Skip to content
Platform hub: modern secure AI-first architecture

An AI-first architecture, identical whether it runs in our cloud or inside your air gap

BigID is cloud native from the ground up, portable to cloud VPCs and data centers, with the AI models decoupled from the platform and the APIs built as MCP interfaces so agents and headless operation are native rather than bolted on. The result is one software build with six ways to run it, from multi-tenant SaaS to a fully air-gapped deployment, and the same discovery, classification, remediation, and AI governance in every one.

Deployment
One build, six ways
Multi-tenant SaaS through fully air-gapped on-premises, with equivalent capability in every mode
Data and AI residency
30+ countries
Data ringfenced inside the country or region your regulator names, on the same platform
Programmatic surface
100% API and MCP
Every function reachable by API and by MCP, with fine grained RBAC on each service
Why architecture became the buying question

Residency law, vendor concentration, and federal isolation all land on the architecture

Data security platforms used to be judged on features and scan coverage. Three things moved that judgement down a layer: regulators started naming the country data has to sit in, boards started treating dependence on a single model provider as concentration risk, and federal and defense programs started requiring isolation the SaaS control plane cannot offer. All three are answered by how the software is built, not by what is on the feature list.

Sovereignty is a deployment question

The moment governance depends on someone else's model, or its logs live in someone else's cloud, the control you set out to keep has already moved. So the air-gapped build has to be the same build.

AI-first is an interface question

APIs became MCP interfaces, with tools rather than read-only reporting. That is what lets an agent operate the platform under a scoped identity instead of screen-scraping a console.

Scale is an elasticity question

Outposts became ephemeral and Kubernetes native, hundreds to a cluster. Capacity follows the workload, so a bursty quarter-end scan does not set the bill for the rest of the year.

See how it is built

Watch it work

Inside BigID's modern secure AI-first architecture

The architecture the product was rebuilt on: cloud native and portable, models decoupled, MCP native interfaces, and outposts that do the processing where the data already sits.

Architecture
In this hub
Start here

Six questions, and where BigID answers each one

Can we run this where our regulator says our data has to stay? Multi-tenant SaaS with full data segregation options, data ringfenced inside any of 30+ countries, and single tenant full isolation set as a deployment flag rather than sold as a different product. Deployment modes → Do we get less product if we deploy in our own VPC, or air-gapped? Self-managed in a VPC and fully air-gapped on-premises run the exact same software as the SaaS, with equivalent discovery, classification, remediation, and AI governance, no phone-home telemetry, and no dependency on a hosted API to function. Deployment modes → What is this costing us while nothing is being scanned? Cloud outposts spin up and down dynamically with the workload, so capacity is not sized and paid for at the peak. On-premises outposts are fully Kubernetes native and scale to hundreds in a cluster. Local processing → Does our data leave our environment to get classified? No copying and no backhaul. Processing stays local to the outpost, side and snapshot scans run behind a security layer that isolates the customer from the vendor, and snapshot scanning does not require admin level access. Local processing → Can our agents operate BigID, or only read from it? MCP plus tools, so the surface goes past reporting into operations, and those tools are extensible by customers and partners. Agentic operation lets business users request and run the work from inside Claude, GPT, Copilot, or Gemini, scoped by their role. API and MCP → What stops an agent from reaching more than it should? Fine grained RBAC on every API and MCP service, so what an application, an agent, or an AI user can see and do is scoped. Agents never reach data directly: they reach it through the identities associated with them, and every interaction is logged. API and MCP →

One build, six ways to run it, nothing dropped on the way down

The usual shape of this is two products: a full SaaS platform, and a smaller on-premises build that trails it by a release or three. That is the version that fails a sovereignty review, because the deployment under the most scrutiny ends up with the least product in it. BigID separates where the software runs from what the software does.

The sovereign build is a different build, so the deployment that needed the most scrutiny is the one running last quarter's product.

The governance tool phones home, and telemetry, logs, or model calls cross the boundary you just spent a year defining.

Scanners are always on, so capacity is sized for the heaviest week of the year and billed for the other fifty-one.

What ships One software build cloud native, fully portable Decoupled AI models swappable, bring your own MCP and API native headless and agent first Ephemeral outposts k8s, hundreds per cluster no phone-home telemetry Same software, every row Six deployment modes Multi-tenant SaaS full data segregation options Residency ringfenced SaaS data held inside 30+ countries Single tenant, full isolation a deployment flag, not a variant Self-managed in your VPC your cloud account, your controls FedRAMP environment public sector and defense Fully air-gapped on-premises zero outbound connectivity customer control increases down the column Identical in all six Discovery and classification same depths, same categories Remediation and AI governance same policies, same RBAC API, MCP, and agent ops same surface, same tools no hosted API dependency to function
The left column is what BigID ships, once: a cloud native build with the AI models decoupled so classification and analysis can run on a model you have already approved, interfaces exposed as MCP and API rather than a console with an integration layer, and outposts that do the processing wherever the data already is. The middle column is the choice you make about where that runs, and the six rows are ordered by how much of it you hold rather than by how much product you get. The right column is what does not move as you go down: a fully air-gapped deployment classifies against the same categories, enforces the same policies, and exposes the same MCP tools as the multi-tenant SaaS.
Capabilities

Where it runs, how it scales, what it exposes, and how it is secured

Seven groups, working outward from the deployment decision: where the software runs, where the processing happens, what the programmatic surface looks like, how it is operated day to day, what comes out of it, what an agent can do with it, and what holds all of that shut.

Same software, any deployment

A cloud-native SaaS that deploys more than one way, so scale and sovereignty are not a trade against each other. The isolation level is configuration; the product is the same product.

  • Multi-tenant SaaS with full data segregation options
  • Data ringfenced within any of 30+ countries for residency and sovereignty requirements
  • Single tenant full isolation is a deployment flag, not a separate product line
  • Self-managed in your own VPC, running the exact same software
  • Fully air-gapped on-premises, running the exact same software
  • FedRAMP environments run that same software too
  • Equivalent discovery, classification, remediation, and AI governance in every mode
  • No phone-home telemetry and no dependency on a hosted API to function
  • AI models decoupled from the platform, so they can be swapped or supplied by you
  • Bring your own AI, so nothing depends on a third party model to do the work

Outposts and local processing

The outpost is where the reading happens, next to the data rather than in a vendor cloud. They were rebuilt as ephemeral in the cloud and Kubernetes native on premises, which is what makes bursty and petabyte workloads survivable.

  • Cloud outposts scale to hundreds per cluster natively
  • In the cloud, outposts spin up and down dynamically, so nothing runs while nothing is being scanned
  • On-premises outposts are fully Kubernetes native and scale to hundreds in a cluster
  • Every outpost supports hundreds of different data sources and formats natively
  • Every outpost can run metadata only scans, smart sampling, configured sampling, and full scans
  • Side and snapshot scans run behind a proprietary security layer that isolates the customer from the vendor
  • Separate local authentication for each local outpost
  • Processing stays local: no data copying and no backhaul

API and MCP architecture

The APIs were rebuilt as MCP interfaces so headless and agent-first operation is native. MCP made BigID accessible to AI; tools on top of it are what make BigID operable by AI, under permissions you set.

  • Full API coverage across the platform
  • Cloud and local MCP server options, so the interface can sit inside your boundary
  • MCP plus tools, so access reaches past reporting into operations
  • Tools are extensible by customers and by partners
  • Fine grained RBAC on every API and MCP service, scoping what applications, agents, and AI users can reach and do
  • Agents reach data through the identities associated with them, never directly
  • Context, not raw data: the interface delivers metadata and classification insight into AI workflows
  • Audit logging on every AI interaction: who accessed what, when, and how

Enterprise-class operations

Scanning is an operational system, not a project, and the tell is whether the settings a change board asks about actually exist. Frequency, blackout windows, exclusions, and triggers are all configuration, with telemetry on the run and an audit trail on the change.

  • Rapid setup and dynamic data discovery
  • Scan configurability including sampling rate, frequency, blackout windows, data exclusions, event triggers, and change based scans
  • No data copying and no backhaul
  • Full localization of processing with outposts
  • Delegated review and remediation, always under RBAC restrictions
  • Full scan and accuracy telemetry, while the scan runs rather than after it fails
  • Integrated agentic ops automation
  • Change audit across configuration and policy

Reporting and investigations

The inventory is only useful where the decisions get made, which is rarely inside the security console. Reporting runs to the warehouse, the BI tool, and the assistant, on the same data the scan produced.

  • In-product reporting, customizable by persona
  • Native support for Tableau and Power BI
  • Integrated ETL into data warehouses and data lakes for advanced analytics
  • Bespoke reporting in plain language through Claude, Copilot, and GPT
  • Breach investigation: impact radius and data flow inference on a real incident
  • Data and access exposure investigation, using the same AI features

Agentic operationalization

The layer that makes the platform operable from the assistant a team already works in. Any agent, all of BigID, with what each one can do bounded by the role and permissions behind it.

  • Business users request and operate the product from inside AI, according to role and permissions
  • Works with Claude, GPT, Copilot, and Gemini
  • AI workflows automate remediation and surface it in the tools people already watch, including Teams
  • Native BigID agents for simplified automation and advanced analysis
  • Simplified third party integration across the rest of the stack

Product security

A data security platform gets read as part of the attack surface, and it should be. This is the group a security architecture review actually opens on: keys, credentials, access paths, and what the vendor can see.

  • FedRAMP authorized
  • Broad certification coverage, including SOC 2, ISO 27001, PCI DSS, and HIPAA
  • Full bring your own keys support
  • No copying and no backhauling of customer data, on premises or in the cloud
  • Review and remediation always RBAC protected
  • Broad password vault support, so scan credentials stay where your team keeps them
  • 100% air-gapped deployment option for AI sovereignty use cases
  • Secure snapshot scanning that does not require admin level access
  • Separate local authentication for local outposts
Verification

Certifications, controls, and the three claims that hold in every deployment mode

True in every deployment mode

  • BigID can govern data and AI entirely inside a customer's own environment, with no dependency on a third-party model to perform the work
  • Runs fully air-gapped, with discovery, classification, and governance operating with zero outbound connectivity required
  • The same architecture runs across cloud, on-premises, private cloud, and air-gapped deployments, with equivalent discovery, classification, remediation, and AI governance in every mode
And from outside
  • Frost & Sullivan 2025 Company of the Year for AI Governance
  • Forrester scored BigID the highest possible 5 out of 5 on secure-by-design commitments in The Forrester Wave™: Sensitive Data Discovery And Classification Solutions, Q2 2026
See the Forrester scoring →

What a security review checks

Certifications and authorizations
  • ✓FedRAMP authorized
  • ✓SOC 2
  • ✓ISO 27001
  • ✓PCI DSS
  • ✓HIPAA
Controls on the architecture itself
  • ✓Bring your own keys
  • ✓Broad password vault support
  • ✓No data copying or backhaul
  • ✓Snapshot scanning without admin access
  • ✓RBAC on every API and MCP service
  • ✓Separate local outpost authentication
  • ✓Audit log on every AI interaction
  • ✓100% air-gapped option
Before the architecture review

What security architects ask first

Is the air-gapped version a cut-down build?

It is the same software. Self-managed in a VPC, FedRAMP, and fully air-gapped on-premises all run the build the multi-tenant SaaS runs, with equivalent discovery, classification, remediation, and AI governance. Isolation is set as a deployment flag, so the sovereign deployment is not a second product on a slower release train.

Does BigID depend on a third party model to do the work?

No, the models are decoupled. They can be swapped, or supplied by you, so classification and analysis run on a model your own review has already approved. Air-gapped deployments operate with zero outbound connectivity required, with no phone-home telemetry and no dependency on a hosted API to function, which is what keeps a governance tool from becoming its own concentration risk.

What does exposing BigID over MCP actually open up?

Only what you scope, and the scoping is per service. There is fine grained RBAC on every API and MCP service, so an application, an agent, or an AI user reaches exactly the functions and data its identity allows. Agents reach data through the identities associated with them rather than directly, the interface delivers context and metadata rather than raw data, every interaction is logged, and the MCP server can run locally so the interface stays inside your boundary.

Take it further

Where to go when the architecture is what your review board is asking about

Trust Center / start here
The BigID Trust Center

Security practices, independent assessments, and responsible AI policy in one place, including SOC 2, ISO 27001, and PCI DSS, plus the certification and vulnerability management detail a security review asks for line by line.

Open the Trust Center →

Tell us where it has to run.

A residency rule, a no-egress network, a model your security team has already approved, or a cluster that has to absorb a quarter-end burst. The architecture walkthrough is scoped to the deployment you actually want, and all six are the same software.

Industry Leadership