AI Blog

Bedrock Knowledge Bases vs Vertex AI Search vs Azure AI Search vs Vectara

You are not buying retrieval quality from a managed knowledge base — you are buying the connector that copies SharePoint's permissions along with its files, and the query path that enforces them per user. Azure's Agents SDK search tool still cannot forward that token, and permission lock-in is the layer that actually holds you.

By Agentic AI Wiki 13 min read

Nobody buys a managed knowledge base for its retrieval quality — the ranking gap between these four closed a while ago, and you could match any of them with pgvector and a reranker in a fortnight. What you are actually buying is the connector that copies SharePoint's permissions along with SharePoint's files, and the query path that enforces them per end user. Which makes the one question worth asking before signing anything: does the agent integration carry the end user's identity, or only the API you evaluated? For one of these four, as of today, the documented answer is no.

At a glance

Four managed services that all sell "RAG without the pipeline", with genuinely different centres of gravity. Three are cloud-native and inherit their cloud's identity plane; one is a standalone platform that will run in your VPC or air-gapped.

ServiceShapeWhere permissions come fromHardest thing to leave
Amazon Bedrock Knowledge BasesManaged or customer-managed, inside BedrockConnector-inherited ACLs, filtered pre-retrieval with a real-time check on topThe IAM and connector wiring
Vertex AI SearchFully managed data stores, Google's search heritageEnd-user identity from your IdP; acl_info metadata for imported dataThe identity-provider integration
Azure AI SearchA search engine you configure, not a pipeline you buyPermission filters via an end-user token on the queryThe index schema and skillset
VectaraFull-pipeline API, SaaS / VPC / on-premisesPlatform RBAC and corpus-level scoping you model yourselfNothing much — which is the product
Four managed knowledge-base services across four axes A matrix comparing Amazon Bedrock Knowledge Bases, Vertex AI Search, Azure AI Search and Vectara on breadth of first-party connectors, query-time permission enforcement, how much of the pipeline you can control, and how portable the result is. Each service is strong on a different axis and none is strong on all four. Where each service leans hardest Connectors Query-time ACLs Pipeline control Portability Bedrock Knowledge Bases Strong Strong Medium Weak Vertex AI Search Strong Strong Weak Weak Azure AI Search Medium Strong in the API, not in the agent tool Strong Medium Vectara Weak Medium Weak Strong Strong Medium Weak Portability here means how much survives a move to another vendor
Each one is strong on a different axis, and none of them is strong on all four.

Amazon Bedrock Knowledge Bases

What it gives you

Six first-party connectors — S3, SharePoint, Confluence, Google Drive, OneDrive and a web crawler — plus custom connectors and direct streaming ingestion for data that does not live in a crawlable system. Metadata filtering supports boolean, string, double and integer attributes, which is enough to express tenancy, document class and recency without inventing a side channel.

The permission story is the differentiator

The managed variant applies ACL filtering before retrieval and layers a real-time access check on top of it, with pre-filtered documents transient for the life of the API call and never surfaced to the model or the user. That second check matters more than it sounds: pre-retrieval filtering against indexed ACLs is only as fresh as your last sync, and a permission revoked this morning is exactly the case where a stale filter is a disclosure. This is the design argued for in permission-aware retrieval — enforce before the search, and late-bind what reaches the prompt.

Where it costs you

The managed variant removes the vector store and therefore the ability to tune it. The customer-managed variant hands it back — OpenSearch Serverless, Aurora, Neptune — and with it the whole ingestion pipeline, which is a different product with a different operational bill. Choosing between them is the real decision, and teams routinely pick managed for the demo and discover they needed customer-managed for the corpus.

Vertex AI Search

What it gives you

The strongest ranking heritage in the group, unsurprisingly, and the least to configure. Data stores ingest from Cloud Storage, BigQuery, websites and first-party connectors; hybrid retrieval and semantic ranking are on by default rather than assembled. If your evaluation is "does it find the right thing with no tuning", this tends to win it.

Access control is identity-first

Google binds retrieval to the end user through your identity provider: the search app identifies who is asking and filters to what they may see. For data you import yourself — Cloud Storage, structured data — you supply an acl_info field in the document metadata and mark the data store as access-controlled. That is a clean model, and it is also a hard dependency: you are wiring your IdP into the retrieval path, which is a security review, not a config change.

Where it costs you

Least pipeline control of the three cloud services. You cannot swap the chunker, and the embedding model is Google's. Pricing is metered per thousand queries with a separate runtime charge, which is predictable but means the cost of a chatty agent that issues four searches per turn is four times the cost of one that issues one — an argument for agentic retrieval discipline that has nothing to do with quality.

Azure AI Search

What it gives you

Strictly speaking this is the odd one out: it is a search engine with an indexer framework, not a managed RAG pipeline. You get more control than the other three combined — custom analyzers, skillsets, your own chunking, vector and hybrid profiles you tune — and correspondingly more to own. For teams who already know what their retrieval should look like, that is the point.

Document-level access control, and the gap

The permission model is genuinely good. Permission metadata is indexed with the content — the SharePoint indexer will ingest ACLs directly — and at query time you pass a user or group token in the x-ms-query-source-authorization header. The service checks that your client holds Search Index Data Reader, then trims results to what that principal may see.

Two paths into the same index, only one of which carries the end user's identity Ingest carries the permissions; the query has to carry the principal Source system SharePoint, Drive, Confluence Connector copies content + ACL metadata Index chunks, vectors, permission keys Retrieval: filter, then rank Path A — direct search call App authenticates the end user, forwards a user or group token on the request Filter has a principal to apply Path B — the agent tool wrapper Agent framework calls the same index with the service identity, no end-user context to pass Filter has nothing to apply — the index is wide open The permission model you evaluated lives on Path A; the agent you are building uses Path B Test the integration you will ship, not the API the documentation describes
The permission model you evaluated lives on the direct path. The agent you are building may not be on it.

Now the part that should change your evaluation. The Azure AI Foundry Agents SDK's built-in Azure AI Search tool does not expose a way to pass that header or any equivalent per-request security context — issue #44454, opened December 2025, still open and awaiting service-team attention at the time of writing. The native SDK supports permission trimming; the agent tool does not. An agent wired up the documented way queries the index as the service, and the filter has no principal to apply.

This is not an Azure-specific moral. It is the general failure mode of buying a permission model and then reaching it through a convenience wrapper: the control was evaluated on one code path and the product ships on another. If you are on Azure and you need trimming today, call the search API directly from your own tool implementation and forward the token yourself. It is not much code, and it is the difference between a retrieval layer that enforces and one that advertises.

Vectara

What it gives you

The only one here that is not a cloud's captive service. Ingestion, embedding via its own Boomerang model, hybrid retrieval, reranking, generation and citations behind a single API, deployable as SaaS, in your VPC, or fully on-premises in an air-gapped environment. For regulated buyers who cannot put the corpus in a hyperscaler's managed service, that deployment range is the entire reason the company is on the shortlist.

The grounding machinery is the real differentiator

Vectara's Hughes Hallucination Evaluation Model scores whether a generated summary is supported by its sources, and the interesting number is latency rather than accuracy: roughly 0.6 seconds on an RTX 3090 against about 35 seconds for a frontier-LLM judge on a 4,096-token context. That is the gap between grounding as an offline eval and grounding as an inline check you can afford on every answer, which is a different product. The Hallucination Corrector API extends it from detection to rewriting the unsupported span.

Where it costs you

Fewest first-party connectors, so more ingestion is yours to build. Permissions are modelled at corpus and platform-RBAC level rather than inherited per document from a source system, which means multi-tenant document-level trimming is design work you own. And it is sold as an enterprise annual contract — reported to start in the low six figures — rather than metered, so it prices out of experiments and into commitments.

What you are actually locked into

Three layers of lock-in, ranked by how hard each is to leave Three columns comparing the vector store, the parser and chunker, and the permission model, on what it costs to replace each one. Vectors are the cheapest to rebuild, parsing is weeks of tuning, and the permission model is the layer that reaches back into your identity provider and your source systems. What it actually costs to leave The vector store The parser and chunker The permission model Re-embed the corpus Re-tune extraction per document type Re-plumb identity and every source Cost: compute, measured in hours Cost: engineering, measured in weeks Cost: a security review, measured in quarters Portable: the corpus is yours Portable: only if you kept the raw source Portable: almost nothing transfers Evaluate the layer that is hardest to leave first — it is not the one the benchmarks measure
The layer everyone benchmarks is the layer that is cheapest to replace.

Vector lock-in is the fear everyone names and the one that barely matters: vectors are derived data, and if you kept your source corpus you can re-embed it into anything over a weekend. The reindexing mechanics are well understood — see reindexing and embedding migrations.

Parser lock-in is worse and less discussed. Every one of these services makes its own decisions about how a PDF becomes text, how a table is linearised, how a scanned page is OCR'd, where a chunk boundary falls. Those decisions determine what your retrieval can possibly find, they are tuned per document type over weeks, and none of that tuning transfers. Keep the raw source of everything you ingest, always, or a migration becomes a re-collection. The trade-offs are in document parsing for RAG, and the enrichment you might layer on top is in contextual retrieval.

Permission lock-in is the one that actually holds you. It reaches back into your identity provider, into each source system's ACL semantics, and into a security review that already passed. Replacing it means redoing all three. If you evaluate one layer properly, evaluate this one — and evaluate it on the integration you will ship.

When to pick which

SituationPickBecause
Enterprise content in SharePoint/Drive/Confluence, per-user answers requiredBedrock KB or Vertex AI SearchConnector-inherited ACLs plus query-time enforcement is the expensive part, and both ship it.
You already know your retrieval design and want to tune itAzure AI SearchMost control in the group — but implement the search call yourself so you can forward the user token.
Corpus cannot leave your network, or must run air-gappedVectaraThe only one with a genuine on-premises path.
Grounding has to be checked inline on every answerVectaraSub-second grounding scoring changes what you can afford to check.
Public documentation, no per-user permissionsNone of them, probablyThe permission plumbing is what you are paying for. Without it, pgvector plus a reranker is cheaper and yours.
Multi-cloud, or you expect to moveVectara, or build itThe three cloud services are deliberately load-bearing on their own identity and storage planes.

The honest default for a team with a normal enterprise corpus: pick the managed service belonging to the cloud your identity provider already trusts, and spend the saved time on evaluation rather than on the pipeline. The failure mode is not choosing wrong — it is choosing on a ranking benchmark and discovering the permission path six months later.

FAQ

Is a managed knowledge base better than building with pgvector?

On retrieval quality, no — a competent hybrid index with a reranker is hard to beat and you control every knob. Managed services win on connectors and permission enforcement, which are the parts that take quarters to build correctly. If your corpus is public or single-tenant, build it.

Do these services support document-level access control?

All four express permissions somehow, but the mechanisms differ sharply: connector-inherited ACLs filtered before retrieval (Bedrock), end-user identity from your IdP plus acl_info on imported data (Vertex), indexed permission metadata plus a per-query user token (Azure), and platform RBAC with corpus scoping you model yourself (Vectara). Test yours with two users and a canary document before you believe any of it.

What is the Azure Foundry Agents SDK gap, exactly?

The built-in Azure AI Search tool in the Agents SDK provides no way to pass x-ms-query-source-authorization or an equivalent per-request security context, so permission trimming that works through the native search SDK does not apply through the agent tool. The workaround is to write your own tool that calls the search API and forwards the token.

How much of my index is portable if I switch?

The vectors: none, and it does not matter. The parsed and chunked text: only if you kept it. The raw source: everything, if you kept that. The permission model: almost nothing. Keep raw sources.

Does contextual retrieval work on top of these?

Only where you control ingestion. Azure's skillsets and Bedrock's custom connectors give you a place to put the preamble; Vertex's managed data stores largely do not. If chunk-level context loss is your main failure mode, that constrains the shortlist.

Further reading

On this wiki:

Sources: