Nobody buys a managed knowledge base for its retrieval quality — the ranking gap between these four closed a while ago, and you could match any of them with pgvector and a reranker in a fortnight. What you are actually buying is the connector that copies SharePoint's permissions along with SharePoint's files, and the query path that enforces them per end user. Which makes the one question worth asking before signing anything: does the agent integration carry the end user's identity, or only the API you evaluated? For one of these four, as of today, the documented answer is no.
At a glance
Four managed services that all sell "RAG without the pipeline", with genuinely different centres of gravity. Three are cloud-native and inherit their cloud's identity plane; one is a standalone platform that will run in your VPC or air-gapped.
| Service | Shape | Where permissions come from | Hardest thing to leave |
|---|---|---|---|
| Amazon Bedrock Knowledge Bases | Managed or customer-managed, inside Bedrock | Connector-inherited ACLs, filtered pre-retrieval with a real-time check on top | The IAM and connector wiring |
| Vertex AI Search | Fully managed data stores, Google's search heritage | End-user identity from your IdP; acl_info metadata for imported data | The identity-provider integration |
| Azure AI Search | A search engine you configure, not a pipeline you buy | Permission filters via an end-user token on the query | The index schema and skillset |
| Vectara | Full-pipeline API, SaaS / VPC / on-premises | Platform RBAC and corpus-level scoping you model yourself | Nothing much — which is the product |
Amazon Bedrock Knowledge Bases
What it gives you
Six first-party connectors — S3, SharePoint, Confluence, Google Drive, OneDrive and a web crawler — plus custom connectors and direct streaming ingestion for data that does not live in a crawlable system. Metadata filtering supports boolean, string, double and integer attributes, which is enough to express tenancy, document class and recency without inventing a side channel.
The permission story is the differentiator
The managed variant applies ACL filtering before retrieval and layers a real-time access check on top of it, with pre-filtered documents transient for the life of the API call and never surfaced to the model or the user. That second check matters more than it sounds: pre-retrieval filtering against indexed ACLs is only as fresh as your last sync, and a permission revoked this morning is exactly the case where a stale filter is a disclosure. This is the design argued for in permission-aware retrieval — enforce before the search, and late-bind what reaches the prompt.
Where it costs you
The managed variant removes the vector store and therefore the ability to tune it. The customer-managed variant hands it back — OpenSearch Serverless, Aurora, Neptune — and with it the whole ingestion pipeline, which is a different product with a different operational bill. Choosing between them is the real decision, and teams routinely pick managed for the demo and discover they needed customer-managed for the corpus.
Vertex AI Search
What it gives you
The strongest ranking heritage in the group, unsurprisingly, and the least to configure. Data stores ingest from Cloud Storage, BigQuery, websites and first-party connectors; hybrid retrieval and semantic ranking are on by default rather than assembled. If your evaluation is "does it find the right thing with no tuning", this tends to win it.
Access control is identity-first
Google binds retrieval to the end user through your identity provider: the search app identifies who is asking and filters to what they may see. For data you import yourself — Cloud Storage, structured data — you supply an acl_info field in the document metadata and mark the data store as access-controlled. That is a clean model, and it is also a hard dependency: you are wiring your IdP into the retrieval path, which is a security review, not a config change.
Where it costs you
Least pipeline control of the three cloud services. You cannot swap the chunker, and the embedding model is Google's. Pricing is metered per thousand queries with a separate runtime charge, which is predictable but means the cost of a chatty agent that issues four searches per turn is four times the cost of one that issues one — an argument for agentic retrieval discipline that has nothing to do with quality.
Azure AI Search
What it gives you
Strictly speaking this is the odd one out: it is a search engine with an indexer framework, not a managed RAG pipeline. You get more control than the other three combined — custom analyzers, skillsets, your own chunking, vector and hybrid profiles you tune — and correspondingly more to own. For teams who already know what their retrieval should look like, that is the point.
Document-level access control, and the gap
The permission model is genuinely good. Permission metadata is indexed with the content — the SharePoint indexer will ingest ACLs directly — and at query time you pass a user or group token in the x-ms-query-source-authorization header. The service checks that your client holds Search Index Data Reader, then trims results to what that principal may see.
Now the part that should change your evaluation. The Azure AI Foundry Agents SDK's built-in Azure AI Search tool does not expose a way to pass that header or any equivalent per-request security context — issue #44454, opened December 2025, still open and awaiting service-team attention at the time of writing. The native SDK supports permission trimming; the agent tool does not. An agent wired up the documented way queries the index as the service, and the filter has no principal to apply.
This is not an Azure-specific moral. It is the general failure mode of buying a permission model and then reaching it through a convenience wrapper: the control was evaluated on one code path and the product ships on another. If you are on Azure and you need trimming today, call the search API directly from your own tool implementation and forward the token yourself. It is not much code, and it is the difference between a retrieval layer that enforces and one that advertises.
Vectara
What it gives you
The only one here that is not a cloud's captive service. Ingestion, embedding via its own Boomerang model, hybrid retrieval, reranking, generation and citations behind a single API, deployable as SaaS, in your VPC, or fully on-premises in an air-gapped environment. For regulated buyers who cannot put the corpus in a hyperscaler's managed service, that deployment range is the entire reason the company is on the shortlist.
The grounding machinery is the real differentiator
Vectara's Hughes Hallucination Evaluation Model scores whether a generated summary is supported by its sources, and the interesting number is latency rather than accuracy: roughly 0.6 seconds on an RTX 3090 against about 35 seconds for a frontier-LLM judge on a 4,096-token context. That is the gap between grounding as an offline eval and grounding as an inline check you can afford on every answer, which is a different product. The Hallucination Corrector API extends it from detection to rewriting the unsupported span.
Where it costs you
Fewest first-party connectors, so more ingestion is yours to build. Permissions are modelled at corpus and platform-RBAC level rather than inherited per document from a source system, which means multi-tenant document-level trimming is design work you own. And it is sold as an enterprise annual contract — reported to start in the low six figures — rather than metered, so it prices out of experiments and into commitments.
What you are actually locked into
Vector lock-in is the fear everyone names and the one that barely matters: vectors are derived data, and if you kept your source corpus you can re-embed it into anything over a weekend. The reindexing mechanics are well understood — see reindexing and embedding migrations.
Parser lock-in is worse and less discussed. Every one of these services makes its own decisions about how a PDF becomes text, how a table is linearised, how a scanned page is OCR'd, where a chunk boundary falls. Those decisions determine what your retrieval can possibly find, they are tuned per document type over weeks, and none of that tuning transfers. Keep the raw source of everything you ingest, always, or a migration becomes a re-collection. The trade-offs are in document parsing for RAG, and the enrichment you might layer on top is in contextual retrieval.
Permission lock-in is the one that actually holds you. It reaches back into your identity provider, into each source system's ACL semantics, and into a security review that already passed. Replacing it means redoing all three. If you evaluate one layer properly, evaluate this one — and evaluate it on the integration you will ship.
When to pick which
| Situation | Pick | Because |
|---|---|---|
| Enterprise content in SharePoint/Drive/Confluence, per-user answers required | Bedrock KB or Vertex AI Search | Connector-inherited ACLs plus query-time enforcement is the expensive part, and both ship it. |
| You already know your retrieval design and want to tune it | Azure AI Search | Most control in the group — but implement the search call yourself so you can forward the user token. |
| Corpus cannot leave your network, or must run air-gapped | Vectara | The only one with a genuine on-premises path. |
| Grounding has to be checked inline on every answer | Vectara | Sub-second grounding scoring changes what you can afford to check. |
| Public documentation, no per-user permissions | None of them, probably | The permission plumbing is what you are paying for. Without it, pgvector plus a reranker is cheaper and yours. |
| Multi-cloud, or you expect to move | Vectara, or build it | The three cloud services are deliberately load-bearing on their own identity and storage planes. |
The honest default for a team with a normal enterprise corpus: pick the managed service belonging to the cloud your identity provider already trusts, and spend the saved time on evaluation rather than on the pipeline. The failure mode is not choosing wrong — it is choosing on a ranking benchmark and discovering the permission path six months later.
FAQ
Is a managed knowledge base better than building with pgvector?
On retrieval quality, no — a competent hybrid index with a reranker is hard to beat and you control every knob. Managed services win on connectors and permission enforcement, which are the parts that take quarters to build correctly. If your corpus is public or single-tenant, build it.
Do these services support document-level access control?
All four express permissions somehow, but the mechanisms differ sharply: connector-inherited ACLs filtered before retrieval (Bedrock), end-user identity from your IdP plus acl_info on imported data (Vertex), indexed permission metadata plus a per-query user token (Azure), and platform RBAC with corpus scoping you model yourself (Vectara). Test yours with two users and a canary document before you believe any of it.
What is the Azure Foundry Agents SDK gap, exactly?
The built-in Azure AI Search tool in the Agents SDK provides no way to pass x-ms-query-source-authorization or an equivalent per-request security context, so permission trimming that works through the native search SDK does not apply through the agent tool. The workaround is to write your own tool that calls the search API and forwards the token.
How much of my index is portable if I switch?
The vectors: none, and it does not matter. The parsed and chunked text: only if you kept it. The raw source: everything, if you kept that. The permission model: almost nothing. Keep raw sources.
Does contextual retrieval work on top of these?
Only where you control ingestion. Azure's skillsets and Bedrock's custom connectors give you a place to put the preamble; Vertex's managed data stores largely do not. If chunk-level context loss is your main failure mode, that constrains the shortlist.
Further reading
On this wiki:
- Permission-Aware Retrieval — why filtering the answer is not a control.
- Contextual Retrieval — the ingestion-side quality lever and what it commits you to.
- Choosing a Vector Database — the build-it-yourself side of the same decision.
- Hybrid Search & Reranking — the quality baseline any managed service has to beat.
- Build vs Buy — how to price the plumbing you are not building.
Sources:
- Amazon Bedrock managed knowledge base documentation
- Document-level access control in Azure AI Search
- azure-sdk-for-python issue #44454 — the Agents SDK permission-trimming gap.
- Vectara release notes