Azure Cosmos DB for MongoDB vCore, now branded Azure DocumentDB, speaks the MongoDB wire protocol and adds vector search. That makes it look like a drop-in: point the Mongo driver at it, create a vector index, run a query. In practice it has its own index rules, its own query shape, its own validator dialect, and a local emulator that is more forgiving than the managed service. We run it as the vector and document tier under Connect, the agentic scheduling platform we built on top of Syncify, and most of what we learned came from the gap between "it speaks Mongo" and "it behaves like Atlas".
This is an operator's field guide: how the tier is provisioned and kept in shape, what the index and validator rules really are, how we emulate it locally and in CI, and a symptom table for when something goes wrong. The retrieval design itself (chunking, re-ingest, scope kinds, ACLs) is covered in the RAG pipeline article. This article is about keeping the database underneath it healthy.
Why the vectors are not in Postgres
Connect's system of record is Postgres: 120+ tables, direct SQL, migrations, foreign keys. Putting embeddings there with an extension would have been the obvious move. We split the tiers instead, along one line: if it needs a transaction across business records, a foreign key or an audit trail, it goes in Postgres. If it is a large, loosely structured document or an embedding, it goes in DocumentDB.
Three things made that split worth a second engine:
- Isolation by connection. Each workspace gets its own database. The chunk store is constructed with one workspace's database handle, so there is no tenant filter for a query to forget.
- The tier carries agent working state too. The Project Info Sheet's durable agent sessions, transcripts, revisions, pending interactions and sandbox leases live in the same per-workspace databases.
- The tier is optional. With no document database configured, the ingest module is null, retrieval degrades to ungrounded answers, and the rest of the platform runs. Keeping vectors out of the system of record means a slow or absent vector store cannot take the core API down.
The cost is real. There are two engines to operate, back up and emulate. Nothing spans them in one transaction, so writes to the document tier are best-effort and repairable rather than atomic with the Postgres write that triggered them. And reporting queries written in SQL can never reach the document tier. They were the right costs for this workload; they might not be for yours.
Provisioning: no migration runner, so make boot idempotent
Mongo has no migration runner, and we did not build one. Instead, every collection, validator and index is ensured idempotently, so a change ships by being deployed. The sequence for one workspace looks like this:
- If the provisioned memo already holds this workspace, return immediately.
- Create the chunk collection with its validator if it is missing; if it exists, re-apply the validator with a collection-modify command.
- Reconcile the unique identity index by name and key (more on why below).
- Create one single-field index per vector filter path.
- Create the compound full-text index.
- Ensure the Info Sheet collections, validators and per-sheet unique indexes the same way, dropping index names that earlier key shapes used.
- Try to create the vector index (HNSW, then IVF), at most once per server.
- Add the workspace to the memo.
The server runs this as a background sweep after it starts listening, eight workspaces at a time. Blocking boot on hundreds of round trips to the document cluster would turn a slow document database into a total outage of an API that mostly does not need it. Because the sweep is backgrounded, a request can arrive for a workspace the sweep has not reached, so every accessor calls the same ensure step on demand rather than assuming the sweep finished. The memo keeps that cheap after the first call.
The sweep reports per-workspace outcomes: "provisioned N of M workspaces", plus one warning if any databases lack a vector index. A single pass/fail or a timeout tells an operator nothing. "N of M" tells them how big the problem is.
Idempotency traps we hit
Writing "create if missing" for a document database sounds trivial. These are the cases that made it not trivial:
- Create races. Two callers can both see "no collection" and both try to create it. The second gets an "already exists" error, which the ensure step treats as success and follows with a validator re-apply.
- Same name, different key. Mongo will not replace an index just because you call create again with a different key. Worse, a create whose name matches an existing index with a different key fails, and if that failure is not handled it takes provisioning down for the whole workspace. So named indexes are reconciled: read the existing index, compare its key field by field and in order, drop it if it differs, then create.
- Renamed indexes. When the Info Sheet moved from one sheet per project to many, the old unique index would have rejected a second sheet. The ensure step drops stale names, tolerating "index not found", before creating the new ones.
- Destructive cutovers. When a workspace's Info Sheet collections are deliberately wiped for a data-model cutover, the code clears that workspace from the provisioned memo and re-runs the ensure step immediately, so indexes are rebuilt in the same request rather than at the next restart.
The validator dialect is not full JSON Schema
Every collection has a $jsonSchema validator, and on this engine the validator does real work. It is also a narrower dialect than you might expect:
- No
additionalProperties. Shape is enforced with the required list and field types only. - No
enum. Lifecycle values (phases, statuses) are validated at the application's schema boundary, not in the database. - Types need care. Counters accept both 32- and 64-bit integers, because drivers pick either depending on the value.
Two validator rules carry more weight than they look. The chunk's project field is required but nullable, so "this chunk is workspace-wide" is always an explicit claim by the writer and a missing field is impossible. And the chunk ACL must have at least one item, because an empty array never matches an $in, which would make the chunk silently unretrievable. The RAG article explains why the scope filter depends on both.
One more limit shapes collection design: the 16 MB document ceiling. The Info Sheet's agent session stores one session item per document, keyed by conversation and sequence, so a long conversation can never grow a single document toward the limit.
Indexes: what each one is for
| Index | Key | Why it exists |
|---|---|---|
| Chunk identity (unique) | workspace, project, source type, source id, chunk index | Upserts are idempotent, and one source can exist in several project scopes |
| One per filter path | workspace, project, embedding model, ACL, source type (each alone) | Filtered vector search requires a single-field index on every filter path |
| Full text | text and title together | Keyword search for spec section and RFI numbers; one text index per collection |
| Vector | embedding, cosine, 1,024 dimensions | Approximate nearest-neighbour search |
| Per-sheet unique keys | project and sheet, plus sequence or item where relevant | Info Sheet sessions, transcripts and revisions |
| Lease expiry | expiry time | Finding stale sandbox leases |
The filter-path rule is the one that surprises people. A compound index over the same fields is not accepted for a filtered vector search. The server reports that the index for the filter path was not found and the query fails outright. That rule is written next to the index loop so nobody "optimises" the five indexes into one.
Choosing the vector index
The vector index is created with a fallback chain:
| Attempt | Kind | Parameters | When it works |
|---|---|---|---|
| 1 | HNSW | m 16, efConstruction 64, cosine, 1,024 dims | Larger cluster tiers; the documentation is inconsistent about which |
| 2 | IVF | one list, cosine, 1,024 dims | From the smallest tiers; honest below roughly ten thousand vectors |
| 3 | None | Exact scan in process | Any server; capped and logged |
HNSW's m of 16 and efConstruction of 64 are moderate defaults that build quickly. IVF with a single list is effectively a flat index behind an index's interface, which is honest for a small corpus; more lists are a tuning step to take on evidence.
Operationally, three rules matter more than the parameters:
- Probe once per server, not per database. Vector index support is a property of the cluster tier. Retrying both index kinds for every workspace means two doomed round trips and a log paragraph per workspace, which buries everything else in the boot log. The first result is cached for the process.
- Log the kind actually built. Dev and production may disagree (the emulator, a small dev cluster and a production tier can each land on a different rung), and nothing else would reveal it. The boot log states the index kind and dimensions once.
- Truncate the failure. A rejected index creation echoes the entire index specification back in the error, more than once. Only the first failure is logged, cut to a few hundred characters.
The dimension count, 1,024, is a single constant kept next to the index definition, and every search rejects a query vector of the wrong length before it reaches the database.
Query shape: cosmosSearch is not $vectorSearch
The single most misleading thing about this engine is that Atlas documentation ranks highly for "Mongo vector search". Atlas uses a $vectorSearch stage. vCore uses $search with a cosmosSearch operator, and the filter goes inside that operator, which is how vCore applies it before ranking. Written from scratch, the shape is:
{ $search: { cosmosSearch: { vector: q, path: "embedding", k: 8,
filter: { embedding_model: { $eq: m },
project_id: { $in: [p, null] } } } } }
{ $project: { text: 1, title: 1, score: { $meta: "searchScore" } } }
Three notes follow:
- The filter grammar is limited. Filtered vector search accepts comparison operators,
$in,$ninand$regex, but not$or. Any "this or that" has to be expressed as$inover values of one field. - Normalise the score. Cosine similarity comes back between -1 and 1. Connect converts it to a distance (one minus the score) so every caller sees the same scale as pgvector's cosine operator, and the exact-scan fallback produces the same numbers.
- Keyword search is separate.
$textruns against the same filter; a missing text index returns nothing and a warning, never a regex fallback.
The exact-scan fallback scores up to 20,000 filtered candidates in process. When it hits that cap it warns that the results are a subset, rather than ranking part of the corpus as if it were all of it.
Local emulation and CI
The open-source DocumentDB project publishes a local container image, and Connect's local compose file runs it alongside Postgres, Azurite and a Key Vault emulator. Details that took some iteration:
- Loopback only. Every local service binds to 127.0.0.1. Nothing on a developer's network can reach the emulator.
- TLS is on. The container serves a self-signed certificate, so the local and test connection strings explicitly allow it. That relaxation lives only in local configuration, never in deployed settings.
- A real health check. The compose health check runs
mongoshinside the container and pings the server. Port-open is not the same as ready; the emulator takes a while to start. - Skip the sample data. The container is started with its init data disabled, so tests start from an empty server.
In CI, the same image runs as a plain container rather than a workflow service, because it needs command arguments the services block cannot express. It boots while dependencies install, a wait loop pings it until healthy, and an environment flag un-skips the document-tier integration suites. Without the flag those suites skip cleanly, so a contributor without Docker can still run the rest.
The integration suite tests behaviour, and its cases read like an operator's checklist: provisioning is idempotent, re-ingesting replaces rather than accumulates, search filters before ranking in whichever mode is available, and the exact fallback ranks the same way the index does.
Where the emulator lies
The local container accepts $or inside a filtered vector search. The managed service does not. So no integration test against the emulator would catch someone rewriting a scope filter with $or. The guard is a plain unit test that builds every scope filter and asserts it uses $in and never $or. When the emulator is more permissive than production, encode the production rule as a test that does not need either.
Symptoms and causes
| Symptom | Likely cause | Where to look |
|---|---|---|
| Vector query fails with "index for filter path not found" | A filter field has no single-field index, or only a compound one | The per-path index loop; add the field there |
| Workspace provisioning fails after a deploy | An index name was reused with a different key | Reconcile by name and key, or drop the stale name first |
| Boot log says no vector index could be created | Cluster tier rejects both HNSW and IVF | Tier choice; retrieval is on the capped exact scan |
| Warning that the exact scan hit its candidate cap | No vector index and a large workspace | Fix the index; results are a subset until then |
| Search returns results that are confidently wrong | Query and chunks embedded by different models, or wrong dimensions | The embedding-model filter and the dimension check |
| Keyword search always empty, with a warning | Text index missing on that workspace | The ensure step's log for that workspace |
| A document insert is rejected | Validator mismatch: wrong integer type, missing required field | The collection's validator; check the writer |
| Works locally, fails in Azure | Emulator accepted something the service rejects | Filter operators and index shapes first |
Why this matters if you're building something similar
- Decide the tier split by consistency needs, not by fashion. Records that need transactions and audit stay in the relational store; embeddings and loose documents can live elsewhere if isolation and optionality are worth a second engine.
- Make provisioning idempotent and background it. Reconcile indexes by name and key, tolerate create races, and never block API boot on a slow optional tier.
- Report outcomes as counts. "N of M workspaces provisioned" and "K databases without a vector index" are actionable. A timeout is not.
- Index every filter path on its own. On vCore, compound indexes do not satisfy filtered vector search.
- Use cosmosSearch, not Atlas examples, and keep filters to the supported operators. Write
$inwhere you would reach for$or. - Treat the emulator as a convenience, not an oracle. Where it is more permissive, pin the production rule with a unit test.
- Log what you actually built. Index kind, dimensions and fallbacks, once per process, so dev and production differences are visible.
Where to go next
- How we built Connect: architecture of an enterprise AI platform, the series overview.
- The RAG pipeline for construction documents, the retrieval design this tier serves.
- Postgres schema design for enterprise SaaS, the other half of the data split.
- Deploying a Node monolith to VMs with CI and rollback, where the local services and CI containers fit.
- Prompt caching strategy for AI agents, what happens to retrieved passages once they reach the prompt.
Planning a vector store for your own documents and not sure which engine fits? Scope it with us.
Frequently asked questions
Why store vectors in Cosmos DB for MongoDB vCore instead of Postgres?
Connect keeps anything that needs transactions, foreign keys or audit in Postgres, and puts embeddings and loosely structured agent documents in DocumentDB. That gives one database per workspace for isolation and keeps the vector tier optional, at the cost of operating a second engine.
Why does my vCore vector query say the index for a filter path was not found?
Filtered vector search requires a single-field index on every field used in the filter. A compound index over the same fields does not satisfy it, so the query fails outright.
Can I use $or in a cosmosSearch filter?
No. The managed service accepts comparison operators, $in, $nin and $regex in a filtered vector search, but not $or. Express alternatives as $in over values of one field. The local emulator accepts $or, so guard this with a unit test.
Should I use HNSW or IVF on DocumentDB?
Try HNSW first, then fall back to IVF if the cluster tier rejects it. IVF works from the smallest tiers and is fine for small corpora. Connect probes once per server, logs which kind it built and falls back to a capped exact scan if neither works.
How do you test DocumentDB vector search locally and in CI?
Connect runs the open-source DocumentDB local container on loopback with a real health check, and runs the same image in CI as a plain container. An environment flag un-skips the integration suites, which test behaviour such as idempotent provisioning and filtering before ranking.
Next step
Have a workflow in mind?
Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
Building a RAG Pipeline for Construction Documents: Chunking, Incremental Re-ingest and Scoped Vector Search
An engineering walkthrough of the retrieval pipeline under Connect: what triggers ingestion, how text is chunked and re-ingested cheaply, how chunks are stored one database per workspace, and how scope filters, index fallbacks and an admin browser keep retrieval correct and inspectable.
Custom applications · October 9, 2026
Postgres Schema Design for an Enterprise AI Platform: 120+ Tables, No ORM, Verified Migrations
A tour of the Postgres schema behind Connect, why it uses direct SQL instead of an ORM, how it splits work with DocumentDB, and the migration discipline that stops the server booting on a mismatched database.
Playbooks · October 9, 2026
Deploying a Node Monolith to VMs: systemd, CI Service Containers and Schema-Safe Rollback
The deployment pipeline behind Connect: three systemd processes from one codebase, Key Vault secrets, emulator-backed CI with a Playwright smoke test, and an honest case for VMs over Kubernetes at this stage.