Build Flows

Playbooks · October 9, 2026 · 11 min read

Architecture Tests: Enforcing Codebase Boundaries So Small Teams and AI Agents Can Move Fast

The tests that govern Connect's structure: import layering, a route manifest checked for equality, catalog parity, agent tool rules, design tokens, migrations and a rehearsed deploy script.

By Charley Forey, founder of Build Flows

Every codebase starts with an architecture diagram and a set of rules: routes call services, services call the database, one module does not reach into another module's tables. The diagram is accurate for about a month. After that, someone under deadline imports a store directly because it is faster, the reviewer misses it, and the rule becomes a suggestion. A year later, the diagram describes a system that no longer exists.

When we built Connect, the AI platform on top of Syncify, we had a small team, a large surface (hundreds of routes, 120+ tables, about 200 MCP tools plus the chat agent's own) and a lot of the code written with AI coding agents in the loop. Rules that live in people's heads do not survive that combination. So we wrote the important ones down as tests. This article walks through the tests that govern Connect's structure, what each one protects, and the real bugs that made us write several of them.

The idea: if a rule matters, a test enforces it

The repo's agent instructions file (the AGENTS.md / CLAUDE.md that every coding agent reads before it touches the code) says it directly: the layering is enforced by a test, not by memory, and a wrong import fails the build. npm run check runs typecheck, the full test suite (over a thousand tests) and the production build, and nothing is handed off until it passes.

That framing changes what a "rule" is. A rule is no longer a paragraph in a wiki. It is a test that names the rule in its title, explains in its doc comment why the rule exists, and produces a failure message that tells the developer (or the agent) what to do instead. Many of Connect's tests are not testing behaviour at all. They test the shape of the code.

Layering: routes, module, store

Each capability under src/server/ (files, schedules, apps, chat, and dozens more) follows the same shape:

routes.ts  ->  module.ts  ->  *-store.ts / external adapters

Routes handle HTTP and validation. The module holds business rules. Stores hold SQL. test/architecture.test.ts walks the server source, reads every relative import with a regular expression, resolves it to a path, and checks a short list of rules:

RuleWhy it exists
Stores import neither routes nor modulesThe dependency arrow only points down
Routes never import a store directlyBusiness rules live in the module, so bypassing it bypasses the rules
A capability never imports another capability's storeCross-capability reads go through the other module, which owns its invariants
platform/ imports nothing above itIt is the floor: config, Postgres, blob storage, signing
The connector library never imports connections/Provider quirks must not leak into the workflow that uses them
Anything with routes or a store has a module.tsNothing gets an HTTP or data surface without a seam in front of it
Agent definitions live under agents/, and the agent runtime imports no agentThe runtime stays generic
The agent-safety capability owns no routes and no storeIt is imported by many agent capabilities, so it must stay a leaf

The cross-capability store rule is the one that pays for itself most often. Connect's authorization hooks check membership through the workspaces module, not its store, and a comment at that call site says why: a feature must not read another feature's tables. Without the test, the shortcut would have been taken many times by now, and every one of them would have been a place where a workspace rule (soft deletes, scoping, audit) was skipped.

The implementation is deliberately low-tech: no dependency-graph library, about a hundred lines of Node. It runs in milliseconds, and the failure names the file, the import and the fix.

The route manifest

src/server/routes.ts begins with MANIFEST: one [method, path, authLevel] row for every route in the API, several hundred in all. test/routes.test.ts builds the real app with its real route wiring and checks the manifest against what actually registered:

  • The registered set equals the manifest. Not "is a subset of". A new route that is not in the manifest fails, and so does a manifest row for a deleted route.
  • Feature-gated routes register at the declared level. The helper that wraps a group of routes in a feature gate tags each route with the gate it registered under, and the test compares that tag with the manifest. We added this after finding a handler registered in a lower-level group than its manifest row declared, a mismatch the earlier method-plus-path comparison could not see.
  • Every ID in a path is classified. Each :param must appear either in the map of project-owned resources (which drives per-project access checks) or in an explicit "not project-owned" list with a written reason. We added this one when we moved project scoping from the URL to the resource: by-ID routes like /notes/:noteId carry no {projectId}, so every ID parameter now has to resolve to its owning project. Both lists are checked in reverse too, so removing a route leaves no stale entry.
  • Every non-public route returns 401 with no session, and a member with no roles gets 403 from every admin and feature-gated route, driven through the real hooks with a stubbed session.

More on the RBAC model this protects in multi-tenant RBAC with a single authorization chokepoint.

RBAC catalog parity and feature packages

Connect's feature catalog exists twice on purpose. The server owns the keys, because route wiring references them and a typo should fail typecheck. The client owns labels, icons and paths, because those are UI concerns. Two lists that agree only by hand will drift, so test/rbac.test.ts reads both files as text and asserts that the key lists match. It reads them as text rather than importing them because the client module pulls in icon libraries and path aliases that do not resolve under the test runner.

The same file pins product decisions as tests: what a new workspace is entitled to, that the seeded roles cascade (Guest is a subset of Member, which is a subset of Admin), that a page whose own data reads are admin-only is not offered at normal level, and that a workspace path resolves to the feature that owns it.

test/feature-packages.test.ts guards the package model. Every feature key must belong to exactly one package, every requires must name a real package, and requirements resolve transitively and stably. The most interesting test checks that each requires edge still matches the wiring it was derived from. If the Schedule Builder stops taking the schedules module as an input, the "Schedule Builder requires Scheduling" edge has become a rule nobody is enforcing, and the test says so.

On the client side, test/client-routes.test.ts makes sure every feature's route module is actually mounted (forgetting one line would 404 a page with no explanation), and that every workspace page is owned by a feature or is explicitly on an un-gated list. Otherwise a new page could render for anyone in the workspace because the route guard did not know which feature it belonged to.

Agent tools: the manifest for the second door

An AI platform has a second entry point into the same capability modules: the agent's tools. test/agent-tools.test.ts treats the tool manifest the way routes.test.ts treats the route manifest:

  • Tool names are unique, and every tool has a description and an owning capability.
  • Every tool is callable by at least one real agent.
  • The chat agent's manifest and its tool implementations are in lockstep, so the manifest never advertises a tool the agent cannot run.
  • Only the reviewed write tools are flagged as direct mutations. The test holds an exact list, with a comment explaining why each one is safe (each writes a draft or a validated state transition, never something client-facing). Any new mutating tool trips the test and gets a deliberate review.
  • Every tool's capability either names a feature or is deliberately un-gated. The capability-to-feature mapping has no table: a tool's capability simply is a feature key where one exists. That is only safe while the set that maps to nothing stays deliberate, because a typo would land there and silently un-gate the tool. The test pins that set to exactly two entries.
  • The agent is offered only the tools its caller can reach. A caller with no feature access gets no tenant-data tools. Reaching one feature does not unlock another.

That last pair matters because, without it, RBAC becomes advisory. A member without schedule access would just ask the assistant. The detail is in agent tool manifest and permissions.

Design system tokens, both directions

DESIGN.md at the repo root has machine-readable frontmatter that lists the colour roles, radii, elevation, motion and type scale. test/design-system.test.ts parses that frontmatter and the CSS theme and asserts they agree in both directions: every colour role in the doc has a CSS variable and a dark-mode value, and every CSS colour variable has a row in the doc. The same goes for radii, shadows, motion and typography.

Other tests in that file check that every theme token has a consumer, that product code uses only the role vocabulary rather than raw colours, that shared primitives own their sizes, that dense record views use semantic tables rather than stacks of cards, and that token pairs keep readable contrast. The design doc cannot quietly go stale, and the CSS cannot grow a token the doc does not describe.

Migrations and comment hygiene

Two smaller families of tests guard things that fail silently.

Migrations. The migration runner sorts by filename and records each file by name. architecture.test.ts requires three-digit, zero-padded, unique prefixes in lower snake case, because two migrations sharing a number run in alphabetical order of whatever follows the number, which is arbitrary. test/migrations.test.ts stops new number collisions (old ones are grandfathered, because renaming a migration breaks every database that already recorded it), refuses any table created by more than one migration, and checks that a database from before a squash is told to re-stamp its ledger rather than rebuild. There is also a check that no source file contains raw control bytes. That came from a real incident: literal NUL characters in one module made grep treat the whole file as binary, so it vanished from every code search without anyone noticing.

Comment hygiene. test/comment-hygiene.test.ts scans comments for pointers a reader cannot follow: references to deleted planning documents, bare migration numbers (ambiguous because numbering restarted after squashes), eval run IDs, and PR numbers. History belongs in git and in the decisions log. A comment should state the reason inline.

Deploy rehearsal

Connect deploys to VMs with a backup-and-restore script. The riskiest lines in it are a handful of mv calls that only ever run on a real host, with elevated privileges, after the service is already stopped. That is the worst possible place to find out one of them is wrong.

test/deploy.test.ts asserts what the script says. test/deploy-rehearsal.test.ts asserts what it does: it runs the real remote script against a temporary directory tree that looks like a host mid-deploy, with sudo, systemctl, curl and npm replaced by shims on the PATH. Then it plays out the scenarios:

ScenarioExpected outcome
Healthy deployStaged build installed, both services restarted
Migration failsPrevious install restored, nothing restarted
Health check failsPrevious install restored and restarted
First deploy, nothing installedSucceeds
First deploy fails healthCandidate kept for forward recovery
Worker inactive while API is healthyRolls back anyway
Restart failsPrevious deploy restored and restarted

Nothing touches the network or a real service. The full pipeline is in deploying a Node monolith to VMs with CI and rollback.

Why tests-as-governance works for small teams and AI agents

Code review is good at judging whether a change is sensible. It is bad at noticing what is missing: the route that was never added to the manifest, the import that crossed a boundary three files down, the token in the CSS that the design doc never mentions. Tests are the opposite. They are bad at judgement and tireless at noticing.

That split suits a small team: nobody has to remember every rule, and newcomers learn them by reading failure messages.

It suits AI coding agents even better. An agent will happily take the shortest path to a passing feature, and an import that reaches straight into another capability's store is the shortest path. A wiki page does not stop it. A failing test with a message like "cross-feature reads go through the other feature's module" does, and the agent fixes it on its own because the instruction is right there. Connect's agent instructions file tells agents to run the check before handing off, so architecture tests become the guardrails the agent works inside rather than a review comment that arrives a day later.

A few things make these tests work in practice:

  • Write the reason into the test. Most of Connect's structural tests have a doc comment that describes the bug that prompted them. The next person to hit the failure knows why the rule exists before they argue with it.
  • Make the failure message the fix. "X imports Y directly; go through module.ts" is an instruction. "assertion failed" is a puzzle.
  • Check both directions. Manifest equals registered routes, catalog A equals catalog B, doc tokens equal CSS tokens. One-way checks let stale entries pile up.
  • Prefer explicit exemption lists over clever inference. "Not project-owned" with a reason per entry, an exact list of mutating tools, an exact set of un-gated capabilities. Adding to the list is easy, but it shows up in the diff.
  • Keep them cheap. Most of these tests read files as text and run in milliseconds. If governance slows the suite down, people will find ways around it.

The cost is real but small: a test file per rule, and occasionally a test that needs adjusting when a rule legitimately changes. In return, the architecture you designed is still the architecture you have a year later.

Why this matters if you're building something similar

  • Pick the three or four rules whose violation would hurt most (layering, auth on every route, no cross-module data access) and write those tests first.
  • Put a route manifest in your API and test it for equality, not inclusion.
  • If you keep two lists in sync by hand, write the test that compares them today.
  • If AI agents write code in your repo, point their instructions at the check command and make the failure messages self-explanatory.
  • Rehearse your deploy script against a fake filesystem before you trust it on a real host.

Where to go next

Want a codebase that stays maintainable as AI agents help write it? Plan your build with us.

Frequently asked questions

What is an architecture test?

A test that checks the shape of the code rather than its behaviour, for example that route files never import a database store directly, or that every API route appears in a manifest with an auth level.

Do you need a special library to enforce layering?

No. Connect's layering test is about a hundred lines of Node that read each server file, extract relative imports with a regular expression, resolve them, and assert the rules. It runs in milliseconds.

Why do architecture tests help with AI coding agents?

Agents take the shortest path to a passing feature, which often crosses a boundary. A failing test with a clear message stops that immediately, and the agent can fix it without a human reviewer spotting it first.

How do you test a deploy script without a server?

Run the real script against a temporary directory tree shaped like a host, with sudo, systemctl, curl and npm replaced by shims on the PATH, and assert the outcome of healthy, failed-migration, failed-health and failed-restart scenarios.

Should every rule become a test?

Start with the few whose violation would hurt most, such as layering, auth on every route and no cross-module data access. Add more when a real bug shows a rule was being broken silently.

Next step

Want help running this playbook?

Bring the report or workflow. We'll help map the work behind it.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning