Build Flows

Playbooks · October 9, 2026 · 11 min read

Deploying a Node Monolith to VMs: systemd, CI Service Containers and Schema-Safe Rollback

The deployment pipeline behind Connect: three systemd processes from one codebase, Key Vault secrets, emulator-backed CI with a Playwright smoke test, and an honest case for VMs over Kubernetes at this stage.

By Charley Forey, founder of Build Flows

Connect is an AI platform with hundreds of routes, about 200 MCP tools, a dozen background job kinds and a React front end. It deploys to Linux virtual machines with rsync, systemd and a shell script. There is no Kubernetes, no container registry and no release artifact. A deploy pushes what is built now, runs database migrations, restarts three services and waits for a readiness check. If anything fails, it rolls back to the previous build, but only after checking that the database still matches it.

That is an unfashionable choice for an enterprise platform, so this article explains what we built, why we think it is the right call at this stage, and what it costs. If you are a small team shipping a serious product, the tradeoffs here may save you months.

One codebase, three processes

Connect is a TypeScript modular monolith: React and Vite on the client, Fastify on the server, Postgres, Azure Blob storage, and an optional DocumentDB tier for vectors and documents. In production it runs as three systemd units built from the same code:

UnitWhat it runsWhy it is separate
connect-serverThe Fastify API and the built React clientRequest latency must not depend on background work
connect-workerThe job loop: syncs, indexing, transcription, updates, evalsA long sync blocks its loop by design; scale by adding processes
connect-buildThe MCP endpoint and the Schedule and App BuildersKeeps long-lived MCP sessions and builder traffic off the main API

Fastify serves the built client from dist/client in production, so the API and the front end ship as one unit with no separate static host. In development, Vite proxies API calls to Fastify instead.

The separation earned its keep early. Schedule Builder turns take minutes of model calls and CPM work. When they ran in-process on the build unit, they starved the event loop until status calls timed out. We moved them into a database-backed job the worker claims, and the leased jobs article covers how that queue works.

The composition root

All three processes, plus the test suite, wire the application the same way: through one buildModules function. It takes a Postgres pool, the validated config, an optional connector registry and an optional document store, and returns every capability module already wired together.

This matters more than it looks. The worker and the server can never disagree about how a module is built. A test can swap in a stub connector registry without hand-assembling dozens of modules. And because the connector registry is built by the same function the worker calls, a connector cannot be offered in the UI without also being executable. Cross-capability glue, such as the regression gate that needs both the agent admin module and the chat module, lives in the composition root, so neither module has to import the other. Architecture tests enforce those boundaries.

Boot is a gate

Every process boots the same way:

  1. Load the local .env file if present.
  2. Fill any still-unset secrets from Azure Key Vault.
  3. Validate config.
  4. Open the Postgres pool and ping it, so a bad connection string fails at boot rather than on the first request.
  5. Verify the schema: the set of applied migrations must exactly equal the set the build ships. The process never migrates itself.
  6. Connect to the document tier if configured, wire modules, and only then listen.

The server exposes /api/health for liveness and /api/ready for readiness. Because the process only binds after its startup checks pass, a 200 from /api/ready means the schema matched and storage was reachable. The deploy uses exactly that as its gate.

Secrets: Key Vault first, environment wins

Secrets come from Azure Key Vault through the VM's managed identity, resolved before config validation runs, so the rest of the process only ever sees a plain environment. Two rules make it pleasant to operate:

  • A vault secret only fills a variable that is unset. The environment wins. During an incident, someone can override a value in the systemd unit without a vault write and a redeploy.
  • A secret missing from the vault is not an error. It falls through to normal validation, so the failure reads "this setting is not set" instead of a Key Vault 404.

Only an explicit allowlist of secrets is ever read from the vault: database and storage connection strings, session and encryption keys, AI and speech service keys, and the platform-owned OAuth client for Procore. Non-secret settings that differ per environment, like the public URL, stay in the host's own config file, so one shared vault never collides across environments. Lookups run in parallel under one overall 30-second timeout, so a slow vault cannot stall a boot forever. Production refuses to start without a vault configured.

For local development, the client talks to Lowkey Vault, an emulator. Its self-signed certificate and relaxed auth challenge are only accepted when the vault host is loopback, so no deployed environment can be pointed at an unverified vault by mistake.

Local development with emulators

A single compose.local.yaml brings up Postgres 18, Azurite for blob storage, DocumentDB Local, and Lowkey Vault, all bound to 127.0.0.1 with health checks. The dev command checks Docker, starts the services, runs migrations, and only then starts the client, server and worker. Running the server directly after pulling a new migration fails on purpose, with a message telling you to migrate.

The point is parity. The same emulators run in CI, so a test that passes locally is testing the same storage semantics CI tests.

CI: real services, then a browser

The CI workflow on GitHub Actions has two jobs.

Check. A Postgres 18 service container (matching local, so CI fails where dev fails) plus Azurite, DocumentDB Local and Lowkey Vault started as plain containers, since they need command overrides the services block cannot express. They boot while npm ci runs. Environment flags then un-skip the integration suites: worker leasing, tenancy, sync, files, retrieval and Key Vault. Then one command runs typecheck, the full test suite of over a thousand tests, and the client and server builds.

Smoke. A separate throwaway Postgres, a production build, and a Playwright run of one end-to-end journey against dist/, not src/. It exercises the same client bundle and server seams production runs, and finishes in about a minute. The broader Playwright pack runs on manual dispatch rather than on every push. Keeping the per-push gate to one fast, stable journey keeps CI signal high, and growing that gate as the UI settles is the next step.

A new push cancels the in-flight run for that branch.

The deploy, step by step

The CLI has commands for the whole lifecycle: dev, stop, build, local (build and run production locally), deploy, ssh, logs (follows all three units' journals, pretty-printed locally), db-shell, db-baseline, restart, env, test, lint and clean-cache. Commands that only read default to the dev host. deploy always requires an explicit environment, because it writes.

Before anything ships

  1. Git guard. The deploy fetches origin/main and refuses if HEAD is not an ancestor of it, if the tree is dirty, or if origin is unreachable. Cannot-verify is treated as a failure, not as clean. The build is stamped with the HEAD commit, so without this guard a host could run code findable on main nowhere. A dev VM once ran four such commits. There is one escape flag for deploying a branch to dev to see it work, and it only downgrades the refusal to a loud warning.
  2. Production confirmation. Prod asks for confirmation unless explicitly skipped.
  3. Build. Prod runs the full test suite by default. Dev deploys skip tests unless asked, since typecheck and build already catch what would stop the server booting there.

On the host

The payload is just dist/, migrations/, package.json and the lockfile. node_modules is never copied over the wire.

  1. Take a directory lock over SSH so two deploys cannot mix payloads.
  2. rsync the payload into a staging directory, with strict host-key checking. SSH and rsync are spawned with argument arrays and no shell, so a hostname with a shell metacharacter cannot become an injection.
  3. Check the Node major version and make sure ffmpeg is present (the worker shells out to it for audio and video keyframes).
  4. Run npm ci --omit=dev in staging, unprivileged. Nothing on the live install has changed yet.
  5. Move the live install into a .previous backup directory, then move the staged build into place. The running processes already have their code loaded, so replacing files under them changes nothing until restart.
  6. Run migrations. After install, before restart. The server refuses to boot on a schema mismatch, so migrating after the restart would fail the health gate and roll back code that was fine.
  7. Restart all three units.
  8. Poll /api/ready for up to 30 seconds, and require every unit to report active. An inactive worker rolls back the deploy even if the API is answering.

Rolling back without lying about the schema

Rollback is where most simple deploy scripts go wrong. Restoring the previous code is easy. Restoring it onto a database that a new migration already changed is how you turn one incident into two.

So before restoring, the script runs a schema check from the new build against the backup's migration inventory. If the database still matches the previous build (because the failing migration rolled back in its transaction, or the failure came after migrations), it restores, restarts and confirms the restored deploy is ready. If the database has moved past what the previous build expects, it keeps both the candidate and the backup in place, prints the logs, and says to deploy a corrected forward build. It will not start old code against a new schema.

There is deliberately no pg_dump in the deploy. The managed database has automated backups and point-in-time restore, and a dump onto the VM's own disk would be a worse copy in a worse place.

The script keeps the design simple on purpose: one backup, overwritten each deploy, and restore is a sequence of moves rather than an atomic swap. A current symlink over dated release directories is the natural next step when we want atomic switches and more than one rollback target.

Rehearsing the rollback for real

A test that only asserts what a deploy script says is weak. The deploy rehearsal test runs the real remote script against a temporary directory tree, with sudo, systemctl, curl and npm shimmed onto the PATH and a staged migrate script that exits with whatever code the test chooses. Live and staged files carry markers, so after each run the test can check which build ended up installed and which systemctl calls happened.

It covers a healthy deploy, a failed migration (restore, no restart), an unhealthy deploy (restore and restart), a first deploy onto an empty host, a failed first deploy (candidate retained), a failed install, an inactive worker, a failed restart, and a compatible rollback that must confirm readiness before reporting recovery. These are a handful of mv calls that otherwise only ever run on a production host, with sudo, after the service is already down. That is the worst place to find out one of them was wrong.

Why VMs and systemd, not Kubernetes

We considered containers and an orchestrator. At this stage, for this team, VMs win on almost every axis that matters.

  • Fewer moving parts. One host, three units, a shell script anyone can read top to bottom. No cluster upgrades, ingress controllers, registries or Helm charts to maintain.
  • Debuggability. journalctl and ssh are the whole observability story for process issues, and the CLI wraps both.
  • The scale fits. Connect serves professional teams at human pace and runs minute-long background jobs. It does not see spiky, elastic consumer traffic. Worker throughput scales by starting more worker processes, which systemd handles fine.
  • Managed services carry the state. Postgres, Blob storage, DocumentDB and Key Vault are managed Azure services with their own backups. The VM is close to stateless.

The costs are real, and we name them:

  • In-process MCP sessions need sticky routing. The MCP server keeps session state in memory and sweeps idle sessions. That assumes one instance, or a load balancer with sticky sessions. Running two connect-build instances behind round-robin would break sessions mid-conversation. The fix is either sticky routing or moving session state into a shared store, and the code comments say so.
  • Restarts are not zero-downtime. Restarting the server drops in-flight requests after a 15-second graceful shutdown. Blue-green would need a second host and a traffic switch.
  • Recovery is a redeploy, not a failover. Each environment runs on one host today. Because the data lives in managed services, replacing a host is a reprovision and a deploy, not a restore, and adding a second host behind a load balancer is a straightforward next step when uptime targets call for it.
  • Host provisioning is partly manual. Node, Chromium for headless app checks, and system libraries are set up outside the deploy (ffmpeg is the one exception the script installs).

None of these are permanent. The process layout, the composition root and the boot gate would carry straight into containers. When traffic, uptime targets or team size justify it, moving to containers is a packaging change, not a redesign.

Why this matters if you're building something similar

  • Split processes by blocking behaviour, not by feature. Anything that can hold an event loop for minutes gets its own process.
  • Wire everything through one composition root. Server, worker and tests built the same way cannot drift apart.
  • Make readiness mean something. Bind only after schema and storage checks pass, then gate deploys on readiness.
  • Migrate between install and restart. And never let the app migrate itself.
  • Check schema compatibility before rolling back. Old code on a new schema is a second outage.
  • Refuse to deploy what is not on main. Stamp builds with the commit, and fail closed when you cannot verify.
  • Rehearse rollback in a test. Shim the system commands and run the real script.
  • Pick the simplest runtime that fits today's load, and write down what would make you change it.

Where to go next

Want a platform your team can actually operate? Plan your build.

Frequently asked questions

Why use VMs and systemd instead of Kubernetes?

At this stage the load fits on a host, state lives in managed Azure services, and a readable shell script with journalctl beats operating a cluster. The process layout carries straight into containers when scale or uptime targets justify it.

When should database migrations run during a deploy?

After the new build is installed and before services restart. The server refuses to boot on a schema mismatch, so migrating after restart would fail the readiness check and roll back good code.

How does the rollback avoid breaking the database?

Before restoring, it checks the database against the previous build's migration inventory. If the schema has moved past what the old code expects, it keeps both builds in place and asks for a forward fix instead of restoring.

Why does the deploy refuse commits that are not on main?

Builds are stamped with the HEAD commit. Without the guard, a host could run code that exists on no shared branch. Unreachable origin also fails closed, and a single flag downgrades the refusal to a warning for dev testing.

What is the catch with in-process MCP sessions?

Session state lives in memory, so it assumes one instance or sticky routing. Scaling the MCP process behind round-robin would break sessions until state moves to a shared store or the load balancer pins clients.

How do you test a deploy script's rollback?

A rehearsal test runs the real remote script against a temporary directory with sudo, systemctl, curl and npm shimmed, then checks which build ended up installed and which restarts happened across nine failure scenarios.

Next step

Want help running this playbook?

Bring the report or workflow. We'll help map the work behind it.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning