Build Flows

Custom applications · October 9, 2026 · 11 min read

Sandboxed AI App Builder: Letting AI Generate Apps Without Your Database

AI-generated apps run in an opaque-origin iframe with no network, reach data only through an SDK bridge that runs named, pre-approved queries against scoped views, and are rendered and reviewed before a person installs them.

By Charley Forey, founder of Build Flows

Every operations team eventually wants a dashboard nobody has built yet: RFIs by trade for two projects, look-ahead activities with weather overlaid, a punch list someone can actually type into. Letting an AI write those apps on demand is now easy. Letting it do that without handing it your database, your session cookies or your other tenants' data is the hard part.

Connect has an App Builder that does exactly this. A user describes what they want, an agent builds a small web app against the workspace's real data, a headless browser checks that it renders, a critic reviews it, and the user decides whether to install it. This article explains the security model that makes that safe: an opaque-origin iframe with no network, a single bridge as the only data path, declared queries that run under a locked-down Postgres role, and a human approval step whenever the agent wants to reach new data.

What a Connect app is

A Connect app is deliberately small. It is either a bundle (a connect-app.json manifest, an entry HTML file, assets, and the names of the queries and data collections it uses) or a single uploaded HTML page. There is no server-side code. The app runs entirely in the browser, inside an iframe, and everything it knows about the workspace arrives through one SDK.

The manifest is the contract. It names:

  • queries: declared, named SQL queries the app is allowed to run
  • aliases: logical data collections the app may read and write
  • adminAliases: the subset of collections only app managers may write

Bundles are immutable. A new change is a new version with its own unguessable id. Old manifests can never be migrated, which is why the manifest carries its own format version.

Bundle upload is the one place untrusted bytes enter the pipeline, so it is hardened: entry count and declared expanded size are checked from the zip's central directory before anything is decompressed, so a zip bomb is refused without running out of memory. Paths, asset counts and manifest size are capped. A single uploaded page is scanned for anything that loads from another origin (a CDN script, a web font, a remote image) and refused with the offending reference named, because the frame would not load it anyway.

The frame: opaque origin, no network

App content is served into an iframe sandboxed with allow-scripts and without allow-same-origin. That puts the app in an opaque origin. It cannot read Connect's cookies, cannot read local storage belonging to Connect, and cannot call Connect's API as the user, because from the browser's point of view it is not Connect.

On top of that, the content routes send a Content Security Policy built around one load-bearing line:

connect-src 'none'

The app cannot make network requests at all. No fetch, no XHR, no WebSocket, no beacon. That is what turns "the app never decides where its data goes" from a guideline into something the browser enforces. The rest of the policy is depth behind that line: default-src 'none', scripts, styles, images and fonts only from Connect's own origin (plus inline and data URIs), no nested frames, no objects, no form posts, no base-tag tricks, and frame-ancestors limited to Connect itself.

Two consequences are worth calling out because they look wrong at first.

The content routes serve app code without a session. An opaque-origin frame cannot send cookies, so a cookie-authenticated asset route would be unreachable by the very frame that needs it. These routes carry only app code, addressed by unguessable version ids. Tenant data never travels this path.

The framing headers are overridden per plugin. The API sets X-Frame-Options: DENY on every response in a hook that runs after the handler, which would stop app content rendering at all. The app content plugin registers its own after-handler hook that relaxes that header for its routes only. Fastify's plugin encapsulation is what makes this safe, and the code carries a warning never to move that hook to the root, where it would strip framing protection from the whole API.

The bridge: the only way out of the frame

If the app cannot talk to the network, how does it get data? Through the SDK, served from three fixed URLs on Connect's origin:

ScriptPurpose
/app-sdk/v1.jsThe bridge client: connect.query.run(name) and connect.data.*
/app-sdk/ui.jsA first-party UI block library
/app-sdk/chart.jsCharting for data the app has already pulled through the bridge

Module imports are governed by script-src, not connect-src, which is why the import works while network calls do not. That is intended.

The protocol is plain browser messaging:

  1. The frame loads the SDK, which announces itself to the parent window every 250 ms. (The parent may not have attached its listener yet when the frame finishes loading. A single announcement is a race.)
  2. The parent shell answers with an init message carrying a context object and a dedicated MessagePort.
  3. The app sends requests over that port: run this named query, list this collection, insert this document.
  4. The shell, which does hold the user's session, calls the authenticated API on the app's behalf and posts the result back.

The SDK is served rather than copied into each bundle on purpose: bundles are immutable, so a vendored copy could never be patched, while a served one gets protocol fixes to every installed app on its next load.

Declared queries: the app names, the server decides

When an app runs a query, it sends exactly one thing: a name. There are no caller parameters. The workspace comes from the session and the project scope comes from the installation. Sorting and filtering are either fixed in the SQL or done in the app's own code. With no caller-supplied value anywhere on the query path, there is nothing to sanitize.

The server checks that the name is one the installed version's manifest declared and an admin granted at install. A compromised app that names a different query gets a not-found, which is also all it should learn.

Then the statement runs under four separate guards:

GuardWhat it stops
Subquery wrapperThe statement is run as SELECT * FROM (<statement>) ... LIMIT n. Bare writes, data-modifying CTEs and stacked statements are all syntax or feature errors in that position.
Read-only transactionA backstop for anything the wrapper reasoning missed, with a 10-second statement timeout
SET LOCAL ROLE to a query roleThe role owns nothing and has SELECT on a named list of views. Sessions, connection secrets and app data are simply not readable.
Scope settingsThe workspace id and project ids are set as transaction-local settings, and the granted views filter on them

The last guard is the one people usually miss. A role decides which tables a statement can reach, never whose rows. A self-join like FROM projects p1 JOIN projects p2 ON true WHERE p1.workspace_id = $1 passes every other check while reading every tenant's projects. Because the grants point at pre-scoped views that filter on the transaction settings, scoping becomes a property of the relation, not of the statement. Forgetting a WHERE clause still returns scoped rows.

Every statement is also bound to two parameters, $1 for the workspace id and $2 for the project ids as a uuid array, and the authoring contract asks for SQL that filters on them. The views make that belt-and-braces rather than the only line of defence. Results are capped at 50,000 rows, applied in SQL so a runaway query never materializes, and the transaction always rolls back because a read has nothing to commit.

Validation is execution

How do you validate a SQL statement before saving it? Not with a parser that will drift from Postgres. Connect runs the statement through the exact production path against a freshly generated workspace id that nothing can own. One round trip catches syntax errors, refused write shapes and ungranted tables, with Postgres's own error message. If that run returns any rows at all, the statement is refused, because rows for an empty workspace mean an unscoped view was granted somewhere.

propose_query: the agent can ask, a super-admin decides

Declaring a query is a super-admin action, because a declared query is a platform artifact. The App Builder agent has a propose_query tool for when no existing query covers what the user needs. A proposal is validated to the same bar as a direct declaration, then stored as pending. The app can already name it, but it returns nothing until a super-admin approves it, and the agent is told to say so to the user. A proposal can never overwrite an existing query name.

That one rule means the model never puts a runnable query into production on its own authority.

App data and the per-app fence

Dashboards often need to store things: a form submission, a status, a note. Each installation gets its own document collections, addressed by alias. The app supplies an alias and a document; the server resolves where it goes. The workspace comes from the path and has already been checked, the installation is filtered by workspace, and the alias must be one the manifest declared and the admin granted.

Incoming documents are checked once before they become rows. The threat model is not a malicious app author; it is that any authenticated user can send arbitrary JSON to the route from devtools. The check caps nesting depth and key counts, rejects bytes Postgres would refuse (NUL characters, lone surrogates) with one consistent message so they cannot become a probing oracle, and rejects __proto__, constructor and prototype keys. It does not try to escape SQL, because the document only ever reaches Postgres as a bound parameter.

The design rule we call the app-data fence: server-side querying of app data happens only through fixed, server-owned predicates. No app-supplied JSON paths, no interpolated update paths, no expression indexes over keys an app controls. That constrains future features, deliberately.

There are also per-installation ceilings (tens of thousands of documents, a few thousand per read), and inserts take an idempotency key so a double-clicked submit becomes one row.

Versions, installations and share links

An installation is a version installed into a workspace with a project scope and a visibility mode (whole workspace or restricted). Grants for queries and collections are re-derived from the manifest at install time, and re-checked then too, because a query can be deleted between publish and install. Changing who may write a collection takes an admin installing a version that says so, not just a publish.

Dashboards can be shared outside the workspace through share links:

  • The raw token is shown once. Only its SHA-256 hash is stored.
  • Links can expire and can be revoked.
  • A link can optionally carry a password, added in a later migration. The password is hashed with scrypt and a per-link salt, compared in constant time, and checked at the single point where a link resolves. A protected link runs nothing until the password is supplied.
  • Share-link access is read-only by construction.

The App Builder agent: build, render, critique, propose

The agent works in a loop that ends with a human decision, not a deploy:

  1. Interview. It asks one or two focused questions: who will use this, which decision it supports, which projects it covers. A build has a small question budget.
  2. Ground in real data. It lists data sources, previews rows from the queries it intends to use, and adapts a similar reference app when one exists. It never invents a query name.
  3. Confirm the plan and scope with the user, especially which projects the app will see.
  4. Build a draft file set and publish a preview version.
  5. Verify. A server-side job renders the preview in headless Chromium through Playwright, inside the same opaque-origin frame users get. The browser may load only that version's assets and the SDK; every other request is blocked and logged. A read-only query bridge serves draft queries, capped per render. After a 10-second budget the job records console errors, failed responses, and whether the frame rendered fewer than a handful of visible characters, which counts as blank. The agent reads this, plus any console errors the live preview reported, through get_preview_logs.
  6. Propose. propose_app is refused while the render check is pending, crashed, blank or showing errors. If it is clean, a critic sub-agent on a cheaper model reviews the files, the allowed query list, the user's request and the render log. It flags blocking defects: external URLs, wrong SDK use, unescaped data in HTML, a broken or blank app, or not doing what was asked. An error verdict sends the agent back to fix. The critic fails open, so an unavailable judge never blocks a build.
  7. Decide. The user accepts, which installs the app, or discards it.

Local CLI and MCP for your own agent

Some teams would rather build apps with the coding agent they already use. A small package provides a CLI (login, init, pull, push, preview) and a stdio MCP server that exposes the same operations as tools: list apps and versions, run declared queries, read and write app data, pull, push and preview. init scaffolds a minimal valid app: a manifest, an index.html that imports the SDK and UI library, and one example query.

Authentication is a personal access token scoped to one workspace, stored in a local config file with owner-only permissions. The MCP server reads it from there, so the agent never handles the token, and a token never grants more than the user's own live access. Pushing goes through the same publish route as the in-product builder, so the same validation applies.

Why this matters if you're building something similar

If you want AI to generate user-facing apps against real data, these are the parts we would keep:

  • Give generated code no network. An opaque-origin sandbox plus connect-src 'none' turns data exfiltration from a code-review problem into a browser-enforced impossibility.
  • Make one bridge the only data path, and keep credentials on the side of the bridge the generated code cannot reach.
  • Let apps send names, not parameters. If nothing caller-supplied reaches SQL, there is nothing to sanitize.
  • Scope in the relation, not the statement. Pre-scoped views plus a role with no ownership beat hoping every query remembers its WHERE.
  • Validate by executing against nothing. It is the only validator that cannot drift from production.
  • Let the agent propose new data access, never grant it. A pending state with human approval is cheap and closes the biggest hole.
  • Check the rendered result, not just the code. A headless render catches blank pages and console errors that no static review will.

Where to go next

Want AI-built dashboards on your own data without opening up your database? Plan your build with us.

Frequently asked questions

Why serve AI-generated apps from an opaque-origin iframe?

Without allow-same-origin the frame is not treated as your application, so it cannot read your cookies or storage or call your API as the user. Combined with connect-src 'none', it cannot send data anywhere at all.

How does an app get data if it has no network access?

It imports an SDK from the host origin and asks the parent window over a MessagePort. The parent shell holds the user's session, calls the authenticated API, and posts the result back.

How do you stop a declared query from reading other tenants' data?

The query runs under a role that can only select from pre-scoped views. Those views filter on transaction-local workspace and project settings taken from the installation, so even a query missing a WHERE clause returns only scoped rows.

Can the AI agent add new SQL queries?

It can propose one. The proposal is validated by executing it against an empty workspace, then stored as pending. The app can reference it, but it returns nothing until a super-admin approves it.

How are generated apps tested before users see them?

A worker renders each preview in headless Chromium through Playwright, records console errors, failed requests and blank renders, and the agent cannot propose the app until that check is clean. A critic sub-agent then reviews the code for blocking defects.

Can I build Connect apps with my own coding agent?

Yes. A local CLI and stdio MCP server let tools like Claude Code or Codex scaffold, preview and publish apps, authenticated with a workspace-scoped personal access token the agent never handles directly.

Next step

Need something built around how your team works?

Describe the users, the workflow, and the systems it touches. We'll tell you whether a custom application makes sense and how we'd build it.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning