Build Flows

AI agents & MCP · October 9, 2026 · 11 min read

Two-Phase AI Agents with Human Approval: How We Built the Connect Info Sheet

A research agent builds a cited Profile, a person approves it through a server route, and a tool-less normalizer converts it into Facts limited to a fixed vocabulary. The pattern applies to any AI that writes to a system of record.

By Charley Forey, founder of Build Flows

Most AI features that fail in production fail at the same point: the moment a model's output stops being a suggestion and becomes a record. A chat answer that is slightly wrong costs a re-read. A wrong value written into the table that drives a schedule, an estimate or a contract costs much more, and you often find out weeks later.

The Project Info Sheet in Connect is the agent that answers "what is being built?" for every project. Its output feeds the Schedule Builder, so it is effectively a system of record. We rebuilt it in September 2026 as a two-phase workflow: a research agent drafts a cited Profile, a person approves it, and only then does a second, locked-down agent turn the Profile into structured Facts. This article walks through how that works in code, what it costs, and why we think "two phases with a human gate in the middle" is the right default for any AI that writes to data other systems trust.

The shape of the workflow

The short version:

PhaseWho actsWhat it can touchOutput
1. ResearchA sandboxed agent plus section workersProject files staged into a sandbox, the Profile documentA Profile: prose findings per Section, each with citations
GateA personThe approval routeAn approved, snapshotted Profile revision
2. NormalizeA tool-less normalizer agentNothing but its input textFacts keyed by a fixed Topic vocabulary, citing only the approved Profile

Phase 1 can be exploratory; Phase 2 must be precise. The gate between them means the exploratory part never writes structured data and the precise part never sees raw documents.

Phase 1: a research agent that builds a cited Profile

The Phase 1 agent runs on the OpenAI Agents SDK as a sandbox agent with filesystem, skills and compaction capabilities. It gets the project's drawings and documents one file at a time through a stage_project_file tool, prepares and renders PDF pages inside its workspace, and looks at drawing crops directly. It works through a fixed set of Sections (project drivers, building scope and so on), and for each Topic in a Section it records a disposition:

  • documented: a note plus at least one citation to a specific file version and locator
  • not_applicable: with a reason
  • explicit_unknown and deferred: written only by the runtime, when a person answers a question that way

That last point matters. The agent's disposition tool accepts only documented and not_applicable. The model cannot decide on its own that something is "unknown" and move on. An unknown has to come from a human answer, so the gaps in the Profile are honest gaps, not places the model gave up.

Section workers

A large drawing set does not fit in one context window, and one agent working through every Section in order is slow. So the parent agent delegates. An investigate_section tool dispatches a fresh section worker with a clean context and a focused brief: which files, which sheets or pages, which questions. Up to three workers run at once. Each worker records dispositions for its own Section only, then returns a report with findings, remaining gaps and suggested questions.

A few rules keep this from turning into chaos:

  • Dispatch returns an identity immediately. The parent consumes results through a wait_for_section_result tool that hands back completions in the order they finished, together with freshly loaded canonical state, so it always merges against the current revision.
  • Workers own disjoint Sections, and their short writes are serialized per project. Research runs in parallel; writes do not.
  • Workers cannot overwrite an assessment that is newer than the one they started from, and they cannot talk to the user or edit the Profile prose. The parent does both.
  • Project Drivers are never delegated. Other Sections stay gated until the parent has resolved them, because job type and building scope change what every other Section needs.
  • If a worker finds two sources that contradict each other, it reports a conflict with bounded, cited claims. An unresolved conflict blocks review even on an optional Topic. Spotting the conflict is model judgment; blocking on it is deterministic code.

Workers have their own turn caps: about 50 model steps for a required Section and 30, at low reasoning effort, for an advisory one. Each step re-sends a growing context, so the cap is a cost ceiling as much as a runaway guard.

The review gate: approval is server code, not an agent tool

The reference design we started from exposed "approve the profile" and "finish normalization" as tools the agent could call. We removed both. In Connect, permissions and tenant scoping live in deterministic server code, and approval is a permission decision.

Approval is an HTTP route. When a person clicks Approve, the server:

  1. Confirms the workflow is waiting for review (or is a re-approval of a Profile edited after approval).
  2. Checks the caller's expected Profile revision against the stored one. If someone else edited it since the page loaded, the request fails with a stale-revision error instead of approving something the person never saw.
  3. Runs the gap check: every required Topic for this job type has a disposition, and no conflict is unresolved.
  4. Marks the Profile approved with a conditional replace, and writes an immutable snapshot of that exact revision.
  5. Seeds the Facts document with status normalizing and the source Profile revision.
  6. Flips the workflow phase last, with a conditional update that fails if the phase moved underneath it.
  7. Starts Phase 2 from the server. The model never starts normalization.

The ordering is deliberate. The phase flip is the commit point, and everything before it is safe to repeat, so a retried approval is idempotent.

Phase 2: a normalizer restricted to a fixed vocabulary

The normalizer is a plain agent with no tools, no sandbox, no file access and no search. It receives three things as text: the approved Profile, that Profile's citations as a numbered list, and the list of fact keys it is allowed to use. It returns structured output against a strict JSON schema.

The allowed keys come from a fixed Topic catalog: every key has the form <subject>.<topicKey>, for the Subjects declared in the Profile and the Topics that apply to its job type. The reference design let the normalizer invent dotted keys. We do not, because invented keys are how a structured store quietly turns into a junk drawer that downstream code cannot rely on.

Before anything is persisted, the server rejects:

  • any key outside the allowed set
  • any citation that did not appear in the approved Profile
  • any output that fails schema validation

A Fact therefore cannot cite a document the human reviewer never saw cited. The provenance chain runs Fact, to Profile citation, to file version and locator, and it stays resolvable from the admin trace viewer.

If the approved Profile has no job type, there is no vocabulary to normalize into, and the run fails up front instead of guessing.

One table for every phase change

The workflow has seven phases: profiling, waiting_for_input, awaiting_project_files, waiting_for_profile_review, normalizing, normalized and normalization_failed, plus a "no workflow yet" starting point. Every legal move between them is one row in a single WORKFLOW_TRANSITIONS table, each with a from, a to and a human-readable trigger. A test asserts the table, and the stores refuse any transition not in it.

An illustrative version of the idea:

const TRANSITIONS = [
  { from: "waiting_for_review", to: "normalizing",   trigger: "approval route, gaps clear" },
  { from: "normalizing",        to: "normalized",    trigger: "output validated" },
  { from: "normalizing",        to: "norm_failed",   trigger: "validation or model failure" },
  { from: "norm_failed",        to: "normalizing",   trigger: "retry, same source revision" },
  // ...every other legal move, and nothing else
]

function assertTransition(from, to) {
  if (!TRANSITIONS.some(t => t.from === from && t.to === to)) throw new IllegalTransition(from, to)
}

"Can a sheet go from X to Y?" is now one grep, and a new path is a reviewed change to one list.

Two transitions the reference design lacked

normalization_failed. In the original design a failed normalization stayed in normalizing. To the client, that looked exactly like a run in progress: an endless spinner with no retry button. We added an explicit failure phase. Failure leaves the approved Profile and its revision untouched and records a user-safe reason. Retry is idempotent for the same source Profile revision, so pressing it twice does not produce two Fact sets.

normalized back to review. If someone edits a Profile after it was approved, the workflow marks it profile_stale and it goes back through review. The previous Facts stay readable while that happens, visibly marked stale, so downstream consumers are not cut off while a correction is in progress.

Three singleton documents per project

The reference kept profile, facts and status in one document keyed by the agent's conversation thread. We split them into three documents, each keyed by project with a unique index:

DocumentRole
WorkflowThe only lifecycle authority: phase, approval metadata, failure reason
ProfilePhase 1 output: Sections, assessments, citations, prose
FactsPhase 2 output: keyed Facts plus the source Profile revision

Each carries its own optimistic revision counter and a hash of the Topic catalog it was written under, so a document from an older catalog is detectable.

Keying by project means one canonical Info Sheet per project however many conversations touch it. Splitting into three makes each phase boundary a single-document conditional update, which keeps phase changes safe on Cosmos DB for MongoDB vCore without multi-document transactions.

Snapshots are taken only at phase boundaries. Per-edit history already lives on our Postgres trace ledger, and storing it twice would create two histories that eventually disagree.

Keeping a long-running agent alive: lease, cap, continuation

Phase 1 can run for a long time on a big drawing set, and the server can restart in the middle. Three mechanisms handle that.

A sandbox lease. A pool keeps one live sandbox per project between turns. After each turn it saves a lease: the serialized sandbox state, the host it lives on, a key describing which project files were selected, and an expiry. On the next turn the pool reuses the live sandbox if there is one, resumes from the lease if it is still valid for the same host and file selection, and otherwise starts fresh. Idle sandboxes close after an hour. Recovery is a resume, not a replay.

A turn cap and a wall-clock kill. Each parent run is capped at 100 SDK turns, and each application turn has a hard timeout that kills the sandbox. Every stop also sweeps up anything the sandbox started, using a marker stamped into each sandbox's environment, and acceptance tests check for leaked processes and workspaces.

Self-driven continuation. When a turn ends with the sheet still in profiling and no pending question, the runtime schedules one more turn itself with an internal "continue from current state" prompt. Each continuation schedules the next when it settles, so profiling reaches the review gate without anyone pressing a button. It stops on any genuine stop: a pending question, missing files, a configured iteration cap, or no progress, measured as no new required Topic assessed since the last self-driven turn. A refused continuation (another turn holds the project, or the phase moved) does not count against the cap.

Only one mutating turn per project is admitted at a time. A second submission gets a 409 and keeps its composer text. That ownership lives in memory, which is correct for our single-process deployment and wrong for overlapping replicas; the decision record says to add owner fencing before that topology ships.

A destructive cutover instead of a migration

The old Info Sheet was an interview over a catalog of items, routing rules and proposals, none of which maps onto Sections, Topics and Facts. A best-effort migration would have produced Profiles with no citations and Facts with no provenance: exactly the data the new design refuses.

So we did not migrate. An explicit cutover command, never a boot-time side effect, wipes the old data per workspace database, and refuses to run against a remote database without an extra flag.

Why this matters if you're building something similar

You do not need the Agents SDK or a scheduling product to use this pattern. If an AI writes to anything another system treats as truth, these are the parts worth copying:

  • Separate exploration from commitment. Let one agent read widely and write prose with citations. Let a different, tool-less step produce structured records. Different jobs, different permissions.
  • Put a person between them, and make approval a server route. If approval is an agent tool, the model can approve its own work. A route can check revisions, run a gap check and commit idempotently.
  • Fix the vocabulary. Reject any key or citation the reviewer could not have seen. A closed vocabulary is what makes structured output safe for downstream code.
  • Write every state change down in one table. It turns "what can happen?" into a list you can read and test.
  • Model failure as a state. If a failed run looks like a running one, users get a spinner and support gets a ticket.
  • Don't migrate data that the new design would reject. A clean cutover is cheaper than defending records with no provenance.
  • Write down your temporary choices. A limitation with a recorded plan is a decision. One nobody recorded is a liability.

The cost is real: two phases are slower, a person has to review, and the state machinery takes code and tests. For a chat assistant that is overkill. For anything feeding a schedule, an estimate or a contract, it is the minimum.

Where to go next

Planning an AI workflow that writes to your own systems? Plan your build with us.

Frequently asked questions

What is a two-phase AI agent workflow?

One agent researches and drafts a human-readable, cited result. A person reviews and approves it. A second, more restricted step then converts only the approved draft into structured records that other systems rely on.

Why not let the agent approve its own work with a tool?

Approval is a permission decision. As a server route it can check that the reviewer saw the current revision, confirm required topics are covered, snapshot the approved version and commit idempotently. None of that is guaranteed if the model can call an approve tool.

How do you stop the normalizer from inventing fields?

It receives the list of allowed keys, built from a fixed Topic catalog for the project's job type and subjects. The server rejects any key outside that set and any citation that did not appear in the approved Profile before anything is saved.

What happens when normalization fails?

The workflow moves to a normalization_failed phase with a user-safe reason. The approved Profile is left untouched, and retry is idempotent for the same source Profile revision, so users get a retry button instead of an endless spinner.

What can the Phase 1 research agent reach?

Only the project files staged into its workspace, one at a time, through a dedicated tool. It holds no Connect credentials, works under resource limits and a per-turn time limit, and cannot approve its own output or write structured Facts.

Next step

Have a workflow in mind?

Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning