Most AI features that fail in production fail at the same point: the moment a model's output stops being a suggestion and becomes a record. A chat answer that is slightly wrong costs a re-read. A wrong value written into the table that drives a schedule, an estimate or a contract costs much more, and you often find out weeks later.
The Project Info Sheet in Connect is the agent that answers "what is being built?" for every project. Its output feeds the Schedule Builder, so it is effectively a system of record. We rebuilt it in September 2026 as a two-phase workflow: a research agent drafts a cited Profile, a person approves it, and only then does a second, locked-down agent turn the Profile into structured Facts. This article walks through how that works in code, what it costs, and why we think "two phases with a human gate in the middle" is the right default for any AI that writes to data other systems trust.
The shape of the workflow
The short version:
| Phase | Who acts | What it can touch | Output |
|---|---|---|---|
| 1. Research | A sandboxed agent plus section workers | Project files staged into a sandbox, the Profile document | A Profile: prose findings per Section, each with citations |
| Gate | A person | The approval route | An approved, snapshotted Profile revision |
| 2. Normalize | A tool-less normalizer agent | Nothing but its input text | Facts keyed by a fixed Topic vocabulary, citing only the approved Profile |
Phase 1 can be exploratory; Phase 2 must be precise. The gate between them means the exploratory part never writes structured data and the precise part never sees raw documents.
Phase 1: a research agent that builds a cited Profile
The Phase 1 agent runs on the OpenAI Agents SDK as a sandbox agent with filesystem, skills and compaction capabilities. It gets the project's drawings and documents one file at a time through a stage_project_file tool, prepares and renders PDF pages inside its workspace, and looks at drawing crops directly. It works through a fixed set of Sections (project drivers, building scope and so on), and for each Topic in a Section it records a disposition:
- documented: a note plus at least one citation to a specific file version and locator
- not_applicable: with a reason
- explicit_unknown and deferred: written only by the runtime, when a person answers a question that way
That last point matters. The agent's disposition tool accepts only documented and not_applicable. The model cannot decide on its own that something is "unknown" and move on. An unknown has to come from a human answer, so the gaps in the Profile are honest gaps, not places the model gave up.
Section workers
A large drawing set does not fit in one context window, and one agent working through every Section in order is slow. So the parent agent delegates. An investigate_section tool dispatches a fresh section worker with a clean context and a focused brief: which files, which sheets or pages, which questions. Up to three workers run at once. Each worker records dispositions for its own Section only, then returns a report with findings, remaining gaps and suggested questions.
A few rules keep this from turning into chaos:
- Dispatch returns an identity immediately. The parent consumes results through a
wait_for_section_resulttool that hands back completions in the order they finished, together with freshly loaded canonical state, so it always merges against the current revision. - Workers own disjoint Sections, and their short writes are serialized per project. Research runs in parallel; writes do not.
- Workers cannot overwrite an assessment that is newer than the one they started from, and they cannot talk to the user or edit the Profile prose. The parent does both.
- Project Drivers are never delegated. Other Sections stay gated until the parent has resolved them, because job type and building scope change what every other Section needs.
- If a worker finds two sources that contradict each other, it reports a conflict with bounded, cited claims. An unresolved conflict blocks review even on an optional Topic. Spotting the conflict is model judgment; blocking on it is deterministic code.
Workers have their own turn caps: about 50 model steps for a required Section and 30, at low reasoning effort, for an advisory one. Each step re-sends a growing context, so the cap is a cost ceiling as much as a runaway guard.
The review gate: approval is server code, not an agent tool
The reference design we started from exposed "approve the profile" and "finish normalization" as tools the agent could call. We removed both. In Connect, permissions and tenant scoping live in deterministic server code, and approval is a permission decision.
Approval is an HTTP route. When a person clicks Approve, the server:
- Confirms the workflow is waiting for review (or is a re-approval of a Profile edited after approval).
- Checks the caller's expected Profile revision against the stored one. If someone else edited it since the page loaded, the request fails with a stale-revision error instead of approving something the person never saw.
- Runs the gap check: every required Topic for this job type has a disposition, and no conflict is unresolved.
- Marks the Profile approved with a conditional replace, and writes an immutable snapshot of that exact revision.
- Seeds the Facts document with status
normalizingand the source Profile revision. - Flips the workflow phase last, with a conditional update that fails if the phase moved underneath it.
- Starts Phase 2 from the server. The model never starts normalization.
The ordering is deliberate. The phase flip is the commit point, and everything before it is safe to repeat, so a retried approval is idempotent.
Phase 2: a normalizer restricted to a fixed vocabulary
The normalizer is a plain agent with no tools, no sandbox, no file access and no search. It receives three things as text: the approved Profile, that Profile's citations as a numbered list, and the list of fact keys it is allowed to use. It returns structured output against a strict JSON schema.
The allowed keys come from a fixed Topic catalog: every key has the form <subject>.<topicKey>, for the Subjects declared in the Profile and the Topics that apply to its job type. The reference design let the normalizer invent dotted keys. We do not, because invented keys are how a structured store quietly turns into a junk drawer that downstream code cannot rely on.
Before anything is persisted, the server rejects:
- any key outside the allowed set
- any citation that did not appear in the approved Profile
- any output that fails schema validation
A Fact therefore cannot cite a document the human reviewer never saw cited. The provenance chain runs Fact, to Profile citation, to file version and locator, and it stays resolvable from the admin trace viewer.
If the approved Profile has no job type, there is no vocabulary to normalize into, and the run fails up front instead of guessing.
One table for every phase change
The workflow has seven phases: profiling, waiting_for_input, awaiting_project_files, waiting_for_profile_review, normalizing, normalized and normalization_failed, plus a "no workflow yet" starting point. Every legal move between them is one row in a single WORKFLOW_TRANSITIONS table, each with a from, a to and a human-readable trigger. A test asserts the table, and the stores refuse any transition not in it.
An illustrative version of the idea:
const TRANSITIONS = [
{ from: "waiting_for_review", to: "normalizing", trigger: "approval route, gaps clear" },
{ from: "normalizing", to: "normalized", trigger: "output validated" },
{ from: "normalizing", to: "norm_failed", trigger: "validation or model failure" },
{ from: "norm_failed", to: "normalizing", trigger: "retry, same source revision" },
// ...every other legal move, and nothing else
]
function assertTransition(from, to) {
if (!TRANSITIONS.some(t => t.from === from && t.to === to)) throw new IllegalTransition(from, to)
}
"Can a sheet go from X to Y?" is now one grep, and a new path is a reviewed change to one list.
Two transitions the reference design lacked
normalization_failed. In the original design a failed normalization stayed in normalizing. To the client, that looked exactly like a run in progress: an endless spinner with no retry button. We added an explicit failure phase. Failure leaves the approved Profile and its revision untouched and records a user-safe reason. Retry is idempotent for the same source Profile revision, so pressing it twice does not produce two Fact sets.
normalized back to review. If someone edits a Profile after it was approved, the workflow marks it profile_stale and it goes back through review. The previous Facts stay readable while that happens, visibly marked stale, so downstream consumers are not cut off while a correction is in progress.
Three singleton documents per project
The reference kept profile, facts and status in one document keyed by the agent's conversation thread. We split them into three documents, each keyed by project with a unique index:
| Document | Role |
|---|---|
| Workflow | The only lifecycle authority: phase, approval metadata, failure reason |
| Profile | Phase 1 output: Sections, assessments, citations, prose |
| Facts | Phase 2 output: keyed Facts plus the source Profile revision |
Each carries its own optimistic revision counter and a hash of the Topic catalog it was written under, so a document from an older catalog is detectable.
Keying by project means one canonical Info Sheet per project however many conversations touch it. Splitting into three makes each phase boundary a single-document conditional update, which keeps phase changes safe on Cosmos DB for MongoDB vCore without multi-document transactions.
Snapshots are taken only at phase boundaries. Per-edit history already lives on our Postgres trace ledger, and storing it twice would create two histories that eventually disagree.
Keeping a long-running agent alive: lease, cap, continuation
Phase 1 can run for a long time on a big drawing set, and the server can restart in the middle. Three mechanisms handle that.
A sandbox lease. A pool keeps one live sandbox per project between turns. After each turn it saves a lease: the serialized sandbox state, the host it lives on, a key describing which project files were selected, and an expiry. On the next turn the pool reuses the live sandbox if there is one, resumes from the lease if it is still valid for the same host and file selection, and otherwise starts fresh. Idle sandboxes close after an hour. Recovery is a resume, not a replay.
A turn cap and a wall-clock kill. Each parent run is capped at 100 SDK turns, and each application turn has a hard timeout that kills the sandbox. Every stop also sweeps up anything the sandbox started, using a marker stamped into each sandbox's environment, and acceptance tests check for leaked processes and workspaces.
Self-driven continuation. When a turn ends with the sheet still in profiling and no pending question, the runtime schedules one more turn itself with an internal "continue from current state" prompt. Each continuation schedules the next when it settles, so profiling reaches the review gate without anyone pressing a button. It stops on any genuine stop: a pending question, missing files, a configured iteration cap, or no progress, measured as no new required Topic assessed since the last self-driven turn. A refused continuation (another turn holds the project, or the phase moved) does not count against the cap.
Only one mutating turn per project is admitted at a time. A second submission gets a 409 and keeps its composer text. That ownership lives in memory, which is correct for our single-process deployment and wrong for overlapping replicas; the decision record says to add owner fencing before that topology ships.
A destructive cutover instead of a migration
The old Info Sheet was an interview over a catalog of items, routing rules and proposals, none of which maps onto Sections, Topics and Facts. A best-effort migration would have produced Profiles with no citations and Facts with no provenance: exactly the data the new design refuses.
So we did not migrate. An explicit cutover command, never a boot-time side effect, wipes the old data per workspace database, and refuses to run against a remote database without an extra flag.
Why this matters if you're building something similar
You do not need the Agents SDK or a scheduling product to use this pattern. If an AI writes to anything another system treats as truth, these are the parts worth copying:
- Separate exploration from commitment. Let one agent read widely and write prose with citations. Let a different, tool-less step produce structured records. Different jobs, different permissions.
- Put a person between them, and make approval a server route. If approval is an agent tool, the model can approve its own work. A route can check revisions, run a gap check and commit idempotently.
- Fix the vocabulary. Reject any key or citation the reviewer could not have seen. A closed vocabulary is what makes structured output safe for downstream code.
- Write every state change down in one table. It turns "what can happen?" into a list you can read and test.
- Model failure as a state. If a failed run looks like a running one, users get a spinner and support gets a ticket.
- Don't migrate data that the new design would reject. A clean cutover is cheaper than defending records with no provenance.
- Write down your temporary choices. A limitation with a recorded plan is a decision. One nobody recorded is a liability.
The cost is real: two phases are slower, a person has to review, and the state machinery takes code and tests. For a chat assistant that is overkill. For anything feeding a schedule, an estimate or a contract, it is the minimum.
Where to go next
- The series overview: How we built Connect
- How the downstream consumer gates on approved data: AI schedule builder build-spec gates
- Pausing and resuming delegated agents: Multi-agent delegation, pause and resume
- Where the per-run traces land: LLM agent tracing on a Postgres ledger
- The governance checklist this pattern implements: Governing AI agents in construction
Planning an AI workflow that writes to your own systems? Plan your build with us.
Frequently asked questions
What is a two-phase AI agent workflow?
One agent researches and drafts a human-readable, cited result. A person reviews and approves it. A second, more restricted step then converts only the approved draft into structured records that other systems rely on.
Why not let the agent approve its own work with a tool?
Approval is a permission decision. As a server route it can check that the reviewer saw the current revision, confirm required topics are covered, snapshot the approved version and commit idempotently. None of that is guaranteed if the model can call an approve tool.
How do you stop the normalizer from inventing fields?
It receives the list of allowed keys, built from a fixed Topic catalog for the project's job type and subjects. The server rejects any key outside that set and any citation that did not appear in the approved Profile before anything is saved.
What happens when normalization fails?
The workflow moves to a normalization_failed phase with a user-safe reason. The approved Profile is left untouched, and retry is idempotent for the same source Profile revision, so users get a retry button instead of an endless spinner.
What can the Phase 1 research agent reach?
Only the project files staged into its workspace, one at a time, through a dedicated tool. It holds no Connect credentials, works under resource limits and a per-turn time limit, and cannot approve its own output or write structured Facts.
Next step
Have a workflow in mind?
Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
How We Built Connect: Architecture of an Enterprise AI Platform
The pillar of our Connect architecture series: a layer-by-layer map of an enterprise agentic AI platform, the three-process shape it runs as, the design principles that kept recurring, and links to every deep-dive article.
Scheduling & P6 · October 9, 2026
Engineering an AI Schedule Builder: Build Specs, Gates and Repair Loops
The engineering behind the Connect Schedule Builder loop: the model edits a declarative Build Spec, code expands and gates it, reviewers and bounded repair passes drive it to done, and approval freezes verified evidence.
AI agents & MCP · October 9, 2026
Multi-Agent Delegation That Pauses and Resumes: How Connect's Chat Hands Work to Sub-Agents
An engineering walkthrough of chat-to-sub-agent delegation in Connect: scope pinned server-side, sub-agent activity streamed into chat, and a one-row-per-conversation pointer that routes the user's next message back to the paused agent, with permission re-checks and stale-pointer recovery.