Every schedule depends on decisions that are not in the drawings. Will excavation use soldier piles and lagging or sheet piling? Is there a tower crane or only a hoist? Does the facade go up from mast climbers or swing stages? How long is the elevator lead time? A senior scheduler answers these from experience, from a phone call with the client, or from a line buried in an email thread. Then the answers disappear into a chat log or someone's memory, and three weeks later nobody can say why the schedule assumed a crane.
Means & Methods is the step in Connect that keeps those answers. It sits between the approved Project Info Sheet (the "what") and the Schedule Builder (the "when"). Its job is to hold the "how" as structured, attributable data: each answer has a value, a confidence level, a rationale, citations, who verified it and when, and a full history of how it changed. This article covers how it is built and the pattern underneath it: capture the expert's reasoning as data, not as chat.
Why not just let the agent ask in chat?
The obvious design is to let the scheduling agent ask the user how the building will go up and carry on. We tried versions of that, and three problems kept coming back:
- Chat has no confidence levels. "Probably sheet piling" and "the client confirmed sheet piling" look the same in a transcript. A schedule built on the first should be treated very differently from one built on the second.
- Chat cannot be handed to the client. Many "how" questions can only be answered by the owner or their construction manager. A chat thread is not something you can send them.
- Chat cannot be audited. When a date slips and someone asks why the schedule assumed a crane, "it came up in a conversation in March" is not an answer.
So Means & Methods became its own artifact with its own store, and the Schedule Builder reads from it instead of re-asking.
The artifact
Migration 005 created one table, means_methods_artifacts, with one row per project's primary Info Sheet. Each row records:
- The source binding: the Info Sheet id and the exact approved Profile and Facts revisions the answers were built against, plus a catalog version for the topic vocabulary.
- A status:
draft,in_revieworverified. - A revision counter used for optimistic concurrency.
- The items: a JSONB map keyed
<subject>.<topicKey>, mirroring how Info Sheet facts are keyed. - A client email draft.
Each item is more than a value:
| Field | Purpose |
|---|---|
value | The chosen method, or null if open |
provenance | Where the answer sits on the confidence ladder (below) |
rationale | Why this answer, in a sentence or two |
question | For an open item, the project-specific question to put to the client |
note | Free-form context an expert adds on top of a chosen option ("sheet piling, tie-backs on the north face only") |
options | The choices offered, from a domain catalog and from the AI, with one marked recommended |
citations | The documents that support the answer |
verifiedBy, verifiedAt | Who signed it off and when |
The topic vocabulary is not a second list. It lives in the same requirements catalog as the Info Sheet, where each topic is tagged by nature: "what" topics belong to the Info Sheet and "how" topics to Means & Methods. One vocabulary means the two features cannot drift apart. Adding "how" topics (we later added dewatering, temporary power, winter protection, concrete curing and permit long-leads) deliberately changes the catalog version, so existing artifacts know they were built against an older list.
Topics are also filtered by project type. A film studio new build should not be asked about traffic phasing on an occupied hospital campus. If the job type is unknown, the full list is used, so nothing is dropped silently.
Choosing, not typing
An empty text box is a bad interface for an expert. Most topics have a short list of methods people actually use, so each carries a catalog of standard options (for support of excavation: open cut, soldier pile and lagging, sheet piling, secant pile wall, slurry wall). The drafter ranks them for the project, marks one recommended with a rationale, and can add project-specific options. The expert picks one, or writes their own, and adds a note.
The confidence ladder
The core of the design is a four-step provenance field:
| Provenance | Meaning | Who can set it |
|---|---|---|
assumed | Inferred from the Info Sheet or documents. Provisional, not reviewed | The AI drafter, or a person |
client_needed | Cannot be answered internally. On the question list for the client | The AI drafter, or a person |
confirmed | An expert at the scheduling consultancy confirmed or entered it | A person only |
client_verified | The client signed off on this exact answer | The client, through a share link |
Three rules make the ladder mean something:
- The AI can only produce the bottom two steps. The parser that reads the drafter's output allows only
assumedorclient_needed, and an item with no value is forced toclient_needed. A confident tone does not change the step. - Nothing automated overwrites the top two steps. Re-running the drafter, or pulling answers from a client's email, skips any item that is already
confirmedorclient_verified. An expert's decision survives every later AI pass. - A client can only answer what was asked. Through a share link, a client can fill in
client_neededitems, which becomeclient_verified. They cannot overwrite something the consultancy already confirmed.
The AI drafter
The drafter runs on pi-agent-core, the runtime Connect uses for its narrower agents. It gets the approved Info Sheet text plus a seed checklist of the project-type topics and their standard options, and three read-only tools:
list_project_files: what evidence exists on the projectsearch_project_documents: semantic search over the project's indexed documents, returning numbered passages to citeread_project_file: the extracted text of one text-like document such as a geotechnical report or spec
Workspace and project are bound on the server. The model only supplies a query or a file id inside the project, so it cannot widen its own scope. Tool failures come back as readable text ("the document index could not be searched") so the agent can carry on instead of the whole turn failing. A deployment without retrieval falls back to drafting from the Info Sheet alone.
The drafter is told to investigate before answering, to drop topics the evidence rules out, and may propose custom topics the catalog does not name if each has a label and a section. It gets up to 24 tool turns and a 16,000-token output budget, because a reasoning model spends part of the budget thinking and JSON cut off mid-object parses to nothing.
Then comes a second pass. A separate model call, on the runtime's child model, rereads the draft against a written quality bar: every assumed value must be supported by a citation, or it becomes a specific client_needed question; items that do not apply are dropped; vague questions are rewritten to name the real project condition; and a clear client_needed is preferred over a confident guess. If this pass fails or returns nothing, the original draft is kept. It can improve a draft but never lose one.
The drafter's system prompt is configurable from the executive admin portal, and every turn, including its tool calls, lands in the tracing ledger as a means_methods turn.
The append-only events log
The items map holds the latest state. History lives somewhere else. Migration 007 added means_methods_events, an append-only log with one row per change:
- action:
save,prepopulate,status,client_emailorclient_submit - actor kind:
user,aiorclient, plus an actor id and label - previous value and next value, so the change can be diffed
- provenance and an optional note
Rows are never updated or deleted. The History panel reads them newest first.
The detail that makes this trustworthy is the write path. Every change to the artifact and its events commit in the same transaction, and the artifact write is conditional on the revision:
BEGIN;
UPDATE artifact SET items = $new, revision = $rev
WHERE key = $key AND revision = $rev - 1; -- 0 rows -> conflict, roll back
INSERT INTO events (...) VALUES (...); -- one row per changed topic
COMMIT;
If two people save at once, one gets a conflict and is asked to reload. An event can never be written without its change, and a change can never be written without its event. That is what turns the log from a best-effort audit into a system of record.
Client shares with no account needed
Migration 008 added means_methods_shares, which lets the consultancy send the open questions straight to the client:
- A user creates a link. The server generates a high-entropy token (the credential in the URL) and a short human-readable code, to be read out or emailed separately as a second factor.
- Only hashes of both are stored. The plaintext is returned once and goes into the email. Creating a new link revokes the old one, and links expire after 30 days.
- The client opens the link and enters the code. Any failure (wrong token, wrong code, expired, revoked) returns the same "invalid or expired" response, so a failed attempt reveals nothing about which part was wrong.
- The client sees only the
client_neededtopics, each phrased as the drafter's specific question where one exists, with the options to choose from. - Submitted answers become
client_verified, each with aclient_submitevent attributed to the client, and the share records when it was opened and submitted.
Some clients reply by email or PDF instead. A user can paste the reply or upload a PDF with selectable text (plain extraction, not OCR; a scan is rejected). The text is screened by the safety layer, and the drafter maps it onto only the topics it addresses. Those answers land as assumed with the note "From client response", not client_verified: pasted text is a lead to review, not a sign-off.
The same hashed-token-plus-code pattern was later reused for job walk reports and client Updates.
Staleness and explicit review
Info Sheets are revised and re-approved. When that happens, the "how" may no longer fit the "what". Because each artifact records the revisions it was built against, the view can flag it as stale, and the portfolio hub counts stale projects alongside open questions, unverified drafts and confirmed items, all derived from the stored items rather than extra columns.
Ordinary saves keep the existing binding. Only an explicit review renews it. To mark the artifact verified, the client sends the revision it displayed and the source revisions it showed the user. If either has changed, the server rejects the request and the user has to reload and review again. A verify event records the person and the exact revisions they reviewed.
The Playwright test e2e/means-methods-review.spec.ts pins this down at phone and desktop widths, in light and dark themes. It loads a stale artifact, confirms with the keyboard, and gets a conflict as if the Info Sheet changed underneath. It then checks that an alert asks for a reload, that the reload names the new revisions to review against, that the second confirm clears the stale state, and that the page never scrolls sideways.
How Means & Methods feeds the Schedule Builder
The Schedule Builder does not read the items map. It asks the Means & Methods module for a rendered block of text, and the module decides what goes in:
- Confirmed and client-verified answers go in under "build to these", with client-verified ones tagged as such and any expert note included.
- Client-needed topics go in under "awaiting client, do not assume".
- Assumed answers do not go in at all. An AI draft nobody has reviewed never reaches the schedule.
- The block opens with a line naming the exact Info Sheet and revisions it applies to, so the builder cannot stretch these answers to another sheet in the same build.
Failures are handled by kind. No Means & Methods at all is fine: the builder falls back to its own interview. A stale artifact, or a mismatch with the Info Sheet revision selected for the build, stops the turn before the model runs. A read that fails for any other reason surfaces a retryable error rather than quietly building without the planning context. Missing input is a gap in the interview, but unreadable input must not silently change the plan.
The same steps are exposed as tools on the Connect MCP server, so a scheduler working from another MCP client follows the same ladder and leaves the same events.
Grading the drafter
Migration 054 switched on automatic evaluation for this agent. The eval harness already ran across every configurable agent kind but skipped any kind without an active rubric, and Means & Methods had never had one, so its turns were going ungraded. The migration widened the allowed agent kinds and seeded a rubric with weighted criteria that match the ladder:
- Grounding: each value is supported by the Info Sheet or marked
client_needed - Provenance:
assumedandclient_neededused correctly, confidence not overstated, locked answers respected - Fit to project: topics and options suit the project type
- Reasoning: rationales are short and options are realistic and distinct
- Completeness: open decisions are raised for the client instead of guessed
Executives can edit or replace the rubric, and only one can be active per kind. The wider harness is covered in LLM evals, golden sets and autopilot.
Why this matters if you're building something similar
- Write down how sure you are. Do not store a bare value. Store how confident it is and who said so. A four-step provenance field does more for trust than any amount of prompt work.
- Cap what the AI can claim. Let the model write only the bottom of the ladder, enforce it in the parser, and never let an automated pass overwrite a human decision.
- Keep current state and history separate, and commit them together. A JSONB map for current state plus an append-only events table, written in one transaction with a revision check, gives you fast reads and a real audit trail.
- Bind derived data to the revision it came from. Then staleness is a comparison, not a guess, and re-approval becomes an explicit review rather than a silent drift.
- Send outside experts a link, not an account. A hashed token, a short human code, one active link and a uniform failure response is a small amount of code that removes the biggest source of delay: waiting on the client.
- Downstream agents should read only reviewed data. The Schedule Builder sees confirmed answers and open questions, never unreviewed drafts. That one filter is the difference between AI-assisted planning and AI-generated guesses.
Where to go next
- The series overview: How we built Connect
- The step before this one: Two-phase AI agent workflow with human approval
- The same share and audit pattern applied to field reports: AI job walks: field capture to schedule actuals
- How corrections feed back into the agents: The AI learning loop from human corrections
- The same pattern for client communication: Versioned client project updates
Want your senior people's judgment captured as data your systems can use? Plan your build with us.
Frequently asked questions
What does means and methods mean in construction scheduling?
It is how a building will be built: excavation support, crane and hoist strategy, facade installation, sequencing and lead times. These decisions drive durations and logic but usually cannot be read from the drawings.
Can the AI confirm its own answers?
No. The drafter can only mark an answer assumed or client_needed. Confirmed requires a person at the consultancy, and client_verified requires the client through a share link.
How does the client answer without a login?
They receive an unguessable link and a separate short code. Both are stored only as hashes, links expire after 30 days, and the client sees and can answer only the open questions.
What happens when the Info Sheet is re-approved?
The artifact is flagged stale because it records the revisions it was built from. The Schedule Builder will not use a stale artifact, and verifying again requires an explicit review against the new revisions.
How is the AI drafter's quality measured?
Its turns are graded automatically against a rubric covering grounding, correct provenance, fit to the project, reasoning and completeness, which executives can edit.
Next step
Thinking about a scheduling pilot?
Start with one exported schedule and an agreed set of checks. We'll discuss your format, method, and approval process.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
How We Built Connect: Architecture of an Enterprise AI Platform
The pillar of our Connect architecture series: a layer-by-layer map of an enterprise agentic AI platform, the three-process shape it runs as, the design principles that kept recurring, and links to every deep-dive article.
AI agents & MCP · October 9, 2026
Two-Phase AI Agents with Human Approval: How We Built the Connect Info Sheet
A research agent builds a cited Profile, a person approves it through a server route, and a tool-less normalizer converts it into Facts limited to a fixed vocabulary. The pattern applies to any AI that writes to a system of record.
Scheduling & P6 · October 9, 2026
Engineering an AI Schedule Builder: Build Specs, Gates and Repair Loops
The engineering behind the Connect Schedule Builder loop: the model edits a declarative Build Spec, code expands and gates it, reviewers and bounded repair passes drive it to done, and approval freezes verified evidence.