Build Flows

Scheduling & P6 · October 10, 2026 · 11 min read

Means & Methods as a system of record: capturing expert reasoning as data, not chat

An engineering walkthrough of Connect's Means & Methods step: one artifact per project, a four-step confidence ladder, an append-only events log, hashed client share links, and a Schedule Builder that only reads reviewed answers.

By Charley Forey, founder of Build Flows

Every schedule depends on decisions that are not in the drawings. Will excavation use soldier piles and lagging or sheet piling? Is there a tower crane or only a hoist? Does the facade go up from mast climbers or swing stages? How long is the elevator lead time? A senior scheduler answers these from experience, from a phone call with the client, or from a line buried in an email thread. Then the answers disappear into a chat log or someone's memory, and three weeks later nobody can say why the schedule assumed a crane.

Means & Methods is the step in Connect that keeps those answers. It sits between the approved Project Info Sheet (the "what") and the Schedule Builder (the "when"). Its job is to hold the "how" as structured, attributable data: each answer has a value, a confidence level, a rationale, citations, who verified it and when, and a full history of how it changed. This article covers how it is built and the pattern underneath it: capture the expert's reasoning as data, not as chat.

Why not just let the agent ask in chat?

The obvious design is to let the scheduling agent ask the user how the building will go up and carry on. We tried versions of that, and three problems kept coming back:

  • Chat has no confidence levels. "Probably sheet piling" and "the client confirmed sheet piling" look the same in a transcript. A schedule built on the first should be treated very differently from one built on the second.
  • Chat cannot be handed to the client. Many "how" questions can only be answered by the owner or their construction manager. A chat thread is not something you can send them.
  • Chat cannot be audited. When a date slips and someone asks why the schedule assumed a crane, "it came up in a conversation in March" is not an answer.

So Means & Methods became its own artifact with its own store, and the Schedule Builder reads from it instead of re-asking.

The artifact

Migration 005 created one table, means_methods_artifacts, with one row per project's primary Info Sheet. Each row records:

  • The source binding: the Info Sheet id and the exact approved Profile and Facts revisions the answers were built against, plus a catalog version for the topic vocabulary.
  • A status: draft, in_review or verified.
  • A revision counter used for optimistic concurrency.
  • The items: a JSONB map keyed <subject>.<topicKey>, mirroring how Info Sheet facts are keyed.
  • A client email draft.

Each item is more than a value:

FieldPurpose
valueThe chosen method, or null if open
provenanceWhere the answer sits on the confidence ladder (below)
rationaleWhy this answer, in a sentence or two
questionFor an open item, the project-specific question to put to the client
noteFree-form context an expert adds on top of a chosen option ("sheet piling, tie-backs on the north face only")
optionsThe choices offered, from a domain catalog and from the AI, with one marked recommended
citationsThe documents that support the answer
verifiedBy, verifiedAtWho signed it off and when

The topic vocabulary is not a second list. It lives in the same requirements catalog as the Info Sheet, where each topic is tagged by nature: "what" topics belong to the Info Sheet and "how" topics to Means & Methods. One vocabulary means the two features cannot drift apart. Adding "how" topics (we later added dewatering, temporary power, winter protection, concrete curing and permit long-leads) deliberately changes the catalog version, so existing artifacts know they were built against an older list.

Topics are also filtered by project type. A film studio new build should not be asked about traffic phasing on an occupied hospital campus. If the job type is unknown, the full list is used, so nothing is dropped silently.

Choosing, not typing

An empty text box is a bad interface for an expert. Most topics have a short list of methods people actually use, so each carries a catalog of standard options (for support of excavation: open cut, soldier pile and lagging, sheet piling, secant pile wall, slurry wall). The drafter ranks them for the project, marks one recommended with a rationale, and can add project-specific options. The expert picks one, or writes their own, and adds a note.

The confidence ladder

The core of the design is a four-step provenance field:

ProvenanceMeaningWho can set it
assumedInferred from the Info Sheet or documents. Provisional, not reviewedThe AI drafter, or a person
client_neededCannot be answered internally. On the question list for the clientThe AI drafter, or a person
confirmedAn expert at the scheduling consultancy confirmed or entered itA person only
client_verifiedThe client signed off on this exact answerThe client, through a share link

Three rules make the ladder mean something:

  1. The AI can only produce the bottom two steps. The parser that reads the drafter's output allows only assumed or client_needed, and an item with no value is forced to client_needed. A confident tone does not change the step.
  2. Nothing automated overwrites the top two steps. Re-running the drafter, or pulling answers from a client's email, skips any item that is already confirmed or client_verified. An expert's decision survives every later AI pass.
  3. A client can only answer what was asked. Through a share link, a client can fill in client_needed items, which become client_verified. They cannot overwrite something the consultancy already confirmed.

The AI drafter

The drafter runs on pi-agent-core, the runtime Connect uses for its narrower agents. It gets the approved Info Sheet text plus a seed checklist of the project-type topics and their standard options, and three read-only tools:

  • list_project_files: what evidence exists on the project
  • search_project_documents: semantic search over the project's indexed documents, returning numbered passages to cite
  • read_project_file: the extracted text of one text-like document such as a geotechnical report or spec

Workspace and project are bound on the server. The model only supplies a query or a file id inside the project, so it cannot widen its own scope. Tool failures come back as readable text ("the document index could not be searched") so the agent can carry on instead of the whole turn failing. A deployment without retrieval falls back to drafting from the Info Sheet alone.

The drafter is told to investigate before answering, to drop topics the evidence rules out, and may propose custom topics the catalog does not name if each has a label and a section. It gets up to 24 tool turns and a 16,000-token output budget, because a reasoning model spends part of the budget thinking and JSON cut off mid-object parses to nothing.

Then comes a second pass. A separate model call, on the runtime's child model, rereads the draft against a written quality bar: every assumed value must be supported by a citation, or it becomes a specific client_needed question; items that do not apply are dropped; vague questions are rewritten to name the real project condition; and a clear client_needed is preferred over a confident guess. If this pass fails or returns nothing, the original draft is kept. It can improve a draft but never lose one.

The drafter's system prompt is configurable from the executive admin portal, and every turn, including its tool calls, lands in the tracing ledger as a means_methods turn.

The append-only events log

The items map holds the latest state. History lives somewhere else. Migration 007 added means_methods_events, an append-only log with one row per change:

  • action: save, prepopulate, status, client_email or client_submit
  • actor kind: user, ai or client, plus an actor id and label
  • previous value and next value, so the change can be diffed
  • provenance and an optional note

Rows are never updated or deleted. The History panel reads them newest first.

The detail that makes this trustworthy is the write path. Every change to the artifact and its events commit in the same transaction, and the artifact write is conditional on the revision:

BEGIN;
UPDATE artifact SET items = $new, revision = $rev
  WHERE key = $key AND revision = $rev - 1;   -- 0 rows -> conflict, roll back
INSERT INTO events (...) VALUES (...);         -- one row per changed topic
COMMIT;

If two people save at once, one gets a conflict and is asked to reload. An event can never be written without its change, and a change can never be written without its event. That is what turns the log from a best-effort audit into a system of record.

Client shares with no account needed

Migration 008 added means_methods_shares, which lets the consultancy send the open questions straight to the client:

  1. A user creates a link. The server generates a high-entropy token (the credential in the URL) and a short human-readable code, to be read out or emailed separately as a second factor.
  2. Only hashes of both are stored. The plaintext is returned once and goes into the email. Creating a new link revokes the old one, and links expire after 30 days.
  3. The client opens the link and enters the code. Any failure (wrong token, wrong code, expired, revoked) returns the same "invalid or expired" response, so a failed attempt reveals nothing about which part was wrong.
  4. The client sees only the client_needed topics, each phrased as the drafter's specific question where one exists, with the options to choose from.
  5. Submitted answers become client_verified, each with a client_submit event attributed to the client, and the share records when it was opened and submitted.

Some clients reply by email or PDF instead. A user can paste the reply or upload a PDF with selectable text (plain extraction, not OCR; a scan is rejected). The text is screened by the safety layer, and the drafter maps it onto only the topics it addresses. Those answers land as assumed with the note "From client response", not client_verified: pasted text is a lead to review, not a sign-off.

The same hashed-token-plus-code pattern was later reused for job walk reports and client Updates.

Staleness and explicit review

Info Sheets are revised and re-approved. When that happens, the "how" may no longer fit the "what". Because each artifact records the revisions it was built against, the view can flag it as stale, and the portfolio hub counts stale projects alongside open questions, unverified drafts and confirmed items, all derived from the stored items rather than extra columns.

Ordinary saves keep the existing binding. Only an explicit review renews it. To mark the artifact verified, the client sends the revision it displayed and the source revisions it showed the user. If either has changed, the server rejects the request and the user has to reload and review again. A verify event records the person and the exact revisions they reviewed.

The Playwright test e2e/means-methods-review.spec.ts pins this down at phone and desktop widths, in light and dark themes. It loads a stale artifact, confirms with the keyboard, and gets a conflict as if the Info Sheet changed underneath. It then checks that an alert asks for a reload, that the reload names the new revisions to review against, that the second confirm clears the stale state, and that the page never scrolls sideways.

How Means & Methods feeds the Schedule Builder

The Schedule Builder does not read the items map. It asks the Means & Methods module for a rendered block of text, and the module decides what goes in:

  • Confirmed and client-verified answers go in under "build to these", with client-verified ones tagged as such and any expert note included.
  • Client-needed topics go in under "awaiting client, do not assume".
  • Assumed answers do not go in at all. An AI draft nobody has reviewed never reaches the schedule.
  • The block opens with a line naming the exact Info Sheet and revisions it applies to, so the builder cannot stretch these answers to another sheet in the same build.

Failures are handled by kind. No Means & Methods at all is fine: the builder falls back to its own interview. A stale artifact, or a mismatch with the Info Sheet revision selected for the build, stops the turn before the model runs. A read that fails for any other reason surfaces a retryable error rather than quietly building without the planning context. Missing input is a gap in the interview, but unreadable input must not silently change the plan.

The same steps are exposed as tools on the Connect MCP server, so a scheduler working from another MCP client follows the same ladder and leaves the same events.

Grading the drafter

Migration 054 switched on automatic evaluation for this agent. The eval harness already ran across every configurable agent kind but skipped any kind without an active rubric, and Means & Methods had never had one, so its turns were going ungraded. The migration widened the allowed agent kinds and seeded a rubric with weighted criteria that match the ladder:

  • Grounding: each value is supported by the Info Sheet or marked client_needed
  • Provenance: assumed and client_needed used correctly, confidence not overstated, locked answers respected
  • Fit to project: topics and options suit the project type
  • Reasoning: rationales are short and options are realistic and distinct
  • Completeness: open decisions are raised for the client instead of guessed

Executives can edit or replace the rubric, and only one can be active per kind. The wider harness is covered in LLM evals, golden sets and autopilot.

Why this matters if you're building something similar

  • Write down how sure you are. Do not store a bare value. Store how confident it is and who said so. A four-step provenance field does more for trust than any amount of prompt work.
  • Cap what the AI can claim. Let the model write only the bottom of the ladder, enforce it in the parser, and never let an automated pass overwrite a human decision.
  • Keep current state and history separate, and commit them together. A JSONB map for current state plus an append-only events table, written in one transaction with a revision check, gives you fast reads and a real audit trail.
  • Bind derived data to the revision it came from. Then staleness is a comparison, not a guess, and re-approval becomes an explicit review rather than a silent drift.
  • Send outside experts a link, not an account. A hashed token, a short human code, one active link and a uniform failure response is a small amount of code that removes the biggest source of delay: waiting on the client.
  • Downstream agents should read only reviewed data. The Schedule Builder sees confirmed answers and open questions, never unreviewed drafts. That one filter is the difference between AI-assisted planning and AI-generated guesses.

Where to go next

Want your senior people's judgment captured as data your systems can use? Plan your build with us.

Frequently asked questions

What does means and methods mean in construction scheduling?

It is how a building will be built: excavation support, crane and hoist strategy, facade installation, sequencing and lead times. These decisions drive durations and logic but usually cannot be read from the drawings.

Can the AI confirm its own answers?

No. The drafter can only mark an answer assumed or client_needed. Confirmed requires a person at the consultancy, and client_verified requires the client through a share link.

How does the client answer without a login?

They receive an unguessable link and a separate short code. Both are stored only as hashes, links expire after 30 days, and the client sees and can answer only the open questions.

What happens when the Info Sheet is re-approved?

The artifact is flagged stale because it records the revisions it was built from. The Schedule Builder will not use a stale artifact, and verifying again requires an explicit review against the new revisions.

How is the AI drafter's quality measured?

Its turns are graded automatically against a rubric covering grounding, correct provenance, fit to the project, reasoning and completeness, which executives can edit.

Next step

Thinking about a scheduling pilot?

Start with one exported schedule and an agreed set of checks. We'll discuss your format, method, and approval process.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning