Every AI product demo eventually gets the question: "Does it learn?" For a scheduling agent the honest answer has to be "yes, but only from things a person has vouched for." A planner who drags a concrete pour from 4 days to 6 has taught the system something. It might be a lesson about how long pours take. It might also mean this floor has twice the slab area, or that the job runs a four-day week in winter. A system that treats every edit as a duration lesson will slowly train itself into nonsense, and nobody will be able to say when it happened.
This article covers the learning loop we built into the Connect Schedule Builder: what it records, what it computes, what it feeds back to the model, and the one narrow, human-gated path by which a correction is allowed to change another project's estimate. It is part of our series on the Connect architecture. The short version is that the system learns plenty on its own from data, and almost nothing on its own from opinions.
Three kinds of signal, three levels of trust
We separate what the system learns from into three tiers, because they deserve very different amounts of trust.
| Signal | Example | Who vouches for it | What it can change |
|---|---|---|---|
| Statistical history | Durations in approved baselines, imported reference schedules, field actuals | The approval itself, or the field | Duration benchmarks, dependency suggestions, mined norms |
| Corrections | A rejected proposal with a reason, a manual Gantt edit, a lost A/B comparison | Nobody yet | The next build prompt for the same project only |
| Promoted knowledge | "Pours on this archetype run about 1.3x what we estimate, reason: crew size" | A named reviewer, versioned and reversible | Every future estimate for that archetype in that workspace |
Most of the engineering is in keeping those tiers from leaking into each other.
Data flows into benchmarks automatically. Corrections reach another estimate only through a reviewer. Dashed pieces are designed, not running yet.
The append-only correction log
The foundation is a single table, schedule_builder_learning. It is append-only by design: rows are never updated or deleted. Each row carries the workspace, project and schedule it belongs to, a kind, a short human summary, an optional free-text reason, who made it and when.
It grew in three steps, and each step fixed a specific weakness.
- The first version recorded two kinds: a rejected proposal (with the reason the user gave) and a manual edit. That was enough to feed corrections back into the prompt, but it said nothing about why a duration changed.
- A
reason_categorycolumn came next, with a fixed vocabulary: scope change, crew size, method change, calendar change, historical slip, estimation error. Older rows keep a null category and are never promotable. An unrecognized category from a client is stored as null rather than rejected, because recording a correction must never fail the action that produced it. - A
payloadjsonb column came last, designed so that distillation never has to parse free-text summaries: the schema is ready for manual edits to carry machine-readable detail (task code, archetype, field, from and to values), and wiring the writers to fill it is the next step. Auser_feedbackkind, from a post-approval rating, joined the log, and anoutcome_variancekind for plan-versus-actual retrospectives is designed in. Thekindcolumn has no database CHECK constraint; allowed values are enforced in the module, as they were for the original two.
Why append-only? Because a correction log is evidence. When an estimate looks wrong six months from now, you want to replay exactly which corrections existed at the time.
One related detail: the A/B tool that grades two builds of the same project prefills a user_feedback correction for the loser but never writes it. Recording is a separate, explicit call, because a read-only comparison should never mutate the learning corpus.
Feeding corrections back into the prompt
The cheapest form of learning is context. When the Schedule Builder starts a turn, it asks for a compact digest of the eight most recent corrections for that project and includes them in its instructions. Each line looks roughly like this:
- Rejected proposal [crew size]: Shortened curtain wall install on levels 3-9 — reason: two crews, not one
- Manual edit [scope change]: Added podium waterproofing activity
The same digest is exposed through the Connect MCP, so an external agent gets the same grounding.
Two choices matter. The digest is scoped to the project, so one job's preferences cannot bleed into another through the prompt. And it is a prompt, not a parameter: context to weigh, not a rule to obey. A wrong correction stops mattering once it scrolls out of the recent window.
The design records a second feedback channel as a next step: reviewer and sequencing findings that recur across at least three distinct builds, distilled into a "you keep getting flagged for X, do Y" block in the same brief. Counting distinct builds rather than raw findings is the point: one build raising the same issue three ways is noise, three builds raising it once each is a habit.
Why we turned off automatic learning
The first version of this loop did the obvious thing. It averaged manual edits and revised schedule imports per activity archetype into a correction_priors insight, and the estimator folded that average into its p50.
We removed that path. The decision record is blunt: an edit is not permission to teach another project. The averages had no reason attached, no human approval, and no way to roll back. A planner who stretched a duration because the scope doubled was, as far as the estimator could tell, saying the work takes twice as long. The existing rows stay stored for audit, and the estimator no longer reads them.
That is the core lesson here. A correction is an opinion about one project, and generalizing it is a judgement call. Judgement calls need a person.
Knowledge promotions: the only correction that changes estimates
What replaced it is schedule_knowledge_promotions. A promotion is a reviewer's deliberate decision to turn a project observation into firm knowledge for one activity archetype. Its rules are enforced in code and in the schema:
- Reviewer-stamped. Only a workspace admin can promote, and the row records who and when. A revert records its own actor and time.
- Duration lessons only. Crew size, method change, historical slip and estimation error are promotable. Scope change and calendar change are refused with an explanation, because neither is a lesson about how long the work itself takes.
- Versioned, not overwritten. Each (workspace, archetype) has a monotonic version. Promoting again supersedes the prior row instead of replacing it, so the table reads as a history.
- At most one active. A partial unique index allows exactly one active promotion per archetype per workspace. The estimator applies one lesson, never a stack of ratios.
- Reversible. A revert flips the status and the estimate stops using it on the next read.
- Workspace-only. There is no firm-wide scope for promotions. One tenant's curated judgement never reaches another tenant's estimate.
In SQL terms, the guard that matters most is small:
CREATE UNIQUE INDEX one_active_lesson
ON knowledge_promotions (workspace_id, archetype)
WHERE status = 'active';
The promotion itself stores a mean ratio (above 1 means the archetype historically ran long), the sample count behind it, the reason category, and a required source summary saying which observation it came from. The decision record is careful about what this is: a mechanism to record and apply a reviewed lesson, not a claim that the lesson is correct.
Learning from history: recomputeBenchmarks
Separately from corrections, the estimating engine learns from history, and this part runs without a human in the loop because its inputs have already been vouched for.
The learning key is the archetype: an activity code with its trailing numeric segments stripped, so CONC.12 and CONC.14 both roll up to CONC. When there is no code, a normalized name (digits and floor words removed, first few words kept) is used as a noisier fallback.
recomputeBenchmarks rebuilds every statistic from scratch from four sources:
- Planned durations from imported schedules in Syncify, the seed corpus.
- Realized actuals from those schedules, which feed both duration learning and the variance projection.
- Approved Connect baselines, which grow the corpus on every approval.
- Reference schedules (prior jobs imported as .xer or auto-registered), for durations and dependency logic, but not for variance, since they are examples rather than this workspace's plan.
Every sample lands in two pools: its own workspace and a global pool across the deployment. Per pool and archetype the engine stores count, mean, standard deviation, p10, p50, p90, min and max. Dependency links are counted per archetype pair with typical lag and link-type mix. Planned and actual distributions are kept apart for comparison.
It runs on two triggers: a worker check on every poll that self-throttles to once a day, and every approval. Both log a run row with the result or error, so a failed learning run is visible.
Note the asymmetry with promotions: raw statistics pool across workspaces as a prior, curated judgement does not.
estimate(): shrinkage toward the prior
When the builder needs a duration, it calls estimate() for the archetype. The interesting problem is small samples. A workspace with two past pours should not trust its own mean over a global pool of hundreds, but a workspace with fifty should.
We use a credibility-weighted blend, in the spirit of James-Stein shrinkage:
w = n_ws / (n_ws + k)
estimate = w * mean_ws + (1 - w) * mean_global
k is the sample count at which the two are weighted equally; we use 5. With 2 local samples the workspace gets about 29% of the weight. With 20 it gets 80%. With 50, about 91%. If one side is missing, the other is used alone.
The estimate also returns a spread (p10, p50, p90 and standard deviation) from the workspace pool when it exists, otherwise from the global pool, plus a source field saying whether the answer was blended, workspace-only or global. The builder is told to cite variance, not a single false-precision number.
Two optional layers ride on top:
- Slip. If the offline miner has a planned-versus-realized history for the archetype with at least five completed samples behind it, the estimate attaches the mean slip ratio and volatility. Below that threshold it says nothing.
- Correction. If, and only if, there is an active promotion for this archetype in this workspace, the p50 and the blended estimate are multiplied by its ratio, and the estimate carries the ratio and sample count as provenance. This is the only correction signal that reaches an estimate.
So a cited duration can always be traced back to its history, its prior, its mined slip and any promotion a named person approved.
suggestDependencies follows the same shape for logic. It merges the most common predecessors and successors for an archetype, workspace history first and the global pool filling pairs the workspace has not seen, with typical lag, sample count, link-type mix and dominant type, capped at eight in each direction.
Mined insights: offline and advisory
A separate offline miner (mining.ts holds the pure core, mine-schedules.ts the I/O) reads a corpus of schedule snapshots, one per revision. Snapshot integer IDs are not stable across revisions, so revisions join on task code. Only activities that completed produce an evolution sample: planned duration at first appearance versus realized actual. Level-of-effort activities are dropped, because as reference content they would look like multi-year tasks and poison the duration benchmarks.
The miner writes insight rows by kind:
| Kind | What it captures |
|---|---|
| Archetype evolution | Planned versus realized duration, slip ratio, volatility |
| WBS playbook | Typical top-level section structure by sector and project type |
| Project growth | Median growth in activity count and finish drift per update |
| Correction priors | Retained for audit, no longer applied to estimates |
The schema is already extended for the next kinds: schedule norms (distributions of structural metrics such as activity count, WBS depth, logic density and FS share), dependency pairs stratified by project type, phase strategy for phased builds, and milestone families. They are designed under the same rule as the rest: advisory and additive, with byte-identical builder behavior when a project type has nothing learned. That is what makes each one safe to ship.
Closing the loop: approval and retrospectives
Approval is where a draft becomes evidence. When a baseline is approved, two best-effort steps run after the commit, and neither can fail the approval:
- A logged benchmark recompute, so the new durations and logic join the pools.
learnFromApproval, which registers the approved revision as a reference schedule keyed to the baseline. Re-approval upserts the same row rather than adding a duplicate.
The next step in the design is the retrospective. The schema already has a table for it: one row per schedule and approved revision, comparing an approved agent-built baseline with that project's field actuals, with a headline verdict (matched archetypes, share within 20% of plan, mean ratio) and per-archetype detail. The plan is to compare distributions rather than pair activities one to one, which holds up when two systems key activities differently, and the admin portal's forecast-accuracy metric already reads from that table.
Once retrospectives are being written, they give a reviewer exactly the evidence a promotion needs. The loop closes, but a human closes it.
Why this matters if you're building something similar
- Classify the why at capture time. A correction without a reason is almost useless for learning. Ask for a category when the edit happens, from a short fixed list, and store null when you do not get one.
- Keep the log append-only. You will want to replay history when an estimate is questioned. Never let a cleanup job delete evidence.
- Let context learn freely and parameters learn carefully. Feeding recent corrections into the prompt is low-risk and reversible by time. Changing a number every future estimate uses deserves a reviewer, a version and an undo.
- Shrink small samples toward a prior. A credibility weight like n / (n + k) is a few lines of code and prevents a workspace with two data points from overriding hundreds.
- Pool data, not judgement. Sharing raw statistics across tenants as a prior is reasonable. Sharing one tenant's curated lessons is not.
- Make every estimate explain itself. Return source, sample count, spread and any applied correction with the number.
- Be willing to turn learning off. Our first automatic path was the obvious design, and it was wrong. Keeping the data while removing it from the estimate path cost very little.
Where to go next
- The Connect architecture pillar for how the learning loop fits the wider platform.
- Engineering the AI schedule builder: Build Specs, gates and repair loops for the build loop that consumes these estimates.
- LLM evals with golden sets and autopilot for how we measure whether the agent is getting better.
- Two-phase AI agent workflows with human approval for the same human-gate pattern applied to agent actions.
- Governing AI agents in construction for the oversight layer around all of this.
If you want a learning loop like this designed around your own data and review process, tell us what you are building.
Frequently asked questions
Should an AI scheduling agent learn automatically from planner edits?
Not directly into estimates. A planner's edit might reflect a scope change, a calendar change or a real duration lesson. We let recent corrections inform the next prompt for the same project, but only a human-reviewed promotion can change estimates for other projects.
What is a knowledge promotion?
A reviewer's deliberate decision to turn an observed duration correction for one activity archetype into firm knowledge. It records who approved it, the ratio and sample count, and the reason category. It is versioned, at most one is active per archetype, it can be reverted, and it never leaves the workspace.
How do you estimate durations when a firm has little history?
We blend the workspace mean with a global prior using a credibility weight of n / (n + k), with k set to 5. With few local samples the prior dominates; with many, local history does. The estimate also returns p10, p50, p90 and where the numbers came from.
Why keep the correction log append-only?
Because it is evidence. When an estimate is questioned later, you need to replay exactly which corrections existed at the time. Rows are never updated or deleted.
How does the system learn from what actually happened in the field?
Realized actuals already feed the duration benchmarks and slip insights. A per-project retrospective is the next step: the schema already holds a retrospective table that compares an approved agent-built baseline with the project's field actuals per activity archetype; once populated, it gives reviewers the evidence a promotion needs.
Next step
Thinking about a scheduling pilot?
Start with one exported schedule and an agreed set of checks. We'll discuss your format, method, and approval process.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
How We Built Connect: Architecture of an Enterprise AI Platform
The pillar of our Connect architecture series: a layer-by-layer map of an enterprise agentic AI platform, the three-process shape it runs as, the design principles that kept recurring, and links to every deep-dive article.
Scheduling & P6 · October 9, 2026
Engineering an AI Schedule Builder: Build Specs, Gates and Repair Loops
The engineering behind the Connect Schedule Builder loop: the model edits a declarative Build Spec, code expands and gates it, reviewers and bounded repair passes drive it to done, and approval freezes verified evidence.
AI agents & MCP · October 9, 2026
LLM Evals in Production: Judges, Golden Sets, Regression Runs and a Gated Autopilot
How we built evaluation into Connect: a rubric-driven LLM judge on every live turn, a deterministic schedule grader, golden items promoted from the ledger, draft-versus-live regression runs, and an autopilot that publishes only through a regression gate.