Ask a large language model for a project's critical path and it will give you one, confidently, and it may even be right. Ask again with the same input and you may get a different answer. For a schedule a contractor will sign, that is disqualifying. Dates, float, health grades and finish probabilities have to come out the same every time, be explainable line by line, and survive an argument with a scheduler who has done this for twenty years.
So in Connect, the language model never does schedule math. Deterministic TypeScript owns the critical path method (CPM) calculation, the structural gut checks, a DCMA-style health score, a grader layered on top of it, and a Monte Carlo simulation of the finish date. The model writes the plan and explains results. This article walks through how each piece works, with a small CPM implementation written from scratch, and why we drew the line where we did. It sits alongside the Schedule Builder engineering article and how we verify XER exports, in our Connect architecture series.
Who does what
The division of labor is strict:
| Layer | Owned by | Output |
|---|---|---|
| What to build and in what order | The model, through a declarative Build Spec | Activity families, logic rules, durations with a basis |
| Expanding the spec into rows | Pure code | Activities, WBS, relationships |
| Dates and float | CPM engine | Early/late dates, total and free float, finish date |
| Is the network valid? | Gut checks | Errors and warnings |
| Is the schedule good? | Health score and grader | 0-100 score, A-F grade, per-metric findings |
| Will it finish on time? | Monte Carlo | Finish percentiles, on-time probability, risk drivers |
| What it means for this project | The model, citing the numbers | Narrative, repair proposals, client updates |
The same functions back the builder's gates, the approval checks, the evaluation harness and the agent tools. When the model reports a health grade, it is reading a number the code computed, not estimating one.
CPM from scratch
The core of CPM fits on a page. Every activity has a duration; every relationship has a type (finish-to-start, start-to-start, finish-to-finish, start-to-finish) and a lag. A forward pass in topological order computes the earliest each activity can start and finish. A backward pass from the project finish computes the latest each can start and finish without moving that finish. The difference is total float.
An illustrative version on a single calendar, with time counted in working days:
order = topological_sort(tasks, deps)
# forward pass: earliest start / finish
for t in order:
t.es = 0
for d in incoming(t):
p = d.pred
bound = { FS: p.ef + d.lag,
SS: p.es + d.lag,
FF: p.ef + d.lag - t.dur,
SF: p.es + d.lag - t.dur }[d.type]
t.es = max(t.es, bound)
t.ef = t.es + t.dur
finish = max(t.ef for t in tasks)
# backward pass: latest start / finish
for t in reversed(order):
t.lf = finish
for d in outgoing(t):
s = d.succ
bound = { FS: s.ls - d.lag,
SS: s.ls - d.lag + t.dur,
FF: s.lf - d.lag,
SF: s.lf - d.lag + t.dur }[d.type]
t.lf = min(t.lf, bound)
t.ls = t.lf - t.dur
t.total_float = t.ls - t.es
# free float: slack before the first successor is pushed
t.free_float = min(gap(d) for d in outgoing(t)) or finish - t.ef if none
Activities with zero total float form the critical path. Free float is a different number: how far an activity can slip before it delays any immediate successor. It is computed per relationship type, not copied from total float.
What production CPM adds
The toy above is correct and nearly useless on a real job. Here is what the production engine adds, and why.
Calendars as working-day indexes. Each calendar (workweek plus holidays) is materialized once into a sorted list of working days, with constant-time lookups from date to index and back. Durations and lags then add as integers, with no weekend arithmetic anywhere in the passes.
Per-activity calendars. A concrete cure can run seven days a week while the trades that follow work five. Each activity runs on its own calendar. Where two calendars meet, the hand-off is a real date: a finish is the end of its last worked day, and the successor starts on its own first working day after that. Lags count on the predecessor's calendar, which is how P6 behaves under the setting the operator's baselines use.
Milestone semantics. A zero-duration activity that something finishes into is a finish milestone and sits at the end of its predecessor's last day, not at the start of the next working day. Getting this wrong shifts every downstream milestone by a day.
Constraints and negative float. The engine supports start-on-or-after and finish-on-or-after as floors on the forward pass, finish-on-or-before and start-on-or-before as ceilings on the backward pass, and as-late-as-possible. A ceiling the logic cannot meet produces negative float, exactly as P6 reports it, and the health check and approval gates surface it. As-late-as-possible consumes only the activity's free float, so it never steals float from work downstream.
Cycles do not crash the render. A half-built schedule often contains a loop. Topological sorting releases the blocked activity with the fewest unplaced predecessors, so only the tie that closes the loop is ignored for ordering and the rest still schedules. The gut check reports the cycle as an error, so it can never be approved.
A calculation version. The engine carries a version number that increments when date, calendar or float semantics change. Approval requires the client to send the version it reviewed. If the engine has changed since, the reviewer has to look again at dates the new engine computed.
Gut checks: is the network valid?
Gut checks ask a narrower question than health: is this schedule structurally sound enough to calculate and export? Errors block; warnings inform.
| Severity | Checks |
|---|---|
| Error | Dangling relationship, self-relationship, milestone with a duration, activity pointing at a missing WBS node, logic cycle (reported as the actual loop of activity codes) |
| Warning | Negative lag, redundant tie between the same pair, non-milestone with zero duration, very long duration, open start or open finish, duplicate activity code, no milestones at all |
Cycle detection is a depth-first search that returns the loop in order, so the message names the loop ("A1000 -> A1010 -> A1020 -> A1000") and tells the user to remove the tie that points backward in time, rather than just saying "cycle detected". Summary level-of-effort rows are exempt from open-end checks, because they span their section by design.
Health: is the schedule good?
The health score is inspired by the DCMA 14-point assessment, applied only where it makes sense. A baseline has no actuals, resources or prior baseline, so the points that need them (invalid dates against a data date, resources, missed tasks, CPLI, BEI) are left out rather than faked. What remains:
| Metric | Target | Weight |
|---|---|---|
| Logic density | at least 1.5 ties per activity | 20 |
| Open ends | at most 5% | 20 |
| Critical path continuity | unbroken from start to finish | 15 |
| Negative float | none | 15 |
| High float (over 44 working days) | at most 5% | 10 |
| Hard constraints | at most 5% | 10 |
| High durations (over 44 working days) | at most 5% | 5 |
| Leads (negative lag) | none | 5 |
| Lags | at most 5% of ties | 5 |
| Relationship types | at least 90% finish-to-start | 5 |
Each metric passes, warns or fails against named constants, and earns full, half or no credit for its weight. The weighted total is a score out of 100, and the bands are A from 90, B from 80, C from 70, D from 55, and F below that. Keeping thresholds as named constants means anyone can read exactly why a schedule got its grade.
Two judgment calls are worth calling out. The finish-to-start share warns rather than fails until it is well below the standard, because the operator's SOP deliberately uses start-to-start with lag for crew flow between floors. And critical path continuity is not "every activity with zero float". Across calendars, a driving activity can legitimately carry a day or two of float, such as a seven-day cure ahead of a five-day successor. So the check walks backward from the latest finish, following driving relationships (a successor starting right after its predecessor's event, allowing for a weekend and the lag), and passes only if the chain reaches a genuine start or a start constraint. That is P6's longest-path idea rather than a float filter.
Alongside this score, Connect also computes a second one: a check-for-check port of the downstream scheduling platform's own quality engine, kept in step with golden outputs from the original. It answers a practical question, which is what the scheduler will see after import.
The grader: health is gameable, outcomes are not
Health measures hygiene, and hygiene can be gamed. Tie every dangling activity to the completion milestone and open ends vanish while the schedule gets no more achievable. Because we use the grade to compare builds and to judge the builder against reference schedules, that loophole mattered.
The grader maps health metrics into four hygiene criteria:
- Logic health: open ends, critical path continuity and negative float, forced to zero by any structural gut-check error.
- Dependencies: logic density, relationship types and lags.
- Duration realism: high durations and high float, since excess float often means a missing tie.
- WBS structure: whether every activity rolls up to a real section, or, against a reference schedule, the share of its WBS codes the produced schedule covers.
Then it adds two outcome criteria when their signals are supplied. On-time credit is the Monte Carlo on-time probability. Requirement coverage is the share of mandatory requirements and required milestones with a reviewed link to an activity. On-time carries more weight than any hygiene criterion because it is the axis hygiene gets gamed against. With the weights we use, a schedule with perfect hygiene but almost no chance of finishing on time cannot get anywhere near a top grade. Without outcome signals, the grade falls back to the health grade. No model call is involved anywhere.
Monte Carlo: how likely is the finish date?
A CPM finish date is one point from a distribution nobody has drawn. The Monte Carlo simulation draws it.
- Choose a distribution for each activity. The estimating engine keeps p10, p50 and p90 durations per activity archetype, learned from approved baselines and prior jobs, preferring the workspace's own history over the firm-wide prior. Where it has a match, that triple becomes a triangular distribution. Where it has none, the default is right-skewed from 85% to 135% of the planned duration, because construction work sometimes finishes a little early and often runs long.
- Sample and run. Each iteration samples every non-milestone duration by inverse transform and runs one forward pass with the same relationship semantics as the CPM engine. The default is a thousand iterations, clamped between one hundred and five thousand.
- Track who drove the finish. After each pass, the simulation walks back from the latest finish along the relationship that actually set each start. The share of iterations an activity sat on that path is its criticality.
- Summarize. Sorting the simulated finishes gives p10, p50, p80 and p90 dates. On-time probability is the share of futures finishing on or before the deterministic CPM date. The working days between that date and p80 are reported as the buffer the plan honestly needs. The top drivers come back ranked.
The random generator is a small seeded PRNG, so a given schedule and seed always produce the same answer. That makes the simulation testable and lets a scheduler rerun it without the numbers wandering. The simulation requires an acyclic network, so callers gate on gut-check errors first.
Scenario analysis reuses the same machinery: the current plan, added crew or phased turnover are derived transformations of the content, never deletions of work, each re-run through CPM and the simulation with its assumptions attached. Every result carries the caveat that the band is an estimate built on the distributions behind it.
Why deterministic code owns the math
We get asked why we don't simply let a capable model compute the schedule. The reasons are practical:
- Reproducibility. The same inputs must produce the same dates, today and in an audit two years from now. Pure functions and a seeded simulation give that; sampling a model does not.
- Auditability. Every grade decomposes into named metrics with named thresholds. A scheduler can disagree with a threshold. Nobody can usefully disagree with a vibe.
- Testability. Pure code runs in unit and property tests in milliseconds. The evaluation harness grades whole corpora deterministically, so a change to a prompt shows up as a change in measured output, not as a feeling.
- Resistance to gaming. When a model is rewarded by a score, it will find the cheapest way to raise it. A deterministic grader with outcome criteria can be hardened against that; a model judging itself cannot.
- Clear accountability. If a date is wrong, there is one function to fix, with a version number that tells reviewers when semantics changed.
The model still does the part it is good at. It reads drawings, proposes sequences, explains why the critical path runs through the hoist, and drafts the client-facing narrative around numbers it did not invent.
Why this matters if you're building something similar
- Draw the line at arithmetic. Let the model propose and explain. Let code compute anything a professional will be held to.
- Make the engine pure and versioned. Pure functions are easy to test and reuse across gates, tools and evaluations. A version number tells reviewers when old evidence needs another look.
- Name every threshold. A grade made of named constants can be argued with and tuned. A black box cannot.
- Pair hygiene metrics with outcome metrics. Any single score becomes a target. Weight the outcome you actually care about above the proxies.
- Seed your randomness. A Monte Carlo result that changes on refresh erodes trust faster than no result at all.
- Do not fake what you cannot measure. Leave out checks that need data you do not have, and say so.
Where to go next
- How we verify Primavera P6 XER exports for how these dates leave Connect intact.
- Schedule Builder engineering: Build Spec and gates for the loop these checks gate.
- LLM evals with golden sets and autopilot for how the grader feeds evaluation.
- The AI learning loop from human corrections for where the duration distributions come from.
- AI CPM scheduling with Primavera P6 for the wider landscape.
To see it working, watch the agentic scheduling demo, or talk to us about your build.
Frequently asked questions
How do the CPM forward and backward passes work?
The forward pass visits activities in topological order and sets each early start to the latest bound imposed by its predecessors, adjusted for relationship type and lag. The backward pass runs in reverse from the project finish and sets each late finish to the earliest bound imposed by its successors. Total float is late start minus early start.
What is the difference between total float and free float?
Total float is how far an activity can slip without moving the project finish. Free float is how far it can slip without delaying any immediate successor. Connect computes free float per relationship type rather than copying total float.
Which DCMA checks apply to a baseline schedule?
Without actuals, resources or a prior baseline, Connect applies the logic, leads, lags, relationship types, hard constraints, high float, negative float, high duration and critical path checks, plus logic density. Checks that need progress or resources are left out rather than faked.
Why not let the LLM calculate the schedule?
Schedule math has to be reproducible, auditable and testable. A model can return different answers to the same input and can learn to raise a score without improving the plan. Pure, versioned code with named thresholds avoids both, while the model handles reading sources, proposing sequences and explaining results.
How is on-time probability calculated?
Each Monte Carlo iteration samples every activity duration from a triangular distribution, using learned p10, p50 and p90 values where available and a right-skewed default otherwise, then runs a forward pass. On-time probability is the share of iterations that finish on or before the deterministic CPM finish date.
Next step
Thinking about a scheduling pilot?
Start with one exported schedule and an agreed set of checks. We'll discuss your format, method, and approval process.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
Scheduling & P6 · October 10, 2026
Verifying Primavera P6 XER Exports Before They Leave
How we write P6 19.12 .xer files and check them before a scheduler downloads them, and why local verification is kept separate from acceptance by the downstream importer.
Scheduling & P6 · October 9, 2026
Engineering an AI Schedule Builder: Build Specs, Gates and Repair Loops
The engineering behind the Connect Schedule Builder loop: the model edits a declarative Build Spec, code expands and gates it, reviewers and bounded repair passes drive it to done, and approval freezes verified evidence.
AI agents & MCP · October 9, 2026
LLM Evals in Production: Judges, Golden Sets, Regression Runs and a Gated Autopilot
How we built evaluation into Connect: a rubric-driven LLM judge on every live turn, a deterministic schedule grader, golden items promoted from the ledger, draft-versus-live regression runs, and an autopilot that publishes only through a regression gate.