A prompt edit that fixes one complaint can quietly break three answers nobody is looking at. Once you have more than one agent in production and more than one person able to change them, you need a way to tell whether a change made things better, and you need to know before it reaches users.
This article covers how we built evaluation into Connect: an LLM judge that scores every live turn, rubrics that executives can edit, golden sets built from real conversations, regression runs that compare a draft config against the live one, a deterministic grader for schedules that doesn't use a model at all, and an autopilot loop that can publish a change on its own, but only through a gate. It also covers what we chose not to do with the numbers, which turned out to be one of the more useful lessons.
Two kinds of grader, on purpose
Connect's agents produce two very different kinds of output. The chat assistant, the Info Sheet and the Updates drafter produce prose and structured answers. A judgment call like "was this grounded, correct and in scope?" needs a model to make it. The Schedule Builder produces a CPM network: activities, durations and logic ties. Whether that's good is mostly a matter of measurement.
So there are two graders:
| LLM judge | Schedule grader | |
|---|---|---|
| Used for | Chat, Info Sheet, Means & Methods, App Builder, Updates | Schedule Builder output |
| Method | Model scores against a rubric | Pure functions over CPM, health metrics and risk |
| Deterministic | No | Yes |
| Cost per grade | One small model call | CPU only |
| Gameable by | Fluent prose that sounds right | Hygiene fixes with no real gain, which is why outcome criteria exist |
Using a model to grade a schedule would be slower, noisier and easier to fool than computing the critical path. Using arithmetic to grade an answer about a drawing set isn't possible. Each kind of output gets the grader that fits it.
The LLM judge
The judge is a small class with three jobs: score a live turn against a rubric, score a candidate answer against a golden ideal, and act as the improvement analyst (covered below).
What it returns
For each turn, the judge returns:
- A score per rubric criterion, from 0 up to the rubric's scale.
- An overall score, which is the weighted mean of the criterion scores. We compute this in code, not by trusting the model's arithmetic.
- A confidence from 0 to 1, the judge's own.
- A grounded share from 0 to 1: the fraction of the answer's factual claims supported by the evidence the agent actually had. An answer with no factual claims scores 1.
- A verdict (pass, weak or fail) and a one- or two-sentence rationale.
The grounded share sits outside the rubric on purpose. Rubrics change. Executives edit them, add criteria and switch the active one. Groundedness is reported on every judged turn whatever rubric is active, so there's always a hallucination signal that's comparable across time.
How we keep it cheap and safe
- Off the request path. A background worker claims un-evaluated ledger rows in batches and grades them. A user never waits on the judge.
- On the smaller model. Grading is an easier task than answering, so the judge runs on the cheaper child deployment, with a token cap, low reasoning effort and a 60-second timeout.
- Unscorable means skipped. The judge must return JSON. We pull out the first object, clamp every score into range and drop unknown keys. If nothing usable comes back, the turn is skipped and retried on a later poll. A judge failure is never a failed user turn.
- Measured outcomes beat prose. For agents whose tools return deterministic facts, we pass those results to the judge in a separate block labeled as measured outcomes, with instructions to weigh them over the narrative. The Info Sheet orchestrator is the clearest case. Its section workers do the drafting, so its own prose is mostly coordination. Without the tool facts (coverage, topic dispositions, open questions, resolved conflicts) the judge consistently under-scored its grounding.
Rubrics are data, not code
Rubrics live in eval_rubrics: a name, a list of criteria with a key, a description and a weight, and a scale. A partial unique index on active rows enforces one active rubric per agent. Executives edit them in the portal, and activating a new one deactivates the old one.
Migration 054 is a good example of why this matters. The eval worker loops over every configurable agent but skips any agent with no active rubric. The Info Sheet and Means & Methods agents had never been seeded one, so their turns had been going ungraded without anyone noticing. The migration seeded a baseline rubric for each and widened the allowed agent kinds on the rubric, golden-item and recommendation tables. The Info Sheet rubric is a fair picture of what we think makes an intake trustworthy:
- Evidence grounding: every derived fact cites the page and section it came from (weight 2)
- Accuracy against the evidence and stated facts (weight 2)
- Gap handling: absent facts are recorded as not found or turned into a targeted question, never invented (weight 1.5)
- Completeness for the job type (weight 1.5)
- Clarity: specific, decision-ready prose (weight 1)
Note what's weighted heaviest. For construction intake, a confident wrong answer is worse than an honest gap.
Golden sets from real turns
A golden item is an input, an ideal output, tags, notes and optionally a workspace. Executives can write them by hand, but the most useful ones come from production. Any ledger row can be promoted to a golden item with one click: the user's text becomes the input, the agent's answer becomes the ideal, and the source workspace is carried over so a regression run answers using the same data.
That last detail matters more than it looks. An agent answering "what's the slab thickness on level 3?" in an empty workspace will fail regardless of how good the prompt is. Pinning the golden item to its source workspace makes the comparison about the config, not the context.
Promoting from the ledger also keeps the set honest. Golden sets written from scratch drift toward the questions the author finds interesting. Ones promoted from real turns reflect what users actually ask, including the awkward ones.
Regression runs: draft versus live
Every agent configuration (system prompt, model, child model, thinking level, enabled tools) is versioned. Edits create a draft. Nothing changes for users until a draft is published, and publishing validates any model override against the live provider first, so a typo in a deployment name fails at publish time instead of in front of a user.
Before publishing, an executive can queue a regression run for the draft:
- The run is inserted as pending against the draft version.
- The background worker claims it and loads both configs, the golden set and the active rubric.
- For each golden item, the worker runs the real agent twice, once under the live config and once under the draft.
- The judge scores each answer against the golden ideal.
- The run stores the live average, the draft average, the delta (draft minus live) and every item with both answers and both scores for drill-down.
The worker does the re-running because the admin module isn't allowed to import the chat runtime. It's a module boundary we hold to across the codebase: governance code observes and configures agents, but it never runs them. Regression runs are currently supported for the agents whose turns can be cleanly replayed from a single input: the chat assistant and the Updates drafter. Multi-step agents like the Info Sheet are judged live, but not replayed.
The review queue shows each pending draft with its delta badge, so the person approving a change sees the evidence beside the diff.
The deterministic schedule grader
The schedule grader takes a schedule's content, optionally a reference schedule, and optionally outcome signals, then returns per-criterion scores, an overall score, the underlying health score, a letter grade from A to F and a rationale built entirely from the metrics.
Hygiene criteria, each derived from standard schedule-health checks:
- Logic health: open ends, critical-path continuity and negative float. Any structural error, like a cycle or a dangling link, zeroes this criterion, because broken logic invalidates everything downstream.
- WBS structure: whether a WBS exists and, if a reference is supplied, how much of the reference's WBS it covers.
- Dependencies: logic density, finish-to-start share and lag hygiene.
- Duration realism: high-duration outliers and high total float, which often points to a missing tie.
Outcome criteria, scored only when their signals are supplied:
- On-time: the Monte Carlo probability of finishing on time.
- Requirement coverage: the share of mandatory contract requirements and required milestones with a reviewed link to the schedule.
The outcome criteria exist because hygiene alone can be gamed. Tie every loose activity to the completion milestone and your open-ends metric improves while the schedule gets no more achievable. So on-time carries more weight than any single hygiene criterion. A schedule with perfect logic and almost no chance of finishing on time can't get a top grade. The hygiene weights also match the recommended schedule rubric in the portal, so the deterministic number and what an executive sees agree.
The same grader backs the reference sets in the portal (upload a standard reference schedule, grade a produced one against it) and a background guardrail that raises an alert when build health, the share of poor grades or on-time probability drops below a floor.
The offline eval CLI: grade mode and build mode
For engineering work there's also a command-line harness with two modes.
- Grade mode runs the deterministic grader, health checks, XER verification and gut checks over a folder of schedules or a workspace's stored schedules. No model and no seeds needed. Every input gets an outcome, including malformed files, so a parse failure counts as a failure rather than disappearing from the denominator.
- Build mode drives the real Schedule Builder over its golden briefs in a seeded project with an approved Info Sheet, then grades what it produced.
Each run writes a JSON artifact and a short markdown report into the repo. Pass a previous run as a baseline and it prints deltas, per case where possible, including cases that are new, missing or changed status. A prompt change is then reviewed as a diff of two committed files, so quality is something you can point to in a pull request.
The harness says plainly that model output varies and that build-mode deltas should only be trusted across about ten or more items.
Why we don't publish the pass rate
The committed baseline in the repo is a grade-mode run over a corpus of schedules mined from historical projects. It measures how good those historical schedules were, including the ones that fail XER verification because their source data had zero-duration activities. The decision log states plainly that this is not a live-generation success rate.
It would be easy to put that number on a slide, and it would be misleading. An offline grade over a historical corpus, or a build-mode run over a handful of golden briefs, tells you whether a change moved things in the right direction. It doesn't tell a customer how often the agent will get their project right. So we use the offline numbers internally, as deltas, and we don't present them as success rates. If you're evaluating AI vendors, ask which kind of number you're looking at.
The improvement loop and the autopilot gate
The last piece connects evaluation back to configuration.
Judge, recommend, draft, regress, gate. Nothing reaches users without a draft-versus-live delta.
The analyst. Using the same judge model, an improvement analyst reads a digest of the recent window: turn count, scored count, average score, the lowest-scoring turns with the judge's rationale, and tool usage. For the Schedule Builder, it also gets measured build quality and a summary of the corrections users made. It proposes up to five recommendations, each categorized as a prompt change, a knowledge document, a model change or general advice.
Applying a recommendation never makes it live. A prompt or model recommendation becomes a draft config version, and a regression run is queued automatically where the agent supports it. A knowledge recommendation becomes a knowledge document that's created disabled. General advice is just acknowledged.
Autopilot is the same loop with a gate instead of a person. It's off by default for every agent. When an executive turns it on, the worker, at most hourly per agent:
- Generates fresh recommendations.
- Applies the open ones in the categories the policy allows. The default is knowledge only, the lowest-risk category.
- Queues a regression for any draft that was created and polls it with a 90-second timeout.
- Asks a pure decision function whether to publish.
- Publishes or holds, and writes an audit event either way.
The decision function is short and fully unit-tested. It holds if autopilot is disabled, if the regression isn't finished, if the delta is below the negative of the policy's regression margin (default zero, meaning the change can't regress at all), or if the agent has already hit its daily change limit (default one). The first rule that fails becomes the stated reason, so a disabled policy never looks like a "would regress" hold. Every action (generated, applied, published, held, error) is appended to an audit table with a reason that's always present.
Why this matters if you're building something similar
- Match the grader to the output. If correctness can be computed, compute it. Save the LLM judge for things only judgment can assess.
- Keep one hallucination signal outside the rubric. Rubrics will change. You need something comparable across those changes.
- Promote golden items from production, with their context. Real questions, pinned to the data they were asked against.
- Gate publishes on a draft-versus-live delta, not an absolute score. Absolute judge scores drift. A paired comparison on the same items, judged in the same run, is much more stable.
- Make autopilot boring. Off by default, lowest-risk category by default, no regression allowed by default, one change a day, and an audit row for everything.
- Don't confuse offline scores with live success rates. Use them as deltas. Be careful who sees them as anything else.
Where to go next
- The series overview: how we built Connect
- The ledger every judged turn comes from: tracing LLM agents in a Postgres ledger
- The portal where rubrics, golden sets and the review queue live: building an executive admin portal for AI agents
- Learning from what users fix by hand: an AI learning loop from human corrections
- The gates inside the schedule build itself: AI schedule builder build-spec gates
Want evaluation built into your agents from day one? Plan your build.
Frequently asked questions
What does the LLM judge return for each turn?
A score per rubric criterion, a weighted overall computed in code, the judge's confidence, a grounded share of supported factual claims, a pass, weak or fail verdict, and a short rationale. Unparseable output is skipped and retried later.
Why grade schedules without an LLM?
Schedule quality is mostly measurable. Logic health, WBS coverage, dependency discipline, duration realism, Monte Carlo on-time probability and requirement coverage can all be computed, which is cheaper, deterministic and harder to fool than a model's opinion.
How does a regression run work?
For each golden item, a background worker runs the real agent under both the live and the draft config, the judge scores both answers against the golden ideal, and the run stores both averages, the delta and every item for drill-down.
Can autopilot change prompts on its own?
Only if an executive enables it and allows that category. By default it is off and limited to knowledge documents, which are created disabled. Any draft must finish a regression within the margin and under the daily change limit before it publishes.
Why don't you publish eval pass rates?
Offline grades over historical corpora or a few golden briefs measure direction of change, not how often the agent gets a customer's project right. We use them internally as deltas and don't present them as live success rates.
Next step
Have a workflow in mind?
Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
How We Built Connect: Architecture of an Enterprise AI Platform
The pillar of our Connect architecture series: a layer-by-layer map of an enterprise agentic AI platform, the three-process shape it runs as, the design principles that kept recurring, and links to every deep-dive article.
AI agents & MCP · October 9, 2026
Tracing LLM Agents in Postgres: Why We Built a Ledger Instead of OpenTelemetry
A deep dive on Connect's agent ledger: two recorders feeding one best-effort sink, per-turn rows with reasoning and cache tokens, a stricter table for MCP tool calls, and the admin views built on plain SQL.
AI agents & MCP · October 9, 2026
Building an Executive Admin Portal for AI Agents: The Control Plane Behind Connect
The build of Connect's executive admin portal: an org-wide surface gated by super-admin plus an operator-domain allowlist, its pages from logs to query approvals, and a read-only admin MCP so executives can ask Claude about agent operations.