Free resource
AI pilot evaluation worksheet
Choose a first job for AI by comparing tasks your team already does. Rate each on value, data readiness, consequence of error, human review, and whether you can test it. Then write down the pilot's boundaries before anyone builds. Your answers stay in your browser.
Background: governing AI agents in construction and what MCP means for construction data.
1. Rate your candidate workflows
List up to 5 tasks people already do. Rate each from 1 to 3. For consequence of error, a lower rating is safer.
Not rated yet
Rule: Rate expected value, data readiness, consequence of error, human review available, evaluation feasibility.
2. Compare side by side
There is deliberately no single score. A high-value task with no way to check the output is not a better pilot than a modest one you can evaluate.
| Workflow | Expected value | Data readiness | Consequence of errorlower is safer | Human review available | Evaluation feasibility | Recommendation |
|---|---|---|---|---|---|---|
| Workflow 1 | Not rated | Not rated | Not rated | Not rated | Not rated | Not rated yet |
3. Scope the pilot you choose
Write the boundaries down before anyone builds. These are the same questions an AI readiness review answers in detail.
One sentence, e.g. "Draft the weekly schedule update for the client."
Start with the least the task needs.
How many known examples with correct answers will you test against?
Who approves output or changes, and what the pilot may never do on its own.
What result on the evaluation set, and in use, would justify the next step?
Your answers stay in this browser.
Frequently asked questions
Why is there no overall score?
The criteria aren't interchangeable. High value cannot make up for an output nobody can check, so each recommendation comes from a stated rule you can read and challenge.
What makes a good first AI pilot?
A task people already do, with reachable and defined data, known examples to test against, a person who reviews the output, and a low cost if it is wrong. Analysis or drafting usually fits better than writing to a system.
Can a pilot change a budget or schedule?
Only if a separately agreed workflow grants that action and meets its validation and approval requirements. Start with read-only analysis or drafts a named person approves.
Is anything I type sent to Build Flows?
No. The worksheet runs in your browser and saves your answers there. Nothing you type is sent anywhere unless you copy the summary and share it yourself.
Next step
Want a second opinion on your first pilot?
An AI readiness review compares your candidate workflows against the data, access, and controls in your environment, and ends with a pilot scope and evaluation rubric.
Prefer email? charley@buildflows.ai
See what the review covers: AI readiness review.