Build Flows

Free resource

AI pilot evaluation worksheet

Choose a first job for AI by comparing tasks your team already does. Rate each on value, data readiness, consequence of error, human review, and whether you can test it. Then write down the pilot's boundaries before anyone builds. Your answers stay in your browser.

Background: governing AI agents in construction and what MCP means for construction data.

1. Rate your candidate workflows

List up to 5 tasks people already do. Rate each from 1 to 3. For consequence of error, a lower rating is safer.

  1. Expected value

    If the pilot works, how much time, risk, or delay does it remove from work people already do?

    Data readiness

    Are the sources the task needs reachable, and are the key fields and identifiers defined?

    Consequence of error(lower is safer)

    What happens if the output is wrong? Lower is safer: a draft someone reads, versus a change to a budget, schedule, or client document.

    Human review available

    Will a qualified person check the output before anyone relies on it?

    Evaluation feasibility

    Do you have known examples with correct answers to test the pilot against?

    Not rated yet

    Rule: Rate expected value, data readiness, consequence of error, human review available, evaluation feasibility.

2. Compare side by side

There is deliberately no single score. A high-value task with no way to check the output is not a better pilot than a modest one you can evaluate.

Ratings and recommendation per candidate workflow
WorkflowExpected valueData readinessConsequence of errorlower is saferHuman review availableEvaluation feasibilityRecommendation
Workflow 1Not ratedNot ratedNot ratedNot ratedNot ratedNot rated yet

3. Scope the pilot you choose

Write the boundaries down before anyone builds. These are the same questions an AI readiness review answers in detail.

One sentence, e.g. "Draft the weekly schedule update for the client."

Start with the least the task needs.

How many known examples with correct answers will you test against?

Who approves output or changes, and what the pilot may never do on its own.

What result on the evaluation set, and in use, would justify the next step?

Your answers stay in this browser.

Frequently asked questions

Why is there no overall score?

The criteria aren't interchangeable. High value cannot make up for an output nobody can check, so each recommendation comes from a stated rule you can read and challenge.

What makes a good first AI pilot?

A task people already do, with reachable and defined data, known examples to test against, a person who reviews the output, and a low cost if it is wrong. Analysis or drafting usually fits better than writing to a system.

Can a pilot change a budget or schedule?

Only if a separately agreed workflow grants that action and meets its validation and approval requirements. Start with read-only analysis or drafts a named person approves.

Is anything I type sent to Build Flows?

No. The worksheet runs in your browser and saves your answers there. Nothing you type is sent anywhere unless you copy the summary and share it yourself.

Next step

Want a second opinion on your first pilot?

An AI readiness review compares your candidate workflows against the data, access, and controls in your environment, and ends with a pilot scope and evaluation rubric.

Prefer email? charley@buildflows.ai

See what the review covers: AI readiness review.