A superintendent walking a floor sees more in twenty minutes than most reports capture in a week. The problem is getting it out of their head and phone and into something the project can use: a report the owner can read, a searchable record the team can query, and progress readings a person has actually confirmed. Photos on a personal camera roll do none of that, and a voice memo nobody transcribes is worse than useless because it feels like it was recorded.
Job Walks is the part of Connect that handles this. A walk is a project-scoped capture session: audio, photos and video taken on site, pushed through transcription and a vision-capable agent, composed into a structured report, and, only after a person confirms it, recorded as schedule actuals. This article walks through how each stage works, where it reuses the recordings pipeline, and the one rule we would not bend: a photo can suggest that an activity is 60% complete, but only a human can make that a fact.
The data model in one migration
The whole feature landed as a single migration (034), folded from several development slices before it ever shipped. It creates four capture tables plus a sink:
| Table | What a row is | Key states |
|---|---|---|
job_walks | One walk, always attached to a project | open, ended, processing, ready, failed |
job_walk_media | One captured artifact (audio, photo or video) | upload status, plus separate transcription and vision statuses |
job_walk_observations | One structured finding from one image | an optional proposed schedule match, and confirmed_at |
job_walk_reports | The composed report for a walk | generating, ready, failed |
schedule_actuals | One human-confirmed progress reading against an activity code | append-only, keyed by project and activity code |
A later migration (060) added two plain-text columns to the walk: who walked it and which client it was for. Both are deliberately free text, not foreign keys. The walker may be a subcontractor or an inspector with no login, and the client is a name on a report, not a tenant. The client name is prefilled from the project when the walk starts and stays editable at any status, because it is a label, not capture data.
One design choice in the media table does a lot of quiet work: each row carries two pipeline statuses, one for transcription and one for vision. Photo and video rows are born with transcription already done, audio rows with vision already done, so each worker drain only sees rows its stage applies to, through partial indexes that match the claim queries.
Capture: built for a bad signal
Job sites have poor connectivity, phones lock, and people navigate away mid-walk. The capture client is built around one assumption: the network will fail, and nothing captured should be lost when it does.
Every photo, clip or audio segment is written to IndexedDB the moment it is captured, next to a manifest of the walk. The database is scoped per user and per workspace, so a shared tablet cannot mix two people's walks. A reload, a dropped signal or a closed tab loses nothing: on the next load the manifest and its blobs are recovered and a banner offers to resume.
A separate uploader drains that buffer whenever the browser reports it is online:
- It runs at most two uploads at a time, so a weak uplink is not swamped.
- A failed item retries with exponential backoff, starting at two seconds and capped at a minute.
- It never blocks capture. A user can keep shooting while earlier items are still uploading.
- The session lives in a module-level store rather than inside a component, so if the superintendent ends the walk and moves to another page, uploads finish and the walk finalizes in the background.
Walks can be guided. A walk carries an editable checklist of checkpoints (up to 40, trimmed and de-duplicated), and each capture is tagged with whichever checkpoint was active. The capture screen shows live coverage per checkpoint and, before ending, what would be skipped. A capture can be re-tagged later if someone shot the east stair under "Level 3".
Media storage: direct to Blob, verified on completion
Media never streams through the application server. Uploads follow the same two-step pattern as recordings (and as files generally, covered in resumable large file uploads to Azure Blob):
- Ticket. The client asks for an upload slot, declaring kind, media type and size. The server checks that the walk is still open, that the media type matches the kind (
image/*for a photo,video/*for a video,audio/*for audio), and that the size is within the ceiling. It inserts the media row inuploadingand returns a short-lived signed URL into a per-workspace container. Small files get a single PUT with a create-only condition. Anything over 256 MB gets a block-upload URL so large videos stage in pieces. - Complete. The client reports it is done. The server checks that the blob actually exists and that its size matches what was declared. A mismatch deletes the blob and the row. Only then does the row flip to
ready.
On completion each media item is also mirrored into the project's Files under a default "Job Walks" folder, stamped with its source surface so Files can show where it came from. The mirror is idempotent and best-effort: a failure there never fails the capture, and a retry will not create a duplicate.
The worker also tidies up. An open walk with no media that is more than a day old is soft-deleted, and media stuck in uploading for half an hour is marked failed so an abandoned upload cannot block the report forever.
Transcription: reusing the recordings pipeline
Audio from a walk goes through the same Azure Speech path that Connect uses for meeting recordings. There is no second transcription stack. The worker claims ready audio rows, pulls the blob, and sends it for transcription with speaker diarization turned on.
Long audio and any video-container audio take a detour first. Anything over a size threshold, or in a video container, is split with ffmpeg into ten-minute, 16 kHz mono WAV windows. Each window is transcribed with its start time passed as an offset, so every speaker turn lands on one shared timeline instead of each window restarting at zero. The text is concatenated and the turns are flattened in order. That is the same segmenter the recordings module uses.
If Speech is not configured, the row fails with a clear message instead of waiting forever.
The difference from recordings is in what happens after the transcript. A meeting recording goes on to an AI summary of the conversation. A walk transcript becomes one input to the walk report, alongside the photos. The other difference is scope: a recording can exist on its own, but a walk is always attached to a project, because its whole purpose is to say something about one job.
The vision-capable job_walk agent
Photos and videos go to a vision pass. It is a single multimodal turn per image through the shared agent runtime, recorded in the telemetry ledger under its own agent kind, job_walk, so its token cost shows up next to every other agent's (see LLM agent tracing on a Postgres ledger). It is not configurable from the executive admin portal. The prompt is short, the output is narrow, and we did not want a tuning surface on it yet.
The model takes images, not video. For a clip, the worker uses ffmpeg to pull one representative frame about a second in, to skip the dark or blurred opening, and falls back to the first frame for very short clips. If no frame can be decoded, that media is marked vision-failed and the report composes without it.
The agent returns strict JSON with between one and four observations per image. Each observation has a caption, a trade, a location if one is legible or can be inferred, a severity (none, defect or safety), a detail sentence for defects and safety issues, and an optional schedule match.
The schedule match is where we were most careful. Before the call, the worker looks up the project's activity list once per project per batch: the approved baseline if there is one, otherwise the only schedule on the project. If there are several schedules and none is approved, the lookup returns nothing and matching is skipped. We would rather propose no match than guess which plan the photo belongs to. Up to 400 activity codes and names go into the prompt with an instruction to cite only those codes.
The parser does not trust the model to follow that instruction:
match = parse(model_output.scheduleMatch)
if match.activityCode not in project_activity_codes:
match = null // the model cannot invent a code
match.progressPct = clamp(round(match.progressPct), 0, 100)
match.confidence = in_range(0, 1) ? match.confidence : 0.5
Malformed JSON yields zero observations rather than an exception. Captions and fields are trimmed and length-capped. The result is written as observation rows with matched_activity_code, matched_confidence and progress_pct. At that point it is still only a proposal.
When no AI is configured, the vision stage fails closed: every photo row reaches a terminal state, and the walk report still composes from whatever there is.
From observations to a report
When the superintendent ends the walk, it moves to ended. The report drain picks up ended walks whose media have all finished processing, and builds the report in two layers.
Deterministic facts. Code, not the model, computes the photo, video and audio counts, the distinct trades seen, the lists of defects and safety findings with their locations, coverage per checkpoint, and the checkpoints with zero captures. These are the numbers people act on, so they come from rows.
A short narrative. The model writes two to four sentences from the title, the superintendent's notes, the transcripts and the facts. The facts are passed in labeled as authoritative, and the prompt tells it to lead with safety, then defects, then progress, and not to introduce any count, location or finding that is not in the input. If the call fails, or no AI is configured, a deterministic summary takes its place ("Captured 14 photos, 2 audio recordings. 1 safety issue flagged..."). The report always ships.
A person can edit the narrative afterwards. They cannot edit the computed facts, which stay tied to the observations.
Once the report is ready, three things fan out, each best-effort so a failure never fails the report:
- An alert,
job_walk.report_ready, goes through the notification system with a dedup key so a retry does not notify twice (see construction alerts across email, push and SMS). - The report body and transcripts are indexed into Connect AI, unless the walk has opted out, so a later question like "what did we see on the east elevation last week?" can be answered with a citation (the pipeline is described in RAG for construction documents).
- A draft client Update is queued with the walk summary as context. It is never auto-published. A person reviews it like any other Update (see versioned client project updates).
A finished walk can also be shared with a client who has no account: an unguessable link plus a short typed code, both stored only as hashes, one active link per walk, a 30-day expiry, and a uniform "invalid or expired" response for any failure. The client sees the report, the photos through short-lived signed URLs, and diarized transcripts. It is the same share pattern used by Means & Methods and Updates, covered in more depth in Means & Methods as a system of record.
The human gate: confirming schedule actuals
On the report page, each observation with a proposed match shows the activity, the model's confidence and its progress reading, with a confirm button. Confirming is the only path from a photo to the schedule, and it is built to be safe under double-clicks, retries and races:
- The server checks that the observation has a match and that schedule actuals are available.
- It flips the observation to confirmed with a conditional update that succeeds only if the observation was not already confirmed. Only the request that wins that update goes on to step 3. A second click gets a conflict.
- It records the actual through the baseline-schedules module, the capability that owns schedule data, rather than writing to its table directly. The row stores the activity code, the progress percentage, who confirmed it and the source observation.
- A partial unique index on the source observation means a replayed insert does nothing. Double-counting is impossible even if step 2 were somehow bypassed.
Un-confirming reverses it: the actual is deleted by its observation and the confirmation is cleared.
The most important property is what this does not touch. Actuals live in their own table, keyed by project and activity code, and are purely additive. The approved baseline stays exactly as it was approved: no duration is changed and no logic is edited because of a photo. Plan and observed progress are kept separate. Nothing reads actuals back into the schedule views yet; showing plan and observed side by side is the next step, and it will be a join, not an edit. The actuals table also has no foreign key back into job walk observations, only a provenance id, so the schedule capability stays independent of the feature that feeds it. Today the only source is job_walk, and the check constraint says so. When other sources such as daily logs are added, the constraint widens and nothing else has to change.
Why this matters if you're building something similar
- Buffer locally before you upload anything. On a job site, assume the network will fail. Writing to IndexedDB at capture time and uploading opportunistically costs a few hundred lines and removes a whole class of "I took the photo, where did it go?" tickets.
- Give each pipeline stage its own status column. Separate transcription and vision statuses, with rows born
donefor stages that do not apply, keep the workers simple and make retrying a failed stage trivial. - Keep the numbers deterministic and let the model write prose. Counts, findings and coverage come from rows. The model writes a summary it is told it cannot contradict, and a deterministic fallback means the report never depends on the model being up.
- Close the vocabulary on anything that touches the plan. The model can only cite an activity code that exists. Validate in the parser, not in the prompt.
- Make AI-to-schedule writes proposals, and make confirmation idempotent. A conditional flip plus a unique index is two lines of SQL and makes double-counting structurally impossible.
- Never let observed progress edit the baseline. Store actuals beside the plan, not in it. Your schedulers will trust the system more when they know a photo cannot quietly move their dates.
- Reuse the pipelines you already have. Walks reuse the recordings transcription path, the Files mirror, the indexing pipeline, the alert dispatcher and the share pattern. The new code is mostly the vision pass and the confirmation gate.
Where to go next
- The series overview: How we built Connect
- The worker these drains run on: Leased background jobs on a Postgres worker
- The share and audit pattern in its original home: Means & Methods as a system of record
- How Connect scores the schedules those actuals sit beside: CPM engine, schedule health and Monte Carlo
- How the job walk fits the wider scheduling workflow: Connect: AI agents for CPM scheduling
Want field capture that feeds your schedule and reports without re-keying? Plan your build with us.
Frequently asked questions
Does the AI update the schedule automatically from site photos?
No. The vision agent proposes an activity match and a progress percentage. Only a person confirming it records an actual, and actuals are stored beside the baseline, never edited into it.
What happens if the site has no signal?
Captures are stored in the browser's IndexedDB immediately and uploaded when the connection returns, with retries and backoff. Capture never waits on the network.
How is video analyzed if the model only takes images?
The worker extracts one representative frame from each clip with ffmpeg and analyzes that still. If no frame can be decoded, the clip is marked failed and the report still composes.
Can the vision model invent an activity code?
No. It is given the project's activity list, and the parser discards any match whose code is not on that list. If the project has several unapproved schedules, matching is skipped rather than guessed.
Where does a finished job walk go?
Media is mirrored into project Files, the report is indexed for AI search, an alert fires, a draft client update is queued for review, and the report can be shared with a client through a code-protected link.
Next step
Need something built around how your team works?
Describe the users, the workflow, and the systems it touches. We'll tell you whether a custom application makes sense and how we'd build it.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
How We Built Connect: Architecture of an Enterprise AI Platform
The pillar of our Connect architecture series: a layer-by-layer map of an enterprise agentic AI platform, the three-process shape it runs as, the design principles that kept recurring, and links to every deep-dive article.
Scheduling & P6 · October 10, 2026
Means & Methods as a system of record: capturing expert reasoning as data, not chat
An engineering walkthrough of Connect's Means & Methods step: one artifact per project, a four-step confidence ladder, an append-only events log, hashed client share links, and a Schedule Builder that only reads reviewed answers.
Workflow automation · October 9, 2026
Leased Background Jobs in Postgres: How Connect's Worker Claims, Renews and Resumes Work
How Connect runs syncs, indexing, transcription and AI jobs from Postgres tables using leases, jittered retries and checkpoints, with no message broker.