Video unavailable? Watch on YouTube or read the written breakdown.
An AI assistant over construction project documents is a search system with a language model on top. Files go through OCR, get split into passages, and are stored as vector embeddings. When someone asks a question, the assistant pulls back the most relevant passages and answers from them, citing the file and passage it used. Done well, it answers "what does the spec say about curing time for slabs on grade?" in seconds and shows you the page. Done badly, it gives a confident answer with no source, from documents the user should never have seen.
This guide covers how retrieval with citations works, what these assistants are good and bad at on drawings, specs and meeting notes, how to set permissions and governance, and how to test one with known questions before anyone relies on it. It draws on two systems we built: the file intelligence layer in Connect, the agentic scheduling platform we built for Syncify (Production), and the file search and tool setup in Construct.Chat agents.
How an AI assistant answers questions from project documents
Most document assistants follow the same pattern, usually called retrieval-augmented generation, or RAG. The model does not "know" your project. It reads a few relevant passages at question time and answers from those.
- Intake. Files land in a known place: a project folder in your document system, SharePoint or Procore, or an upload area in the assistant itself. In Connect, every project has a predefined folder structure, so drawings, specs, client documents and transcripts always land somewhere the agents expect.
- OCR and text extraction. Native PDFs have a text layer. Scanned pages, photographed field documents and many older drawing sets do not, so they go through optical character recognition. OCR output also keeps page numbers, which you need later for citations.
- Chunking. The text is split into passages: a spec section, a paragraph of meeting notes, a block of general notes from a drawing sheet. Each passage keeps metadata: project, file, page, section, upload date and revision.
- Embedding. Each passage is turned into a vector, a list of numbers that represents its meaning. Passages about the same thing end up close together even when they use different words. See vector embeddings in our glossary.
- Retrieval. At question time, the question is embedded the same way and the closest passages are pulled back, often combined with plain keyword search so exact terms like spec section numbers and product names still match.
- Answer with citations. The model gets the question plus the retrieved passages, writes an answer, and names the file and passage behind each claim. If the passages do not support an answer, it should say so.
One design choice matters more than it looks: keep the original files intact next to the index. In Connect, agents use the vectors to find context and identify exactly which file the information came from, but they still pull in the files themselves. That way a citation always opens the real document, and you can re-index when the embedding approach changes without losing anything.
If you are building on Microsoft's stack, Microsoft documents the same pattern in its retrieval-augmented generation overview for Azure AI Search and its vector search overview.
What we built: Connect file intelligence and Construct.Chat file search
Connect helps schedulers turn drawings, specs and client answers into CPM schedules. Its file intelligence layer is the foundation for every agent in the platform:
- Each project has a predefined folder structure for uploads.
- Files are processed with OCR and stored as vector embeddings, so agents can find and cite the exact file and passage.
- Meeting recordings and job walks are transcribed and attached to the project, so what was said in a client meeting is as searchable as the spec.
- Access to project data follows the users assigned to that project.
- Each workspace has its own vector knowledge base, which admins can add to from the workspace's data sources and re-index from the admin portal.
On top of that sits the Info Sheet agent, which reads the project's files and transcripts and produces a normalized "what is being built" record. Questions the documents cannot answer go to the Means and Methods agent, which sends them to the client by email or secure link. Answers come back marked client verified, with an audit trail. That split is the useful lesson: the assistant answers what the documents support and routes the rest to a person who knows. You can see it in the Connect walkthrough and read the design decisions in how Connect's specialized agents standardize CPM scheduling.
Construct.Chat agents show the same idea in a general-purpose agent builder. A narrow agent gets a chosen model plus optional capabilities, including file search over uploaded documents (RAG), code execution and web search, alongside MCP tools for live systems like Procore. The Construct.Chat demo shows an agent with read-only Procore tools and a log of every tool call. Pairing document search with live system data is where these assistants get most useful: the spec says what was required, and Procore says what was submitted.
What document assistants are good and bad at
Be honest about this before you start, because the failure modes are predictable.
| Task | How well it works | Why |
|---|---|---|
| Finding a requirement in specs ("what's the testing frequency for concrete cylinders?") | Strong | Specs are structured text with section numbers, and retrieval handles them well |
| Summarizing meeting notes or job walk transcripts | Strong | Plain language, clear dates, and a transcript is already text |
| Finding the RFI or submittal that addressed a topic | Strong, if the log is in the index | Short, well-labeled documents with clear subjects |
| Comparing two revisions of a spec section | Moderate | Works when both revisions are indexed and labeled; fails when superseded files are mixed in without revision metadata |
| Reading text on drawings (general notes, schedules, title blocks) | Moderate | OCR picks up the text, but small fonts, rotated text and dense sheets lose accuracy |
| Interpreting graphical drawing content (dimensions, clashes, what a detail shows) | Weak | Text retrieval cannot see geometry; this needs model-based tools or a person |
| Quantities and takeoffs | Weak | Needs measurement, not retrieval; treat any number as unverified |
| Questions the documents don't answer | Risky | Without a rule to say "not found," a model fills the gap with general knowledge |
Two failure patterns cause most of the trouble:
- Stale or superseded documents. If addendum 3 changed a requirement and both the original and the addendum are indexed with no revision metadata, the assistant may cite the old one. Fix it at the source: tag revisions at intake, mark superseded files, and have the assistant prefer the latest revision and say when sources conflict.
- Answers without evidence. A fluent answer with no citation is worse than no answer, because people believe it. Make "no citation, no answer" a hard rule in the instructions, and check it in testing.
Governance and permissions for project document AI
Construction documents are full of things not everyone should see: bid numbers, subcontractor pricing, owner contracts, HR material that ended up in a project folder. A document assistant can leak all of it in a single answer. Governance has to be built in from the first index, not added after a pilot.
Permissions follow the project. The assistant should only retrieve passages from projects the user is assigned to. In Connect, access to project files and transcripts is set by project membership, and workspaces add roles (admin, member, guest) and feature flags on top. The key point: filter at retrieval time, using the user's identity, before passages reach the model. Telling the model "don't mention confidential documents" is not a control.
Least privilege for tools. If the assistant also calls live systems through MCP, the tool list is its permission boundary. The Construct.Chat demo agent only has list and show tools for Procore, so it can read but not create, update or delete. Start read-only. See least privilege and our guide to governing AI agents in construction.
Log everything. Record each question, which passages were retrieved, which tools were called and what the answer was. Connect's admin portal keeps a full log of every agent's tool calls, and Construct.Chat sends each tool call to a dashboard with its request, response and status. Logs are how you answer "why did it say that?" and how you find bad documents in the index.
People approve what leaves the building. An assistant can draft an answer to an owner's question; a person sends it. Connect follows this throughout: agents draft, schedulers and clients verify. See human in the loop.
Controlled changes. Prompt, model and tool changes change answers. In Connect, agent prompts, models and tools are version controlled, and changes an agent proposes go to a review queue for an admin to approve. Autopilot can be turned on, but it is a deliberate choice.
Hosting and data residency are decided during discovery, typically in your own cloud tenant or a managed environment. Our governance page covers how we handle credentials, access and audit across builds.
How to evaluate a document assistant with known questions
You cannot judge an assistant by trying a few questions in a demo. Build a test set first, run it before anyone relies on the assistant, and run it again after every change.
- Collect real questions. Pull 30 to 50 questions people actually asked on past projects: from RFIs, emails to the project engineer, and pre-construction meetings. Include easy lookups, cross-document questions, and some the documents cannot answer.
- Record the expected source. For each question, note the correct answer and the file and page or section it comes from. This becomes your golden set.
- Score more than the answer. For each run, check: did it retrieve the right passage, cite it correctly, answer accurately, and say "not found" when it should have? A correct answer with the wrong citation is still a failure.
- Use a written rubric. Define what pass, partial and fail mean for each criterion, so two reviewers grade the same way. Connect's admin portal uses grading rubrics and golden reference sets for exactly this, so agent quality is tracked over time rather than judged by feel.
- Rerun on every change. New documents, a different chunking approach, a model upgrade or a prompt change can all move results. Rerun the set and compare before releasing.
- Measure time against a baseline. Time how long it takes today to answer a sample of the questions by searching documents, then how long it takes with the assistant, including checking the citation. That is the honest time-saved number. We do not quote one until it has been measured on your documents.
The unanswerable questions matter most. An assistant that says "the documents don't cover this, here is who to ask" on those questions is one people will trust on the others.
Pilot readiness checklist
Before you start a pilot, check that you can say yes to each of these:
- The document set is defined: which projects, which document types, and which are out of scope.
- Files have a consistent folder structure and revision labels, and superseded documents are marked.
- Scanned documents have been tested through OCR on a sample, including some drawing sheets.
- Access rules are written down: who can query which projects, and how that maps to your existing project membership.
- Confidential folders (bids, contracts, HR) are excluded or permissioned separately.
- The assistant must cite file and passage for every answer and say when the documents don't support one.
- Any connection to live systems starts with read-only tools.
- Every question, retrieval and answer is logged, and someone owns reviewing the logs.
- A golden set of real questions with expected sources exists, and a pass threshold is agreed before the pilot starts.
- A baseline time-to-answer has been measured, so results can be compared after.
- The source of truth stays in your document system; the assistant only reads from it.
If you want a structured way to work through these, our AI pilot worksheet walks through scope, data, access and measures for a first agent.
Build it or buy it?
Many document platforms now ship a built-in assistant, and if your documents already live in one system and its assistant respects your permissions and cites its sources, start there. A custom build makes sense when the answers you need cross systems (specs plus Procore records plus meeting transcripts), when you need your own permission model, logging and test set, or when the assistant is one step in a larger workflow, the way Connect's Info Sheet feeds scheduling. More on how we approach this is on the AI agents capability page and the document and field automation page, with the specific pattern described in AI assistant for project documents. Related patterns include meeting and job walk transcription and invoice and document OCR extraction.
Where to go next
- To scope a document assistant against your own files, access rules and questions, the AI readiness review has fixed scope and price, agreed after discovery, or you can plan your build.
- See file intelligence, citations, rubrics and golden sets in production in the Connect work example and the Connect walkthrough.
- Work through scope, data and measures for a first agent with the AI pilot worksheet.
- Want to talk it through first? Start a conversation.
Frequently asked questions
Can AI read construction drawings and specifications?
It can read the text in them, including scanned pages through OCR, and find passages relevant to a question. It is strongest on specs, narratives, RFIs and meeting notes. It is weak on graphical drawing content such as dimensions and details, so treat those answers as unverified.
How do I stop an AI assistant from making things up?
Require a citation to the file and passage for every answer, and instruct the assistant to say when the documents do not support an answer. Then test it with questions the documents cannot answer and check that it declines rather than guessing.
How do permissions work for an AI document assistant?
Retrieval should be filtered by the user's identity and project membership, so passages from projects they cannot access never reach the model. Confidential folders such as bids and contracts should be excluded or permissioned separately, and every question and answer should be logged.
How do you evaluate an AI assistant for project documents?
Build a golden set of real questions with the expected answer and source, score retrieval, citation accuracy and answer accuracy against a written rubric, and rerun the set after every change to documents, prompts or models. Measure time to answer against a baseline before claiming savings.
Does an AI document assistant replace Procore or SharePoint?
No. Your document system stays the source of truth and the assistant reads from it. Keeping original files intact next to the index means every citation opens the real document.
Next step
Have a workflow in mind?
Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
Scheduling & P6 · October 8, 2026
How Specialized AI Agents Standardize CPM Scheduling
A walkthrough of Connect, the agentic platform we built on Syncify, where one agent per step turns drawings, specs and client answers into validated, versioned CPM schedules.
AI agents & MCP · October 8, 2026
Governing AI Agents in Construction: A Checklist for IT and Leadership
A practical checklist for putting AI agents into construction systems safely, covering credentials, least privilege, read-first rollout, tool routing, telemetry, evaluation and human approval, drawn from our own builds.
AI agents & MCP · January 16, 2026
Building AI Agents for Construction with MCP Tools and Procore
A walkthrough of Construct.Chat: building a Procore financials agent with MCP tools, auditing every tool call, and the structure of a 735-tool Procore MCP server.


