Build Flows

Custom applications · October 9, 2026 · 11 min read

Resumable Large File Uploads to Azure Blob for Construction Files

An engineering walkthrough of Connect's Files layer: signed direct-to-blob uploads, server-side sessions that let multi-gigabyte uploads pause and resume, idempotent ZIP expansion, provenance and the handoff to indexing.

By Charley Forey, founder of Build Flows

Construction files are big and awkward. A drawing set arrives as one 150 MB PDF, a model package is several gigabytes, and the plan room hands over everything as a single Drawings.zip. Then the person uploading it is on a jobsite trailer's Wi-Fi, closes the laptop halfway through, and expects it to still be there in the morning.

When we built the Files layer for Connect, the AI platform we built on top of Syncify, we treated this as an engineering problem in its own right, not a form with a file input. Everything the agents do downstream (reading drawings, citing specs, building schedules) depends on files landing intact, being findable, and being indexed. This article walks through how uploads work: bytes going straight from the browser to Azure Blob, server-side sessions that make large uploads pausable and resumable, versions and folders, ZIP expansion, provenance, and the handoff to indexing.

The shape of the problem

We had four requirements that pulled against each other:

  • The API server should never carry file bytes. Proxying gigabytes through Node ties up memory and connections on the box serving the app.
  • A failed upload must not look like a file. If the browser dies halfway, nobody should see a zero-byte entry they can click on.
  • Large uploads must survive a pause, a reload or a dropped connection. Restarting a 6 GB upload from zero is not acceptable.
  • Nothing the browser holds should be able to overwrite existing data. An upload URL that leaks into a log should not become a way to replace someone's drawings.

The design that satisfies all four is direct-to-blob upload with short-lived signed URLs, a database row that tracks the upload's state, and a server-side "complete" step that only trusts what it can observe in storage.

Step one: a ticket, not a byte stream

An upload starts with a small JSON request to the API: display name, size, media type, optional project and folder. The server validates it, creates the file and version rows in one transaction, and returns an upload ticket. The ticket holds a signed URL pointing at exactly one blob, plus the headers the browser must send.

The sequence looks like this:

  1. Browser posts file metadata to the API.
  2. API creates a file_assets row and a file_versions row in status uploading, links any projects, and points the file at the new version.
  3. API signs a short-lived SAS URL (15 minutes) for the blob <fileId>/<versionId> in the workspace's container.
  4. Browser PUTs the bytes directly to Azure Blob.
  5. Browser calls complete on the API.
  6. API calls HEAD on the blob, reads the real size and ETag from storage, and only then marks the version ready.

A few details in that list do real work.

One container per workspace. The storage account's own boundary matches the tenancy boundary, so a query bug can't mix tenants' bytes.

Blob names are derived, never chosen. The name is always <assetId>/<versionId>, both fresh server-minted UUIDs. Versions are immutable, so each version has exactly one blob.

Two independent locks against overwrite. For a normal upload, the SAS grants create permission only, not write. On top of that, the browser must send If-None-Match: *, which tells storage to refuse the PUT if a blob already exists. Either mechanism alone would stop a leaked URL from replacing data. We use both because the cost is one header.

The listing filters out unfinished uploads. The file row exists before any bytes do, because the ticket needs it. The file browser only lists files whose current version is ready, and ready can only be set by the complete step, which reads size and ETag from storage. A database check constraint ties the ready status to those observed values being present. A closed tab, an expired SAS or a CORS failure leaves a row nobody sees, not a broken file. (A scheduled sweep to tidy those invisible rows is the next step.)

CORS is checked at boot, and a failure is fatal. Direct upload silently breaks without the right CORS rule, so the server ensures it at startup and refuses to start otherwise.

Step two: large files become staged blocks

A single PUT to Azure Blob has a size ceiling, and it is all-or-nothing. Above roughly 4 GiB, Connect switches to Azure's block model:

  • Put Block uploads one chunk (we use 100 MiB blocks) under a block id. Staged blocks are stored but not yet part of the blob.
  • Put Block List commits an ordered list of block ids, and the blob comes into existence in one atomic step.

For this path the server signs a SAS with create and write, because staging blocks can't be done under create-only. The protection moves to the commit: the browser still commits with If-None-Match: * against a fresh blob name nothing else points at. The overall cap is 20 GiB per file.

The trick that makes resume cheap is deterministic block ids. The browser derives each block's id from its index, zero-padded and base64-encoded. A resumed upload doesn't need to remember which random ids it used. It recomputes the same list, skips the blocks before the checkpoint, stages the rest, and commits the full ordered list.

In illustrative TypeScript:

const blockId = (i: number) => btoa(String(i).padStart(12, "0"))

for (let i = Math.floor(startOffset / blockSize); i < totalBlocks; i++) {
  await putBlock(url, blockId(i), file.slice(i * blockSize, (i + 1) * blockSize))
  await saveProgress(sessionId, { staged: i + 1, nextOffset: (i + 1) * blockSize })
}
await putBlockList(url, range(totalBlocks).map(blockId))

Step three: a session row that outlives the tab

Deterministic ids tell the browser how to resume. It still needs to know where to resume from, and that has to survive a reload. Migration 055 added a file_upload_sessions table for that.

ColumnWhy it exists
file_id, version_id, blob_nameWhich version and blob this upload is filling
mode, block_size_bytes, total_bytesEnough to recompute the block plan exactly
staged_block_ids (jsonb)Which blocks storage already holds
next_offsetThe byte to resume from
statusactive, paused, completed, aborted, expired
created_by, expires_atWho can resume it, and until when

After each block lands, the browser patches the session with the new staged list and offset. That write is best-effort: if it fails, the worst case is re-staging one block on resume. The upload itself doesn't fail.

Resume runs like this:

  1. The user reopens Files and sees a "resume your upload" list of their active or paused sessions in this workspace.
  2. They pick the file again. Browsers can't keep a file handle across a reload, so the user has to re-select it. The client checks that its size matches total_bytes and refuses a mismatch.
  3. The API returns the session view with a freshly signed upload URL. The original SAS expired long ago, so resuming never depends on an old credential.
  4. The client stages from next_offset, commits, and calls complete as usual.

Sessions are scoped to the user who created them, so another member can't see or resume your upload. The same user can pick it up on a different machine. The session's lifetime is seven days, which deliberately matches how long Azure keeps uncommitted blocks. A session that outlived its blocks would promise a resume storage can no longer deliver.

One accepted limitation: files under the block threshold use a single PUT, which restarts from zero on failure. We judged one simple path for those sizes worth more than resume support.

Versions, folders and project links

Once a file is ready, it takes part in the rest of the model.

Versions. Replacing a file adds a new version with its own ticket, and restoring an earlier version just repoints the file. A replacement must keep the original name, so "replace" can't become "rename". A connector-sourced file can't be replaced by hand, because the next sync would revert it. Anyone other than the owner or an admin gets "not found" instead of "forbidden", so the response doesn't confirm the file exists.

Project links. A file can belong to several projects. A member restricted to certain projects sees it if at least one of its projects is in reach. A file with no links is workspace-level.

Default folders (migration 063). Every project now opens with the same system folders: Drawings, From Client, Schedules, Logistics, Meetings and Job Walks. They are ordinary folder rows with a non-null default_key, and a partial unique index allows one folder per key per project. They can't be renamed, moved or deleted. When a file belongs to exactly one project and the caller didn't choose a folder, it auto-files by origin: a job walk capture goes to Job Walks, a generated schedule export goes to Schedules. Anything else stays at the project root for a person to file.

For existing projects, the migration adopted a matching root folder (an existing "Drawings") instead of duplicating it, then moved only routable single-project files. The folder list lives in both the migration and a TypeScript constant, and a test asserts they agree.

ZIP packages become individual, indexed files

Drawing packages usually arrive as one ZIP. Stored as-is, that ZIP is opaque: the agents' drawing and document tools need one PDF per sheet set to read. Migrations 048 to 050 added archive expansion.

When a ZIP version becomes ready, it's queued for expansion. A worker downloads it, plans the extraction, and registers each supported member as its own project file. The member is linked to the ZIP's projects, catalogued and indexable. Planning is a pure function of the archive's bytes, so the whole policy can be unit-tested in memory:

  • Zip-slip is rejected. Absolute paths, drive letters, .. segments and NUL bytes are skipped.
  • Junk is dropped. That covers __MACOSX, .DS_Store, Thumbs.db, AppleDouble files and dotfiles.
  • Only readable types are kept: PDF, DOCX, XLSX, images, CSV, XML, text and XER. Nested archives aren't recursed into.
  • Zip bombs are capped: 500 members, 500 MiB uncompressed in total, 50 MiB per member. If a cap stops extraction early, the run says so in a warning instead of silently dropping files.

Idempotence comes from an index, not from careful code. Each member row stores its parent archive's id and its path inside the archive, under a unique partial index. The expander inserts the asset row first, with ON CONFLICT DO NOTHING, before writing any blob, so re-running expansion is a no-op. That paid off twice. Migration 049 backfilled older ZIPs by simply re-queuing them. Migration 050 re-queued archives that had failed under the strict parser, once a salvage path existed that walks local file headers to recover the intact members of a truncated archive. Neither backfill could duplicate a member.

The member keeps a foreign key to its parent ZIP with ON DELETE SET NULL on that column. Deleting the ZIP leaves its extracted drawings readable, because by then they are files in their own right.

Provenance without coupling (migration 040)

Originally a file only knew "upload" or "connector", so a manual upload, an Info Sheet attachment, a job walk photo and a chat drop all looked the same. Migration 040 added two columns:

  • source_surface: which surface created the file (files_manual, connection, info_sheet, job_walk, chat, and later archive), enforced by a check constraint.
  • source_ref: a UUID pointing at the thing it came from, such as an info sheet, a job walk or a conversation.

source_ref is deliberately not a foreign key. It points into a different table depending on the surface, and a real FK would tie the files table to every capability that can create a file. We accepted a reference the database can't check in exchange for keeping Files independent of every capability that feeds it. The only surface a browser may claim is chat, and the upload route whitelists it with a UUID conversation reference. Every other surface is stamped on the server. The UI uses this to label a file's source and to link a chat drop back to its thread.

Text extraction, limits and the handoff to indexing

When the upload completes, the server queues the version for cataloguing, clears any index built on the previous version, enqueues ZIP expansion if it applies, and (if the workspace has opted into auto-indexing) enqueues an index job. That last step never throws. A knowledge-base hiccup must not fail an upload the user just watched finish.

Indexing extracts text by type. PDFs go page by page, so a large drawing set never becomes one giant string. DOCX goes through mammoth. XLSX goes through a bounded sheet reader with caps on sheets, rows, columns and cell length. Text-like types are decoded directly, and anything containing NUL bytes is rejected as binary. Image-only PDF pages yield no text from this step.

The limits are guards against runaway work, not quality gates:

LimitValuePurpose
Bytes loaded to index500 MiBStops a multi-GB blob being pulled into worker memory
Pages5,000Bounds extraction time
Chunks200,000Bounds embedding cost
Chunk size / overlap1,600 / 200 charactersFits the embedding model's input window
Batch200 chunksKeeps memory flat and requests inside limits

An index job is a leased, resumable background job, the same model as connector runs. It renews a two-minute lease after each batch, records processed_units, and checks for cancellation. Chunking is deterministic, so a worker that picks up an abandoned job rebuilds the identical chunk list and carries on from the checkpoint without re-embedding earlier chunks. If a file's project links changed during indexing, the job throws its output away and re-enqueues, so a chunk is never visible to the wrong project.

Why this matters if you're building something similar

  • Keep bytes off your API. Signed, single-blob, short-lived URLs let storage do what it's good at. Your server only handles metadata and verification.
  • Make "ready" something you observe, not something you're told. Promote a file only after reading size and ETag from storage. The client's word that the upload finished isn't enough.
  • Put resume state on the server and derive the rest. Deterministic block ids plus a session row with an offset is a very small amount of state, and it survives reloads and device changes.
  • Match your retention to the platform's. A seven-day session matches seven days of uncommitted blocks. Promising more than storage keeps is a bug that only shows up later.
  • Get idempotence from a unique index. If re-running a job is always safe, backfills and retries stop being risky operations.
  • Be deliberate about which references are foreign keys. A polymorphic reference that isn't an FK is a fair trade when the alternative couples one table to every feature.

Where to go next

Have a file-heavy workflow that needs to be built properly? Plan your build.

Frequently asked questions

How do you resume a large upload to Azure Blob Storage after a page reload?

Stage the file as blocks with Put Block, using block ids derived from each block's index, and record staged blocks and the next byte offset in a server-side session. On resume, the user re-selects the file, the server issues a fresh signed URL, the client stages the remaining blocks and commits the full ordered list with Put Block List.

Why upload directly from the browser to blob storage instead of through the API?

Proxying gigabytes through an application server ties up memory and connections and couples upload throughput to the app. A short-lived SAS URL scoped to one blob lets storage receive the bytes while the server only creates records and verifies the result.

How long can a paused Azure Blob block upload be resumed?

Azure keeps uncommitted blocks for about seven days, so Connect's upload sessions expire after seven days. A session that outlived its staged blocks would promise a resume that storage can no longer honour.

How do you stop a leaked upload URL from overwriting files?

Blob names are fresh server-minted ids per immutable version, normal uploads use create-only SAS permission, and the client must send If-None-Match: * so storage refuses to write over an existing blob.

How are ZIP files of drawings handled?

A worker expands each ZIP into individual project files, rejecting path traversal, skipping junk and unsupported types, and capping member count and size. A unique index on the parent archive and member path makes re-running expansion safe, and damaged archives are salvaged where possible.

Next step

Need something built around how your team works?

Describe the users, the workflow, and the systems it touches. We'll tell you whether a custom application makes sense and how we'd build it.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning