Design a file storage service: uploads, scans, downloads
Design a file storage service where bytes never pass through the API: scoped signed uploads, a metadata state machine, an async scan queue, and edge downloads.
The incident report reads the same every time. A customer uploads a 2 GB video export through
POST /files, the request runs for eleven minutes, and at minute ten the load balancer's idle
timeout closes the connection. The customer retries, two more customers start uploads of their
own, and the API is now holding several long transfers while memory climbs and the health check
in the same process starts timing out. The storage system behind all of this is idle. It was never
the bottleneck. The application server was, because someone decided file bytes should travel
through it.
Everything in this design follows from one rule: the API decides who may store what, and object storage moves the bytes.
The naive design posts a multipart form to the API, which reads the stream and writes it to storage. It works in the demo and fails at the third concurrent large upload, for mechanical reasons:
- Every byte crosses the application server twice, in from the client and out to storage. A 2 GB upload is 4 GB of traffic on a machine sized for JSON.
- The request holds a connection, a body buffer, and a request slot for the whole transfer. Fifty clients on slow mobile links occupy fifty slots for an hour.
- Timeouts are sized for API calls, not transfers, and a failure at 99 percent restarts from zero because a proxied stream has nothing to resume.
Back-of-the-envelope: 1,000 uploads an hour at 200 MB each is 200 GB an hour through the API, roughly 450 Mbit/s sustained in each direction, before any product traffic. Object storage absorbs that without noticing. An API fleet would have to be scaled for bandwidth it should never carry.
Direct-to-storage upload splits the operation into a control call and a data transfer. The client
asks the API for permission. The API records a metadata row, then asks the storage provider to sign
a URL that permits exactly one operation: a PUT to one object key, with a maximum size, a fixed
content type, and an expiry a few minutes out. The client uploads straight to storage with that
URL. The API never sees the bytes.
Microsoft's architecture guidance calls this the valet key pattern: a token granting limited, time-boxed access to one resource. Every major object store has a signing primitive; the header names differ and the shape does not.
public async Task<UploadGrant> BeginUploadAsync(
Guid ownerId, string displayName, string contentType, long declaredSize,
CancellationToken cancellationToken)
{
_limits.EnsureAllowed(contentType, declaredSize);
var fileId = Guid.NewGuid();
var objectKey = $"tenants/{ownerId}/{fileId:N}"; // never the user's file name
_db.Files.Add(new StoredFile
{
Id = fileId,
OwnerId = ownerId,
DisplayName = Path.GetFileName(displayName),
ContentType = contentType,
DeclaredSize = declaredSize,
ObjectKey = objectKey,
Status = FileStatus.Pending
});
await _db.SaveChangesAsync(cancellationToken);
// The signature covers key, method, content type, and maximum length.
var signed = _storage.SignPut(new SignPutRequest(
Key: objectKey,
ContentType: contentType,
MaxContentLength: declaredSize,
ExpiresIn: TimeSpan.FromMinutes(15)));
return new UploadGrant(fileId, signed.Url, signed.Headers, signed.ExpiresAt);
}Three details in that method carry most of the system's security. The object key is generated by the server and never derived from the user's file name, which closes the path traversal family of bugs before it opens. The size and content type are part of the signature, so a 10 MB image slot cannot receive a 4 GB archive. And the URL expires in minutes, so a leaked grant is a small problem instead of a permanent one.
One 2 GB upload, two paths
Through the API, or straight to storage
Proxied through the API
every byte crosses the app server twice
multipart/form-data - client → API 2 GB in
body buffered, one request slot held for 11 min
- API → storage 2 GB out
the same bytes again, on the API's network interface
the API is the bottleneck; storage is idle
Direct to storage
the API only ever touches metadata
signed PUT, 15 min - 01 POST /files 300 B
client → API: row saved as pending; signed PUT for one key, one content type, one max size, 15 min
- 02 PUT signed-url 2 GB
client → storage: the API is not on this path
- 03 POST /files/{id}/finalize 100 B
client → API: HEAD the object, check length and hash, pending → uploaded
- 04 file.uploaded key only
queue → scanner: verdict sets available or quarantined; a scanner outage only grows the queue
2 GB moved once; the API served about 400 bytes
files.status
The object in the bucket is not the file. The row in your database is the file, and the object is one of its attributes. That row moves through a small set of states, and every operation in the service is a guarded transition between them:
CREATE TABLE files (
id uuid PRIMARY KEY,
owner_id uuid NOT NULL,
object_key text NOT NULL UNIQUE,
display_name text NOT NULL,
content_type text NOT NULL,
declared_size bigint NOT NULL,
actual_size bigint,
sha256 bytea,
status text NOT NULL CHECK (status IN
('pending', 'uploaded', 'scanned', 'available', 'quarantined', 'deleted')),
legal_hold boolean NOT NULL DEFAULT false,
created_at timestamptz NOT NULL DEFAULT now(),
updated_at timestamptz NOT NULL DEFAULT now()
);
-- finalize: only a pending row can become uploaded, so a second call is a no-op
UPDATE files
SET status = 'uploaded', actual_size = $2, sha256 = $3, updated_at = now()
WHERE id = $1 AND status = 'pending';pending means a grant was issued and nothing is known about the bytes. uploaded means the
server verified what landed. scanned, then available or quarantined, are set by the scanning
pipeline. Only available rows can be downloaded or listed, so a half-uploaded or infected object
is invisible to the product even though it exists in the bucket.
Completion needs a signal: either the storage provider emits an object-created event, or the
client calls POST /files/{id}/finalize. Accept both and make them idempotent; the
WHERE status = 'pending' guard turns the second signal into a no-op. In neither case does the
server take the client's word for it. It issues a HEAD request against the object, compares
length and content type with what was declared, verifies a content hash for anything it will
serve to other users, and moves a mismatch straight to quarantined.
The upload is finished when the server has verified the object, not when the client stops sending. Until then the file does not exist as far as the product is concerned.
Malware scanning, image re-encoding, thumbnailing, and text extraction are slow, and they fail
independently of the upload. Running them inline in finalize makes the client wait on a scanner
it cannot see and ties your upload success rate to the scanner's uptime.
So finalize commits the state change and an outbox row in one transaction, and a worker pool
consumes a file.uploaded queue. Each worker fetches the object from storage, scans it, and writes
the verdict: available or quarantined, plus the scanner version so a later engine update can
re-scan old files. If the scanner is down, the queue grows and uploads keep succeeding; when it
recovers, the backlog drains. That is queue-based load leveling
applied to a job nobody should be waiting on. The metric to alert on is the age of the oldest
uploaded row, and a file that fails scanning three times goes to a dead-letter lane and a human.
Downloads follow the same rule in reverse. GET /files/{id}/download checks authorization and
status, then redirects to a signed download URL, again scoped to one object and expiring in
minutes. The bytes flow from storage, or the CDN in front of it, and the API's involvement ends at
the redirect.
GET /files/7c1e.../download
-> 302 Found
Location: https://cdn.example.com/o/tenants/42/7c1e...?exp=1758470400&sig=...
(the edge serves 2 GB; the API served 300 bytes)
GET https://cdn.example.com/o/tenants/42/7c1e...?exp=1758470400&sig=...
Range: bytes=1048576-2097151
-> 206 Partial Content
Content-Range: bytes 1048576-2097151/2147483648Put a CDN in front of storage for anything read more than once. Signed URLs and
edge caching cooperate: the edge validates the signature and caches by
object key, so the origin sees one fetch per object per region. Serve Content-Disposition from
the sanitized display name in the metadata row, never from anything the object carries.
Range requests matter more than they look: video players seek, PDF viewers load page by page, and
a resumed download asks for the tail.
RFC 9110 section 14 defines Range and
Content-Range; every object store and CDN implements them, and a proxying API would have had to
implement them by hand.
Above a few hundred megabytes, one PUT is fragile: a dropped connection restarts everything.
Multipart upload fixes that. The API opens a multipart session with storage and hands the client
signed URLs for numbered parts (say 32 MB each); the client uploads them in any order and in
parallel, each part returns an ETag, and a final complete call lists the ETags so storage can
stitch the object together. Resumability falls out of that shape: a client that dies at part 40 of
64 asks the API which parts storage has recorded and uploads only the missing ones. A lifecycle
rule on the bucket aborts sessions older than a day, so abandoned uploads do not accumulate as
billable fragments.
Deletion, legal hold, quota enforcement, and re-scanning belong on a separate surface from the
user-facing API: different routes, different credentials, ideally a different deployment. Deletion
is a transition to deleted plus a scheduled object removal after a retention window, because a
hard delete on request is how a compromised account erases the evidence of its own compromise.
legal_hold = true blocks that transition in the state machine, not in a UI check. Quotas are
enforced twice: at grant time against the declared size, and at finalize against the actual size.
When a quarantine or a quota breach needs to reach a person, the fan-out belongs in a
notification system, not in the finalize handler.
Every decision above is a topology decision: which component may talk to which, and which path bytes are never allowed to take. Katabench's System Design Studio has a "Build a File Vault" series that grows this design on a canvas one constraint at a time: "Store the First File", "Upload Directly to Storage", "Scan Before You Share", "Download from the Edge", and "Guard the Vault". Each challenge is graded structurally and deterministically against authored rules, so an uploader that pushes bytes through the API, or a scan verdict that never reaches the metadata store, fails for the specific bypass it contains. The tracks overview shows where system design sits beside the code tracks. The mistake in a file service is drawn long before it is coded, and the canvas is the cheapest place to catch it.
Practice what you just read
- Open the Pro kata: Store the First File
System Design Studio Easy Katabench Pro
Store the First File
The MVP stuffs every upload into a database row and backups now take all night. Give file bytes an object store and keep the database for metadata.
More like this: System Design Studio →