Skip to content
Katabench
Try free
8 min read The Katabench team

Dead-Letter Queues: Diagnose Failures and Replay Safely

Use a dead-letter queue to isolate poison messages, diagnose failed deliveries, and replay repaired work without duplicating effects or hiding lost jobs.

The queue is busy, workers are running, and one invoice never arrives. Its message has been received repeatedly. Every attempt throws the same parsing exception. Adding another worker gives the bad message another process to fail in.

A dead-letter queue, usually shortened to DLQ, holds messages that the normal consumer cannot finish under its delivery policy. Moving failed work there lets healthy work continue and preserves evidence for investigation. It does not repair the message or finish the business operation.

The useful design question is therefore bigger than "where do failed messages go?" Someone must recognize the failure, decide whether the work is still valid, and resume it without repeating an effect that already happened.

Healthy work continues while failed work gets a recovery path

Normal delivery

Queue → validate → durable effect → complete

A repeated operation is recognized before applying the effect again.

Failed delivery

Invalid input or exhausted attempts → DLQ

Preserve the payload, operation ID, and diagnostic reason.

  1. 1. Diagnose
    Find the cause and decide whether the work is still valid.
  2. 2. Repair
    Fix the handler or approve a payload correction.
  3. 3. Replay
    Return to the queue with the same business operation ID.
Replay returns work to the idempotent consumer. Quarantine alone does not complete the operation.

A dead-letter queue is a recovery inbox. If nobody owns its contents, it is a quieter form of losing work.

Separate a poison message from a temporary failure

A database connection failure may disappear on another attempt. A message containing an invalid invoice identifier will not become valid because the worker waits longer. A schema version the consumer does not understand needs a compatible consumer or a deliberate transformation.

The term poison message describes input that repeatedly prevents successful processing under the current handler. The fault may be in the payload, the handler, or the relationship between them. Calling the message poison does not prove that the producer is wrong.

Failure Useful first response What changes before recovery
Temporary dependency outage Bounded retries with backoff Dependency becomes available
Invalid required field Quarantine with a reason Correct input or business decision
Unsupported schema version Quarantine and investigate rollout Compatible handler or approved conversion
Handler bug on valid input Limit repeated attempts Fix and test the handler
Completion acknowledgment lost Expect possible redelivery Idempotent handling recognizes prior work

Keep transport retries distinct from application attempts. An SDK retrying a network operation is not necessarily another invocation of your business handler. Record enough context to tell the difference before tuning a retry count.

Azure Service Bus, for example, can move messages to its DLQ after the configured maximum delivery count is exceeded. Applications can also dead-letter explicitly and attach a reason and description. Those are broker-specific mechanisms implementing the broader quarantine idea. Microsoft: Service Bus dead-letter queues

Reject known-invalid input before performing effects

Suppose the message requests creation of an invoice. It carries a business operation identifier that remains stable across delivery attempts:

JSON
{
  "operationId": "ed688dfa-4415-4425-a11b-981482535d29",
  "invoiceId": "cff96a5f-6236-4514-a6ed-af1db8b9d439",
  "schemaVersion": 1,
  "currency": "EUR"
}

The consumer's request model represents those fields explicitly:

C#
public sealed record InvoiceRequested(
    Guid OperationId,
    Guid InvoiceId,
    int SchemaVersion,
    string Currency);

Validate the schema and required fields before creating the invoice. This excerpt handles an unreadable payload in an Azure.Messaging.ServiceBus processor callback. It assumes the processor has AutoCompleteMessages = false, so successful processing must explicitly complete the message.

C#
InvoiceRequested? request;
try
{
    request = args.Message.Body.ToObjectFromJson<InvoiceRequested>(
        new JsonSerializerOptions(JsonSerializerDefaults.Web));
}
catch (JsonException)
{
    await args.DeadLetterMessageAsync(
        args.Message,
        deadLetterReason: "InvalidJson",
        deadLetterErrorDescription: "Invoice request could not be parsed.",
        cancellationToken: args.CancellationToken);
    return;
}

if (request is null || request.OperationId == Guid.Empty)
{
    await args.DeadLetterMessageAsync(
        args.Message,
        deadLetterReason: "MissingOperationId",
        deadLetterErrorDescription: "A stable operation identifier is required.",
        cancellationToken: args.CancellationToken);
    return;
}

// Validate the remaining fields and schema version before processing.
// Commit the idempotent business operation, then complete the message.

The excerpt intentionally stops before business processing. The missing success path matters: completing before the durable effect can lose work, while completing afterward can produce a redelivery if the acknowledgment fails. The SDK method accepts the custom reason and description shown here. Microsoft: DeadLetterMessageAsync

Do not surround the whole handler with a catch that labels every exception InvalidJson. A database timeout and a programming bug require different responses. Cancellation during shutdown also needs careful handling; it is not evidence that the payload is malformed.

Use compact reason codes that operators can group. Keep secrets and unnecessary personal data out of diagnostic descriptions. Store sensitive payloads under an explicit access and retention policy instead of copying their contents into every log line.

Investigate the failure class before replaying

Start with the reason code, consumer version, schema version, and operation identifier. Then ask whether the failures began with a producer deployment, a consumer deployment, or a dependency incident. A collection of messages failing at the same field is a different investigation from unrelated messages whose locks expired during a worker restart.

Preserve the original payload and relevant metadata before repair. Record what changed and who authorized the business correction. If the fix changes an amount or recipient, it is a business decision as well as an operational action.

An unknown currency may indicate invalid input, but it could also mean a legitimate currency was added before all consumers understood it. Replaying into the same incompatible handler simply rebuilds the same backlog. First reproduce one failure in a controlled environment, then verify the proposed repair against that example.

Some messages should never be replayed. A time-sensitive request may have expired in business terms, or a customer may already have canceled the operation through another path. Record a final disposition and reconcile the user-visible state instead of treating an empty DLQ as the only goal.

Make duplicates safe before opening the replay valve

A worker can commit an invoice and crash before completing its queue message. The broker cannot infer whether the database commit occurred, so later delivery must be safe. Microsoft's settlement guidance explicitly describes completion failures and the need to recognize repeated work. Microsoft: Message settlement

For a local database effect, record the operation and create the invoice atomically. This PostgreSQL example assumes processed_operations.operation_id has a unique constraint. The placeholders are bound parameters:

SQL
WITH accepted AS (
    INSERT INTO processed_operations (operation_id)
    VALUES (@operation_id)
    ON CONFLICT (operation_id) DO NOTHING
    RETURNING operation_id
)
INSERT INTO invoices (id, operation_id, currency)
SELECT @invoice_id, operation_id, @currency
FROM accepted;

A repeated operation inserts no row into accepted, so it creates no additional invoice. If the invoice insertion fails, the statement rolls back its operation marker too. PostgreSQL documents the conflict handling and RETURNING behavior used by this statement. PostgreSQL: INSERT

Scope operation identifiers to the consumer when several independent consumers legitimately process the same event. Retain deduplication records long enough for the supported replay window. If identifiers can be reused with different payloads, store and check a payload fingerprint or another invariant rather than silently treating conflicting requests as duplicates.

This statement does not make an external payment or email atomic with the database. External effects need their own idempotency contract or a durable handoff. The idempotency article develops that boundary in more detail.

Replay as a controlled transfer

Choose a bounded group of messages with the same diagnosed cause. Repair or upgrade the handler, then replay a small sample and inspect its business outcomes. Increase the rate only while live traffic and dependency capacity remain healthy. Recovery traffic consumes the same resources as new work.

When a replay tool sends a replacement and then removes the DLQ copy, sending first avoids deleting the only durable copy before acceptance. A crash between send and removal can still produce a duplicate. Preserve the stable business operation ID and rely on the consumer's idempotency boundary; do not generate a fresh operation ID merely to make the broker accept the replay.

Transport message IDs and business operation IDs need not be identical. Broker duplicate-detection settings may influence which transport ID a replay uses. Record the original transport ID and a replay identifier for auditability while preserving the business identity.

Ordering also needs an explicit decision. Quarantining an early event and processing later events can leave the replayed event stale. A version check, per-entity sequencing rule, or reconciliation job may be necessary. A DLQ cannot restore business ordering just by putting a message back.

Measure recovery, not only queue depth

Monitor new dead-letter arrivals by reason and the age of unresolved work. Track replay outcomes and repeated dead-lettering of the same operation. Assign an owner and a response expectation that matches the business effect: a delayed analytics record and a missing paid invoice have different urgency.

The queue-based load leveling guide addresses normal backlog and worker capacity. A DLQ adds a separate recovery workflow. A healthy main queue does not establish that all accepted work completed.

To practice the design, trace one operation through a crash after its database commit, then through a replay after a handler fix. Katabench's track guide describes the System Design Studio, while how it works explains the broader feedback model. Use that same habit here: name the failure, follow the durable state, and prove what the user sees after recovery.

Practice System Design

More like this: System Design Studio →

Get new puzzles and .NET tips in your inbox

A short note when fresh kata land, plus the C# and performance tricks behind the grading. No spam, unsubscribe anytime.