Postgres as a job queue: how it works, and when it is enough

A job queue needs three things from its store: a claim that two workers cannot both win, a way to notice a worker that died, and a place to keep what happened. Postgres provides all three with features it has had for years.

Guides

The claim: FOR UPDATE SKIP LOCKED

The core of a Postgres queue is one query. Select the oldest job that has not been claimed, lock its row for update, and skip any row another transaction already holds. Two workers running the query at the same instant get two different rows or, when only one is left, one of them gets nothing and tries again. There is no polling of a shared counter and no lost update.

Sturdle runs the claim in a transaction that also marks the job processing and increments its attempt counter, so a claimed job is visible as such to every other worker the moment the transaction commits.

SELECT s.job_id
FROM sturdle.job_streams s
WHERE s.queue_name = $1
  AND s.acknowledged = false
  AND (s.consumer_name IS NULL OR s.consumer_name = '')
ORDER BY s.timestamp ASC
FOR UPDATE OF s SKIP LOCKED
LIMIT 1;

UPDATE sturdle.jobs
SET status = 'processing', started_at = now(), attempts = attempts + 1
WHERE id = $1;
The claim in Sturdle's Postgres adapter, trimmed to its shape

Ordering, and the per-queue lock

SKIP LOCKED alone gives no guarantee about counts: two workers can both read that zero jobs hold a concurrency key before either commits. Sturdle takes a transaction-scoped advisory lock per queue around the claim. It costs one lock per claim, it is released at commit, and it makes the per-key limit exact.

Ordering is by enqueue time within a queue; a priority moves a job ahead of older ones. Queues are just a column, so a billing queue and an email queue can be served by different worker pools against one database.

Leases, heartbeats and stale jobs

A worker that crashes mid-job holds nothing: the row lock died with its connection. What remains is a job marked processing that nobody is processing. Sturdle's worker renews a lease every thirty seconds while a job runs; a sweeper releases jobs whose lease is older than five minutes back to the queue, or dead-letters them if they have used their attempts. A job that resumes is treated as a retry: its completed steps replay, its unfinished step runs again.

Dispatcher.heartbeatIntervalMs = 30_000; // the lease is renewed every 30 s
adapter.releaseStaleJobs(queueName, 300_000); // stale after 5 min by default
The two knobs: the heartbeat and the stale threshold

State next to your data

Because the queue is a schema in the same database, everything about a job is a query away from the rows it changes: which invoice a failed charge belonged to, how many jobs a customer has waiting, what the p95 was last night. Backups, replicas and access control are the ones you already run. Nothing leaves the region your database is in.

The cost is that job traffic shares the database's capacity. For most product workloads, thousands of jobs an hour, that is a rounding error; for a firehose of millions, it is a real budget to plan.

When Postgres is enough, and when it is not

Postgres is enough when your jobs are the ones a product generates: emails, webhooks, syncs, reports, billing runs, ordered work per customer. It is the wrong tool when you need sub-millisecond dispatch, when a single queue must absorb millions of messages an hour, or when the database is already the bottleneck of the application. In those cases a dedicated broker earns its operational cost.

Sturdle's engine is storage-agnostic behind a small adapter interface. The Postgres adapter is the one that exists today; it is the one we recommend starting with, because it is the one you already run.

Try it on the database you already run

Early access is open. No credit card, nothing to provision.

Join early access