How Sturdle works
An engine that runs in your process and keeps every job, step and log line in a schema of your own Postgres. This page walks through what it does today, mechanism by mechanism, and names what is still planned.
A schema of your own database
On the first connect the adapter creates a dedicated schema, sturdle by default, and applies its versioned SQL migrations inside a session-level advisory lock, so several instances can boot at the same time and only one applies each migration. Six tables hold everything: jobs, the queue streams, delayed jobs, per-job metadata such as steps and logs, distributed locks, and the migration bookkeeping.
Teams whose DBAs review every change can export the DDL without touching a database and apply it through their own tooling; the adapter then finds nothing left to migrate.
import { migrationsSql } from "@sturdle/postgres";
for (const { id, sql } of migrationsSql("sturdle")) {
console.log(id, sql);
}Claiming work without contention
A worker claims the oldest unacknowledged job in its queue with FOR UPDATE SKIP LOCKED, so two workers never wait on each other for the same row, and marks it processing in the same transaction. A transaction-scoped advisory lock serializes claims per queue, which is what makes the per-key concurrency count below consistent across workers.
A running job renews its lease every thirty seconds. If a process dies, the job goes stale and is released back to the queue after five minutes by default, or dead-lettered when it has already used its attempts.
SELECT s.job_id
FROM sturdle.job_streams s
JOIN sturdle.jobs j ON j.id = s.job_id
WHERE s.queue_name = $1
AND s.acknowledged = false
AND (s.consumer_name IS NULL OR s.consumer_name = '')
ORDER BY s.timestamp ASC
FOR UPDATE OF s SKIP LOCKED
LIMIT 1Retries, backoff and the dead letter
A failed attempt is retried until the job's maxAttempts, three by default, is used up. Each retry waits an exponential backoff of one second doubled per attempt, plus up to a second of jitter, capped at five minutes. An error named NonRetriableError skips the retries and dead-letters the job at once.
A dead-lettered job keeps its payload, its attempt history and the reason, and the task's onDeadLetter hook runs with them, which is where an alert or a manual replay starts.
// min(1s * 2^(attempt - 1) + jitter, 5 min)
export function calculateExponentialBackoff(attempt: number) {
const baseDelayMs = 1000;
const maxDelayMs = 5 * 60 * 1000;
const jitterMs = Math.random() * 1000;
const exponentialDelay = baseDelayMs * Math.pow(2, attempt - 1);
return Math.min(exponentialDelay + jitterMs, maxDelayMs);
}Durable steps and replay
A step is a named unit of work inside a job. When a step completes, its result is stored with the job; when the job runs again on a later attempt, or resumes after a sleep, every completed step returns its stored result instead of running. The step's name is its identity, so it must be stable across deploys.
The console shows a replayed step as cached, with no duration: the work happened once. This is what lets a job charge a card exactly once even though the call after it timed out twice.
func: async ({ event, step }) => {
const invoice = await step.run("Fetch invoice", () =>
fetchInvoice(event.data),
);
const charge = await step.run("Charge card", () =>
chargeCard(invoice),
);
await step.run("Record ledger entry", () =>
recordLedger(invoice, charge),
);
await step.sendEvent("Send receipt email", {
name: "email/send",
data: { customerId: invoice.customerId, chargeId: charge.id },
});
}Sleeping for minutes or days
step.sleep and step.sleepUntil suspend the job: the handler unwinds, the worker is freed, and the job is re-enqueued at the wake time from the delayed_jobs table. On wake, the completed steps replay and the handler continues after the sleep, in the same attempt.
await step.run("Send welcome", () => sendWelcome(email));
await step.sleep("Wait three days", 3 * 24 * 60 * 60 * 1000);
await step.run("Send follow-up", () => sendFollowUp(email));Cron with a distributed lock
A task with a cron expression is scheduled by every instance, and a distributed lock in the distributed_locks table makes sure exactly one instance enqueues each tick. The lock has a two-minute lifetime and is never released early: its expiry is the guarantee, so a crashed scheduler cannot hold the schedule hostage.
JobRegistry.register({
name: "Nightly ledger report",
cron: "0 2 * * *",
func: async ({ step }) => {
await step.run("Build report", () => buildLedgerReport());
},
});Concurrency per task and per key
A number limits how many jobs of a task run at once, engine-wide. An object with a limit and a key limits per value of a payload field: a limit of one keyed on customerId lets jobs for different customers run in parallel while jobs for the same customer wait their turn. The count is taken at claim time, inside the per-queue lock, so it never overshoots.
// at most 4 of these run at once, engine-wide
concurrency: 4
// at most 1 per customer, any number of customers in parallel
concurrency: { limit: 1, key: "customerId" }Payloads validated before they run
A task may declare a zod schema for its event payload. The engine validates the payload before the handler runs, so a malformed event fails as a validation error with the issues listed, rather than deep inside a step.
import { z } from "zod";
JobRegistry.register({
name: "Sync invoice to ledger",
event: "billing/invoice.sync",
payloadSchema: z.object({
invoiceId: z.string(),
customerId: z.string(),
}),
func: async ({ event, step }) => {
// ...
},
});One telemetry hook, no vendor SDK
The engine never talks to an error tracker or a hosting platform. It emits engine-level events, job.error, job.dropped, job.dead-lettered, job.assignment-error, hook.error, engine.error, to one onEvent listener, and you forward them to whatever your stack already uses. A listener that throws is logged and never propagates.
const engine = new Engine({
databaseAdapter,
onEvent: (event) => {
if (event.type === "job.dead-lettered") {
errorTracker.capture(event);
}
},
});Connection poolers
Claims use transaction-scoped advisory locks and are safe behind a transaction-pooling proxy such as PgBouncer in transaction mode. Migrations use a session-level lock and are not: run migrate() once, out of band, against a direct or session-mode connection, and let the pooled instances find nothing left to apply.
What is planned
Everything above runs today, in production, inside the product Sturdle was extracted from. What follows does not exist yet and is labelled as such on every page, with the tier it lands in: free under MIT, or Sturdle Pro, a flat licence per organisation. There is no hosted service on the roadmap.
- Preview APIA plug-and-play SDK: defineJob, a single route file and npx sturdle init. Free.
- PlannedAn installable console application; today the console's components are what this site embeds. Free.
- PlannedFlow control: rate limits, throttling, batches and workflows. Sturdle Pro.
- PlannedAlerts that understand retries: one alert after retries are exhausted, with the context to fix it. Sturdle Pro, to validate.
import { defineJob, start } from "sturdle";
import { z } from "zod";
export const welcome = defineJob({
id: "send-welcome",
event: "user/created",
schema: z.object({ email: z.string().email() }),
retries: 5,
run: async ({ event, step }) => {
await step.run("email", () => sendEmail(event.data.email));
await step.sleep("wait", "3d");
await step.run("follow-up", () => sendFollowUp(event.data.email));
},
});
// Reads DATABASE_URL, creates its schema on first run, starts the workers.
await start({ jobs: [welcome] });Questions engineers ask
- Why Postgres and not Redis?
- Because the database is already there, already backed up and already in your region, and because job state next to the data it changes can be queried, joined and audited with the tools you have. Redis-backed queues are faster at very high throughput; if you push millions of jobs an hour, measure first.
- How many jobs can it run?
- We publish no benchmark yet. The claim is a single indexed SELECT with SKIP LOCKED per job, so throughput scales with your Postgres and your worker count; the honest answer is to run the demo workload against your own database before you commit.
- Does it run on serverless?
- Not today. Workers run in a long-lived Node process, the one you already deploy. Platforms without one are not supported for now.
- Is a step run exactly once?
- A completed step is never run again for that job. A step that crashed mid-way, before its result was stored, runs again on the next attempt, so a step's own side effect should be idempotent, as with any durable-execution engine.
- Does it need my ORM?
- No. The adapter uses the pg driver directly and owns its schema; your ORM never sees Sturdle's tables unless you want it to.
- What is free?
- The engine, the Postgres adapter and the console, under MIT, and everything they do today. Sturdle Pro, flow control and retry-aware alerts, is planned as a flat licence per organisation; the Pro page shows the offer and takes licence requests. Nothing is on sale yet, and nothing will meter steps.
Run the demo job on your own Postgres
Early access is open. No credit card, nothing to provision.