Durable steps in TypeScript: memoize, sleep, resume
A background job that calls three external services will, some day, fail after the second call. Durable steps are how a job survives that: each step's result is stored as it completes, so a retry replays what was done and runs only what was not.
The problem a plain retry cannot solve
Retrying a whole job is easy: catch the error, enqueue again. The trouble is what the job already did. If it charged a card and then timed out sending the receipt, a plain retry charges the card again. Making every call idempotent by hand works, but it is exactly the kind of bookkeeping that gets skipped under deadline.
Memoized steps
A step is a named function inside the job. When it completes, its result is stored with the job, keyed by the step's name and the attempt. When the job runs again, whether after a crash, a retry or a sleep, a step whose result exists returns it immediately and its function is not called. The handler is written as straight-line code; durability comes from the runtime.
The name is the identity. Rename a step and it will run again on the next attempt of an in-flight job; reorder steps and the stored results still match by name. Keep names stable, and keep the result JSON-serializable, since that is how it is stored.
func: async ({ event, step }) => {
const invoice = await step.run("Fetch invoice", () =>
fetchInvoice(event.data),
);
const charge = await step.run("Charge card", () =>
chargeCard(invoice),
);
await step.run("Record ledger entry", () =>
recordLedger(invoice, charge),
);
await step.sendEvent("Send receipt email", {
name: "email/send",
data: { customerId: invoice.customerId, chargeId: charge.id },
});
}What happens on failure
When a step throws, the attempt fails and the job is scheduled for a retry with exponential backoff. On the next attempt every completed step replays and the failed step runs again, from the start of that step. This is why a step's own side effect should be safe to repeat: a step that crashed after the charge succeeded but before its result was stored will charge again. Keep each step to one external effect and give it an idempotency key where the provider supports one.
An error named NonRetriableError ends the job at once: no backoff, straight to the dead letter, with the payload, the attempt history and the reason preserved for the onDeadLetter hook.
Sleeping without a worker
step.sleep and step.sleepUntil suspend the job. The handler unwinds, the worker is released, and the job is re-enqueued at the wake time from a delayed-jobs table. When it wakes, its completed steps replay and execution continues after the sleep. A follow-up email three days later costs a row, not a process.
await step.run("Send welcome", () => sendWelcome(email));
await step.sleep("Wait three days", 3 * 24 * 60 * 60 * 1000);
await step.run("Send follow-up", () => sendFollowUp(email));Child jobs with sendEvent
step.sendEvent enqueues another job from inside a step and is memoized like any other step, so a retry of the parent never enqueues the child twice. The console shows the child next to its parent, with its own attempts and logs.
What you see in the console
Each attempt is a lane on the execution timeline: the steps that ran, the one that failed, the backoff wait, and on the next attempt the replayed steps drawn as cached, with no duration. The steps table and the log stream share that selection, so a failed run reads top to bottom without leaving the page.
Try it on the database you already run
Early access is open. No credit card, nothing to provision.