Quick overview
This workflow implements a checkpoint-and-resume system for a multi-stage pipeline, storing stage outputs in Postgres and sending failure/recovery alerts to Slack, with manual, webhook, and scheduled triggers to start runs and automatically resume failed or stalled executions.
How it works
- Starts a new run via Manual Trigger, receives a POST request via webhook, or runs every 10 minutes on a schedule to sweep for resumable runs.
- Ensures the required Postgres tables for runs, checkpoints, and run events exist.
- For scheduled sweeps, queries Postgres for failed runs whose retry time has passed or running runs whose lease expired, then calls this workflow’s webhook to resume each run.
- For manual/webhook runs, loads the run state and completed checkpoints from Postgres, restores prior stage outputs into context, and determines the first stage that still needs to run.
- Acquires a lease-based lock in Postgres to ensure only one execution owns the run, and exits if the run is already completed or currently locked by another execution.
- Executes the pipeline stages in order (validate order, reserve inventory, charge payment, create shipment, send confirmation), saving a Postgres checkpoint and event after each successful stage.
- On stage failure, retries the stage inline with exponential backoff when allowed, otherwise marks the run as failed or dead-lettered in Postgres and sends a Slack alert, and on completion marks the run completed and optionally posts a Slack recovery notification when resuming from checkpoints.
Setup
- Add Postgres credentials for the database where the workflow can create and update the workflow_runs, workflow_checkpoints, and workflow_run_events tables.
- Add Slack credentials and set the target Slack channel in the configuration (and update the Slack nodes if you want a different channel behavior).
- Update the selfWebhookUrl value in the configuration to this workflow’s production webhook URL ending in /webhook/checkpoint-run so the sweeper can trigger resumes.
- If you want crash notifications, enable the Error Trigger node and set this workflow as the Error workflow in your n8n instance settings.