Outworx
All case studies
Steadhold logo

Backend-as-a-Service

Steadhold

A developer asks for a project and gets an isolated PostgreSQL 17 database with auto-generated APIs, auth, row-level security and object storage. The interesting part is what happens when the machine doing the work dies halfway through.

p50 create, 20 consecutive runs
2.46 s
kill points proven to resume clean
11
failures across the measured runs
0
Built byAbdalla Mohamed, founder of Outworx. Designed and built end to end — control plane, data nodes, dashboard, operations.StackTypeScript · PostgreSQL 17 · Docker · PgBouncer · PostgREST · pgBackRest · Next.js · Prometheus · Grafana · Loki

The problem

Provisioning a database is a distributed transaction in disguise. One request has to reserve capacity, create a volume, start a container, wait for Postgres to become healthy, create roles and schemas, start a pooler and an API layer, mint keys, register a route and enable backups — and most of those steps touch a system that can fail independently of the row that claims they happened.

If a worker dies between starting a container and recording it, a naive retry starts a second one. This class of bug never shows up in development, because in development nothing dies.

The design

The API never provisions. It commits intent — a row and a job in one transaction — and returns 202. Workers run an eleven-step saga where every step is written as a question before an action: does this already exist? If so, verify it and continue; if not, create it. That rule is what makes the saga resumable.

Desired state lives in the control plane and actual state on the node. A reconciliation loop closes the gap, deliberately asymmetrically: containers are restarted or stopped automatically, but an orphaned volume is only ever reported. Nothing that holds data is deleted by a machine.

Proving it, rather than claiming it

The worker was killed with SIGKILL at each of the eleven steps, and every run was checked for duplicate containers, volumes, credentials and leaked capacity bookings. Twenty create-and-delete cycles leave nothing behind.

Backups get the same standard: each one is restored into a scratch cluster and checked for completed recovery, page checksums, pg_amcheck on the indexes, and row counts. To trust the verifier, it is deliberately fed a corrupted backup and expected to reject it and name the check that caught it.

Security and operations

Organisations have three enforced roles, and every mutating call writes to an append-only audit table held by a database trigger — a test fails if any mutating route is unaudited. Each project gets its own ES256 keypair published over JWKS, secrets are envelope-encrypted, and deletion takes a final backup first with a seven-day recovery window. Point-in-time restore replays WAL into a fresh project and never touches the original.

Free 15-min onboarding call

Tell us what you're building. We'll bring the team.

Outsourced engineering, design, consulting — or kicking the tires on one of our products. Book a 15-minute call, or send us a message and we reply within one business day.

No credit card · No upfront commitment · NDA on request

Chat with Outworx on WhatsApp