Skip to content

How TinyConductor is checked

A process engine is only useful if you can trust it with work. TinyConductor promises four things about every request it confirms:

  • Nothing is lost. When the API confirms a command (a started instance, a completed job, a published message, a cancellation and so on), the change is stored and takes effect, also if an engine node or the database fails right after the answer.
  • Nothing is repeated. A confirmed command takes effect exactly once. A command whose answer was lost takes effect at most once when the client repeats it with the same Idempotency-Key.
  • No job is handed to two workers at once. A job belongs to one worker until that worker finishes it, gives it back, or its timeout runs out.
  • Nothing gets stuck. Every process instance either finishes or is visibly waiting for something: a job, a timer, a message, a user task, an incident or a child instance.

This page explains how these promises are tested. The same test runs against every mode (Embedded, Bundled and Cluster) and is a release gate.

The test has three parts.

  1. A workload that keeps a ledger. A test client starts process instances and runs job workers, a message publisher, a user-task clerk and an incident resolver, the way an application would. It writes down every request it sends and every answer it gets, and every job each of its workers receives, in an append-only log. The workload’s process covers jobs that complete, fail and retry, fail into an incident and are resolved, and throw a business error; messages published before and after the instance waits for them; message correlation; a timer; a user task that is assigned and completed; cancellation after changing variables; and a call to a child process.
  2. Faults while it runs. While the workload runs, a fault runner breaks things on a schedule drawn from a seed, so every run can be repeated exactly (see below).
  3. An independent checker. After the workload has finished and the engine is idle, a separate checker reads everything the engine wrote: the full record stream, the search and history tables, and each instance as the API reports it. It compares all of it with the client’s log. The checker shares no code with the engine, so a mistake in the engine cannot hide itself.

Any violation fails the run. The report names the exact instance, job, message or record involved, and the seed that reproduces the run.

PromiseWhat is checked
Confirmed commands take effect onceFor every command the client got a success for, the matching records are in the stream exactly once: one creation per start, one completion per job, one stored message per publish, and so on. Each command also sets a variable with a unique name, and that variable must appear exactly once. A command the engine refused must have had no effect.
Unanswered commands take effect at most onceCommands whose answer was lost (network error, timeout, server error) and were repeated with the same Idempotency-Key appear at most once.
JobsEach job is created once and ends at most once (completed, failed with no retries left, error thrown or canceled). No job is handed to a second worker while the first worker still holds it (before its timeout). A completion from a worker whose job was already handed to another worker is refused. No job is handed out again after its completion was confirmed.
Element lifecycleEvery element follows the legal order (activating, activated, completing, completed; or terminating, terminated), with no step twice, and starts only inside a scope that is running. A finished instance leaves nothing running behind it.
Nothing stuckOnce the workload has finished, every instance is finished or waiting for something real. No job of a type that workers serve is left open, no timer is left overdue, and no message is left published but unmatched next to an instance waiting for it.
Keys and positionsNo key and no record position is used twice, and within one instance the stream never puts an effect before its cause or goes back in time.
Tenant isolationEvery record of an instance belongs to the instance’s tenant. A second tenant’s session probes the first tenant’s instances, jobs, incidents, user tasks, variables and audit log by key and by search, and tries to change them. Each probe must find nothing and change nothing.
History and searchOnce the history has caught up, the search and history tables agree with the record stream for every instance, element, incident and user task. So does every instance read back through the API.
MessagesA message is correlated at most once to each waiting instance. A published message is stored once per messageId.

Each run draws its fault schedule from a seed, so the same seed repeats the same faults at the same points of the workload.

FaultWhat happens
Node crashAn engine node is killed without warning and started again, sometimes after its hold on its work has run out, so another node takes over.
Paused nodeA node is frozen for a few seconds, or for longer than its hold on its work. The second case is a node that wakes up believing it still owns work that another node has taken over.
Database crashPostgreSQL is killed and started again on the same storage.
Database failoverThe PostgreSQL primary is killed, a synchronous standby is promoted, and the proxy in front of the database is switched to it.
Slow or lossy networkThe database’s network gets added delay or packet loss for a while.
Network partitionThe database is cut off completely for up to 15 seconds.
Crash during a deploymentA node is killed while a new process version is being deployed.
Clock skewOne node’s clock jumps forward or back (by a few seconds, by more than a node’s hold on its work, or by a minute) for up to 30 seconds while it keeps running.

Cluster nodes share the work of a database between them. Each node owns a share of the work under a time-limited hold (a lease) stored in the database. When the hold changes hands, a counter goes up, and every write checks that counter, so a node that lost its hold can never write again. This protocol is described in a formal model and checked with the TLC model checker, which tries every order in which the steps can happen within small limits: crashes, pauses, restarts, database commits whose outcome the node cannot see, clean hand-overs and takeovers. The checker confirms that:

  • at any time, at most one node writes for a given counter value;
  • no confirmed change is ever lost;
  • no write from a node that lost its hold is ever accepted;
  • with fair scheduling and a bounded number of faults, the work always finds an owner again and every request gets an answer.

The model assumes the database itself works correctly and loses nothing it committed. Through a database failover, that needs synchronous replication, as described in Cluster mode.

Two kinds of repetition are part of the contract and are reported, but do not fail the test:

  • A job handed out again after its timeout. If a worker does not finish or give back a job before its timeout, for example because its completion request failed during a database failover, the job is handed to a worker again. The work then runs twice. Job workers should be written so that running a job twice is safe, and completions should be sent before the timeout runs out.
  • A start repeated without an idempotency key. If a start request’s answer is lost and the client sends the start again without an Idempotency-Key, a second instance is created. Send an Idempotency-Key with every command you might repeat.
  • On every change to the engine: a quick run of about five minutes per mode, with faults, must finish with no violations.
  • Before every release: a 24-hour run with faults throughout must finish with no violations.