Skip to content

How the engine runs your processes

This page explains how a durable TinyConductor engine keeps your processes in memory and saves every change to PostgreSQL. You don’t need to know this to use TinyConductor. It helps when you size a deployment, read the metrics, or plan an upgrade.

Bundled and Cluster mode work as described here: Bundled on one node, Cluster on several nodes that share the partitions. Embedded mode runs the same engine with one partition, in memory only, and saves nothing.

The engine splits its work into partitions. A partition is one engine thread with its own share of memory. It runs the process instances of the tenants placed on it, one command at a time, in the order they arrive.

  • A deployment has 24 partitions by default (CONDUCTOR_PARTITIONS). The number is fixed when the database is created.
  • Each tenant has one home partition. All of a tenant’s instances, jobs, timers and messages live there, so a tenant’s work never has to be coordinated across partitions.
  • A partition holds many tenants. Partitions that hold no tenant cost almost nothing.
  • Because each partition is its own thread, a deployment with several busy tenants uses several processor cores.

When a tenant is created, the engine places it on the least-loaded partition: the one with the fewest tenants, then the one that has been least busy, then the lowest number. The tenant stays there. The placement is saved in PostgreSQL, so it survives restarts.

Bundled and Cluster mode start with one tenant, default, and you add more in the console or through the tenant API (see Tenants); each is placed on a home partition as described above.

Tenants on the same partition take turns. The engine measures how much processing time each tenant’s commands use and gives each tenant its fair share. So a tenant that sends a flood of requests slows down only itself: the other tenants on its partition wait for at most one short turn. While other tenants are waiting, one tenant can fill at most half of a save batch.

If a fault is caused by one tenant (for example a command that crashes while it runs, or saved data that fails its checks), only that tenant is quarantined. Its requests answer 503 with the problem type urn:bpm:error:tenant-unavailable, and the other tenants on the partition keep running. GET /v2/tenants/{tenantId} shows the quarantine, with its reason and since when.

Once the cause is fixed, an operator with the tenant:admin permission releases the tenant with DELETE /v2/tenants/{tenantId}/quarantine. The tenant then reloads from its saved state. The call answers 204 and can be repeated safely.

The engine answers a request only after the change it made is saved in PostgreSQL.

  1. The partition applies the command in memory.
  2. It writes one frame to PostgreSQL: the command, together with the records it produced. A frame always belongs to one tenant.
  3. Frames that are ready at the same moment are saved together, in one commit (CONDUCTOR_COMMIT_BATCH caps how many). Batching only happens when commits are slower than the engine, so it adds no waiting time.
  4. Once the commit is done, the engine answers the request.

A tenant’s frames are numbered without gaps. Keys and record positions are never reused, also not after a restart.

Every 5 seconds by default (CONDUCTOR_SNAPSHOT_INTERVAL), the engine saves a segment for each tenant: a copy of the tenant’s current state. Frames that are older than every tenant’s latest segment, and that the exporter has already read, are removed. A tenant with no recent activity gets a fresh segment first, so old frames never pile up.

Searches, history, the audit log and the analytics views read from PostgreSQL query tables. The exporter fills them from the saved frames, shortly after each commit. So a search can be a moment behind the change you just made. The console waits until it sees its own changes.

The exporter keeps track of each tenant separately. If one tenant’s query tables fall behind, the engine holds back that tenant’s new instances until the exporter catches up. Other tenants are not affected.

Behind byWhat happens to new instances of that tenant
half the limit: 500,000 records or 30 secondsThey are slowed down. The engine logs export lag alarm is firing
the limit: 1,000,000 records (about 20 seconds of work at full rate) or 60 secondsThey are refused with 503, urn:bpm:error:backpressure and Retry-After: 1
back under half the limitThey are admitted again. The engine logs export lag alarm cleared

The limits are CONDUCTOR_EXPORTER_MAX_LAG and CONDUCTOR_EXPORTER_MAX_LAG_AGE_MS.

Each partition is held under a lease in PostgreSQL, which names the engine that owns it. The lease is renewed every second or so, and an owner that stops renewing loses it after CONDUCTOR_PARTITION_LEASE_MS (10 seconds by default).

Every engine has an owner identity. In Bundled mode it is always tinyconductor-bundled, because one engine serves one database. When an engine restarts with the same identity, it takes its partitions back at once, without waiting for the lease to run out, and anything the old process still tried to save is refused. Then each partition loads every tenant’s latest segment and replays the frames after it. A restart takes a few seconds.

Two engines must never run with the same owner identity at once. The Bundled image and the Kubernetes chart make sure of that: the Bundled container runs one engine next to its own PostgreSQL, and in Cluster mode the chart’s StatefulSet never runs two pods with the same name.

If the connection to PostgreSQL breaks while a commit is in flight, the engine cannot know whether the commit went through. It answers the waiting requests with 503 and urn:bpm:error:partition-unavailable, reloads the partition from PostgreSQL, and carries on. Retry such a request with the same idempotency key: if the first attempt was saved, the retry returns its result and does nothing twice.

In Cluster mode several engines, called nodes, share one database and its partitions. Each node has its own name (CONDUCTOR_CLUSTER_NODE_NAME), and each partition is owned by exactly one node at a time.

The nodes share the partitions fairly, by how busy each partition is. A node that starts takes free partitions; when a node joins or leaves, the others hand partitions over automatically, at most one per second, so the cell never moves much work at once. Each node keeps at least its equal share (the number of partitions divided by the number of nodes, rounded down). A clean hand-over takes a second or two for the partition being moved. A node that owns no partition still answers every request, by forwarding it.

A client can send any request to any node. If the tenant’s partition is owned by another node, the node forwards the request to the owner over an internal port and returns the owner’s answer. Nodes find each other through the addresses they publish (CONDUCTOR_NODE_ADVERTISE_URL) and sign every forwarded request with a key made from a shared secret (CONDUCTOR_NODE_SECRET), which never travels itself. A node never forwards a forwarded request a second time. Without internal TLS, forwarded traffic is not encrypted: keep the internal port on a network that only the nodes can reach, or turn on mutual TLS (see The internal port).

If a partition changes owner while a request is on its way, the node looks up the new owner and tries once more. If there is no owner at that moment (for example while a lease is being taken over), the request answers 503 with urn:bpm:error:partition-unavailable and a Retry-After header.

A tenant that needs more than its share of a shared partition can be moved to a partition of its own, and back. An operator starts the move; the tenant’s work is paused for a moment while its state moves, and nothing is lost. See Moving a busy tenant.

What happensHow long its partitions are unavailable
The node stops cleanly (a rolling update, a scale-down)A second or two: it finishes the requests it accepted, saves, and releases its leases so another node takes them at once
The node restarts in place with the same owner identityA few seconds: it takes its own partitions back at once
The node crashes or loses its networkAbout 10 to 12 seconds: the other nodes wait until its leases run out (CONDUCTOR_PARTITION_LEASE_MS), then take the partitions over

In every case the new owner loads each tenant’s saved state and carries on. Nothing that was acknowledged is lost or done twice, and a node that has lost a partition can no longer save anything for it, even if it is still running.

Data written by preview builds before 0.1.0 is not carried over. On such a volume the Bundled container refuses to start and says what to do: stop the container, remove its data volume (docker volume rm tinyconductor-pgdata, or whichever volume you mounted at /var/lib/postgresql/data), and start the new image with a fresh volume. On Kubernetes, delete the release’s volume claim (data-<release>-tinyconductor-0) before you install the new chart.

In Cluster mode migrate and the nodes refuse such a database and list what they found, for example:

database check failed: this database holds data from a preview build or an earlier TinyConductor release (schema bpm_jobs, table public.bpm_m6_schema_migrations), and this release cannot read it. Upgrading needs a fresh database: nothing is carried over, so deploy your process models again afterwards, and recreate users and tenants that do not come from configuration. To reset, create a new database and point the database URL at it, or remove the old data from this one: DROP SCHEMA bpm_jobs CASCADE; DROP TABLE public.bpm_m6_schema_migrations;

A database of a preview build names the migrations it recorded instead, and the statements remove every schema the engine created. Create a new database for the cell, or run the statements the message names, then start the nodes again.

Running instances and history are not carried over. Deploy your process models again. Create your tenants again in the console or through the tenant API, and the API clients of your programs in the console or through the clients API. Single sign-on works as before, and what your configuration provides (local users, static tokens, a first-start tenant seed) is applied again.

From 0.1.0 on, the database is upgraded forward. A Bundled engine updates the database layout itself when a newer version starts on the same volume. In Cluster mode, tinyconductor migrate of the new version updates it (the Helm chart runs it as a Job before every upgrade); the nodes only check that it was done. To upgrade:

  1. Back up the database.
  2. Stop the old version. In Cluster mode, stop every node of the cell: rolling upgrades, with old and new nodes running side by side, are not supported yet.
  3. In Cluster mode, run migrate of the new version as the owner login (see Preparing the database).
  4. Start the new version on the same database (in Cluster mode, all nodes of the cell).

An older version cannot be started again on a database that a newer version has upgraded; restore the backup instead.

The engine publishes metrics for each part described here. See Observability for the full list, and the example alerts and dashboard.

QuestionMetric
Does every partition have an owner?tinyconductor_partition_owned
Are leases renewed?tinyconductor_partition_lease_renew_failures_total
Does each node’s clock agree with the database?tinyconductor_clock_offset_seconds
How long does a save take, and how many frames share one?tinyconductor_partition_commit_seconds, tinyconductor_partition_commit_batch_frames
Does a tenant wait for its turn?tinyconductor_tenant_queue_wait_seconds
Which tenant uses the most processing time?tinyconductor_tenant_engine_seconds_total
Is a tenant quarantined?tinyconductor_tenant_quarantined
How far behind are searches and history?tinyconductor_exporter_lag_records, tinyconductor_exporter_lag_seconds, tinyconductor_exporter_lag_frames
Is a tenant being slowed because of that?tinyconductor_exporter_backpressure_total
Has the export of a tenant stopped?tinyconductor_exporter_quarantined