Skip to content

Performance and scaling

This page explains how each deployment mode uses CPU, memory and PostgreSQL, how each one scales, which settings change its capacity, and how to measure it yourself. To size an installation, start with Hardware requirements: the measured footprint of each mode and three starting sizes. TinyConductor is pre-release. The benchmark results below were measured with realistic processes on a single test machine; the page states the machine and the conditions next to the numbers. They show how each mode behaves, not what your hardware will do.

The three modes run the same engine and the same API. They differ only in durability, latency, capacity and tenancy.

EmbeddedBundledCluster
Intended useProcess tests and CIOne team’s production; the quickstartShared, multi-tenant production
Where state livesIn process memory onlyIn memory, in partitions, with every change saved to PostgreSQLThe same, with the partitions shared out between the nodes of a cell
Cost of one acknowledged writeNo storageOne frame in one PostgreSQL commit; frames that are ready together share a commitThe same, plus one hop between nodes when the request reaches a node that does not own the tenant’s partition
Where reads come fromMemoryReads by key from memory; searches from PostgreSQL query tables that the exporter fillsThe same; reads by key from the owning node’s memory
TimeVirtual clockWall clockWall clock
How it scalesOne processOne node, verticallySeveral nodes per cell (one PostgreSQL), and more cells for more tenants

Embedded mode is an in-memory engine with a virtual clock and no storage. It is fast because it never waits for a disk or a network, and its capacity is the memory and CPU of the one process that hosts it. It does not persist anything, so it is not a production mode and there is nothing to scale out. Timers fire only when the test advances the clock. A finished instance leaves the engine’s working state, so requests do not slow down as instances accumulate. The record stream and the history views that answer queries about finished instances still stay in memory, so a long-running Embedded engine slowly uses more memory. For long test runs, reset the engine between suites.

A Bundled deployment is one engine process and one PostgreSQL database in the same container or pod. The engine keeps its state in memory, split into partitions (24 by default). Each partition is one thread that runs the tenants placed on it, one command at a time, and saves each change as one frame before it answers. Frames that are ready together share one commit. The exporter copies saved records into the query tables on its own connection, and no request waits for it. How the engine runs your processes explains each part.

What this means for capacity:

  • Latency is bounded by the commit. An acknowledged create costs one frame in one commit. Storage with a fast, honest fsync matters more than CPU. Batching helps when commits are slower than the engine.
  • One tenant runs on one core; several tenants use several. A tenant’s commands run in its home partition, one at a time, so one busy tenant is bounded by one core. Tenants on other partitions run in parallel. Scale Bundled up with faster cores and faster storage.
  • Memory follows the running instances. A finished instance leaves the engine’s memory; its records live in the query tables. A restart loads each tenant’s latest saved state and replays only the frames saved after it, so restart time does not grow with the history.
  • No horizontal scaling. One engine owns one database. If one team outgrows one node, that is the point to move to Cluster mode.
  • Durability and recovery are PostgreSQL’s: backups, point-in-time recovery and a synchronous replica protect against machine loss. See Bundled mode.

A Cluster deployment is organised in cells. A cell is one PostgreSQL database (CONDUCTOR_CELL_ID names it) and every engine node attached to it. Each tenant lives in exactly one cell; instances, call activities and signal broadcasts never cross cells.

clients, job workers, console
|
load balancer
(any node can serve any request of the cell)
+--------------------+--------------------+
| | |
+---------------+ +---------------+ +---------------+
| engine node 1 |<-->| engine node 2 |<-->| engine node 3 |
| partitions | | partitions | | partitions |
| 0-7 | | 8-15 | | 16-23 |
+-------+-------+ +-------+-------+ +-------+-------+
| forwarding between nodes: internal port |
+--------------------+--------------------+
|
+-------------------------------+
| PostgreSQL primary (cell 1) |
| leases, saved frames, tenant |
| state, query tables |
+-------------------------------+
cell 2 = its own PostgreSQL and its own nodes, serving other tenants

How the nodes of one cell share work (the details are in How the engine runs your processes):

  • Partitions, not instances, are owned. The cell has 24 partitions by default. Each tenant lives in one of them, and each partition is owned by one node at a time, under a lease in PostgreSQL. The nodes share the partitions fairly and hand them over one at a time when a node joins or leaves.
  • State lives in memory. The owning node runs the tenant’s instances in memory and saves each change as one frame before it answers. Frames that are ready together share one commit. Reads of a live instance, job or task by its key come from the owner’s memory; searches and history come from the query tables, which any node reads.
  • Any node answers. A node that does not own the tenant’s partition forwards the request to the owner over the internal port.
  • Fenced leases. Every save names the lease it was made under. A node that has lost a partition can no longer save for it, even if it is still running.
  • Failover. A node that stops cleanly hands its partitions over at once. A node that restarts with the same identity takes its partitions back at once. A node that crashes loses its partitions when their leases run out, after CONDUCTOR_PARTITION_LEASE_MS (10 seconds by default), and the others take them over.

What scales with more nodes, and what does not:

Scales out with nodesShared by the whole cell
HTTP handling, authentication and request validationThe PostgreSQL primary: every commit and every flush
Engine work: each partition runs on its own thread on its ownerThe exporter’s writes into the query tables
Reads by key, answered from memorySearches, answered from the query tables

One tenant runs on one core. A tenant’s commands run in its home partition, one at a time. A tenant that needs more than its share of a shared partition can get a partition of its own (see Moving a busy tenant). Different tenants run in parallel, on the same node or on different nodes.

PostgreSQL is the part every node of a cell shares. Once the primary is saturated, the way to grow is more cells: tenants in a new cell with its own PostgreSQL. Tenants are placed in a cell by configuration; clients address the cell that serves their tenant.

These are the results of the project’s benchmark, measured on 2026-09-27 on a single test machine, with every mode running the partition engine described above. They show how each mode behaves with realistic processes and where each one runs out of capacity. They are not a sizing guide: every number depends on the machine, the storage, the network path to PostgreSQL and the load of the clients. Measure your own processes on your own hardware before you plan capacity (see Measuring it yourself).

HostApple M5 Max (arm64), 18 CPU cores, 128 GiB memory, macOS 26.6.2, on mains power, kept awake during the runs
ContainersDocker Engine 29.1.3 in a Linux virtual machine (Rancher Desktop, Alpine Linux 3.23, kernel 6.6) with 6 virtual CPUs and 9.7 GiB memory. The CPU, disk and network of this virtual machine are measured in Reference environment
TinyConductorRelease build of the engine at the code of 2026-09-27, running natively on macOS (one process per node)
PostgreSQL16 (Alpine image) in a container inside the virtual machine, data on an in-memory filesystem (tmpfs, 3 GiB; 5 GiB for loan approval), fsync on, default settings except max_connections=300. A fresh database for every point. The engine logs in as its own role, which is neither a superuser nor exempt from row-level security
Path from engine to databaseThe engine on macOS reaches PostgreSQL in the virtual machine through the container runtime’s port forwarding, which adds about 0.6 ms to every database round trip
Benchmark clientThe project’s benchmark client (Node.js 26.3.1, one process) on the same host, calling the REST API over HTTP on the loopback interface
Other work on the machineNo builds or other benchmarks ran. The one-minute load average before a point was between 1.6 and 6.0 (the virtual machine’s own services and the previous point); a point that starts above 6 is discarded and repeated. No point was discarded for load or for sleep
Durations10 s warm-up, 30 s measured window, up to 15 s for instances started in the window to finish. Two runs per point; ranges show the spread

tmpfs makes a PostgreSQL flush almost free, which favours the durable modes; the port forwarding makes every database round trip slower than on a Linux host. The single benchmark client on the same host adds noise of its own. Read the results as the shape of each mode’s behaviour, not as its limits.

Each scenario is a BPMN process model, described below. Job workers are simulated by the benchmark client (see Workload settings).

Straight-through (perf-straight-through), three service tasks in a row:

start → step 1 (job) → step 2 (job) → step 3 (job) → end

Order fulfilment (perf-order-fulfilment), about twelve steps, with an order of 5–20 KB (JSON with 20–80 line items, customer and address) as variables:

start → validate order (job)
→ order type? ── express (30 %) → prioritise order (job) ─┐
└─ standard ────────────────────────────────┴→ shipping class (DMN decision table)
→ ┌ reserve stock (job) ┐
└ charge payment (job)┘ (in parallel)
→ wait for the payment confirmation (message, correlated by order id)
└ after 5 s without it (2 % of orders): escalate payment (job) → end
→ ship (job) → notify customer (send task, job) → end

The benchmark’s payment provider publishes the confirmation 5–50 ms after the “charge payment” job completes; for 2 % of the orders it never does, so the timer on the wait fires and the order escalates.

Loan approval (perf-loan-approval, with the called process perf-loan-document-check), about fourteen steps:

start → credit score (job, with input and output mappings)
→ risk class (DMN decision table)
→ review by 3 reviewers (multi-instance service task, in parallel)
→ document check (call activity: start → verify documents (job) → end)
→ clerk approval (user task, completed by a simulated clerk)
→ extra checks (inclusive gateway): fraud check (job) if the amount is over 50,000,
collateral check (job) if the risk is high,
otherwise neither
→ disburse (job) → end

In 1 % of the loans the credit-score job fails with no retries left, which raises an incident; the benchmark finds it through the incident search, gives the job a retry, and resolves the incident. The simulated clerk polls the user-task search every 200 ms and completes each task through the user-task API.

Tiny shapes (jobs and messages, kept for comparison with earlier runs): start → service task → end, and start → message catch → service task → end, with the client itself creating, publishing, activating and completing, and no simulated work.

SettingValue
Load modelClosed loop: N clients each start one instance and wait until it has finished, then start the next (reported as “clients”). Open loop for the latency points: starts at a fixed rate whatever the engine does
Clients128 and 512 (tiny shapes: 32, 128 and 512)
Job workersPer job type: 4 worker processes, each with up to 128 jobs in flight, activating with a 1 s long poll (like the standard job-worker clients)
Simulated work per job10–50 ms, uniformly distributed
Payment confirmationPublished 5–50 ms after the charge job; 2 % never published (timer escalation after 5 s)
Clerk2 per tenant, polling every 200 ms, 10–50 ms per task
Incidents1 % of loans, resolved by the benchmark
PayloadOrders 5–20 KB; loans about 1 KB; straight-through about 100 bytes
Instance dataEvery instance’s data and its path (express or not, escalation, incident) is drawn from a fixed seed and the instance’s number, so every run sees the same instances
TenantsOne tenant. With two Cluster nodes, the benchmark sends every request to both nodes in turn; the node that does not own the tenant’s partition forwards it to the one that does
  • Throughput (instances/s): instances that finished inside the 30 s window, divided by 30. An instance has finished when the completion of its last job (“notify”, “escalate payment”, “disburse”, “step 3”) has been acknowledged.
  • End-to-end duration: from sending the create request to the acknowledgement of that last job’s completion, for instances started inside the window. It includes the simulated work (for an order about 160 ms of job work and payment delay on the critical path), the waits for workers and the 5 s timer of escalated orders.
  • p50, p95, p99: the value below which 50 %, 95 % and 99 % of the measurements fall (nearest rank).
  • Create: create request to acknowledgement.
  • Job pickup: from the acknowledgement of the instance’s previous step to the moment a worker receives the next job. It includes the engine’s work between the two (gateways, decisions, mappings) and any wait for a free worker.
  • Engine CPU per instance: CPU time used by the engine processes during the run, divided by the instances that finished (including warm-up and the finishing period). With two nodes it is the sum of both. Database CPU per instance: the same for the PostgreSQL container.
  • Database log per instance: PostgreSQL write-ahead log bytes written during the run, divided by the finished instances.
  • Engine memory: the highest resident memory of the engine process(es), sampled every 2 seconds.

With a closed loop, the clients wait for their instances; once the engine is saturated, more clients only make each instance wait longer. So the highest rate is the saturation throughput, and the durations at 512 clients are mostly queueing.

Tiny shapes (jobs and messages):

ModeClientsInstances/sEnd-to-end p50 / p95 / p99 (ms)Create p50 / p95 (ms)Engine CPU per instanceDatabase CPU per instanceDatabase log per instance
Embedded325,186–5,2316 / 8 / 101.6 / 3.00.35 ms——
Embedded1285,233–5,30123 / 31–32 / 35–376.4 / 9.0–9.30.38 ms——
Embedded5124,897–4,898100 / 131 / 138–13929.9–30.0 / 36.30.41 ms——
Bundled321,903–1,90416 / 22 / 263.8 / 6.10.56 ms0.70 ms29 KiB
Bundled1283,139–3,20539–40 / 51–52 / 56–599.2–9.4 / 13.6–13.90.54–0.55 ms0.59–0.60 ms17 KiB
Bundled5123,577–3,596136–137 / 182 / 208–21235.8–36.0 / 48.00.54 ms0.53 ms15–16 KiB
Cluster, 1 node321,899–1,90216 / 22–23 / 263.8 / 6.1–6.20.55 ms0.71 ms30 KiB
Cluster, 1 node1283,125–3,17039–40 / 52 / 58–599.3–9.5 / 13.8–13.90.54 ms0.60–0.61 ms17–18 KiB
Cluster, 1 node5123,583–3,612136–137 / 182 / 20235.9–36.0 / 47.60.53–0.54 ms0.53 ms16 KiB
Cluster, 2 nodes321,759–1,76618 / 24 / 273.9 / 6.2–6.31.1 ms0.73–0.74 ms31–32 KiB
Cluster, 2 nodes1282,897–2,90843 / 56 / 639.5–9.6 / 14.81.2 ms0.65 ms18 KiB
Cluster, 2 nodes5123,335–3,354147–149 / 195–196 / 215–22137.5–37.8 / 49.7–50.11.2 ms0.56 ms16 KiB

Straight-through (three jobs):

ModeClientsInstances/sEnd-to-end p50 / p95 / p99 (ms)Create p50 / p95 (ms)Engine CPU per instanceDatabase CPU per instanceDatabase log per instance
Embedded1281,308–1,30998 / 131 / 142–1431.5 / 2.9–3.00.83 ms——
Embedded5122,998–3,028167–169 / 229–232 / 247–2505.3–5.4 / 9.9–10.20.62–0.63 ms——
Bundled128959–963132–133 / 171 / 187–1883.4–3.5 / 6.11.1 ms1.4 ms58–60 KiB
Bundled5121,378–1,396342–349 / 608–610 / 715–7169.0–10.0 / 25.8–28.00.96–0.97 ms1.1 ms45–46 KiB
Cluster, 1 node128957–962132–133 / 171–172 / 187–1893.5 / 6.11.0–1.1 ms1.4 ms62–63 KiB
Cluster, 1 node5121,385–1,401344 / 594–622 / 707–7279.5–10.0 / 27.7–29.10.96 ms1.1 ms47 KiB
Cluster, 2 nodes128896142 / 185 / 204–2073.7 / 6.3–6.42.2 ms1.5 ms61–62 KiB
Cluster, 2 nodes5121,148–1,153411–413 / 772–795 / 966–9796.7–6.8 / 20.2–21.22.2 ms1.3 ms51–52 KiB

At 128 clients Embedded is limited by the clients, not the engine: each instance spends about 90 ms in simulated job work, so 128 clients cannot start more than about 1,300 instances per second.

Order fulfilment:

ModeClientsInstances/sEnd-to-end p50 / p95 / p99 (ms)Create p50 / p95 (ms)Engine CPU per instanceDatabase CPU per instanceDatabase log per instanceEngine memory (peak)
Embedded128367346–347 / 452–453 / 504–5099.3 / 20.2–20.74.1 ms——2.4 GiB
Embedded512359–3711,344–1,391 / 1,728–1,791 / 1,850–1,89988.6–91.3 / 152.1–153.74.0–4.2 ms——3.3–3.6 GiB
Bundled128267–269374–381 / 515–520 / 5,253–5,25410.5–11.0 / 26.3–28.56.0–6.1 ms4.4–4.5 ms204 KiB803–998 MiB
Bundled512292–2971,572–1,606 / 2,568–2,739 / 5,930–5,99787.6–94.0 / 183.5–195.25.7 ms3.7–3.8 ms194 KiB1.1–1.4 GiB
Cluster, 1 node128267–268378–379 / 514–517 / 5,242–5,25110.7–11.1 / 27.5–27.86.0 ms4.5–4.6 ms212 KiB1.0–1.1 GiB
Cluster, 1 node512288–2941,592–1,613 / 2,842–2,864 / 5,975–6,10380.5–87.5 / 194.1–200.75.6 ms3.8 ms203–204 KiB1.2–1.3 GiB
Cluster, 2 nodes128265–267378–382 / 526–530 / 5,246–5,2499.9–10.5 / 26.2–27.68.9 ms4.7 ms212 KiB909–1,015 MiB (2 nodes)
Cluster, 2 nodes512284–2861,510–1,536 / 3,399–3,551 / 5,884–5,91232.7–34.5 / 114.0–130.58.6 ms4.0–4.1 ms206 KiB1.5–1.6 GiB (2 nodes)

The p99 of about 5 s in Bundled and Cluster is the 2 % of orders that wait for the payment timer. In Embedded the clock is virtual, so no order escalates there. The engine memory is the orders in flight with their 5–20 KB documents, plus the records waiting for the query tables; a finished order leaves the engine’s memory.

Loan approval:

ModeClientsInstances/sEnd-to-end p50 / p95 / p99 (ms)Create p50 / p95 (ms)Engine CPU per instanceDatabase CPU per instanceDatabase log per instanceIncidents raised / resolved within the run
Embedded128553–557228–230 / 289–295 / 329–3421.9 / 3.7–3.93.0–3.1 ms——229–230 / 229–230
Embedded512683–687746–749 / 912–918 / 972–99020.7–21.9 / 48.8–49.42.8 ms——276–277 / 276–277
Bundled128329–331297–298 / 366–367 / 403–4084.4–4.5 / 8.84.2 ms4.7 ms232–235 KiB130 / 112–114
Bundled512413–4181,056–1,061 / 1,630–1,770 / 1,934–2,13119.5–19.8 / 63.4–64.73.8–3.9 ms4.1–4.2 ms202–204 KiB171–174 / 108
Cluster, 1 node128328–330297 / 365–367 / 406–4134.4 / 8.8–8.94.1–4.2 ms4.7–4.8 ms241–245 KiB130 / 110–112
Cluster, 1 node512414–4191,056–1,060 / 1,622–1,704 / 2,001–2,26620.4–20.6 / 67.8–68.13.8 ms4.2 ms210–211 KiB173–174 / 108–109
Cluster, 2 nodes128278–317307–351 / 377–461 / 422–5554.8–6.2 / 9.3–12.67.1–8.0 ms5.1–5.8 ms242–244 KiB108–125 / 100–108
Cluster, 2 nodes512362–3891,151–1,238 / 1,772–2,029 / 2,172–2,56016.8–18.4 / 52.9–56.96.8–7.2 ms4.6–4.9 ms217–222 KiB146–156 / 104–106

No request failed. In Embedded every raised incident was found and resolved and every loan had finished before the end of the run. In Bundled and Cluster the benchmark finds an incident through the incident search, which reads the query tables; at these loan rates the query tables were 7–38 seconds behind, so some incidents were found only after the run’s 15 s finishing period and their loans were still waiting. Run again with 60 s to finish, every incident was resolved and every loan finished. See Known limits.

Order fulfilment with starts at a fixed rate (40 s window, every start on time, no errors). The rates are those of the earlier measurements: about half and about 80 % of Bundled’s highest closed-loop rate, and the rates that were half and 80 % of Cluster’s earlier maximum:

ModeRate (instances/s)Create p50 / p95 (ms)Job pickup p50 / p95 (ms)Job completion p50 / p95 (ms)Message publish p50 / p95 (ms)End-to-end p50 / p95 (ms), orders not escalated
Bundled1353.7 / 5.4–5.52.1–2.9 / 4.6–6.44.1–4.6 / 6.8–7.33.0 / 4.8191–192 / 249
Bundled2154.6–5.4 / 10.6–33.44.1–7.9 / 14.0–1666.5–8.5 / 16.6–71.54.1–5.0 / 10.8–29.9225–247 / 294–1,031
Cluster, 2 nodes8.56.9–7.0 / 9.7–9.91.8–2.4 / 3.1–4.94.0–6.3 / 4.8–8.34.0–4.1 / 5.6–6.1199–200 / 254
Cluster, 2 nodes145.8 / 8.3–8.51.5–2.0 / 3.0–4.53.3–5.4 / 4.8–8.43.5 / 5.2–5.4192 / 244–247
Cluster, 2 nodes1354.0–4.1 / 5.7–7.62.4–3.5 / 4.8–11.24.6–5.0 / 7.3–11.33.2–3.3 / 5.0–6.9195–200 / 255–272
Cluster, 2 nodes2155.0–7.0 / 11.2–53.94.5–13.5 / 15.0–4717.3–11.7 / 16.8–1024.4–6.3 / 11.4–52.4230–280 / 299–2,247

Job pickup and completion show the range over the order’s job types (the escalation job, which waits for the 5 s timer, is left out). About 160 ms of every order’s end-to-end duration is the simulated work and payment delay, so the engine adds about 30–70 ms to an order at these rates, in Bundled and in Cluster alike.

At 215 orders per second, one of the two runs in each mode had a pause of a few seconds in which creates took up to 96 ms (Bundled) and 233 ms (Cluster); the high p95 values at that rate come from those pauses. The other run in each mode stayed at about 225–230 / 295–300 ms end to end.

With one tenant, a second node does not add capacity: the tenant’s commands run in its one partition on one node, and the other node forwards the requests it receives to that node. Forwarding costs about 3–10 % of throughput and, per instance, about as much engine CPU again on the forwarding node (the tables above). More nodes add capacity when there are several tenants, spread over the partitions of the cell. Several cells and a database failover were not measured again in this round.

One Cluster node with 3,721 tenants, measured twice:

Until ready
New database (every tenant is placed and its configuration saved)4.4–4.5 s
Restart on the same database13.4–14.3 s

A restart takes about three times as long as a start on a new database: the node reads each tenant’s remembered idempotency keys back one tenant at a time, even for tenants that have none. See Known limits.

Connections each node holds to PostgreSQL, with 24 partitions and one tenant, idle and at the highest count seen while 512 clients ran the tiny shapes against all nodes:

SetupIdleUnder load
Bundled3031
Cluster, 1 node3031
Cluster, 2 nodes18 and 1718 and 19

One tenant with about 70,000 finished instances (the tiny job shape), and one with about 136,000 (Cluster, one node), 20 requests one at a time after 3 warm-ups; Bundled and Cluster (one node) gave the same times:

Requestp50 (ms), 70,000p95 (ms), 70,000p50 / p95 (ms), 136,000
Process-instance search, first page of 50, no sort9.0–9.310.8–10.911.1 / 12.3
Process-instance search sorted by start date, first page of 509.0–9.510.8–10.911.1 / 12.2
Process-instance search, page of 50 starting in the middle of the matches17.1–17.526.2–27.121.1 / 21.8
Decision-definition search (2 decisions deployed)2.73.82.7 / 3.8
Variable search for one instance3.03.8–3.93.0 / 4.1
Variable read by its key2.52.8–2.92.5 / 3.2

Sorting and paging run in the database, so a sorted search or a page far from the start costs about as much as the first page and grows only a little with the number of matches.

The query tables are filled after the engine has answered, and at the highest rates they fall behind: after a 20-second burst at the tiny shapes’ highest rate, they needed another 35 seconds to hold every finished instance, and 72 seconds after a 40-second burst.

Jobs and messages, 64 clients, then a clean stop and a restart after each step. Without idempotency keys, 45 s of load per step:

Instances so farInstances/s during the stepCreate p50 / p99 (ms)Engine memory before the stopRestart until readyEngine memory after the restartSaved tenant stateDatabase size
118,2562,6265.5 / 10.3102 MiB0.66 s102 MiB235 KB1.9 GB
227,2722,4215.8 / 11.2111 MiB0.96 s100 MiB268 KB3.9 GB

With an idempotency key on every create, activation and completion, 15 s of load per step:

Instances so farInstances/s during the stepCreate p50 / p99 (ms)Engine memory before the stopRestart until readyEngine memory after the restartSaved tenant stateRemembered keys
37,2792,4815.5 / 15.5146 MiB0.62 s119 MiB251 KB83,519 (32 MB)
72,4952,3435.8 / 16.3149 MiB0.60 s143 MiB262 KB162,720 (62 MB)
106,4462,2596.0 / 17.9208 MiB0.61 s184 MiB287 KB239,293 (90 MB)
140,0472,2366.0 / 17.8214 MiB0.68 s188 MiB278 KB314,930 (119 MB)

Memory and restart time no longer grow with the history: the engine keeps only the running instances, and a restart loads the tenant’s latest saved state and replays the few frames after it. Idempotency keys are stored in their own table, not in the saved state; the engine keeps the keys of the last 10 minutes in memory (CONDUCTOR_CLAIM_WINDOW) and a compact filter for older ones, so memory grows during the first 10 minutes of heavy keyed load and should then stop growing. In a continuous 10-minute run with a key on every request (about 4,200 keyed requests per second, 1.1 million instances) the engine reached 540 MiB, still growing when the run ended at the 10-minute mark, and 750 MiB after a restart; the saved tenant state stayed at 190–320 KB, and no request failed. The database grows by about 17 KB per small instance, most of it the record stream that the query tables and the history are built from; a third step without keys did not fit the benchmark’s 5 GiB in-memory database. See Known limits.

Exporting traces, metrics and logs of every request to a local collector made no measurable difference to throughput or CPU per instance; see Overhead of tracing. That was measured on the earlier engine and has not been repeated.

  • Embedded runs about 4,900–5,300 small instances, 3,000 straight-through, 360–370 orders and 550–690 loans per second. Its limit is the engine’s one partition, which here also keeps every finished record in memory for the query side (1.8–6.5 GiB after these runs).
  • Bundled runs up to about 3,600 small instances, 1,400 straight-through, 270–300 orders and 330–420 loans per second, with create latency of a few milliseconds below saturation. Its limit is the tenant’s one partition (one core). Memory stays at the instances in flight. At these rates the query tables that searches read fall behind (see Known limits).
  • Cluster runs the same as Bundled on one node, because it is the same engine: about 3,600 small instances, 1,400 straight-through, 270–295 orders and 330–420 loans per second for one tenant. PostgreSQL is lightly loaded (about 4–5 ms of database CPU and 200–250 KiB of write-ahead log per order or loan). One tenant does not get faster with more nodes; more tenants do.

The benchmark harness that produced these results is part of the engineering test suite and is not shipped. To measure your own processes on your own hardware, see Measuring it yourself.

Most of the database is the finished instances, which are kept 30 days by default (per tenant; see How long history is kept) and then removed completely. A small instance (a few elements, small variables) takes about 20 KB: 16 KB of records and 4 KB of searchable history. Instances with more steps and larger variables take several times more; measure yours by dividing the growth of the database by the instances started.

Instances per dayKept 1 dayKept 7 daysKept 30 days
100,000about 2 GBabout 14 GBabout 60 GB
1,000,000about 20 GBabout 140 GBabout 600 GB

The figures are for small instances; multiply by your own size per instance. Add room for the change log, the saved tenant state, PostgreSQL’s own write-ahead log and one more day of history (the history is removed a whole day at a time). At a steady load the database stays about this size once the first retention period has passed. See Retention and disk.

  • Connections. A node holds one connection for the log of each partition it owns, one for its snapshot writer, one more while it maintains the reporting tables, and its pool (16 by default, shared by requests, searches, identity and the exporter). A Bundled engine, or a Cluster node that owns all 24 partitions, therefore opens at most 42. In a cell the log connections are shared out between the nodes, so the cell needs at most partitions + nodes × (pool size + 2): 24 + 18 × nodes with the defaults, all direct or through a pooler in session mode. Idle, a node holds about one connection per partition it owns plus 5: 29 on one node with 24 partitions, 17 on each of two nodes. With one busy tenant the counts barely rise (30, and 17–18 on each of two nodes; see Database connections); many busy tenants use more of the pools. In Cluster mode set max_connections to cover them plus your tools, backups and monitoring: PostgreSQL’s default of 100 is tight above two or three nodes. The Bundled image sets max_connections to 100; if you raise CONDUCTOR_PARTITIONS, raise it too (CONDUCTOR_POSTGRES_OPTIONS="-c max_connections=200").
  • Storage. Every acknowledged write waits for a PostgreSQL commit, so commit latency (the fsync of the WAL device) is the floor of every write’s latency in both durable modes. Use low-latency storage with honest flushes; do not turn fsync or synchronous_commit off.
  • Network. In Cluster mode every save is one round trip to PostgreSQL, and a forwarded request adds one round trip between two nodes. Keep the nodes and the database in the same zone.
  • CPU and memory. The exporter’s writes, index maintenance and vacuum run on the primary. Engine memory follows the instances that are still running.

Per-tenant limits protect the cell from one tenant’s burst. They are exact, because each tenant’s work runs in one partition; see Tenant quotas. A tenant’s weight sets its share of its partition’s processing time when several tenants are busy at once.

The engine never stops to shed load; it answers with a typed error that says what to do:

AnswerWhyWhat the client does
429 urn:bpm:error:quota-exhaustedThe tenant reached one of its limitsWait the Retry-After delay and retry
503 urn:bpm:error:backpressureThe tenant’s searches are too far behindWait the Retry-After delay and retry
503 urn:bpm:error:partition-unavailableThe tenant’s partition has no owner at the moment (a failover or a move)Wait the Retry-After delay and retry, with the same idempotency key

In each case nothing was changed, except that a partition-unavailable answer after a lost connection may come from a saved command: retrying with the same idempotency key is always safe. Add some jitter to the retries. A steady stream of these answers means a tenant needs higher limits or a partition of its own, or the cell needs a larger database.

The engine is built for human and agent cadence: many active instances, each moving every few seconds to days. It is not built for payment-grade volume in a single cell. Plan for more cells, not bigger ones, when the volume grows.

Every setting below is read at startup. They apply to Bundled and Cluster mode.

VariableDefaultEffect on performance
CONDUCTOR_PARTITIONS24How many partitions share the work. Each is one thread on its owner; tenants on different partitions run in parallel. Fixed when the database is created
CONDUCTOR_COMMIT_BATCH256The most frames saved together in one commit
CONDUCTOR_SNAPSHOT_INTERVAL5000Milliseconds between two saved copies of a tenant’s state. Shorter means faster restarts and earlier removal of old frames, at the price of more writes
CONDUCTOR_EXPORTER_BATCH256Frames the exporter reads per pass
CONDUCTOR_EXPORTER_MAX_LAG, CONDUCTOR_EXPORTER_MAX_LAG_AGE_MS1000000, 60000How far a tenant’s searches may fall behind before its new instances are refused; they are slowed from half of it
CONDUCTOR_PARTITION_LEASE_MS10000How long a crashed node’s partitions stay unavailable before another node takes them over. Shorter means faster failover, at the price of more renewals and a higher risk that a slow node loses its partitions
CONDUCTOR_CLAIM_WINDOW10mHow long an idempotency key is kept in memory for an exact check; older keys are checked with a compact filter and, on a match, one database read. Shorter means less memory under heavy keyed load, at the price of more database reads for late retries
CONDUCTOR_DEDUP_RETENTION24hHow long an idempotency key is remembered at all. The key table in the database holds every key for this long

GET /metrics on the metrics port (9090; see Scraping /metrics) exposes the measures that matter for capacity, among them:

  • tinyconductor_command_seconds: from request to acknowledgement, per kind of command;
  • tinyconductor_partition_commit_seconds and tinyconductor_partition_commit_batch_frames: time and size of each save;
  • tinyconductor_tenant_queue_wait_seconds: how long a tenant waits for its turn;
  • tinyconductor_tenant_engine_seconds_total: processing time per tenant, to find a tenant that needs a partition of its own;
  • tinyconductor_exporter_lag_seconds: how far searches are behind;
  • bpm_engine_jobs_available and bpm_engine_jobs_activated: the job backlog.

See Observability.

The figures on this page describe the reference environment. To learn what your installation does, measure it with your own processes:

  1. Compare your hardware with the reference environment: the CPU, disk, PostgreSQL commit latency and network figures in Comparing with your hardware show which of the numbers above you can expect to reach, and which will be lower or higher.
  2. Run the engine from the released image or download, in the mode and topology you plan to use, with PostgreSQL on the storage you plan to use.
  3. Deploy your own process models, and drive them with your own job workers or a load tool of your choice through the REST API: start instances at a fixed rate (for latency) or with a fixed number of concurrent clients (for the highest rate).
  4. Watch the measures listed under What to watch, and the CPU, memory and disk figures listed under Measuring your own installation.
  5. Record the hardware, the storage and the machine load with every result. In-memory storage for PostgreSQL makes the numbers look better, and a busy machine makes them look worse; neither transfers to production.

To find the highest rate your setup sustains, raise the number of concurrent clients (for example 16, 64, 256) until the rate stops rising; then measure latency at about half and at 80 % of that rate.

  • All measurements come from a single test machine, with PostgreSQL on an in-memory filesystem behind the container runtime’s port forwarding; none has been taken on production-like hardware.
  • Idempotency keys cost memory and database space. The engine remembers each key for a day (CONDUCTOR_DEDUP_RETENTION): the keys of the last 10 minutes in memory (CONDUCTOR_CLAIM_WINDOW), older ones in a compact filter, and every key in a database table. With a key on every create, activation and completion at about 4,200 requests per second, the engine used about 540 MiB after 10 minutes (about 110 MiB without keys), the key table grew by about 380 bytes per key (1 GB after 10 minutes), and throughput was about 10 % lower than without keys. A shorter CONDUCTOR_CLAIM_WINDOW lowers the memory.
  • Searches can lag behind the engine. The exporter that fills the query tables keeps up at full load on the test machine (a backlog of about 0.2–0.3 seconds), but on slower storage or a busy database it can fall behind; searches, the incident and task lists and the history then lag by that much. Reads by key of a live instance, job or task are always current. Watch tinyconductor_exporter_lag_seconds; when a tenant’s backlog reaches half of CONDUCTOR_EXPORTER_MAX_LAG or CONDUCTOR_EXPORTER_MAX_LAG_AGE_MS, its new instances are slowed, and at the limit refused.
  • Throughput drifts down under sustained full load: by 12–22 % over 5–10 minutes in these runs, with or without idempotency keys, while the change log and the database grow.
  • History takes most of the database. About 20 KB per small instance, more for larger variables, kept 30 days by default; see History retention.
  • A Cluster node with thousands of tenants restarts more slowly than it first starts: 13–14 seconds with 3,721 tenants, against 4.5 seconds on a new database, because it reads each tenant’s remembered idempotency keys back one tenant at a time.
  • One tenant runs on one core: its home partition. A tenant that needs more than one partition can hold is not supported yet.
  • Embedded keeps every record of finished instances in memory, because tests read them: after the benchmark’s runs that was 1.8–6.5 GiB.
  • The Kubernetes chart offers no autoscaler.
  • PostgreSQL is a single point of contention per cell; more cells mean more databases, and tenants are placed in cells by configuration. Tenants do not move between cells.
  • Bundled mode runs on one node; it cannot be scaled out.
  • Principals with resource-level grants (rather than an action on every resource) cost an extra lookup on instance and user-task requests; the effect has not been measured.