Performance and scaling
This page explains how each deployment mode uses CPU, memory and PostgreSQL, how each one scales, which settings change its capacity, and how to measure it yourself. To size an installation, start with Hardware requirements: the measured footprint of each mode and three starting sizes. TinyConductor is pre-release. The benchmark results below were measured with realistic processes on a single test machine; the page states the machine and the conditions next to the numbers. They show how each mode behaves, not what your hardware will do.
What each mode is designed for
Section titled “What each mode is designed for”The three modes run the same engine and the same API. They differ only in durability, latency, capacity and tenancy.
| Embedded | Bundled | Cluster | |
|---|---|---|---|
| Intended use | Process tests and CI | One team’s production; the quickstart | Shared, multi-tenant production |
| Where state lives | In process memory only | In memory, in partitions, with every change saved to PostgreSQL | The same, with the partitions shared out between the nodes of a cell |
| Cost of one acknowledged write | No storage | One frame in one PostgreSQL commit; frames that are ready together share a commit | The same, plus one hop between nodes when the request reaches a node that does not own the tenant’s partition |
| Where reads come from | Memory | Reads by key from memory; searches from PostgreSQL query tables that the exporter fills | The same; reads by key from the owning node’s memory |
| Time | Virtual clock | Wall clock | Wall clock |
| How it scales | One process | One node, vertically | Several nodes per cell (one PostgreSQL), and more cells for more tenants |
How each mode scales
Section titled “How each mode scales”Embedded: one process
Section titled “Embedded: one process”Embedded mode is an in-memory engine with a virtual clock and no storage. It is fast because it never waits for a disk or a network, and its capacity is the memory and CPU of the one process that hosts it. It does not persist anything, so it is not a production mode and there is nothing to scale out. Timers fire only when the test advances the clock. A finished instance leaves the engine’s working state, so requests do not slow down as instances accumulate. The record stream and the history views that answer queries about finished instances still stay in memory, so a long-running Embedded engine slowly uses more memory. For long test runs, reset the engine between suites.
Bundled: one node, vertical scaling
Section titled “Bundled: one node, vertical scaling”A Bundled deployment is one engine process and one PostgreSQL database in the same container or pod. The engine keeps its state in memory, split into partitions (24 by default). Each partition is one thread that runs the tenants placed on it, one command at a time, and saves each change as one frame before it answers. Frames that are ready together share one commit. The exporter copies saved records into the query tables on its own connection, and no request waits for it. How the engine runs your processes explains each part.
What this means for capacity:
- Latency is bounded by the commit. An acknowledged create costs one
frame in one commit. Storage with a fast, honest
fsyncmatters more than CPU. Batching helps when commits are slower than the engine. - One tenant runs on one core; several tenants use several. A tenant’s commands run in its home partition, one at a time, so one busy tenant is bounded by one core. Tenants on other partitions run in parallel. Scale Bundled up with faster cores and faster storage.
- Memory follows the running instances. A finished instance leaves the engine’s memory; its records live in the query tables. A restart loads each tenant’s latest saved state and replays only the frames saved after it, so restart time does not grow with the history.
- No horizontal scaling. One engine owns one database. If one team outgrows one node, that is the point to move to Cluster mode.
- Durability and recovery are PostgreSQL’s: backups, point-in-time recovery and a synchronous replica protect against machine loss. See Bundled mode.
Cluster: cells, nodes and partitions
Section titled “Cluster: cells, nodes and partitions”A Cluster deployment is organised in cells. A cell is one PostgreSQL
database (CONDUCTOR_CELL_ID names it) and every engine node attached to it. Each
tenant lives in exactly one cell; instances, call activities and signal
broadcasts never cross cells.
clients, job workers, console | load balancer (any node can serve any request of the cell) +--------------------+--------------------+ | | | +---------------+ +---------------+ +---------------+ | engine node 1 |<-->| engine node 2 |<-->| engine node 3 | | partitions | | partitions | | partitions | | 0-7 | | 8-15 | | 16-23 | +-------+-------+ +-------+-------+ +-------+-------+ | forwarding between nodes: internal port | +--------------------+--------------------+ | +-------------------------------+ | PostgreSQL primary (cell 1) | | leases, saved frames, tenant | | state, query tables | +-------------------------------+
cell 2 = its own PostgreSQL and its own nodes, serving other tenantsHow the nodes of one cell share work (the details are in How the engine runs your processes):
- Partitions, not instances, are owned. The cell has 24 partitions by default. Each tenant lives in one of them, and each partition is owned by one node at a time, under a lease in PostgreSQL. The nodes share the partitions fairly and hand them over one at a time when a node joins or leaves.
- State lives in memory. The owning node runs the tenant’s instances in memory and saves each change as one frame before it answers. Frames that are ready together share one commit. Reads of a live instance, job or task by its key come from the owner’s memory; searches and history come from the query tables, which any node reads.
- Any node answers. A node that does not own the tenant’s partition forwards the request to the owner over the internal port.
- Fenced leases. Every save names the lease it was made under. A node that has lost a partition can no longer save for it, even if it is still running.
- Failover. A node that stops cleanly hands its partitions over at once.
A node that restarts with the same identity takes its partitions back at
once. A node that crashes loses its partitions when their leases run out,
after
CONDUCTOR_PARTITION_LEASE_MS(10 seconds by default), and the others take them over.
What scales with more nodes, and what does not:
| Scales out with nodes | Shared by the whole cell |
|---|---|
| HTTP handling, authentication and request validation | The PostgreSQL primary: every commit and every flush |
| Engine work: each partition runs on its own thread on its owner | The exporter’s writes into the query tables |
| Reads by key, answered from memory | Searches, answered from the query tables |
One tenant runs on one core. A tenant’s commands run in its home partition, one at a time. A tenant that needs more than its share of a shared partition can get a partition of its own (see Moving a busy tenant). Different tenants run in parallel, on the same node or on different nodes.
PostgreSQL is the part every node of a cell shares. Once the primary is saturated, the way to grow is more cells: tenants in a new cell with its own PostgreSQL. Tenants are placed in a cell by configuration; clients address the cell that serves their tenant.
Benchmark results
Section titled “Benchmark results”These are the results of the project’s benchmark, measured on 2026-09-27 on a single test machine, with every mode running the partition engine described above. They show how each mode behaves with realistic processes and where each one runs out of capacity. They are not a sizing guide: every number depends on the machine, the storage, the network path to PostgreSQL and the load of the clients. Measure your own processes on your own hardware before you plan capacity (see Measuring it yourself).
Test machine and conditions
Section titled “Test machine and conditions”| Host | Apple M5 Max (arm64), 18 CPU cores, 128 GiB memory, macOS 26.6.2, on mains power, kept awake during the runs |
| Containers | Docker Engine 29.1.3 in a Linux virtual machine (Rancher Desktop, Alpine Linux 3.23, kernel 6.6) with 6 virtual CPUs and 9.7 GiB memory. The CPU, disk and network of this virtual machine are measured in Reference environment |
| TinyConductor | Release build of the engine at the code of 2026-09-27, running natively on macOS (one process per node) |
| PostgreSQL | 16 (Alpine image) in a container inside the virtual machine, data on an in-memory filesystem (tmpfs, 3 GiB; 5 GiB for loan approval), fsync on, default settings except max_connections=300. A fresh database for every point. The engine logs in as its own role, which is neither a superuser nor exempt from row-level security |
| Path from engine to database | The engine on macOS reaches PostgreSQL in the virtual machine through the container runtime’s port forwarding, which adds about 0.6 ms to every database round trip |
| Benchmark client | The project’s benchmark client (Node.js 26.3.1, one process) on the same host, calling the REST API over HTTP on the loopback interface |
| Other work on the machine | No builds or other benchmarks ran. The one-minute load average before a point was between 1.6 and 6.0 (the virtual machine’s own services and the previous point); a point that starts above 6 is discarded and repeated. No point was discarded for load or for sleep |
| Durations | 10 s warm-up, 30 s measured window, up to 15 s for instances started in the window to finish. Two runs per point; ranges show the spread |
tmpfs makes a PostgreSQL flush almost free, which favours the durable modes; the port forwarding makes every database round trip slower than on a Linux host. The single benchmark client on the same host adds noise of its own. Read the results as the shape of each mode’s behaviour, not as its limits.
The scenarios
Section titled “The scenarios”Each scenario is a BPMN process model, described below. Job workers are simulated by the benchmark client (see Workload settings).
Straight-through (perf-straight-through), three service tasks in a row:
start → step 1 (job) → step 2 (job) → step 3 (job) → endOrder fulfilment (perf-order-fulfilment), about twelve steps, with an
order of 5–20 KB (JSON with 20–80 line items, customer and address) as
variables:
start → validate order (job) → order type? ── express (30 %) → prioritise order (job) ─┐ └─ standard ────────────────────────────────┴→ shipping class (DMN decision table) → ┌ reserve stock (job) ┐ └ charge payment (job)┘ (in parallel) → wait for the payment confirmation (message, correlated by order id) └ after 5 s without it (2 % of orders): escalate payment (job) → end → ship (job) → notify customer (send task, job) → endThe benchmark’s payment provider publishes the confirmation 5–50 ms after the “charge payment” job completes; for 2 % of the orders it never does, so the timer on the wait fires and the order escalates.
Loan approval (perf-loan-approval, with the called process
perf-loan-document-check), about fourteen steps:
start → credit score (job, with input and output mappings) → risk class (DMN decision table) → review by 3 reviewers (multi-instance service task, in parallel) → document check (call activity: start → verify documents (job) → end) → clerk approval (user task, completed by a simulated clerk) → extra checks (inclusive gateway): fraud check (job) if the amount is over 50,000, collateral check (job) if the risk is high, otherwise neither → disburse (job) → endIn 1 % of the loans the credit-score job fails with no retries left, which raises an incident; the benchmark finds it through the incident search, gives the job a retry, and resolves the incident. The simulated clerk polls the user-task search every 200 ms and completes each task through the user-task API.
Tiny shapes (jobs and messages, kept for comparison with earlier runs): start → service task → end, and start → message catch → service task → end, with the client itself creating, publishing, activating and completing, and no simulated work.
Workload settings
Section titled “Workload settings”| Setting | Value |
|---|---|
| Load model | Closed loop: N clients each start one instance and wait until it has finished, then start the next (reported as “clients”). Open loop for the latency points: starts at a fixed rate whatever the engine does |
| Clients | 128 and 512 (tiny shapes: 32, 128 and 512) |
| Job workers | Per job type: 4 worker processes, each with up to 128 jobs in flight, activating with a 1 s long poll (like the standard job-worker clients) |
| Simulated work per job | 10–50 ms, uniformly distributed |
| Payment confirmation | Published 5–50 ms after the charge job; 2 % never published (timer escalation after 5 s) |
| Clerk | 2 per tenant, polling every 200 ms, 10–50 ms per task |
| Incidents | 1 % of loans, resolved by the benchmark |
| Payload | Orders 5–20 KB; loans about 1 KB; straight-through about 100 bytes |
| Instance data | Every instance’s data and its path (express or not, escalation, incident) is drawn from a fixed seed and the instance’s number, so every run sees the same instances |
| Tenants | One tenant. With two Cluster nodes, the benchmark sends every request to both nodes in turn; the node that does not own the tenant’s partition forwards it to the one that does |
How each number is measured
Section titled “How each number is measured”- Throughput (instances/s): instances that finished inside the 30 s window, divided by 30. An instance has finished when the completion of its last job (“notify”, “escalate payment”, “disburse”, “step 3”) has been acknowledged.
- End-to-end duration: from sending the create request to the acknowledgement of that last job’s completion, for instances started inside the window. It includes the simulated work (for an order about 160 ms of job work and payment delay on the critical path), the waits for workers and the 5 s timer of escalated orders.
- p50, p95, p99: the value below which 50 %, 95 % and 99 % of the measurements fall (nearest rank).
- Create: create request to acknowledgement.
- Job pickup: from the acknowledgement of the instance’s previous step to the moment a worker receives the next job. It includes the engine’s work between the two (gateways, decisions, mappings) and any wait for a free worker.
- Engine CPU per instance: CPU time used by the engine processes during the run, divided by the instances that finished (including warm-up and the finishing period). With two nodes it is the sum of both. Database CPU per instance: the same for the PostgreSQL container.
- Database log per instance: PostgreSQL write-ahead log bytes written during the run, divided by the finished instances.
- Engine memory: the highest resident memory of the engine process(es), sampled every 2 seconds.
With a closed loop, the clients wait for their instances; once the engine is saturated, more clients only make each instance wait longer. So the highest rate is the saturation throughput, and the durations at 512 clients are mostly queueing.
Results by scenario
Section titled “Results by scenario”Tiny shapes (jobs and messages):
| Mode | Clients | Instances/s | End-to-end p50 / p95 / p99 (ms) | Create p50 / p95 (ms) | Engine CPU per instance | Database CPU per instance | Database log per instance |
|---|---|---|---|---|---|---|---|
| Embedded | 32 | 5,186–5,231 | 6 / 8 / 10 | 1.6 / 3.0 | 0.35 ms | — | — |
| Embedded | 128 | 5,233–5,301 | 23 / 31–32 / 35–37 | 6.4 / 9.0–9.3 | 0.38 ms | — | — |
| Embedded | 512 | 4,897–4,898 | 100 / 131 / 138–139 | 29.9–30.0 / 36.3 | 0.41 ms | — | — |
| Bundled | 32 | 1,903–1,904 | 16 / 22 / 26 | 3.8 / 6.1 | 0.56 ms | 0.70 ms | 29 KiB |
| Bundled | 128 | 3,139–3,205 | 39–40 / 51–52 / 56–59 | 9.2–9.4 / 13.6–13.9 | 0.54–0.55 ms | 0.59–0.60 ms | 17 KiB |
| Bundled | 512 | 3,577–3,596 | 136–137 / 182 / 208–212 | 35.8–36.0 / 48.0 | 0.54 ms | 0.53 ms | 15–16 KiB |
| Cluster, 1 node | 32 | 1,899–1,902 | 16 / 22–23 / 26 | 3.8 / 6.1–6.2 | 0.55 ms | 0.71 ms | 30 KiB |
| Cluster, 1 node | 128 | 3,125–3,170 | 39–40 / 52 / 58–59 | 9.3–9.5 / 13.8–13.9 | 0.54 ms | 0.60–0.61 ms | 17–18 KiB |
| Cluster, 1 node | 512 | 3,583–3,612 | 136–137 / 182 / 202 | 35.9–36.0 / 47.6 | 0.53–0.54 ms | 0.53 ms | 16 KiB |
| Cluster, 2 nodes | 32 | 1,759–1,766 | 18 / 24 / 27 | 3.9 / 6.2–6.3 | 1.1 ms | 0.73–0.74 ms | 31–32 KiB |
| Cluster, 2 nodes | 128 | 2,897–2,908 | 43 / 56 / 63 | 9.5–9.6 / 14.8 | 1.2 ms | 0.65 ms | 18 KiB |
| Cluster, 2 nodes | 512 | 3,335–3,354 | 147–149 / 195–196 / 215–221 | 37.5–37.8 / 49.7–50.1 | 1.2 ms | 0.56 ms | 16 KiB |
Straight-through (three jobs):
| Mode | Clients | Instances/s | End-to-end p50 / p95 / p99 (ms) | Create p50 / p95 (ms) | Engine CPU per instance | Database CPU per instance | Database log per instance |
|---|---|---|---|---|---|---|---|
| Embedded | 128 | 1,308–1,309 | 98 / 131 / 142–143 | 1.5 / 2.9–3.0 | 0.83 ms | — | — |
| Embedded | 512 | 2,998–3,028 | 167–169 / 229–232 / 247–250 | 5.3–5.4 / 9.9–10.2 | 0.62–0.63 ms | — | — |
| Bundled | 128 | 959–963 | 132–133 / 171 / 187–188 | 3.4–3.5 / 6.1 | 1.1 ms | 1.4 ms | 58–60 KiB |
| Bundled | 512 | 1,378–1,396 | 342–349 / 608–610 / 715–716 | 9.0–10.0 / 25.8–28.0 | 0.96–0.97 ms | 1.1 ms | 45–46 KiB |
| Cluster, 1 node | 128 | 957–962 | 132–133 / 171–172 / 187–189 | 3.5 / 6.1 | 1.0–1.1 ms | 1.4 ms | 62–63 KiB |
| Cluster, 1 node | 512 | 1,385–1,401 | 344 / 594–622 / 707–727 | 9.5–10.0 / 27.7–29.1 | 0.96 ms | 1.1 ms | 47 KiB |
| Cluster, 2 nodes | 128 | 896 | 142 / 185 / 204–207 | 3.7 / 6.3–6.4 | 2.2 ms | 1.5 ms | 61–62 KiB |
| Cluster, 2 nodes | 512 | 1,148–1,153 | 411–413 / 772–795 / 966–979 | 6.7–6.8 / 20.2–21.2 | 2.2 ms | 1.3 ms | 51–52 KiB |
At 128 clients Embedded is limited by the clients, not the engine: each instance spends about 90 ms in simulated job work, so 128 clients cannot start more than about 1,300 instances per second.
Order fulfilment:
| Mode | Clients | Instances/s | End-to-end p50 / p95 / p99 (ms) | Create p50 / p95 (ms) | Engine CPU per instance | Database CPU per instance | Database log per instance | Engine memory (peak) |
|---|---|---|---|---|---|---|---|---|
| Embedded | 128 | 367 | 346–347 / 452–453 / 504–509 | 9.3 / 20.2–20.7 | 4.1 ms | — | — | 2.4 GiB |
| Embedded | 512 | 359–371 | 1,344–1,391 / 1,728–1,791 / 1,850–1,899 | 88.6–91.3 / 152.1–153.7 | 4.0–4.2 ms | — | — | 3.3–3.6 GiB |
| Bundled | 128 | 267–269 | 374–381 / 515–520 / 5,253–5,254 | 10.5–11.0 / 26.3–28.5 | 6.0–6.1 ms | 4.4–4.5 ms | 204 KiB | 803–998 MiB |
| Bundled | 512 | 292–297 | 1,572–1,606 / 2,568–2,739 / 5,930–5,997 | 87.6–94.0 / 183.5–195.2 | 5.7 ms | 3.7–3.8 ms | 194 KiB | 1.1–1.4 GiB |
| Cluster, 1 node | 128 | 267–268 | 378–379 / 514–517 / 5,242–5,251 | 10.7–11.1 / 27.5–27.8 | 6.0 ms | 4.5–4.6 ms | 212 KiB | 1.0–1.1 GiB |
| Cluster, 1 node | 512 | 288–294 | 1,592–1,613 / 2,842–2,864 / 5,975–6,103 | 80.5–87.5 / 194.1–200.7 | 5.6 ms | 3.8 ms | 203–204 KiB | 1.2–1.3 GiB |
| Cluster, 2 nodes | 128 | 265–267 | 378–382 / 526–530 / 5,246–5,249 | 9.9–10.5 / 26.2–27.6 | 8.9 ms | 4.7 ms | 212 KiB | 909–1,015 MiB (2 nodes) |
| Cluster, 2 nodes | 512 | 284–286 | 1,510–1,536 / 3,399–3,551 / 5,884–5,912 | 32.7–34.5 / 114.0–130.5 | 8.6 ms | 4.0–4.1 ms | 206 KiB | 1.5–1.6 GiB (2 nodes) |
The p99 of about 5 s in Bundled and Cluster is the 2 % of orders that wait for the payment timer. In Embedded the clock is virtual, so no order escalates there. The engine memory is the orders in flight with their 5–20 KB documents, plus the records waiting for the query tables; a finished order leaves the engine’s memory.
Loan approval:
| Mode | Clients | Instances/s | End-to-end p50 / p95 / p99 (ms) | Create p50 / p95 (ms) | Engine CPU per instance | Database CPU per instance | Database log per instance | Incidents raised / resolved within the run |
|---|---|---|---|---|---|---|---|---|
| Embedded | 128 | 553–557 | 228–230 / 289–295 / 329–342 | 1.9 / 3.7–3.9 | 3.0–3.1 ms | — | — | 229–230 / 229–230 |
| Embedded | 512 | 683–687 | 746–749 / 912–918 / 972–990 | 20.7–21.9 / 48.8–49.4 | 2.8 ms | — | — | 276–277 / 276–277 |
| Bundled | 128 | 329–331 | 297–298 / 366–367 / 403–408 | 4.4–4.5 / 8.8 | 4.2 ms | 4.7 ms | 232–235 KiB | 130 / 112–114 |
| Bundled | 512 | 413–418 | 1,056–1,061 / 1,630–1,770 / 1,934–2,131 | 19.5–19.8 / 63.4–64.7 | 3.8–3.9 ms | 4.1–4.2 ms | 202–204 KiB | 171–174 / 108 |
| Cluster, 1 node | 128 | 328–330 | 297 / 365–367 / 406–413 | 4.4 / 8.8–8.9 | 4.1–4.2 ms | 4.7–4.8 ms | 241–245 KiB | 130 / 110–112 |
| Cluster, 1 node | 512 | 414–419 | 1,056–1,060 / 1,622–1,704 / 2,001–2,266 | 20.4–20.6 / 67.8–68.1 | 3.8 ms | 4.2 ms | 210–211 KiB | 173–174 / 108–109 |
| Cluster, 2 nodes | 128 | 278–317 | 307–351 / 377–461 / 422–555 | 4.8–6.2 / 9.3–12.6 | 7.1–8.0 ms | 5.1–5.8 ms | 242–244 KiB | 108–125 / 100–108 |
| Cluster, 2 nodes | 512 | 362–389 | 1,151–1,238 / 1,772–2,029 / 2,172–2,560 | 16.8–18.4 / 52.9–56.9 | 6.8–7.2 ms | 4.6–4.9 ms | 217–222 KiB | 146–156 / 104–106 |
No request failed. In Embedded every raised incident was found and resolved and every loan had finished before the end of the run. In Bundled and Cluster the benchmark finds an incident through the incident search, which reads the query tables; at these loan rates the query tables were 7–38 seconds behind, so some incidents were found only after the run’s 15 s finishing period and their loans were still waiting. Run again with 60 s to finish, every incident was resolved and every loan finished. See Known limits.
Latency below the maximum
Section titled “Latency below the maximum”Order fulfilment with starts at a fixed rate (40 s window, every start on time, no errors). The rates are those of the earlier measurements: about half and about 80 % of Bundled’s highest closed-loop rate, and the rates that were half and 80 % of Cluster’s earlier maximum:
| Mode | Rate (instances/s) | Create p50 / p95 (ms) | Job pickup p50 / p95 (ms) | Job completion p50 / p95 (ms) | Message publish p50 / p95 (ms) | End-to-end p50 / p95 (ms), orders not escalated |
|---|---|---|---|---|---|---|
| Bundled | 135 | 3.7 / 5.4–5.5 | 2.1–2.9 / 4.6–6.4 | 4.1–4.6 / 6.8–7.3 | 3.0 / 4.8 | 191–192 / 249 |
| Bundled | 215 | 4.6–5.4 / 10.6–33.4 | 4.1–7.9 / 14.0–166 | 6.5–8.5 / 16.6–71.5 | 4.1–5.0 / 10.8–29.9 | 225–247 / 294–1,031 |
| Cluster, 2 nodes | 8.5 | 6.9–7.0 / 9.7–9.9 | 1.8–2.4 / 3.1–4.9 | 4.0–6.3 / 4.8–8.3 | 4.0–4.1 / 5.6–6.1 | 199–200 / 254 |
| Cluster, 2 nodes | 14 | 5.8 / 8.3–8.5 | 1.5–2.0 / 3.0–4.5 | 3.3–5.4 / 4.8–8.4 | 3.5 / 5.2–5.4 | 192 / 244–247 |
| Cluster, 2 nodes | 135 | 4.0–4.1 / 5.7–7.6 | 2.4–3.5 / 4.8–11.2 | 4.6–5.0 / 7.3–11.3 | 3.2–3.3 / 5.0–6.9 | 195–200 / 255–272 |
| Cluster, 2 nodes | 215 | 5.0–7.0 / 11.2–53.9 | 4.5–13.5 / 15.0–471 | 7.3–11.7 / 16.8–102 | 4.4–6.3 / 11.4–52.4 | 230–280 / 299–2,247 |
Job pickup and completion show the range over the order’s job types (the escalation job, which waits for the 5 s timer, is left out). About 160 ms of every order’s end-to-end duration is the simulated work and payment delay, so the engine adds about 30–70 ms to an order at these rates, in Bundled and in Cluster alike.
At 215 orders per second, one of the two runs in each mode had a pause of a few seconds in which creates took up to 96 ms (Bundled) and 233 ms (Cluster); the high p95 values at that rate come from those pauses. The other run in each mode stayed at about 225–230 / 295–300 ms end to end.
Cluster: one node or two
Section titled “Cluster: one node or two”With one tenant, a second node does not add capacity: the tenant’s commands run in its one partition on one node, and the other node forwards the requests it receives to that node. Forwarding costs about 3–10 % of throughput and, per instance, about as much engine CPU again on the forwarding node (the tables above). More nodes add capacity when there are several tenants, spread over the partitions of the cell. Several cells and a database failover were not measured again in this round.
Starting a node with many tenants
Section titled “Starting a node with many tenants”One Cluster node with 3,721 tenants, measured twice:
| Until ready | |
|---|---|
| New database (every tenant is placed and its configuration saved) | 4.4–4.5 s |
| Restart on the same database | 13.4–14.3 s |
A restart takes about three times as long as a start on a new database: the node reads each tenant’s remembered idempotency keys back one tenant at a time, even for tenants that have none. See Known limits.
Database connections
Section titled “Database connections”Connections each node holds to PostgreSQL, with 24 partitions and one tenant, idle and at the highest count seen while 512 clients ran the tiny shapes against all nodes:
| Setup | Idle | Under load |
|---|---|---|
| Bundled | 30 | 31 |
| Cluster, 1 node | 30 | 31 |
| Cluster, 2 nodes | 18 and 17 | 18 and 19 |
Searches over many finished instances
Section titled “Searches over many finished instances”One tenant with about 70,000 finished instances (the tiny job shape), and one with about 136,000 (Cluster, one node), 20 requests one at a time after 3 warm-ups; Bundled and Cluster (one node) gave the same times:
| Request | p50 (ms), 70,000 | p95 (ms), 70,000 | p50 / p95 (ms), 136,000 |
|---|---|---|---|
| Process-instance search, first page of 50, no sort | 9.0–9.3 | 10.8–10.9 | 11.1 / 12.3 |
| Process-instance search sorted by start date, first page of 50 | 9.0–9.5 | 10.8–10.9 | 11.1 / 12.2 |
| Process-instance search, page of 50 starting in the middle of the matches | 17.1–17.5 | 26.2–27.1 | 21.1 / 21.8 |
| Decision-definition search (2 decisions deployed) | 2.7 | 3.8 | 2.7 / 3.8 |
| Variable search for one instance | 3.0 | 3.8–3.9 | 3.0 / 4.1 |
| Variable read by its key | 2.5 | 2.8–2.9 | 2.5 / 3.2 |
Sorting and paging run in the database, so a sorted search or a page far from the start costs about as much as the first page and grows only a little with the number of matches.
The query tables are filled after the engine has answered, and at the highest rates they fall behind: after a 20-second burst at the tiny shapes’ highest rate, they needed another 35 seconds to hold every finished instance, and 72 seconds after a 40-second burst.
Bundled restart time and memory
Section titled “Bundled restart time and memory”Jobs and messages, 64 clients, then a clean stop and a restart after each step. Without idempotency keys, 45 s of load per step:
| Instances so far | Instances/s during the step | Create p50 / p99 (ms) | Engine memory before the stop | Restart until ready | Engine memory after the restart | Saved tenant state | Database size |
|---|---|---|---|---|---|---|---|
| 118,256 | 2,626 | 5.5 / 10.3 | 102 MiB | 0.66 s | 102 MiB | 235 KB | 1.9 GB |
| 227,272 | 2,421 | 5.8 / 11.2 | 111 MiB | 0.96 s | 100 MiB | 268 KB | 3.9 GB |
With an idempotency key on every create, activation and completion, 15 s of load per step:
| Instances so far | Instances/s during the step | Create p50 / p99 (ms) | Engine memory before the stop | Restart until ready | Engine memory after the restart | Saved tenant state | Remembered keys |
|---|---|---|---|---|---|---|---|
| 37,279 | 2,481 | 5.5 / 15.5 | 146 MiB | 0.62 s | 119 MiB | 251 KB | 83,519 (32 MB) |
| 72,495 | 2,343 | 5.8 / 16.3 | 149 MiB | 0.60 s | 143 MiB | 262 KB | 162,720 (62 MB) |
| 106,446 | 2,259 | 6.0 / 17.9 | 208 MiB | 0.61 s | 184 MiB | 287 KB | 239,293 (90 MB) |
| 140,047 | 2,236 | 6.0 / 17.8 | 214 MiB | 0.68 s | 188 MiB | 278 KB | 314,930 (119 MB) |
Memory and restart time no longer grow with the history: the engine keeps
only the running instances, and a restart loads the tenant’s latest saved
state and replays the few frames after it. Idempotency keys are stored in
their own table, not in the saved state; the engine keeps the keys of the
last 10 minutes in memory (CONDUCTOR_CLAIM_WINDOW) and a compact filter for
older ones, so memory grows during the first 10 minutes of heavy keyed
load and should then stop growing. In a continuous 10-minute run with a
key on every request (about 4,200 keyed requests per second, 1.1 million
instances) the engine reached 540 MiB, still growing when the run ended at
the 10-minute mark, and 750 MiB after a restart; the saved tenant state
stayed at 190–320 KB, and no request failed. The database grows by about
17 KB per small instance, most of it the record stream that the query
tables and the history are built from; a third step without keys did not
fit the benchmark’s 5 GiB in-memory database. See
Known limits.
Tracing overhead
Section titled “Tracing overhead”Exporting traces, metrics and logs of every request to a local collector made no measurable difference to throughput or CPU per instance; see Overhead of tracing. That was measured on the earlier engine and has not been repeated.
What limits each mode
Section titled “What limits each mode”- Embedded runs about 4,900–5,300 small instances, 3,000 straight-through, 360–370 orders and 550–690 loans per second. Its limit is the engine’s one partition, which here also keeps every finished record in memory for the query side (1.8–6.5 GiB after these runs).
- Bundled runs up to about 3,600 small instances, 1,400 straight-through, 270–300 orders and 330–420 loans per second, with create latency of a few milliseconds below saturation. Its limit is the tenant’s one partition (one core). Memory stays at the instances in flight. At these rates the query tables that searches read fall behind (see Known limits).
- Cluster runs the same as Bundled on one node, because it is the same engine: about 3,600 small instances, 1,400 straight-through, 270–295 orders and 330–420 loans per second for one tenant. PostgreSQL is lightly loaded (about 4–5 ms of database CPU and 200–250 KiB of write-ahead log per order or loan). One tenant does not get faster with more nodes; more tenants do.
How to reproduce
Section titled “How to reproduce”The benchmark harness that produced these results is part of the engineering test suite and is not shipped. To measure your own processes on your own hardware, see Measuring it yourself.
Capacity planning
Section titled “Capacity planning”History retention
Section titled “History retention”Most of the database is the finished instances, which are kept 30 days by default (per tenant; see How long history is kept) and then removed completely. A small instance (a few elements, small variables) takes about 20 KB: 16 KB of records and 4 KB of searchable history. Instances with more steps and larger variables take several times more; measure yours by dividing the growth of the database by the instances started.
| Instances per day | Kept 1 day | Kept 7 days | Kept 30 days |
|---|---|---|---|
| 100,000 | about 2 GB | about 14 GB | about 60 GB |
| 1,000,000 | about 20 GB | about 140 GB | about 600 GB |
The figures are for small instances; multiply by your own size per instance. Add room for the change log, the saved tenant state, PostgreSQL’s own write-ahead log and one more day of history (the history is removed a whole day at a time). At a steady load the database stays about this size once the first retention period has passed. See Retention and disk.
PostgreSQL
Section titled “PostgreSQL”- Connections. A node holds one connection for the log of each
partition it owns, one for its snapshot writer, one more while it
maintains the reporting tables, and its pool (16 by default, shared by
requests, searches, identity and the exporter). A Bundled engine, or a
Cluster node that owns all 24 partitions, therefore opens at most 42. In a
cell the log connections are shared out between the nodes, so the cell
needs at most partitions + nodes × (pool size + 2): 24 + 18 × nodes with
the defaults, all direct or through a pooler in session mode.
Idle, a node holds about one connection per partition it owns plus 5:
29 on one node with 24 partitions, 17 on each of two nodes. With one busy
tenant the counts barely rise (30, and 17–18 on each of two nodes; see
Database connections); many busy tenants use
more of the pools. In Cluster mode set
max_connectionsto cover them plus your tools, backups and monitoring: PostgreSQL’s default of 100 is tight above two or three nodes. The Bundled image setsmax_connectionsto 100; if you raiseCONDUCTOR_PARTITIONS, raise it too (CONDUCTOR_POSTGRES_OPTIONS="-c max_connections=200"). - Storage. Every acknowledged write waits for a PostgreSQL commit, so
commit latency (the
fsyncof the WAL device) is the floor of every write’s latency in both durable modes. Use low-latency storage with honest flushes; do not turnfsyncorsynchronous_commitoff. - Network. In Cluster mode every save is one round trip to PostgreSQL, and a forwarded request adds one round trip between two nodes. Keep the nodes and the database in the same zone.
- CPU and memory. The exporter’s writes, index maintenance and vacuum run on the primary. Engine memory follows the instances that are still running.
Quotas
Section titled “Quotas”Per-tenant limits protect the cell from one tenant’s burst. They are exact,
because each tenant’s work runs in one partition; see
Tenant quotas. A tenant’s weight sets its share of its
partition’s processing time when several tenants are busy at once.
When a tenant or a node is at capacity
Section titled “When a tenant or a node is at capacity”The engine never stops to shed load; it answers with a typed error that says what to do:
| Answer | Why | What the client does |
|---|---|---|
429 urn:bpm:error:quota-exhausted | The tenant reached one of its limits | Wait the Retry-After delay and retry |
503 urn:bpm:error:backpressure | The tenant’s searches are too far behind | Wait the Retry-After delay and retry |
503 urn:bpm:error:partition-unavailable | The tenant’s partition has no owner at the moment (a failover or a move) | Wait the Retry-After delay and retry, with the same idempotency key |
In each case nothing was changed, except that a partition-unavailable
answer after a lost connection may come from a saved command: retrying with
the same idempotency key is always safe. Add some jitter to the retries. A
steady stream of these answers means a tenant needs higher limits or a
partition of its own, or the cell needs a larger database.
Workload class
Section titled “Workload class”The engine is built for human and agent cadence: many active instances, each moving every few seconds to days. It is not built for payment-grade volume in a single cell. Plan for more cells, not bigger ones, when the volume grows.
Tuning settings
Section titled “Tuning settings”Every setting below is read at startup. They apply to Bundled and Cluster mode.
| Variable | Default | Effect on performance |
|---|---|---|
CONDUCTOR_PARTITIONS | 24 | How many partitions share the work. Each is one thread on its owner; tenants on different partitions run in parallel. Fixed when the database is created |
CONDUCTOR_COMMIT_BATCH | 256 | The most frames saved together in one commit |
CONDUCTOR_SNAPSHOT_INTERVAL | 5000 | Milliseconds between two saved copies of a tenant’s state. Shorter means faster restarts and earlier removal of old frames, at the price of more writes |
CONDUCTOR_EXPORTER_BATCH | 256 | Frames the exporter reads per pass |
CONDUCTOR_EXPORTER_MAX_LAG, CONDUCTOR_EXPORTER_MAX_LAG_AGE_MS | 1000000, 60000 | How far a tenant’s searches may fall behind before its new instances are refused; they are slowed from half of it |
CONDUCTOR_PARTITION_LEASE_MS | 10000 | How long a crashed node’s partitions stay unavailable before another node takes them over. Shorter means faster failover, at the price of more renewals and a higher risk that a slow node loses its partitions |
CONDUCTOR_CLAIM_WINDOW | 10m | How long an idempotency key is kept in memory for an exact check; older keys are checked with a compact filter and, on a match, one database read. Shorter means less memory under heavy keyed load, at the price of more database reads for late retries |
CONDUCTOR_DEDUP_RETENTION | 24h | How long an idempotency key is remembered at all. The key table in the database holds every key for this long |
What to watch
Section titled “What to watch”GET /metrics on the metrics port (9090; see
Scraping /metrics) exposes the measures that matter for capacity, among them:
tinyconductor_command_seconds: from request to acknowledgement, per kind of command;tinyconductor_partition_commit_secondsandtinyconductor_partition_commit_batch_frames: time and size of each save;tinyconductor_tenant_queue_wait_seconds: how long a tenant waits for its turn;tinyconductor_tenant_engine_seconds_total: processing time per tenant, to find a tenant that needs a partition of its own;tinyconductor_exporter_lag_seconds: how far searches are behind;bpm_engine_jobs_availableandbpm_engine_jobs_activated: the job backlog.
See Observability.
Measuring it yourself
Section titled “Measuring it yourself”The figures on this page describe the reference environment. To learn what your installation does, measure it with your own processes:
- Compare your hardware with the reference environment: the CPU, disk, PostgreSQL commit latency and network figures in Comparing with your hardware show which of the numbers above you can expect to reach, and which will be lower or higher.
- Run the engine from the released image or download, in the mode and topology you plan to use, with PostgreSQL on the storage you plan to use.
- Deploy your own process models, and drive them with your own job workers or a load tool of your choice through the REST API: start instances at a fixed rate (for latency) or with a fixed number of concurrent clients (for the highest rate).
- Watch the measures listed under What to watch, and the CPU, memory and disk figures listed under Measuring your own installation.
- Record the hardware, the storage and the machine load with every result. In-memory storage for PostgreSQL makes the numbers look better, and a busy machine makes them look worse; neither transfers to production.
To find the highest rate your setup sustains, raise the number of concurrent clients (for example 16, 64, 256) until the rate stops rising; then measure latency at about half and at 80 % of that rate.
Known limits
Section titled “Known limits”- All measurements come from a single test machine, with PostgreSQL on an in-memory filesystem behind the container runtime’s port forwarding; none has been taken on production-like hardware.
- Idempotency keys cost memory and database space. The engine
remembers each key for a day (
CONDUCTOR_DEDUP_RETENTION): the keys of the last 10 minutes in memory (CONDUCTOR_CLAIM_WINDOW), older ones in a compact filter, and every key in a database table. With a key on every create, activation and completion at about 4,200 requests per second, the engine used about 540 MiB after 10 minutes (about 110 MiB without keys), the key table grew by about 380 bytes per key (1 GB after 10 minutes), and throughput was about 10 % lower than without keys. A shorterCONDUCTOR_CLAIM_WINDOWlowers the memory. - Searches can lag behind the engine. The exporter that fills the query
tables keeps up at full load on the test machine (a backlog of about 0.2–0.3
seconds), but on slower storage or a busy database it can fall behind;
searches, the incident and task lists and the history then lag by that
much. Reads by key of a live instance, job or task are always current.
Watch
tinyconductor_exporter_lag_seconds; when a tenant’s backlog reaches half ofCONDUCTOR_EXPORTER_MAX_LAGorCONDUCTOR_EXPORTER_MAX_LAG_AGE_MS, its new instances are slowed, and at the limit refused. - Throughput drifts down under sustained full load: by 12–22 % over 5–10 minutes in these runs, with or without idempotency keys, while the change log and the database grow.
- History takes most of the database. About 20 KB per small instance, more for larger variables, kept 30 days by default; see History retention.
- A Cluster node with thousands of tenants restarts more slowly than it first starts: 13–14 seconds with 3,721 tenants, against 4.5 seconds on a new database, because it reads each tenant’s remembered idempotency keys back one tenant at a time.
- One tenant runs on one core: its home partition. A tenant that needs more than one partition can hold is not supported yet.
- Embedded keeps every record of finished instances in memory, because tests read them: after the benchmark’s runs that was 1.8–6.5 GiB.
- The Kubernetes chart offers no autoscaler.
- PostgreSQL is a single point of contention per cell; more cells mean more databases, and tenants are placed in cells by configuration. Tenants do not move between cells.
- Bundled mode runs on one node; it cannot be scaled out.
- Principals with resource-level grants (rather than an action on every resource) cost an extra lookup on instance and user-task requests; the effect has not been measured.