Skip to content

Hardware requirements

This page says how much CPU, memory and disk TinyConductor needs, for the engine and for PostgreSQL, and how to size an installation. Every number on it was measured with the released images.

All measurements ran in the reference environment described below: a Linux virtual machine with 6 virtual CPUs, measured on its own so that you can compare it with your hardware. No other benchmark or build ran during a measurement. Use the numbers as a starting point, compare the reference environment with your own (see Comparing with your hardware), measure your own processes (see Measuring your own installation), and leave headroom.

What you runWhat it costs
Running instances (started, not yet finished)Engine memory. Every running instance is kept in memory, with its variables. A small instance takes about 12–17 KB; one with 50 KB of variables about 0.5 MB, roughly ten times the size of its variables
Instances per secondCPU of the engine and of PostgreSQL, the database’s write-ahead log (WAL) and the disk writes. Every acknowledged request waits for one PostgreSQL commit
Finished instancesDisk only. A finished instance leaves the engine’s memory; its history stays in PostgreSQL until the tenant’s history retention removes it
How long history is keptDisk: most of the database is the finished instances of that period, 30 days by default
TenantsLittle: 1,000 idle tenants add about 15–25 MiB of engine memory and a few hundredths of a core
Idempotency keysEngine memory: about 200 bytes per key for its first 10 minutes, then under 3 bytes per key for the rest of the day. A database table keeps every key for a day, stored by the hour, and removes each hour whole

The engine’s memory therefore follows the instances that are waiting (for a job worker, a message, a timer or a person), not the history, and its CPU follows the rate at which instances move.

Idle, one minute after start, with no instances. “Engine memory” is the resident memory of the engine process; for Bundled, the container holds the engine and PostgreSQL.

ModeTenantsEngine memoryPostgreSQL memoryCPUDatabase connectionsStart until ready
Embedded119 MiB—under 0.01 cores—under 0.1 s
Bundled154 MiB200 MiB for the whole container0.01 cores300.7 s
Bundled10059 MiB280 MiB for the whole container0.02 cores410.8 s
Bundled1,00070 MiB350 MiB for the whole container0.08 cores411.4 s
Cluster node138 MiB166 MiB0.01 cores310.4 s
Cluster node10042 MiB237 MiB0.02 cores410.5 s
Cluster node1,00060 MiB292 MiB0.09 cores411.2 s

Embedded holds one tenant, and it keeps the records of finished instances in memory for tests to read, so it grows with the history: it is not a production mode. The Cluster node owns all 24 partitions of its cell. PostgreSQL memory is the container’s memory, including the files it has cached; its own share (without the cache) was 80–190 MiB.

Instances of a process that waits on a job nobody completes, created one after the other, then measured after the engine had settled:

Running instancesVariables per instanceEngine memory (Bundled)Engine memory (Cluster node)
0—60 MiB40 MiB
10,000under 100 bytes411 MiB397 MiB
50,000under 100 bytes909 MiB878 MiB
1,00050 KB1.2 GiB
5,00050 KB2.8 GiB
10,00050 KB5.0 GiB (6.5 GiB peak)

To estimate the engine’s memory:

engine memory ≈ 100 MiB + running instances × (15 KB + 10 × the size of their variables)

Keep variables small: store large documents elsewhere (for example with Documents) and keep a reference in the instance. 50,000 instances with 50 KB each would need about 25 GiB, which did not fit the test machine and was not measured.

After a burst of creates the engine holds more memory than it needs at rest: 50,000 instances created in 9 seconds across 1,000 tenants left the engine at 1.6 GiB; after a restart the same instances took 0.9–1.0 GiB. Allow about 1.5 times the estimate as the memory limit.

A process with one service task, run at a fixed rate for 5 minutes: every instance is created, its job is activated and completed (three requests), and the instance finishes. Latency is measured by the client, over HTTP on the same machine.

The rates below are the rates tested, not the most the engine can take. A process with more steps costs more per instance: count roughly one of these instances per service task (one job each) when you estimate your own load. All tenants of one database share its capacity, and a single tenant runs in one partition.

ModeInstances/s (and jobs completed/s)Engine CPUPostgreSQL CPUCreate p50 / p99Job completion p50 / p99Engine memoryWAL written
Bundled500.11 cores0.28 cores2.9 / 3.9 ms2.7 / 3.4 ms63 MiB2.0 MB/s
Bundled2000.24 cores0.99 cores2.5 / 4.4 ms2.5 / 4.2 ms71 MiB7.7 MB/s
Bundled5000.42 cores1.19 cores2.3 / 6.2 ms2.4 / 6.1 ms138 MiB18.6 MB/s
Cluster node500.11 cores0.29 cores3.0 / 3.9 ms2.7 / 3.3 ms47 MiB2.0 MB/s
Cluster node2000.24 cores0.99 cores2.4 / 4.5 ms2.5 / 4.2 ms55 MiB7.2 MB/s
Cluster node5000.41 cores1.18 cores2.3 / 6.2 ms2.4 / 6.0 ms83 MiB17.2 MB/s

Every start was on time and no request failed. On the engine’s own measurement (tinyconductor_command_seconds), a create took 0.6–0.8 ms on average from its arrival to its durable answer, and a commit to PostgreSQL 0.3–0.4 ms.

Per 100 instances per second of this one-job process, the engine used 0.08–0.22 cores and PostgreSQL 0.24–0.6 cores (the lower figures at the higher rates, where more work shares each commit). Each additional job, message or user task in a process adds about as much again. PostgreSQL made 2.6–3 commits per instance (1,330 per second at 500 instances per second), each one a flush of the WAL.

These rates are far below the most one tenant can do on this machine (see Performance and scaling): the sizes below leave room for bursts and longer processes.

Three starting points. The “Measured at that rate” row is what the runs above used with a one-job process; the resources include headroom for slower server cores, processes with more steps, and bursts.

SmallMediumLarge
Typical loadup to 50 instances/s at peak, 100,000 a day, 10,000 runningup to 200 instances/s at peak, 1 million a day, 50,000 runningup to 500 instances/s at peak, 5 million a day, 100,000 running
Measured at that rate50/s: engine 0.11 cores, PostgreSQL 0.3 cores, create p99 3.9 ms200/s: engine 0.24 cores, PostgreSQL 1.0 core, create p99 4.5 ms500/s: engine 0.42 cores, PostgreSQL 1.2 cores, create p99 6.2 ms
Engine (per Cluster node, or the engine in Bundled)1 CPU, 1 GiB memory2 CPUs, 2 GiB4 CPUs, 4 GiB
PostgreSQL CPU and memory2 CPUs, 4 GiB4 CPUs, 8 GiB8 CPUs, 16 GiB
PostgreSQL disk for finished instances kept 30 days (records and searchable history)60 GB600 GB3 TB (keep 7 days: 700 GB)
PostgreSQL disk speed1,000 IOPS, low-latency flushes3,000 IOPS6,000 IOPS or more, and 25 MB/s of WAL
Bundled pod (engine and PostgreSQL together)1–2 CPUs, 4 GiB, the disk above4 CPUs, 8 GiBuse Cluster mode

Memory assumes small variables; add 10 × the variable size per running instance for larger ones (see Memory per running instance). The disk figures follow from Retention and disk; add room for the WAL and for PostgreSQL’s own maintenance. The commands of one tenant run on one core: the measured 500 instances per second of one tenant used less than half of it, but a tenant with long processes can fill it, and then only more tenants (or a partition of its own) spread the load over more cores.

The Kubernetes chart’s default resources are the Small size: 1 CPU and 1 GiB requested with a 2 GiB limit for a Cluster node, and 1 CPU and 2 GiB requested with a 4 GiB limit for the Bundled pod (engine and PostgreSQL).

The runs above used PostgreSQL 16 with its default settings. Every acknowledged request waits for PostgreSQL to flush its write-ahead log (WAL) to disk, so what matters most is how much WAL the database writes and how quickly the disk flushes it. For a Cluster, use these settings, scaled to PostgreSQL’s memory:

shared_buffers = 1GB # about a quarter of PostgreSQL's memory
effective_cache_size = 3GB # about three quarters of it
wal_buffers = 16MB
max_wal_size = 4GB # fewer checkpoints under steady writes
checkpoint_timeout = 15min
wal_compression = lz4
max_connections = 200

The Bundled image applies the write settings itself: shared_buffers=256MB, wal_buffers=16MB, max_wal_size=2GB, checkpoint_timeout=15min, wal_compression=lz4 and max_connections=100. Change any of them with CONDUCTOR_POSTGRES_OPTIONS, for example CONDUCTOR_POSTGRES_OPTIONS="-c shared_buffers=1GB -c max_wal_size=4GB" (the Helm chart’s postgres.serverOptions); what you set there wins. Leave room for the WAL on the data volume: up to about max_wal_size plus 1 GB.

Why they matter over hours: with PostgreSQL’s defaults (128 MB shared buffers, 4 MB WAL buffers, max_wal_size of 1 GB), a Bundled container at 100 instances per second with a key on every request and retention removing finished instances checkpointed every 37–75 seconds. More than half of the WAL (57 %) was full copies of pages written again after each checkpoint, and the database wrote 700–850 MB of WAL a minute. With the settings above the same load wrote 480–570 MB a minute, a third less. Create p99 per 10 minutes then stayed at 10–13 ms from minute 10 to minute 60, against 8 ms in the first 10 minutes.

At 500 instances per second in a 5-minute run they made little difference (the same CPU and latency, 4 % less WAL): a short run hardly checkpoints.

Keep fsync and synchronous_commit on: every acknowledged request relies on the commit.

A flush of the WAL waits for the disk. When something else keeps that disk busy, every commit waits longer, and so does every acknowledgement. In the reference environment, a full scan of a 2.4 GB table every 30 seconds on the same disk stretched single WAL flushes to 0.45–0.7 s. Create p99 went from under 25 ms to 175–240 ms in the minutes the scans ran.

  • Run reports, exports and ad-hoc queries that read whole tables against a replica, not against the database the engine writes to.
  • Count rows with the planner’s estimate (pg_class.reltuples) or with pg_stat_user_tables, not count(*), in monitoring that runs every few seconds.
  • Give PostgreSQL a disk of its own, or one whose flush latency does not depend on other workloads. Measuring your hardware the same way shows how to check its flush latency.

Each of the engine’s connections is a PostgreSQL process. After hours of load each engine connection’s process held 70–77 MB, of which PostgreSQL itself accounted for about 13 MB; the rest is memory it used once and kept. The processes give it back only when the connection closes. Plan for:

PostgreSQL memory ≈ shared_buffers + 80 MB × the engine’s connections + the file cache you want

A three-node Cluster holds about 75 connections (see Connections): in a 6-hour run their processes reached 3.0 GiB, so plan 3–3.5 GB for them beyond shared_buffers, plus the file cache. The PostgreSQL of a Bundled container reached 1.3 GiB in the same run. A PostgreSQL container limited too tightly keeps the connections’ memory and loses its file cache instead, and then reads from disk what it would have found in memory.

A node holds one connection per partition it owns, for that partition’s log, one for its snapshot writer, one more while it maintains the reporting tables, and its pool (CONDUCTOR_STORAGE_POOL_SIZE, 16 by default). For a cell of N nodes the log connections are shared out, so the cell needs at most:

max_connections ≥ partitions + N × (pool size + 2) + what your tools and backups use (20 or more)

that is 24 + 18 × N connections for the engine with the defaults. A node that owns all 24 partitions, as a Bundled engine does, opens at most 42; measured on one node with 24 partitions: 30–31 connections with one tenant, 32–33 under load and 41 with 100 or more tenants. For three nodes the engine needs at most 78, about 100 with your tools: raise PostgreSQL’s default of 100. The Bundled image sets 100, which is enough for its one node; if you raise CONDUCTOR_PARTITIONS or the pool size, raise it by the same number. A node refuses to start when the server cannot hold the 42 (with the defaults) it may open alone.

Connect the engine directly or through a connection pooler in session mode: it holds session advisory locks and listens for notifications, so a pooler in transaction mode does not work.

The WAL grew by 35–41 KB per instance of the one-job process: 2 MB/s at 50 instances per second, 7–8 MB/s at 200 and 17–19 MB/s at 500. Size the WAL disk, WAL archiving and any replica’s network for that, and for more per instance when processes have more steps or larger variables.

A finished small instance added about 20–24 KB to the database:

PartPer finished instanceRemoved by the history retention?
Records of the instance (the record stream)about 16 KBYes, with the day it was written, once that day is older than the tenant’s retention (30 days by default)
Searchable history: the instance, its elements, variables, incidents, jobs and user tasksabout 4 KBYes, together with its records
Change log of the partitions0.7–4 KBYes, within minutes, once the records are saved
Idempotency keys (only when clients send them)about 350 bytes per key (the key table and its indexes)Yes, with the hour in which they expire, a day after they were made (CONDUCTOR_DEDUP_RETENTION)

The history is stored by day. A day is removed whole once it is older than the longest retention of any tenant, so the database keeps up to one day more than the retention (with retentions under eight days, an eighth of the longest retention). What is still needed then (instances still running or finished less than their retention ago) is set aside and removed on its own later. So a steady load of R instances per second needs about:

disk ≈ R × (retention + 1 day) × 20 KB (16 KB of records and 4 KB of searchable history) + keys of the last day

For example, 1 million small instances a day, kept 30 days, need about 600 GB: 480 GB for the records and 120 GB for the searchable history. Once the first retention period has passed, the database stays about that size at a steady load: the retention removes as much each day as the new instances add. The reports are the exception: their daily and per-element summaries are kept for good, a few rows per process, day, element and job type, however many instances they count. Instances with more steps and larger variables take several times more; measure yours (see below).

The idempotency keys are the one part that grows for a whole day by design: every key is kept for 24 hours, then removed. A client that sends a key with every request at K requests per second needs about K × 86,400 × 350 bytes for them, 30 GB a day at 1,000 keyed requests per second; after the first day that size stays level.

Removing a day gives its space back to the operating system at once. The rows set aside because they were still needed (instances that run longer than about a day, and deployed definitions) are removed one by one when their time comes; PostgreSQL does not give that space back but reuses it for new rows. With processes that run for weeks or months most of the history is set aside, and the database then behaves as if there were no days: the space settles at what the retention period holds.

Checked in the reference environment with a one-hour Bundled run at 50 instances per second, a key on every request (about 150 new keys a second) and a 10-minute retention: from minute 25 on the records table stayed at 1.05–1.13 GB and 1.30–1.31 million rows, the searchable history at 380–410 MB and the change log of the partitions between 145 and 415 MB (it is emptied every few minutes); only the key table grew, by 3.3 MB a minute.

This was also checked with a one-hour run at 200 instances per second and a 15-minute retention, before the retention also removed the searchable history: after the first 15 minutes, the database grew by about 8 KB per instance instead of 22 KB. The rest was the searchable history (4 KB, which the retention now removes as well), the idempotency keys of the day, and space of deleted records that PostgreSQL had not reused yet.

Removing history by day was measured in the reference environment with a one-hour Bundled run at 100 instances per second, a key on every request and a 10-minute retention (days of one minute), against the same run removing every instance one by one: from minute 20 on the database wrote 4.4 MB of WAL a second instead of 5.0 MB (12 % less; full-page copies fell from about 400 to 50 a second), create p99 was 5.5 ms instead of 6.8 ms, and the database ended at 1.75 GB instead of 2.11 GB. Processes that are still running when their day is removed cost one extra copy of their history, once: with every instance waiting 15 minutes (longer than the retention; 50 instances per second for 20 minutes), the database wrote 62 KB of WAL per instance instead of 43 KB, about the size of an instance’s history once more.

A restart loads each tenant’s latest saved state and replays the few changes saved after it. With 1,000 tenants and 50,000 running instances:

ModeStart until readyPeak engine memory during the restartEngine memory after
Bundled (engine and PostgreSQL restart together)5.0–5.1 s0.9 GiB (1.2 GiB for the container)0.9 GiB
Cluster node (engine only)5.1–5.2 s1.1 GiB1.0 GiB

A restart needs no more memory than the running instances themselves.

Bundled, 200 instances per second of the one-job process, with a 15-minute history retention so that the retention also ran:

RunEngine memory after 15 minAfter 30 minAfter 60 minGrowth after the first 15 min
No idempotency keys (v2 API)147 MiB151 MiB—0.2–0.3 MiB a minute
No idempotency keys (v1 API)137 MiB151 MiB—0.6 MiB a minute
A key on every request (about 370 new keys a second)146 MiB158 MiB167 MiB after 50 min0.4 MiB a minute

Throughput stayed at 200 instances per second and the median create at 2.4–2.5 ms from the first minute to the last, in every run. Finished instances leave memory: the engine levels off at about 150–160 MiB, with keys or without.

The engine remembers every idempotency key for a day (CONDUCTOR_DEDUP_RETENTION, 24 hours): exactly for the first 10 minutes (CONDUCTOR_CLAIM_WINDOW), about 200 bytes per key, then in a compact filter, about 2.4 bytes per key. So with keys the engine’s memory grows slowly for the first day and then stays level. At 60 keyed requests per second that is about 7 MiB for the last 10 minutes and 12 MiB for the rest of the day; at the 370 keys a second above, about 40 MiB and 75 MiB. Runs longer than an hour were not measured.

On Linux the engine fixes one setting of the system’s memory allocator when it starts (the size from which a block gets memory of its own, 128 KiB). Without it, the allocator keeps memory the engine has already given back, and under a steady keyed load the engine’s memory grew by about 2 MiB a minute for hours. This holds for the images and the release archive alike; there is nothing to set. A value you set yourself (GLIBC_TUNABLES=glibc.malloc.mmap_threshold=…) is used instead.

Every figure on this page was measured in one Linux virtual machine that ran the containers (engine, PostgreSQL and, in Bundled mode, both in one container). The virtual machine is described and measured below, with the same tools you can run on your own hardware (see Comparing with your hardware).

Reference environment
Virtual machineLinux virtual machine of Rancher Desktop 1.24 (Apple Virtualization framework): Alpine Linux 3.23, kernel 6.6.137, arm64
Container runtimeDocker Engine 29.1.3 in that virtual machine
CPU6 virtual CPUs, arm64 (ARMv8, vendor Apple; the virtual machine reports no model name), one thread per core. The host is an Apple M5 Max with 18 cores
Memory23.4 GiB for the virtual machine (the host has 128 GiB)
StoragePostgreSQL’s data on a named Docker volume on the virtual machine’s disk, which is a disk image on the host’s internal SSD
NetworkEngine and PostgreSQL in separate containers on one Docker network (Cluster mode); in Bundled mode the engine reaches PostgreSQL over a Unix socket inside its container
PostgreSQL16, default settings except where PostgreSQL says otherwise; fsync and synchronous_commit on

Measured on 2026-09-27 in three runs, with no other benchmark or build running (the virtual machine’s own idle services kept running); where the runs differed, the range is shown:

MeasurementToolResult
CPU, one threadsysbench cpu, 10 s11,500–12,000 events/s
CPU, all 6 virtual CPUssysbench cpu, 10 s45,000–46,800 events/s
Memory bandwidth, one threadsysbench memory, 1 MiB blocks, write48–49 GiB/s
Memory bandwidth, all virtual CPUssysbench memory, 1 MiB blocks, write104–107 GiB/s
Write of 8 KiB followed by fsyncfio, one job, 30 s1,700–6,500 writes/s; fsync p50 0.09–0.24 ms, p99 0.6–6.8 ms
Random writes, 4 KiB, queue depth 32fio, direct I/O, 30 s22,600–36,900 IOPS
Sequential writes, 1 MiB blocksfio, direct I/O, 30 s2,100–4,700 MB/s
One commit (a single-row insert per transaction), 1 clientpgbench, 60 s0.18–0.29 ms per commit (3,400–5,500 commits/s)
Standard pgbench transactions, 1 clientpgbench (TPC-B-like, scale 10, simple protocol), 60 s2,100–3,200 transactions/s, 0.31–0.48 ms each
Standard pgbench transactions, 8 clientsthe same, 8 clients5,100–9,700 transactions/s, 0.83–1.57 ms each
Round trip, container to the PostgreSQL containerping; SELECT 1 through pgbench0.03 ms (ping); 0.02 ms (SELECT 1)

The numbers above are what the virtual machine delivers to its containers, which is what the engine and PostgreSQL used. The disk is the host’s internal SSD: its flushes are fast but vary from run to run (the fsync p99 above), and the host’s own caching sits underneath. Server storage, and network storage in particular, usually has slower but steadier flushes; compare your commit latency with the figure above.

Three figures decide most of the difference between the reference environment and yours:

FigureWhat it drivesHow to scale the numbers on this page
Commit latency (the single-row insert above; fsync latency of your PostgreSQL disk)Acknowledgement latency and the highest write rate of one tenant. Every acknowledged request waits for one PostgreSQL commit, and the commands of one tenant are saved by one partition, one commit after the other (commands that are ready together share a commit)Add the difference in commit latency to every acknowledgement. The highest rate of one tenant falls roughly in proportion to the commit latency once commits are no longer shared: with commits 5 times slower than here, expect about a fifth of the per-tenant rate. Spread busy tenants over more partitions
CPU speed (one-thread sysbench cpu)Engine and PostgreSQL CPU per 100 instances per second (see CPU and latency per instance rate)Multiply the cores on this page by the reference one-thread result divided by yours. A core half as fast needs twice the cores for the same rate
Memory per running instanceEngine memory (see Memory per running instance)Does not depend on the speed of the hardware: use the formula as it is. Memory bandwidth changes only how fast a restart loads the state

Also compare the network round trip between the engine and PostgreSQL: each commit is at least one round trip, so a database 1 ms away adds about 1 ms to every acknowledgement and lowers the per-tenant rate accordingly. Random write IOPS and sequential write throughput matter for the exporter, checkpoints and the write-ahead log at high rates; compare them with the Sizes table.

As a rule of thumb: start from the size for your load, scale its CPU by the one-thread CPU ratio, keep its memory, and check that your commit latency and round trip leave room for the acknowledgement latency you need. Then measure your own processes (see Measuring your own installation).

These commands run the same checks with public container images (debian:bookworm-slim for sysbench and fio, postgres:16 for pgbench). They need Docker, take about seven minutes, and remove everything they create at the end. Run them on the machine (or in the Kubernetes node or virtual machine) where PostgreSQL will run, when it is otherwise idle.

CPU, memory and disk. The disk tests write to a named Docker volume; to test another disk, replace tc-bench-disk with a directory on it (-v /path/on/that/disk:/data):

Terminal window
docker volume create tc-bench-disk
docker run --rm -v tc-bench-disk:/data debian:bookworm-slim sh -c '
apt-get update -qq && apt-get install -y -qq sysbench fio >/dev/null
sysbench cpu --threads=1 --time=10 run | grep "events per second"
sysbench cpu --threads=$(nproc) --time=10 run | grep "events per second"
sysbench memory --threads=1 --memory-block-size=1M --memory-total-size=100G --time=10 run | grep transferred
cd /data
fio --name=fsync --rw=write --bs=8k --size=256m --ioengine=sync --fsync=1 \
--runtime=30 --time_based | grep -A 8 "fsync/fdatasync"
fio --name=randwrite --rw=randwrite --bs=4k --size=1g --ioengine=libaio --direct=1 \
--iodepth=32 --runtime=30 --time_based | grep "IOPS="
fio --name=seqwrite --rw=write --bs=1m --size=2g --ioengine=libaio --direct=1 \
--iodepth=8 --runtime=30 --time_based | grep "bw="
rm -f fsync.* randwrite.* seqwrite.*'
docker volume rm tc-bench-disk

PostgreSQL commit latency and network round trip, with PostgreSQL 16 on its default settings and pgbench in a second container, as the engine would connect:

Terminal window
docker network create tc-bench
docker run -d --name tc-bench-pg --network tc-bench -e POSTGRES_PASSWORD=bench \
-v tc-bench-pgdata:/var/lib/postgresql/data postgres:16
sleep 10
PGB="docker run --rm --network tc-bench -e PGPASSWORD=bench postgres:16"
$PGB pgbench -h tc-bench-pg -U postgres -i -s 10 -q postgres
$PGB pgbench -h tc-bench-pg -U postgres -M simple -c 1 -j 1 -T 60 postgres
$PGB pgbench -h tc-bench-pg -U postgres -M simple -c 8 -j 8 -T 60 postgres
docker exec tc-bench-pg psql -U postgres -c 'CREATE TABLE commit_test (v int)'
$PGB sh -c "echo 'INSERT INTO commit_test VALUES (1);' > /tmp/c.sql &&
pgbench -h tc-bench-pg -U postgres -M simple -n -c 1 -T 60 -f /tmp/c.sql postgres"
$PGB sh -c "echo 'SELECT 1;' > /tmp/s.sql &&
pgbench -h tc-bench-pg -U postgres -M simple -n -c 1 -T 10 -f /tmp/s.sql postgres"
docker rm -f tc-bench-pg
docker network rm tc-bench
docker volume rm tc-bench-pgdata

Read tps and latency average from each pgbench run: the insert run is the commit latency, the SELECT 1 run the round trip to the database. To measure a PostgreSQL you already run (a managed service, for example), point the pgbench container at it with -h, -p, -U and a database of its own, from a machine where the engine will run.

Run your own processes at your expected rate on the hardware you plan to use, and watch:

  • Engine memory: the container’s memory (docker stats, or the pod’s container_memory_working_set_bytes). Divide its growth by the running instances to get your memory per instance.
  • CPU: engine and PostgreSQL containers, in cores, at your rate.
  • Latency: tinyconductor_command_seconds (from arrival to durable answer) and tinyconductor_partition_commit_seconds (one commit) on GET /metrics; see Observability.
  • WAL: SELECT pg_current_wal_lsn() before and after a run; the difference in bytes, divided by the seconds.
  • Disk per instance: SELECT pg_database_size(current_database()) before and after a run, divided by the finished instances, once tinyconductor_exporter_lag_seconds is back near zero.
  • Connections: SELECT count(*) FROM pg_stat_activity WHERE datname = current_database().

Then scale: memory with the running instances, CPU and WAL with the rate, disk with the rate and the retention.

  • The numbers come from one reference environment, a virtual machine on a workstation; none was taken on server hardware or network storage.
  • Large variables cost about ten times their size in engine memory while the instance runs.
  • Memory with idempotency keys was measured for 50 minutes, not for the full day the engine remembers them; the day’s total is computed from the per-key cost (see Memory over time).
  • Space comes back to the disk when a whole day is removed. Instances removed one by one (those set aside because they ran longer than about a day, and those of a tenant whose retention is shorter than another tenant’s) leave space that PostgreSQL reuses but does not give back: after you shorten such a retention, the files stay as large as before until a VACUUM FULL in a maintenance window.
  • Tenants and credentials are provisioned through the API and the console: tenants through the tenant API or the console’s Tenants area, people through single sign-on, programs as API clients. The first-start seed (CONDUCTOR_TENANT_CONFIGS_JSON) and static tokens (CONDUCTOR_AUTH_CREDENTIALS_JSON) are meant for a few entries; a large seed is not the way to provision many tenants.