Skip to content

Cluster mode

Use Cluster mode to run one shared production engine for many teams or customers (tenants). Cluster mode runs the same engine and serves the same API as Embedded and Bundled mode; the modes differ only in durability, speed, capacity and tenancy.

  • Nodes and one database. A Cluster is one or more engine processes, called nodes, that share one PostgreSQL database you operate, and serve many tenants. Each tenant’s data is kept apart in the database (PostgreSQL row-level security), and each tenant has its own quotas.
  • Partitions. The work is split into partitions, 24 by default. Each partition is served by one node at a time. The nodes share the partitions out between them, and when a node stops, the others take its partitions over. Any node answers any request: if the tenant’s partition is on another node, it forwards the request there.
  • Tenants. Each tenant is placed in one partition, usually next to other tenants. A busy tenant can be moved to a partition of its own, and moved back (see Moving a busy tenant).
  • The cell. One such deployment, one database with its nodes, is called a cell. It has an id, CONDUCTOR_CELL_ID. A cell serves many tenants; it is not one per tenant. You run a second cell, with its own database and nodes, only when you need more capacity than one database gives. Every tenant belongs to exactly one cell.
cell 1 (one PostgreSQL database, CONDUCTOR_CELL_ID=1)
├── node 1 ── partitions 1–12
│ partition 1: tenants acme, globex
│ partition 2: tenant initech
│ …
└── node 2 ── partitions 13–24
partition 13: tenant umbrella (a partition of its own)
…

How the engine runs your processes explains partitions, failover, placement and forwarding in more detail.

Before the first node starts, and before nodes of a new release start, apply the release’s database schema with the migrate command of the same image or program. It connects as the database’s owner login (see The database logins), creates or updates the engine’s schemas, gives the engine’s own login the rights it needs, and exits:

Terminal window
docker run --rm --env-file cluster.env \
-e CONDUCTOR_STORAGE_OWNER_USER=tinyconductor_owner \
-e CONDUCTOR_STORAGE_OWNER_PASSWORD='<owner password>' \
registry.tinyfactory.ai/tinyblox/tinyconductor:0.1.0 migrate

cluster.env holds the nodes’ settings (below): migrate reads the database connection from them exactly as a node does, so it reaches the same database and knows the engine’s login. migrate exits with 0 when the schema is up to date. Running it again changes nothing, and it is safe while nodes of the previous release keep serving. It refuses a database that a newer release has migrated, and says so. tinyconductor migrate --help lists its settings. The Helm chart runs it as a Job before every install and upgrade.

A Cluster node is the tinyconductor program with --mode cluster. It is published as the container image registry.tinyfactory.ai/tinyblox/tinyconductor (for linux/amd64 and linux/arm64) and as a download for Linux, macOS and Windows (see Embedded mode for the list). The image starts the Embedded test engine unless you pass --mode cluster, so always pass it: a node started without it keeps nothing.

Terminal window
docker run -d --name tinyconductor-node-1 -p 8080:8080 \
--env-file cluster.env \
registry.tinyfactory.ai/tinyblox/tinyconductor:0.1.0 --mode cluster

cluster.env holds the settings below. The downloaded program takes the same settings from its environment: tinyconductor --mode cluster.

On startup the engine connects to PostgreSQL, checks that the database schema is the one of its release, sets up its tenants (see Tenants), and only then accepts requests. A node never changes the schema itself: when migrate has not run yet, or not for this release, the node does not start and its log says to run it. If a setting is missing or invalid, the engine does not start either.

To check the settings before a rollout, run tinyconductor config check --mode cluster with the same environment (in the image: docker run --rm --env-file cluster.env registry.tinyfactory.ai/tinyblox/tinyconductor:0.1.0 config check --mode cluster). It reads every setting exactly as a start would, without connecting to anything, prints what it found (never a secret) and exits 0, or exits 1 naming the setting that is wrong. A CONDUCTOR_* variable the engine does not know is refused there, and logged as a warning at a start.

VariableDefaultMeaning
CONDUCTOR_STORAGE_HOST, CONDUCTOR_STORAGE_PORT(required), 5432The PostgreSQL server: a host name or address, or the directory of its Unix socket
CONDUCTOR_STORAGE_DB_NAME, CONDUCTOR_STORAGE_USER(required)The cell’s database and the engine’s login (see The database logins)
CONDUCTOR_STORAGE_PASSWORDunsetThe login’s password. Keep it in a secret store, or give the path of a mounted secret file in CONDUCTOR_STORAGE_PASSWORD_FILE instead
CONDUCTOR_DATABASE_URLunsetInstead of the five settings above: one connection URL, postgres://user:password@host:5432/database. Set the URL or the CONDUCTOR_STORAGE_* connection settings, not both; the engine refuses both. CONDUCTOR_DATABASE_URL_FILE names a file that holds it
CONDUCTOR_STORAGE_TLS_MODEverify-full for a database on another hostdisable, prefer, require, verify-ca or verify-full; set it here or as sslmode in the URL, not both. A mode weaker than verify-full for another host is refused (see Connecting over TLS). A local database (a Unix socket, localhost or a loopback address) keeps PostgreSQL’s default
CONDUCTOR_STORAGE_TLS_ALLOW_UNVERIFIEDfalsetrue accepts a TLS mode weaker than verify-full for a database on another host; the engine then warns at every start. For test setups only
CONDUCTOR_STORAGE_CA_FILEunsetA PEM file with the certificate authority that signed PostgreSQL’s server certificate. Without a TLS mode, it makes the connection verify-full
CONDUCTOR_STORAGE_POOL_SIZE16Connections in each node’s pool, 6–1000. The cell needs at most CONDUCTOR_PARTITIONS + nodes × (pool size + 2) connections (see Database connections)
CONDUCTOR_STORAGE_STATEMENT_TIMEOUT_SECONDS30The longest one statement of a node’s pool may run; 0 sets no bound. Schema updates, history retention and the reporting tables’ maintenance are not bounded by it
CONDUCTOR_STORAGE_CONNECT_TIMEOUT_SECONDShalf of CONDUCTOR_PARTITION_LEASE_MSThe longest a node waits for a database connection, opening one if need be; at most half the partition lease
CONDUCTOR_STORAGE_OWNER_USERunsetmigrate only: the owner login it connects as, on the server and database of the settings above. Unset: migrate connects as the engine’s login
CONDUCTOR_STORAGE_OWNER_PASSWORDunsetThe owner login’s password, or the path of a mounted secret file in CONDUCTOR_STORAGE_OWNER_PASSWORD_FILE instead. Give it to migrate only, never to the nodes
CONDUCTOR_STORAGE_AUTO_MIGRATEfalsetrue lets a node apply pending schema changes itself when it starts, as the owner login when CONDUCTOR_STORAGE_OWNER_USER is set (otherwise as its own login, which then must own the database). Meant for a single node; with several nodes, run migrate instead
CONDUCTOR_DATA_DIRunsetA directory only the engine’s user can read. The first node that creates the console administrator (with local users on) writes its password once to the file first-boot-secrets there, never to the log. Unset, the file goes to a new private directory under the system’s temporary directory, and the log says which
CONDUCTOR_SERVER_LISTEN_HOST, CONDUCTOR_SERVER_HTTP_PORT127.0.0.1, 8080The address (an IP address) and port of the API and the console. --bind host:port overrides both
CONDUCTOR_SERVER_METRICS_PORT9090The port /metrics is served on, on the same address, without a credential (see Observability). CONDUCTOR_SERVER_METRICS_ENABLED=false serves none
CONDUCTOR_SERVER_PUBLIC_URLunsetWhere clients reach the engine, host[:port] without a scheme, for example behind a load balancer. Document links name this host; unset, the host each request was sent to
CONDUCTOR_CELL_ID(required)This cell’s id, 1–8191. Every node of the cell uses the same id
CONDUCTOR_CLUSTER_NODE_NAME(required)Stable name of this node, without leading or trailing spaces. Partitions are leased to it, so give every node its own, and keep it when the node restarts: a node that comes back with the same name takes its partitions back at once
CONDUCTOR_AGENT_USAGE_AUTH_SECRET(required)The key the engine checks agent usage reports with: at least 32 random bytes. Keep it in a secret store, or give the path of a mounted secret file in CONDUCTOR_AGENT_USAGE_AUTH_SECRET_FILE instead
CONDUCTOR_INITIAL_TENANTdefaultThe id of the one tenant the cell starts with on its very first start (see Tenants)
CONDUCTOR_TENANT_CONFIGS_JSONunsetOptional: a few tenants, with their limits, to start with instead, read on the very first start only. Create and change tenants through the console or the tenant API (see Tenants)
CONDUCTOR_AUTH_CREDENTIALS_JSONrequired unless OIDC or local users are configuredStatic bearer tokens, each for one tenant: the bootstrap credential of a new cell, and for tests. Give people single sign-on and programs API clients (see Credentials). A token for a tenant that does not exist yet is left out with a warning in the log; restart the engine after creating the tenant to use it
CONDUCTOR_PARTITIONS24How many partitions the cell has. Fixed when the database is created; a different value later is refused
CONDUCTOR_PARTITION_LEASE_MS10000How long a partition’s lease lasts without renewal, 1000–300000 milliseconds. A node renews ten times per lease. It is also how long the others wait before they take over the partitions of a node that stopped without warning, and how long a node keeps retrying when PostgreSQL does not answer
CONDUCTOR_MAX_CLOCK_SKEW_MS1000How far this node’s clock may differ from PostgreSQL’s, in milliseconds. It must be positive and below half of CONDUCTOR_PARTITION_LEASE_MS. A node whose clock is further off gives up its partitions and reports not ready until its clock is back in range. Run NTP on every node
CONDUCTOR_INTERNAL_BIND_ADDR(required)The address this node’s internal port listens on, where the other nodes forward requests to it, for example 10.0.0.5:9600 (the node’s private address). There is no default: a node without it does not start. Never put it on a public network (see The internal port)
CONDUCTOR_NODE_ADVERTISE_URL(required)The address other nodes use to reach this node’s internal port: a URL with host and port and no path, http://node-1:9600, or https://node-1:9600 with internal TLS
CONDUCTOR_NODE_SECRET(required)A secret of at least 32 random bytes, the same on every node of the cell. Every forwarded request is signed with a key made from it; the secret itself is never sent. Keep it in a secret store, or give the path of a mounted secret file in CONDUCTOR_NODE_SECRET_FILE instead
CONDUCTOR_NODE_SECRET_PREVIOUSunsetWhile you change the node secret: the previous secret, still accepted (see Changing the node secret); CONDUCTOR_NODE_SECRET_PREVIOUS_FILE names a file that holds it
CONDUCTOR_INTERNAL_TLS_CERT_FILE, CONDUCTOR_INTERNAL_TLS_KEY_FILE, CONDUCTOR_INTERNAL_TLS_CA_FILEunsetOptional TLS for the internal port: this node’s certificate (PEM, with its chain), its private key, and the certificate authority that signs the certificates of every node of the cell. Set all three or none. See The internal port
CONDUCTOR_COMMIT_BATCH256The most frames saved together in one commit, 1–4096
CONDUCTOR_SNAPSHOT_INTERVAL5000Milliseconds between two saved copies of a tenant’s state
CONDUCTOR_EXPORTER_BATCH256How many frames the exporter reads per pass
CONDUCTOR_EXPORTER_MAX_LAG, CONDUCTOR_EXPORTER_MAX_LAG_AGE_MS1000000, 60000How far a tenant’s searches may fall behind (records, milliseconds) before its new instances are refused; they are slowed from half of it
CONDUCTOR_DEDUP_RETENTION24hHow long an idempotency key is remembered (units ms, s, m, h, d)
CONDUCTOR_CLAIM_WINDOW10mHow long a remembered key is kept exactly in memory; older keys are kept in a compact filter and looked up in the database only when the filter matches
CONDUCTOR_CLAIM_FILTER_FALSE_POSITIVE0.01Share of new keys that cost one extra database lookup because the filter matched by chance
CONDUCTOR_CLAIM_FILTER_ROTATION1hTime span each filter covers; a filter is dropped once all its keys have expired
CONDUCTOR_EVALUATE_BOUNDARY_EVENT_CORRELATION_KEY_IN_ACTIVITY_SCOPEfalseWhere a message boundary event’s correlation key is evaluated. false uses the scope around the activity, so the activity’s input mappings are not visible (the default in compatible engines). true uses the activity’s own scope (see Compatibility for the equivalent setting in other engines). A boundary on an embedded sub-process always uses the sub-process scope. Read by every mode, including Embedded and Bundled; restart to change it
CONDUCTOR_MAX_NAME_FIELD_LENGTH256The longest id, name, job type and resource name a deployment may use, and the longest message name and correlation key a running instance may evaluate, in characters. A longer one rejects the deployment, or raises an incident on the waiting element. The default is the same as in compatible engines. Read by every mode, including Embedded and Bundled; restart to change it

How the engine runs your processes explains partitions, leases, forwarding and the exporter. The metrics and log settings (CONDUCTOR_METRICS_TENANT_LABEL_LIMIT, CONDUCTOR_LOGGING_FORMAT, CONDUCTOR_LOGGING_LEVEL and the standard OTEL_* variables) in Observability, and the identity settings (OIDC, local users, seeding, the token endpoint) in Identity and access, and the document store settings (CONDUCTOR_DOCUMENT_*; a cell of more than one node needs the S3 store) in Documents.

Every secret setting (CONDUCTOR_STORAGE_PASSWORD, CONDUCTOR_DATABASE_URL, CONDUCTOR_NODE_SECRET, CONDUCTOR_NODE_SECRET_PREVIOUS, CONDUCTOR_AGENT_USAGE_AUTH_SECRET, CONDUCTOR_DOCUMENT_LINK_SECRET, CONDUCTOR_OIDC_CLIENT_SECRET, CONDUCTOR_SIGNING_KEY_ENCRYPTION_KEY) also takes the path of a file that holds it, in the same name with _FILE added, for example a mounted Kubernetes Secret: such a file shows in no process listing. Set one of the two, not both; the file’s trailing line break is not part of the secret, and an empty or unreadable file stops the start with the variable’s name.

The engine reads every setting once, at startup. The console’s configuration view shows the values as they were read then; restart the engine to change them. Tenants are the exception: you manage them while the engine runs (see Tenants).

The engine uses two PostgreSQL logins on its database:

  • The owner login owns the database and the engine’s schemas, and everything in them. Only migrate uses it. It needs membership in the role analytics_view_owner, which owns the analytics views.
  • The engine’s login is CONDUCTOR_STORAGE_USER (or the login of CONDUCTOR_DATABASE_URL), which every node connects as. It reads and writes the engine’s tables and nothing else: it cannot create, change or drop tables, schemas or views. migrate grants it these rights every time it runs; when the login does not exist yet and the owner login may create roles, migrate creates it with its password from CONDUCTOR_STORAGE_PASSWORD.

Neither login may be a superuser, have BYPASSRLS, or be a member of a role that is or has either: with any of these, row-level security does not apply to it. The engine’s login must not be a member of the owner login either. An administrator sets them up once per cell, for example:

CREATE ROLE tinyconductor_owner LOGIN PASSWORD '<owner password>';
CREATE ROLE tinyconductor LOGIN PASSWORD '<engine password>';
CREATE DATABASE tinyconductor OWNER tinyconductor_owner;
GRANT analytics_view_owner TO tinyconductor_owner;

analytics_view_owner and analytics_reader (the role you grant to reporting tools, see Analytics) belong to the whole PostgreSQL server, not to one database: every cell on the server shares the same two roles. Neither can log in. An administrator creates them once per server, before the first migrate:

CREATE ROLE analytics_view_owner NOLOGIN NOINHERIT;
CREATE ROLE analytics_reader NOLOGIN NOINHERIT;

Instead, the owner login can have the CREATEROLE attribute on a server where these two roles do not exist yet: migrate then creates them and makes the owner login a member. When several cells share a server, every other cell’s owner login needs the GRANT.

You can run with one login for both, as Bundled mode does: give CONDUCTOR_STORAGE_USER the owner login and leave CONDUCTOR_STORAGE_OWNER_USER unset. The nodes then hold a login that can change the schema, which two logins avoid. The engine checks its login at every start and, if it is a superuser or can bypass row-level security, writes a warning to its log that names the login, the role it is or can act as, and the attribute; it starts all the same. Keep a separate login for your own administration. If a right is missing, migrate or the node stops, and its log names the step that failed and the database’s reason, such as the GRANT to run.

What row-level security protects, and what it does not: it keeps each tenant’s rows apart when a query in the engine misses a tenant filter, and it keeps out other database logins, such as reporting tools, that are granted only the analytics role. The engine’s login cannot turn it off or change a table, but it does choose the tenant of each statement, so it can read any tenant’s rows. Treat both passwords like the data itself: keep them in a secret store, give the engine’s password to the nodes only and the owner’s password to migrate only, and require TLS for their connections (see Connecting over TLS).

A database on another host is reached over TLS that checks the server (verify-full) unless you set a TLS mode: the engine refuses a server whose certificate or name does not match. Set CONDUCTOR_STORAGE_CA_FILE to the certificate authority (a PEM file) that signed the server’s certificate, unless it chains to a public authority.

A weaker mode for a database on another host (disable, prefer, require, or verify-ca, which does not check the server’s name) stops the start, with a message that names the mode and the host. To accept it anyway, for a test setup, set CONDUCTOR_STORAGE_TLS_ALLOW_UNVERIFIED=true: the engine then starts and warns at every start. A database on the same machine (a Unix socket, localhost or a loopback address) keeps PostgreSQL’s default mode; the Bundled image’s database is one. On the PostgreSQL side, accept the engine’s login only over TLS (hostssl lines in pg_hba.conf).

Each node holds one connection for the log of each partition it owns, one for its snapshot writer, one more while it maintains the reporting tables, and its pool (CONDUCTOR_STORAGE_POOL_SIZE, 16 by default). The cell’s partitions are shared out between the nodes, so a cell of n nodes needs at most:

CONDUCTOR_PARTITIONS + n × (CONDUCTOR_STORAGE_POOL_SIZE + 2) connections: 24 + 18 × n with the defaults

plus what your tools, backups and monitoring use. A node checks at start that the server’s max_connections, less its reserved slots, holds the most it can open (CONDUCTOR_PARTITIONS + CONDUCTOR_STORAGE_POOL_SIZE + 2, when it owns every partition), and refuses to start otherwise. tinyconductor config check prints both figures for your settings.

Every connection must be a session of its own: the engine holds session advisory locks and listens for PostgreSQL notifications. Connect it directly or through a connection pooler in session mode, never through one in transaction mode.

A new cell starts with one tenant, default (or the id you give in CONDUCTOR_INITIAL_TENANT), with no limits. You add, change, suspend and reactivate tenants while the cell runs, in the console’s Tenants area or through the tenant API. Nothing restarts.

The engine keeps its tenants in its PostgreSQL database, and that list is what the cell serves. A new tenant is placed on the least-loaded partition and can be used immediately, through any node: a node that has not heard of it yet looks it up on the first request that names it. A new tenant, a changed limit or retention, a suspension and a reactivation reach every node within moments. Members and administrators added to an existing tenant work on every node as soon as the request that added them is answered. If a node misses a change of limits, retention or suspension, it catches up when it reads the whole list again, every 30 seconds.

If a create request fails or its answer is lost, send the same request again: it finishes the tenant and answers 201, as the first one would have. A create that names an existing tenant with other values answers 409. A suspension or reactivation is saved first; if it answers 503, repeat it, or wait: the engine applies the saved state by itself within 30 seconds. A change of limits that answers 503 is saved too, but the node running the tenant has not heard of it yet: repeat the request.

Tenants are managed by the administrators and tenant operators of the starting tenant; a tenant’s own administrators cannot see or change any tenant (see Who administers tenants).

ToConsoleAPI (permission tenant:admin)
list tenantsTenantsPOST /v2/tenants/search with {}
create a tenantTenants → New tenantPOST /v2/tenants
change its name, limits or history retentionTenants → the tenant → SettingsPUT /v2/tenants/{tenantId}
suspend it (it refuses all work)Tenants → the tenant → More actions → Suspend tenant…POST /v2/tenants/{tenantId}/suspension
reactivate itTenants → the tenant → Reactivate…POST /v2/tenants/{tenantId}/activation
name its administratorTenants → New tenant → Administrator, or the tenant’s Settings → Administratorsadmins in POST /v2/tenants or PUT /v2/tenants/{tenantId} (how)
add or remove membersTenants → Add member / RemovePUT/DELETE /v2/tenants/{tenantId}/users/{username} and its siblings
export a suspended tenant to a fileTenants → the tenant → Export…POST /v2/tenants/{tenantId}/export
import a file as a new tenantTenants → Import tenant…POST /v2/tenants/{tenantId}/import
purge a suspended tenant and everything it holdsTenants → the tenant → More actions → Purge tenant…POST /v2/tenants/{tenantId}/purge/confirmation, then POST /v2/tenants/{tenantId}/purge

For example, create a tenant with a few limits and 7 days of history, then raise one limit:

Terminal window
curl -X POST https://bpm.example.com/v2/tenants \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"tenantId": "acme", "name": "Acme",
"quotas": {"activeInstanceLimit": 10000, "jobsPerSec": 1000},
"retention": {"recordRetentionMicros": 604800000000}}'
curl -X PUT https://bpm.example.com/v2/tenants/acme \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"quotas": {"jobsPerSec": 2000}}'

A change names only what it changes; everything else stays. The limits are listed in Tenant quotas. A quota you leave out of a new tenant is no limit (the tenant API leaves it out of quotas, and the console shows No limit); its size limits (limits) are the engine’s own until you lower them, and its history is kept for 30 days.

A tenant you no longer need is removed with a purge, which needs a confirmation (see below). DELETE /v2/tenants/{tenantId} answers 409 and says so.

To let a program manage tenants for you, give its API client the built-in tenant-operator role: it may manage tenants and nothing else. See Automating tenant management.

You can copy one tenant to a file, bring that file back as a tenant on this cell or on another one, and remove a tenant with everything it holds. The three work together: export a tenant before you purge it, and you can restore it later from the file.

Export. Suspend the tenant first: only a suspended tenant is exported, so its running instances and its history are read at one and the same moment. The export answers one file (<tenant>-<time>.tenant.jsonl.gz) with everything the tenant holds:

  • its settings, limits and history retention, and its cluster variables;
  • its users (with the issuer and subject they sign in with), groups, roles, mapping rules, grants and API clients. API clients come without their secrets, and the file holds no password or token: local users who sign in with a password are not exported, so create them again after an import;
  • its audit;
  • its deployed processes, decisions and forms;
  • its running instances with their timers, jobs, message subscriptions and user tasks, and the answers kept for repeated requests;
  • its history within its retention, its analytics and its documents.

The file is gzip-compressed JSON lines. It names its format and version, the engine version that wrote it and the tenant it came from, and carries a SHA-256 hash of each section and of the whole file. The file is sent while it is written; a download that breaks off is incomplete, and an import refuses it.

Terminal window
curl -X POST https://bpm.example.com/v2/tenants/acme/suspension \
-H "Authorization: Bearer $TOKEN"
curl -X POST https://bpm.example.com/v2/tenants/acme/export \
-H "Authorization: Bearer $TOKEN" -o acme.tenant.jsonl.gz

The file holds the tenant’s data: keep it as safely as the database. Anyone with tenant administration in the starting tenant, including an API client with the tenant-operator role, can export a tenant.

Import. An import creates a tenant from the file: under its own id, or another one you name in the path. The id must be new, or the id of a purged tenant. The engine must be the same version as the one that wrote the file, and the file must be whole and unchanged; otherwise the import answers 400. Everything is stored in one step, or nothing is.

Terminal window
curl -X POST https://bpm.example.com/v2/tenants/acme/import \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/gzip' \
--data-binary @acme.tenant.jsonl.gz

The imported tenant starts suspended. Give its API clients new secrets, then reactivate it: its running instances go on where they stopped. Timers that came due meanwhile fire at once, jobs can be activated again, and messages are correlated as before. Process instance, job and other keys stay the same, so a worker or a link that names one still finds it, and a request repeated with the idempotency key it had before the export gets its first answer. Importing needs tenant administration on every tenant, as creating a tenant does.

Purge. A purge removes a suspended tenant and everything it holds, in every partition it ever lived in: its instances, records, history, documents, analytics, users, groups, grants, API clients, audit and settings, and the bindings of reporting tools’ database logins to it. Afterwards nothing in the database names the tenant, and its id can be used again, for a new tenant or to import its file. The starting tenant’s audit keeps who purged which tenant and when. A purge cannot be undone, so export the tenant first if you may need it again. The starting tenant itself cannot be purged.

A reporting tool’s database login bound to the tenant in analytics.reader_login (see Analytics) loses its binding, so it reads nothing of a new tenant given the same id. The login itself is a database role the administrator created; drop it when it is no longer needed.

A purge needs a confirmation, valid for ten minutes:

Terminal window
curl -X POST https://bpm.example.com/v2/tenants/acme/purge/confirmation \
-H "Authorization: Bearer $TOKEN"
# {"tenantId":"acme","confirmation":"…","expiresAt":"…"}
curl -X POST https://bpm.example.com/v2/tenants/acme/purge \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"confirmation": "…"}'

In the console, Purge tenant… asks you to type the tenant’s id instead. Once a purge has started, the tenant cannot be reactivated. If the purge answers 503, send the same request again: it continues where it stopped, also after the confirmation expired. GET /v2/tenants/{tenantId}/purge says whether a purge waits for its confirmation (CONFIRMING) or has started and not finished (PURGING).

Tenants are provisioned through the console or the tenant API, as described above. If you want the very first start to begin with a few tenants already in place, you can seed them in CONDUCTOR_TENANT_CONFIGS_JSON. The first one listed becomes the starting tenant (the tenant of the first administrator), unless CONDUCTOR_INITIAL_TENANT names another of them. Each entry needs only its tenantId and this cell (cellId); a quota it leaves out is no limit:

[{"tenantId": "acme", "cellId": 1,
"activeInstanceLimit": 10000, "jobsPerSec": 1000, "feelBudget": 100000,
"timerArmRate": 1000, "broadcastAdmission": 100, "mailboxBound": 10000,
"retention": {"recordRetentionMicros": 2592000000000}}]

retention is optional: 30 days by default. limits is optional too; left out, an entry gets the same size limits as a tenant created through the tenant API: {"maxDeploymentBytes": 4194304, "maxVariablesBytes": 33554432, "maxBatchSize": 1000} (4 MiB, 32 MiB and 1,000; see Size limits). An entry that gives limits gives all three.

The variable is read only on the very first start, when the database has no tenants yet. After that the engine ignores it, says so once in its log, and keeps the tenants as you manage them in the console or through the API. You can remove the variable after the first start. A large seed is not the way to provision many tenants: create them through the tenant API, for example from a provisioning script (see Automating tenant management).

Each tenant has its own history retention: how long its finished instances are kept, 30 days unless you set another value. Once an instance ended longer ago than that, it is removed completely: its records (element timeline, variable history, audit entries) and its searchable history (the instance, its element instances, variables, incidents, jobs and user tasks). Searches no longer list it, and reading it by its key answers 404, as for a key that never existed. An instance started by a call activity is removed together with the top-level instance that started it, once that one ended longer ago than the retention.

Running instances are never affected. Records that belong to no instance are removed by the same retention, counted from their own time: a published message once its time to live ended longer ago than the retention (a message that can still be correlated always stays), and a signal, a decision evaluated on its own or a refused request once they are older than the retention. Deployed processes, decisions and forms are kept. The reports under Analytics keep counting removed instances: they are daily and per-element summaries, and the retention does not change them.

The history is stored by the day it was written: an instance, the instances its call activities started and everything recorded about them belong to the day the instance started. Once a whole day is older than the longest retention of any tenant, the database removes that day in one step instead of instance by instance: it writes far less, and the disk space comes back at once. What that day still holds that must stay (instances still running, instances that ended less than their tenant’s retention ago, deployed definitions, messages that can still be correlated) is first set aside, once, and removed on its own when its time comes. So an instance can be kept up to one day longer than its retention; when every tenant keeps its history for less than eight days, the days are shorter (an eighth of the longest retention, at least a minute), and so is that delay. A tenant whose retention is shorter than another tenant’s has its instances removed one by one when they are due, until their day is removed.

Set it in the console (Tenants → the tenant → Settings → History kept for (days)) or through the API, in microseconds:

Terminal window
curl -X PUT https://bpm.example.com/v2/tenants/acme \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"retention": {"recordRetentionMicros": 604800000000}}'
  • 7 days: 604800000000, to save database space;
  • 1 day: 86400000000;
  • 90 days: 7776000000000, to look further back.

The cleanup, which runs about once a minute, uses the new value on its next round; no restart is needed. The database space this takes is estimated under Capacity planning. For history you need longer than that, copy the analytics views into your own reporting store (see Analytics).

A new tenant is placed on the least-loaded shared partition. A tenant that needs more than its share can get a partition of its own. An operator with the tenant:admin permission moves a tenant with POST /v2/tenants/{tenantId}/move, with the body {"dedicated":true} (a partition of its own) or {"partition":7} (a shared partition). The answer is 202 with a moveId. Follow the move with GET /v2/tenants/{tenantId}/moves/{moveId}: its step goes from REQUESTED through FROZEN, IMPORTED and SWITCHED to DONE. While it runs, the tenant’s requests wait up to two seconds, then answer 503 with Retry-After. A second move of the same tenant while one runs answers 409.

A move whose next step cannot be taken yet, for example because the partition the tenant moves into is not served for a moment after a database restart, waits and tries again after 1, 2, 4, then at most 8 seconds. It goes on by itself once the partition is served again. While it waits, the move’s answer has a blocked object: since (when it started waiting), retries and reason. The metric tinyconductor_tenant_move_blocked_seconds shows how long the longest-waiting move has waited (see Observability).

GET /api/v1/console/partitions (permission operations:read) lists the partitions, the nodes that own them and how many tenants each holds. A person or client of a tenant sees only its own tenant by name: where it is placed, and whether it is among the tenants that keep a shared partition busiest. Only a service client sees every tenant id, the busiest tenants of the whole cell and the nodes’ internal addresses.

Terminal window
export CONDUCTOR_STORAGE_HOST=db.internal
export CONDUCTOR_STORAGE_DB_NAME=bpm
export CONDUCTOR_STORAGE_USER=bpm
export CONDUCTOR_STORAGE_PASSWORD_FILE=/run/secrets/bpm-db-password
export CONDUCTOR_STORAGE_CA_FILE=/etc/tinyconductor/db-ca.crt
export CONDUCTOR_SERVER_LISTEN_HOST=0.0.0.0
export CONDUCTOR_CELL_ID=1
export CONDUCTOR_CLUSTER_NODE_NAME="cell-1-node-1"
export CONDUCTOR_INTERNAL_BIND_ADDR='10.0.0.5:9600'
export CONDUCTOR_NODE_ADVERTISE_URL='http://10.0.0.5:9600'
export CONDUCTOR_NODE_SECRET='<32 or more random bytes>'
export CONDUCTOR_AGENT_USAGE_AUTH_SECRET='<32 or more random bytes>'
export CONDUCTOR_AUTH_CREDENTIALS_JSON='[{"token":"<secret>","subject":"ops","tenantId":"default","kind":"service","actions":["*"]}]'
tinyconductor --mode cluster

The static token is the cell’s bootstrap credential; see Credentials for what to use once the cell runs. The internal port listens on the node’s private address (10.0.0.5 here); see The internal port.

GET /readyz answers 200 once the node has loaded the cell’s partition map, so it knows which node owns which partition and can serve or forward every request, and while PostgreSQL answers it within 2 seconds (see Observability). GET /livez tells an orchestrator such as Kubernetes that the process is alive; it never reads the database, so a database outage takes the node out of service without restarting it. The container image’s built-in health check uses /readyz. The image already sets CONDUCTOR_SERVER_LISTEN_HOST=0.0.0.0, because a container’s published port reaches the container’s own interface, not its loopback. /metrics listens on port 9090 of the same address, without a credential: publish it only to your metrics scraper.

A cell can run several engine nodes against the same PostgreSQL. They share the cell’s partitions: each partition is owned by one node at a time, and any node answers any request by forwarding it to the owner when needed. Every node needs its own CONDUCTOR_CLUSTER_NODE_NAME and its own CONDUCTOR_NODE_ADVERTISE_URL; everything else is the same. The nodes may start at the same time, once migrate has prepared the database.

When a node shuts down cleanly, it hands its partitions over at once. When it stops without warning, the other nodes take its partitions over once its leases run out, after CONDUCTOR_PARTITION_LEASE_MS (10 seconds by default). Until then the tenants on those partitions answer 503 with Retry-After; nothing acknowledged is lost. See How the engine runs your processes.

This example runs two nodes of cell 1 next to a PostgreSQL container. The PostgreSQL container creates the two logins (the owner, which owns the database, and the engine’s login; neither is a superuser or able to bypass row-level security) and serves TLS with a certificate from a small authority of your own; the nodes check it (verify-full, the default for a database on another host). The migrate service prepares the database as the owner login, and the nodes start once it has finished. In production, point the CONDUCTOR_STORAGE_* settings at the PostgreSQL you operate, set CONDUCTOR_STORAGE_CA_FILE to its authority, and drop the postgres service.

x-tinyconductor-node: &node
image: registry.tinyfactory.ai/tinyblox/tinyconductor:0.1.0
command: ["--mode", "cluster"]
restart: unless-stopped
stop_grace_period: 30s
depends_on:
migrate:
condition: service_completed_successfully
volumes:
- ./tls/ca.crt:/etc/tinyconductor/db-ca.crt:ro
environment: &node-env
CONDUCTOR_STORAGE_HOST: postgres
CONDUCTOR_STORAGE_DB_NAME: tinyconductor
CONDUCTOR_STORAGE_USER: tinyconductor
CONDUCTOR_STORAGE_PASSWORD: ${ENGINE_PASSWORD}
CONDUCTOR_STORAGE_CA_FILE: /etc/tinyconductor/db-ca.crt
CONDUCTOR_CELL_ID: "1"
CONDUCTOR_INTERNAL_BIND_ADDR: 0.0.0.0:9600
CONDUCTOR_NODE_SECRET: ${NODE_SECRET}
CONDUCTOR_AGENT_USAGE_AUTH_SECRET: ${AGENT_USAGE_SECRET}
CONDUCTOR_AUTH_CREDENTIALS_JSON: >-
[{"token":"${TINYCONDUCTOR_TOKEN}","subject":"ops","tenantId":"default","kind":"service","actions":["*"]}]
services:
migrate:
image: registry.tinyfactory.ai/tinyblox/tinyconductor:0.1.0
command: ["migrate"]
depends_on:
postgres:
condition: service_healthy
volumes:
- ./tls/ca.crt:/etc/tinyconductor/db-ca.crt:ro
environment:
CONDUCTOR_STORAGE_HOST: postgres
CONDUCTOR_STORAGE_DB_NAME: tinyconductor
CONDUCTOR_STORAGE_USER: tinyconductor
CONDUCTOR_STORAGE_PASSWORD: ${ENGINE_PASSWORD}
CONDUCTOR_STORAGE_CA_FILE: /etc/tinyconductor/db-ca.crt
CONDUCTOR_STORAGE_OWNER_USER: tinyconductor_owner
CONDUCTOR_STORAGE_OWNER_PASSWORD: ${OWNER_PASSWORD}
postgres:
image: postgres:16
# PostgreSQL wants its key owned by its own user: copy the pair in first.
command:
- bash
- -c
- >-
install -o postgres -g postgres -m 0600 /tls/server.crt /tls/server.key /var/lib/postgresql/
&& exec docker-entrypoint.sh postgres -c ssl=on
-c ssl_cert_file=/var/lib/postgresql/server.crt
-c ssl_key_file=/var/lib/postgresql/server.key
environment:
POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
OWNER_PASSWORD: ${OWNER_PASSWORD}
ENGINE_PASSWORD: ${ENGINE_PASSWORD}
configs:
- source: engine-login
target: /docker-entrypoint-initdb.d/engine-login.sh
volumes:
- pgdata:/var/lib/postgresql/data
- ./tls/server.crt:/tls/server.crt:ro
- ./tls/server.key:/tls/server.key:ro
healthcheck:
test: ["CMD", "pg_isready", "-U", "postgres", "-d", "tinyconductor"]
interval: 2s
retries: 30
node-1:
<<: *node
environment:
<<: *node-env
CONDUCTOR_CLUSTER_NODE_NAME: cell-1-node-1
CONDUCTOR_NODE_ADVERTISE_URL: http://node-1:9600
ports:
- "127.0.0.1:8081:8080"
node-2:
<<: *node
environment:
<<: *node-env
CONDUCTOR_CLUSTER_NODE_NAME: cell-1-node-2
CONDUCTOR_NODE_ADVERTISE_URL: http://node-2:9600
ports:
- "127.0.0.1:8082:8080"
configs:
engine-login:
content: |
psql -v ON_ERROR_STOP=1 -U postgres <<SQL
CREATE ROLE tinyconductor_owner LOGIN PASSWORD '$$OWNER_PASSWORD' NOSUPERUSER NOBYPASSRLS;
CREATE ROLE tinyconductor LOGIN PASSWORD '$$ENGINE_PASSWORD' NOSUPERUSER NOBYPASSRLS;
CREATE DATABASE tinyconductor OWNER tinyconductor_owner;
CREATE ROLE analytics_view_owner NOLOGIN NOSUPERUSER NOCREATEDB NOCREATEROLE NOINHERIT NOREPLICATION NOBYPASSRLS;
CREATE ROLE analytics_reader NOLOGIN NOSUPERUSER NOCREATEDB NOCREATEROLE NOINHERIT NOREPLICATION NOBYPASSRLS;
GRANT analytics_view_owner TO tinyconductor_owner;
SQL
volumes:
pgdata:

Create the certificate authority and the database’s certificate (for the name postgres, which the nodes connect to), put the secrets in a .env file next to compose.yaml (Compose reads it automatically), and start the cell:

Terminal window
mkdir -p tls
openssl req -x509 -newkey rsa:2048 -nodes -days 365 -subj /CN=tinyconductor-db-ca \
-addext basicConstraints=critical,CA:TRUE -addext keyUsage=critical,keyCertSign,cRLSign \
-keyout tls/ca.key -out tls/ca.crt
openssl req -newkey rsa:2048 -nodes -subj /CN=postgres \
-keyout tls/server.key -out tls/server.csr
printf 'subjectAltName=DNS:postgres\nextendedKeyUsage=serverAuth\n' > tls/server.ext
openssl x509 -req -in tls/server.csr -CA tls/ca.crt -CAkey tls/ca.key -CAcreateserial \
-days 365 -extfile tls/server.ext -out tls/server.crt
cat > .env <<EOF
POSTGRES_PASSWORD=$(openssl rand -hex 16)
OWNER_PASSWORD=$(openssl rand -hex 16)
ENGINE_PASSWORD=$(openssl rand -hex 16)
NODE_SECRET=$(openssl rand -hex 32)
AGENT_USAGE_SECRET=$(openssl rand -hex 32)
TINYCONDUCTOR_TOKEN=$(openssl rand -hex 24)
EOF
docker compose up -d
curl -s 127.0.0.1:8081/readyz
curl -s 127.0.0.1:8082/readyz

Both nodes serve the same API for the same tenants. Put a load balancer in front of them and send Authorization: Bearer <TINYCONDUCTOR_TOKEN> with every request. The internal port 9600 is only for the nodes themselves: don’t publish it. Inside each container it listens on 0.0.0.0, which is the container’s own network interface on the Compose network, not the host’s. Without internal TLS, requests on it are signed but not encrypted, so keep that network to the nodes alone, or turn on internal TLS (see The internal port).

Any node answers any request: when another node runs the tenant, the node forwards the request to it over the internal port. What protects that port:

  • Signed requests. Every forwarded request is signed with a key made from CONDUCTOR_NODE_SECRET, over the request, its body, the node it is meant for and the time it was sent. The secret itself never travels. A node refuses a request without a valid signature, a request sent more than two minutes before or after the time on its own clock, and a request it has already received, so a request someone recorded cannot be sent again. Run NTP on every node.
  • Checked callers. The node that runs a forwarded request checks again that the caller may act in the tenant the request names, and refuses it otherwise.
  • An address you choose. The port listens only on CONDUCTOR_INTERNAL_BIND_ADDR; there is no default. Use the node’s private address, and let only the other nodes reach it (a firewall, a security group or a Kubernetes NetworkPolicy).
  • Optional TLS. Without TLS, the forwarded requests and their data travel unencrypted between the nodes, and the node logs a warning at every start. On a network that is not private to the cell, turn on mutual TLS: create a certificate authority for the cell, give every node a certificate from it that names the host of its CONDUCTOR_NODE_ADVERTISE_URL, set the three CONDUCTOR_INTERNAL_TLS_* files on every node, and change every CONDUCTOR_NODE_ADVERTISE_URL to https://. Each node then accepts connections only from nodes with a certificate from that authority. All nodes of a cell use TLS, or none do.

You can change CONDUCTOR_NODE_SECRET without stopping the cell, in two rounds of restarts, one node at a time:

  1. On every node, set CONDUCTOR_NODE_SECRET to the new secret and CONDUCTOR_NODE_SECRET_PREVIOUS to the old one, and restart the node. A node on the new secret still accepts the old one, and when a node that has not been restarted yet refuses the new secret, it signs again with the old one, so forwarding goes on during the round.
  2. Once every node runs with the new secret, remove CONDUCTOR_NODE_SECRET_PREVIOUS from every node and restart them one at a time again. The old secret is then refused everywhere.

The tinyconductor Helm chart with postgres.bundled: false runs the nodes of one cell as a Kubernetes StatefulSet on the PostgreSQL you operate. It passes --mode cluster and gives each pod a stable CONDUCTOR_CLUSTER_NODE_NAME from its pod name. Before every install and upgrade it runs migrate as a Job with the owner login. It probes /livez and /readyz, serves /metrics on a Service of its own that only your scraper’s namespace may reach, and takes the database passwords, the node secret, the credentials and the identity settings from Secrets. The internal port is reachable only from the release’s own pods, and forwarding.tls turns on mutual TLS for it with a certificate from a Secret. Kubernetes covers the database setup, installing, scaling and upgrades.

Every request runs in one tenant. A person or client that belongs to one tenant needs nothing more. One that may act in several tenants names the tenant of each request in the X-Tenant-Id header. Row-level security keeps every tenant’s data apart in PostgreSQL (see The database logins for what it does and does not protect against). Permissions, roles and memberships are set per tenant. See Identity and access.

  • People sign in to the console with your identity provider over single sign-on (OIDC), or as local users where there is none. See Identity and access.
  • Programs (job workers, scripts, connector runtimes, provisioning pipelines) each get an API client: create it in the console (Access → Clients) or with POST /v2/clients, give it the roles or grants it needs, and let it fetch short-lived tokens from the engine’s token endpoint. See Tokens for SDKs and connectors.
  • Static tokens (CONDUCTOR_AUTH_CREDENTIALS_JSON) are for the first start of a cell, before any client or sign-on exists, and for tests. Once single sign-on and API clients are in place, keep only the static tokens you need for recovery; with single sign-on or local users configured, the variable may be left out.

Users, groups, roles, mapping rules, grants and API clients are the same on every node. A change answered by one node is seen by the next request to any node: a new client gets a token from any node, and a deleted client or a removed grant is refused by every node at once. Each node checks with PostgreSQL that its copy is current before it answers a request (one small read, shared by requests that arrive together); if PostgreSQL cannot be reached, the node answers 503.

Cluster mode serves:

  • every operation of the v1 and v2-compatible API;
  • TinyConductor Console at /console, including the Tenants and Credentials areas;
  • GET /metrics (Prometheus text or OpenMetrics with exemplars), the audit log, the history and analytics reads, and the identity administration API.

The trace context of instances and published messages is saved with the commands, so a trace continues across message correlation and call activities, and across restarts, failovers and tenant moves. A request forwarded to another node stays in the same trace. Timers start new traces. See Observability.

It refuses only the Embedded-mode clock and record routes. They answer 501 urn:bpm:error:mode-not-supported and name the mode. The API reference overview lists the differences per mode.

The engine saves every accepted command in PostgreSQL before it confirms it. Timers follow the real (wall) clock. Backups, point-in-time recovery and replication are ordinary PostgreSQL operations on the database you run.

Data written by preview builds before 0.1.0 is not carried over: migrate and the nodes refuse such a database and name the reset. From 0.1.0 on, migrate of a newer version upgrades the database forward; run it before you start the new version, and stop every node of the cell before you start the new version, because rolling upgrades are not supported yet. A node of an older version refuses a database that a newer version has migrated. See Upgrading.

A cell depends on its one PostgreSQL database. To survive the loss of the database machine, run PostgreSQL with a standby and a tool that promotes it (for example Patroni, CloudNativePG or your cloud’s managed failover), and give the engine one address that always leads to the current primary: a virtual IP, a DNS name the tool updates, or a TCP proxy. The engine does not pick or promote a new primary itself.

  • Requests fail for a moment. While the database is away, the tenants’ requests answer 503 with urn:bpm:error:partition-unavailable and a Retry-After header. Clients should retry, with the same idempotency key.
  • The nodes carry on by themselves. A partition keeps retrying while its lease is valid, and reloads its saved state when the database answers again. No restart is needed. If the outage lasts longer than the lease (CONDUCTOR_PARTITION_LEASE_MS, 10 seconds by default), the nodes claim the partitions again once the database is back.
  • Confirmed work is kept. Every command the API confirmed, such as a started instance, a completed job or a published message, was saved in the database before the answer went out. A confirmed job completion is never timed out or handed to another worker.
  • Zero loss needs synchronous replication. With asynchronous replication, work the old primary confirmed but had not yet sent to the standby is lost with it, including commands the API already confirmed. The engine cannot notice this. If confirmed work must survive the loss of the primary, use synchronous replication (synchronous_commit = on and synchronous_standby_names, or your tool’s equivalent).
  • An answer can be lost even when the work was saved. A request that was in progress when the connection dropped may have been saved even though the client got an error. Retrying with the same idempotency key is always safe; a retried start without a key creates a second instance.
  • The old primary must stay down. Stopping a former primary from coming back as a second writable database is the failover tool’s job.