Kubernetes
TinyConductor ships two Helm charts:
| Chart | What it runs | Use it for |
|---|---|---|
tinyconductor | The engine, as a StatefulSet. With postgres.bundled: true (the default): one pod with the engine and its own PostgreSQL, and one persistent volume (Bundled mode). With postgres.bundled: false: the engine nodes of one Cluster mode cell (one database and its nodes, serving many tenants), on a PostgreSQL database you operate | One team’s production (Bundled); shared, multi-tenant production (Cluster) |
tinyconductor-script-worker | The JavaScript script worker, pointed at the engine | Script tasks written in JavaScript |
Both charts are published to the registry as OCI charts, at
oci://registry.tinyfactory.ai/tinyblox/charts, under the release’s version, and
are also release downloads, tinyconductor-0.1.0.tgz and
tinyconductor-script-worker-0.1.0.tgz, each with a .sha256 file next to
it. From the registry (log in with the credentials that come with your
access):
helm registry login registry.tinyfactory.aihelm show chart oci://registry.tinyfactory.ai/tinyblox/charts/tinyconductor --version 0.1.0From the downloads, check them before you use them:
shasum -a 256 -c tinyconductor-0.1.0.tgz.sha256 tinyconductor-script-worker-0.1.0.tgz.sha256The commands on this page install from these files. Each chart checks its values against a schema before it renders anything, so a typo or a missing setting stops the install with a message that names it.
The defaults are the safe ones:
- authentication is always on, and no chart can turn it off;
- secrets are never written into your values: you create the Secrets, and the charts only name them;
- pods run as a non-root user, on a read-only file system, with every Linux capability dropped;
- a NetworkPolicy admits traffic to the engine only from the pods of its own namespace, and from the peers you add; its metrics port only from your metrics scraper’s namespace.
Before you start
Section titled “Before you start”You need:
- Kubernetes 1.30 or later (the pods pause before they stop with the
built-in
sleepaction); - Helm 3 or Helm 4 (the charts are tested with Helm 3.22 and Helm 4.2);
- the published images, which the charts use by default:
registry.tinyfactory.ai/tinyblox/tinyconductor-bundled(Bundled mode),registry.tinyfactory.ai/tinyblox/tinyconductor(Cluster nodes) andregistry.tinyfactory.ai/tinyblox/tinyconductor-script-worker. The image tag is the chart’sappVersion; setimage.tag, or betterimage.digest, to pin a build. If you mirror the images, setimage.bundledRepositoryandimage.repository.
The examples below install into a namespace called tinyconductor, with the
release name tinyconductor. The chart then names its StatefulSet and Service
tinyconductor, and its pods tinyconductor-0, tinyconductor-1 and so on.
kubectl create namespace tinyconductorWhy a StatefulSet
Section titled “Why a StatefulSet”The engine runs as a StatefulSet in both variants. A StatefulSet gives each pod a stable name that it keeps when it restarts, and Kubernetes never runs two pods with the same name at the same time. The engine takes its owner identity from that name. So the same owner identity never runs twice at once, which is what lets the engine hand its work safely from one pod to the next.
A headless Service (tinyconductor-headless) gives each pod a stable DNS name
next to the ordinary Service (tinyconductor) that clients use.
Bundled mode
Section titled “Bundled mode”Bundled mode is the chart’s default. The chart runs exactly one pod, with the engine and PostgreSQL inside it, and one persistent volume for PostgreSQL’s data. PostgreSQL listens only on a socket inside the pod, never on the network.
The pod meets Kubernetes’ restricted policy: it runs as the image’s own
non-root user (uid and gid 65532, with fsGroup: 65532 so the persistent
volume is writable), with a read-only root file system, no privilege
escalation, all capabilities dropped and the RuntimeDefault seccomp
profile. The image runs the same way under Docker. The database superuser
logs in only with its password, kept on the volume, and the engine runs in
a kernel sandbox (Landlock) that keeps it away from that password and
PostgreSQL’s files. So code running inside the engine cannot act as the
superuser; it is limited to the engine’s own login, which is not a
superuser and cannot bypass row-level security. The sandbox needs Linux
5.13 or later on the nodes; where a node’s kernel lacks it, the engine logs
so at every start. See Bundled mode
for what the sandbox covers.
The persistent volume also holds PostgreSQL’s socket, so give it a storage
class backed by a block device (the usual ReadWriteOnce classes are), not
a network file share.
For database administration, kubectl exec into the pod and run psql: it
connects as the superuser with the password from the volume.
kubectl -n tinyconductor exec -it tinyconductor-0 -- psqlhelm install tinyconductor ./tinyconductor-0.1.0.tgz --namespace tinyconductorkubectl -n tinyconductor rollout status statefulset/tinyconductorThe first-start token
Section titled “The first-start token”With no credentials configured, the engine creates an API token on its first
start. It never writes it to its log. It writes it once to the file
first-boot-secrets on the data volume, readable only by the engine’s user,
and stores only a digest of it, so the token cannot be read back later. The
first start also creates the console administrator admin and writes its
password to the same file. Read the file right after the first start, keep
both in your secret manager, and delete the file:
kubectl -n tinyconductor exec tinyconductor-0 -- cat /var/lib/postgresql/data/tinyconductor/first-boot-secretskubectl -n tinyconductor exec tinyconductor-0 -- rm /var/lib/postgresql/data/tinyconductor/first-boot-secretsUse them to set the installation up, then move to lasting credentials:
- People: turn on single sign-on with
oidc.*(see Identity and access), so that nobody depends on the generated password. - Programs: turn on the engine’s token endpoint with
auth.tokenIssuer.issuer, then create an API client for each job worker, script or scraper in the console (Access → Clients) or withPOST /v2/clients(see Tokens for SDKs and connectors).
If your installation must start with a fixed token instead of the printed
one (for example in an automated test environment), put a credential
document in a Secret and set auth.credentials.existingSecret before the
first start; the stored first-start token is then ignored:
kubectl -n tinyconductor create secret generic tinyconductor-credentials \ --from-file=credentials.json=./credentials.jsonhelm install tinyconductor ./tinyconductor-0.1.0.tgz --namespace tinyconductor \ --set auth.credentials.existingSecret=tinyconductor-credentialsTo replace a lost first-start token, see Bundled mode.
Key values
Section titled “Key values”| Value | Default | Meaning |
|---|---|---|
postgres.bundled | true | Bundled mode. replicaCount must stay 1, and the other postgres server settings (host, password, ca, caFile) must stay empty |
publicUrl | none | Where people and programs reach the engine, as host or host:port without https://, usually the ingress host |
tenants.initialTenant | default | The one tenant the engine starts with. Add and change tenants afterwards in the console or through the tenant API (Tenants) |
tenants.config or tenants.existingConfigMap | none | Optional: a few tenants to seed on the very first start only. Provision tenants through the console or the tenant API |
agentUsage.existingSecret | none | Optional: a Secret with the key agent usage reports are checked with |
persistence.size | 20Gi | Size of PostgreSQL’s volume |
persistence.storageClassName | the cluster default | Use storage with fast, honest flushes: every acknowledged request waits for one |
persistence.retainOnDelete | true | Keep the volume when the release is removed |
auth.* | the first-start token and local users | Static tokens from a Secret, OIDC and the other identity settings, as in Cluster mode |
partitions.count | 24 | The cell’s partitions, fixed when the database is created; for a new database only |
settings | none | Any other non-secret setting this chart does not model yet |
postgres.serverOptions | none | Extra PostgreSQL server options, for example -c shared_buffers=512MB |
postgres.sharedMemorySize | 256Mi | Shared memory for PostgreSQL |
resources | 1 CPU and 2 GiB requested; 4 GiB limit | The container runs the engine and PostgreSQL. The default is the Small size of Hardware requirements; the engine’s memory grows with the instances that are still running |
startupProbe.failureThreshold | 120 (10 minutes) | A start creates or checks the schema and loads every partition before the engine is ready |
Back up the volume with volume snapshots, or with PostgreSQL tools plus a
copy of the engine’s own directory on it (tinyconductor/, the secrets the
engine generated on its first start; see
Operations). There is no
disruption budget: one pod cannot stay available during a node drain.
Cluster mode
Section titled “Cluster mode”Database setup
Section titled “Database setup”In Cluster mode the chart does not install PostgreSQL. Use PostgreSQL 16 from one of:
- a managed service (Amazon RDS or Aurora, Google Cloud SQL, Azure Database for PostgreSQL and similar);
- the CloudNativePG operator, an open-source (Apache-2.0) operator that
runs PostgreSQL with replicas and automatic failover inside Kubernetes.
helm/cloudnative-pg-cluster.yamlin the examples download of the release (tinyconductor-0.1.0-examples.tar.gz) is a ready example for one cell; install the operator first, as its own documentation describes; - a PostgreSQL server you already run.
Give the engine its own database and two logins on it:
- the owner login (
postgres.owner) owns the database and the engine’s schemas. Only the chart’s migration Job uses it: before every install and upgrade it runstinyconductor migrate, which creates or updates the schemas and gives the engine’s login its rights; - the engine’s login (
postgres.user) is the one the nodes use. It reads and writes the engine’s tables and cannot create, change or drop anything, so a node never changes the schema.
Neither may be a superuser, have BYPASSRLS, or be a member of a role that
has either: with any of these, PostgreSQL’s row-level security does not
apply to it. The engine checks its login at every start and logs a warning
that names the login and the attribute if it is such a role. Row-level
security keeps the tenants apart if a query misses its tenant filter, and
keeps out other logins such as reporting tools. Keep each password in a
Secret: the owner’s for the migration Job only, the engine’s for the nodes
only. See Cluster mode for the full
list of rights.
The engine also uses two roles without login rights for the analytics
views: analytics_view_owner owns the views, and analytics_reader is the
role you grant to reporting tools. These roles belong to the whole
PostgreSQL server. Create them once per server, as an administrator (on a
managed service, its admin user), and make the owner login a member of
analytics_view_owner:
CREATE ROLE analytics_view_owner NOLOGIN NOINHERIT;CREATE ROLE analytics_reader NOLOGIN NOINHERIT;CREATE ROLE tinyconductor_owner LOGIN PASSWORD '…';CREATE ROLE tinyconductor LOGIN PASSWORD '…';CREATE DATABASE tinyconductor OWNER tinyconductor_owner;GRANT analytics_view_owner TO tinyconductor_owner;If several cells share one server, create the two analytics roles once and
run the GRANT for each cell’s owner login. Alternatively, give the owner
login the CREATEROLE attribute on a server where the two roles do not
exist yet: the first migration then creates them and makes the owner a
member. If the owner has neither right, the migration Job fails with a
message that names the GRANT to run.
The CloudNativePG example creates the database, both logins and the analytics roles.
Connections. Each node holds one connection for the log of each
partition it owns, one for its snapshot writer, one more while it maintains
the reporting tables, and its pool (postgres.poolSize, 16 by default). The
cell’s 24 partitions are shared out between the nodes, so the cell needs at
most
partitions + replicas × (pool size + 2), that is 24 + 18 × replicas with the defaults
connections. Set max_connections to cover that plus your tools, backups and
monitoring: PostgreSQL’s default of 100 is tight above three nodes. A node
refuses to start when the server cannot hold the most it can open alone
(24 + 18 with the defaults). Every connection must be a session of its own
(the engine holds session advisory locks and listens for notifications):
connect directly or through a connection pooler in session mode, never in
transaction mode. The chart prints the figure for your release when you
install it.
Latency. Every acknowledged request waits for a PostgreSQL commit. Keep the database in the same zone as the nodes, on storage with fast and honest flushes. See Performance and scaling.
Secrets
Section titled “Secrets”Create the Secrets the chart refers to. Only their names go into your values.
# The password of the engine's database login. The server, database and# login are plain values (postgres.host, postgres.dbName, postgres.user).kubectl -n tinyconductor create secret generic tinyconductor-db \ --from-literal=password='…'
# The password of the owner login (postgres.owner.user), for the migration# Job only.kubectl -n tinyconductor create secret generic tinyconductor-db-owner \ --from-literal=password='…'
# The certificate authority that signed the database's certificate, unless# it chains to a public one (postgres.ca in the values). The engine checks# the server's certificate and name (postgres.tlsMode: verify-full).kubectl -n tinyconductor create secret generic tinyconductor-db-ca \ --from-file=ca.crt=./db-ca.crt
# At least 32 random bytes the nodes sign forwarded requests with.kubectl -n tinyconductor create secret generic tinyconductor-node-secret \ --from-literal=secret="$(openssl rand -hex 32)"
# At least 32 random bytes that authenticate agent usage reports.kubectl -n tinyconductor create secret generic tinyconductor-agent-usage \ --from-literal=secret="$(openssl rand -hex 32)"
# A bootstrap token for the first start, before single sign-on or API# clients exist.kubectl -n tinyconductor create secret generic tinyconductor-credentials \ --from-file=credentials.json=./credentials.jsoncredentials.json is the credential document described in
Bundled mode; every
entry names an existing tenant (default on a new cell). A cell needs this
bootstrap token, single sign-on or local users to start. If you configure
single sign-on (oidc.*) from the first install, you can leave the
token out. With CloudNativePG, the operator creates the owner login’s
password Secret itself: use tinyconductor-db-app with the key password
for postgres.owner.password, the operator’s -rw Service as
postgres.host, and create the engine’s password Secret as the example’s
header shows.
Install
Section titled “Install”A minimal values.yaml for one cell with two nodes:
publicUrl: tinyconductor.example.compostgres: bundled: false host: db.internal dbName: tinyconductor user: tinyconductor password: existingSecret: tinyconductor-db owner: user: tinyconductor_owner password: existingSecret: tinyconductor-db-owner ca: existingSecret: tinyconductor-db-careplicaCount: 2cell: id: 1forwarding: nodeSecret: existingSecret: tinyconductor-node-secretagentUsage: existingSecret: tinyconductor-agent-usageauth: credentials: existingSecret: tinyconductor-credentialsThe cell starts with one tenant, default. Once it runs, add your tenants
in the console’s Tenants area or through the tenant API, and set their
limits there; see Tenants. To automate that, give
an API client the tenant-operator role
(Automating tenant management).
Give job workers and scripts API clients of their own (the metrics scraper
needs none: the metrics port takes no credential)
(set auth.tokenIssuer.issuer to turn on the token endpoint, with
auth.tokenIssuer.keyEncryptionKey.existingSecret naming a Secret with the
key that encrypts the stored signing keys:
kubectl -n tinyconductor create secret generic tinyconductor-signing-kek --from-literal=key="$(openssl rand -base64 32)"),
and people single sign-on; see Credentials.
helm install tinyconductor ./tinyconductor-0.1.0.tgz \ --namespace tinyconductor -f values.yamlkubectl -n tinyconductor rollout status statefulset/tinyconductorThe install first runs the migration Job (tinyconductor-migrate) and
creates the nodes only when it succeeded. If it fails, helm install fails
too; read its log with
kubectl -n tinyconductor logs job/tinyconductor-migrate, fix the cause (a
missing right, a wrong password) and run the install again.
Try it from your machine:
kubectl -n tinyconductor port-forward service/tinyconductor 8080:8080curl -s localhost:8080/readyzInternal TLS
Section titled “Internal TLS”The nodes forward requests to each other on the internal port. The
NetworkPolicy lets only the release’s own pods reach it and every request
is signed, but without TLS the requests and their data cross the cluster
network unencrypted. Turn on mutual TLS unless that network is private to
you. Put a certificate for
*.tinyconductor-headless.tinyconductor.svc (the headless Service of the
release, in its namespace), its key and the CA that signed it in one
Secret, for example with a cert-manager Certificate from a CA Issuer of
its own, and name it:
forwarding: nodeSecret: existingSecret: tinyconductor-node-secret tls: enabled: true existingSecret: tinyconductor-internal-tlsEvery pod then presents that certificate and accepts only peers with a
certificate from the same CA. Enable or disable it for all pods at once
(helm upgrade); a cell where some nodes use TLS and others do not cannot
forward between them.
Changing the node secret
Section titled “Changing the node secret”- Create a Secret with a new node secret. Set
forwarding.nodeSecret.existingSecretto it andforwarding.previousNodeSecret.existingSecretto the old one, and runhelm upgrade. The pods restart one by one and keep forwarding: a pod on the new secret accepts the old one, and signs again with the old one when a pod not yet restarted refuses the new one. - Once the rollout is complete, remove
forwarding.previousNodeSecret.existingSecretand runhelm upgradeagain. Then delete the old Secret.
Key values
Section titled “Key values”| Value | Default | Meaning |
|---|---|---|
postgres.bundled | true | Set it to false for Cluster mode |
postgres.host, postgres.port, postgres.dbName, postgres.user | (required), 5432, (required), (required) | The PostgreSQL server, the engine’s database and the engine’s login, which the nodes connect as |
postgres.password.existingSecret, postgres.password.key | (required), password | The Secret with the engine’s login’s password |
postgres.owner.user, postgres.owner.password.existingSecret, postgres.owner.password.key | none, none, password | The owner login and the Secret with its password. Only the migration Job uses them. Empty: the Job connects as postgres.user, which then must own the database |
postgres.tlsMode | verify-full | TLS to the server: disable, prefer, require, verify-ca or verify-full. verify-full checks the server’s certificate and name. The engine refuses a weaker mode for a server on another host unless postgres.tlsAllowUnverified is true (test setups only) |
postgres.autoMigrate | the engine’s own default | true lets a node apply pending schema steps itself at start instead of relying on the migrate Job; meant for a single node |
postgres.ca.existingSecret, postgres.ca.key | , ca.crt | The Secret with the certificate authority (PEM) that signed PostgreSQL’s server certificate, mounted as a file. Or postgres.caFile: the path of a file you mount yourself with extraVolumes |
postgres.poolSize, postgres.connectTimeoutSeconds, postgres.statementTimeoutSeconds | 16, half the partition lease (5 s), 30 | Connections in each node’s pool (6 to 1000), the longest wait for a connection (at most half the partition lease), and how long one statement may run (0: no bound; schema updates, retention and maintenance are not bounded) |
migrate.enabled | true | Runs tinyconductor migrate as a Job before every install and upgrade |
migrate.backoffLimit, migrate.activeDeadlineSeconds, migrate.resources | 1, 900, 100m CPU and 128 MiB requested | How often the Job retries, how long it may run, and its container’s resources |
publicUrl | none | Where people and programs reach the cell, as host or host:port without https://, usually the ingress host |
replicaCount | 1 | Engine nodes of this cell |
cell.id | 1 | This cell’s id, 1 to 8191 |
forwarding.nodeSecret.existingSecret, forwarding.nodeSecret.key | (required), secret | The Secret with the node secret: every request forwarded between the nodes is signed with a key made from it, and the secret itself is never sent |
forwarding.previousNodeSecret.existingSecret, forwarding.previousNodeSecret.key | none, secret | While you change the node secret: a Secret with the previous one, still accepted (Changing the node secret) |
forwarding.port | 9600 | The internal port the nodes forward requests on. Only this release’s pods may reach it |
forwarding.tls.enabled, forwarding.tls.existingSecret | false, none | Mutual TLS on the internal port, with the node certificate, its key and the CA from a Secret (keys forwarding.tls.certKey, keyKey, caKey: tls.crt, tls.key, ca.crt). Without it, forwarded requests travel unencrypted inside the cluster network (Internal TLS) |
agentUsage.existingSecret, agentUsage.key | (required), secret | The Secret with the agent usage key |
tenants.initialTenant | default | The one tenant a new cell starts with. You add and change tenants afterwards in the console or through the tenant API (Tenants) |
tenants.config | none | Optional: a few tenants to seed on the very first start only, as a list. A tenant without cellId gets cell.id; one that names another cell stops the install. Provision tenants through the console or the tenant API |
tenants.existingConfigMap, tenants.existingConfigMapKey | A ConfigMap you manage that holds that starting list instead | |
auth.credentials.existingSecret | Static tokens, for bootstrap and tests. Cluster mode needs these, single sign-on or local users to start | |
oidc.* | off | Single sign-on for people: enabled, issuer, clientId, clientSecret.existingSecret, redirectUri (https://<host>/auth/callback), audience, scopes, displayName, alias (the provider’s short name in usernames such as alice@corp, default sso), usernameClaim, displayNameClaim, tenants (see Who a person is) |
oidc.providers | none | Further identity providers, each with its own alias, issuer, clientId, clientSecret.existingSecret and the same settings as the first provider (Several identity providers) |
auth.defaultProvider | The identity provider whose people a plain username such as alice names: local or the OIDC alias. Required when both single sign-on and local users are on | |
auth.mappingRules, auth.roles | none | Rules from identity-provider claims to groups and roles, and extra roles (Identity and access) |
auth.localUsers.* | off | Local users; their password hashes come from a Secret |
auth.init.existingSecret | A Secret with a seeding document for users, groups, roles, grants and clients, mounted as a file | |
auth.tokenIssuer.issuer | The public URL that turns on the engine’s own token endpoint for SDKs and connectors | |
auth.tokenIssuer.keyEncryptionKey.existingSecret, .key | , key | The Secret with 32 random bytes, base64-encoded, that encrypt the stored token signing keys. Required in Cluster mode with auth.tokenIssuer.issuer; in Bundled mode the engine generates one on its volume when it is empty |
auth.session.insecureCookie | false | true drops Secure from the session cookie. Loopback development over plain HTTP only |
auth.auditRetentionDays | 30 | Days a refused request stays in the audit log (Audit log API) |
partitions.* | the engine’s own defaults | The partition engine, for a new database only: count (24), leaseMs (10000), maxClockSkewMs (1000), commitBatch (256), snapshotIntervalMs (5000), dedupRetention (24h), claimWindow (10m), claimFilterFalsePositive (0.01), claimFilterRotation (1h), exporterBatch (256), exporterMaxLag (1000000), exporterMaxLagAgeMs (60000) (How the engine runs your processes) |
maxNameFieldLength | 256 | Longest id, name, job type and resource name a deployment may use, and longest message name and correlation key a running instance may evaluate |
boundaryEventCorrelationInActivityScope | false | true evaluates a message boundary event’s correlation key in the activity’s own scope instead of the scope around it (Compatibility) |
logging.level, logging.format | info, json | The log level and line format. logLevel (RUST_LOG), a full filter, wins over logging.level when set |
metrics.tenantLabelLimit | 100 | Distinct tenants that keep their own metric label before the rest fold into other |
settings | none | Any other non-secret setting this chart does not model yet. Secrets and the settings the chart manages are refused here |
otel.* | no export | OpenTelemetry export (see Observability) |
resources | 1 CPU and 1 GiB requested; 2 GiB limit | The Small size of a node in Hardware requirements; raise it for Medium and Large loads |
podDisruptionBudget.maxUnavailable | 1 | With two or more nodes, at most one node is drained at a time |
topologySpread | spread over hosts and zones | Keeps the nodes on different machines where it can |
service, ingress, httpRoute | a ClusterIP Service | How the API is exposed (Exposure) |
networkPolicy.* | on | Who may connect, and where the nodes may connect to |
metrics.enabled, metrics.port | on, 9090 | /metrics on a port of its own and a Service of its own (<release>-tinyconductor-metrics, port metrics), without a credential; the ingress never routes it |
networkPolicy.metrics.scraperNamespace, networkPolicy.metrics.from | observability, none | The only namespace (and further peers) the NetworkPolicy admits to the metrics port |
metrics.vmServiceScrape, metrics.serviceMonitor, metrics.prometheusRule | off | Scrape objects for the VictoriaMetrics operator and the Prometheus Operator, and the alert rules (Observability) |
preStopSleepSeconds | 5 | A pause before shutdown, so Services and load balancers stop sending new requests first. It must be shorter than terminationGracePeriodSeconds |
documents.store | none | The document store: memory, local (Bundled only, on a volume of its own, documents.local.size, default 5Gi) or s3. Without one the document operations answer 501. With more than one replica only s3 is accepted (Documents) |
documents.s3.* | bucket (required), prefix, region, endpoint (for MinIO and other S3-compatible services) and forcePathStyle | |
documents.s3.credentials.existingSecret | A Secret with the access key id and secret access key (keys accessKeyId and secretAccessKey). Leave it empty to use the pod’s own cloud identity through serviceAccount.annotations | |
documents.maxUploadBytes, documents.defaultTtl, documents.sweepInterval, documents.linkBaseUrl | 10 MiB, none, 60s, the request’s host | The upload limit, the default expiry, how often expired documents are deleted, and the public address of engine links |
documents.linkSecret.existingSecret | A Secret with the key that signs links of the memory and local stores; without it links end when the pod restarts |
Every value is described in the chart’s values.yaml.
How a node starts
Section titled “How a node starts”Each node takes its name (CONDUCTOR_CLUSTER_NODE_NAME) from its pod name (cell-1-tinyconductor-0,
cell-1-tinyconductor-1 and so on), and publishes its address for forwarding
through the headless Service. On startup a node checks that the database
schema is the one of its release, loads the cell’s partition map and only
then reports ready. A node never changes the schema: the migration Job did
that before the nodes were created or rolled. A node that finds the schema
not migrated for its release does not start, and its log says to run
tinyconductor migrate.
All nodes start together. They share the partitions between them. If a start fails, the node’s log names the step that failed and the database’s reason, for example a missing right.
Health checks use the engine’s own endpoints:
- a startup check on
/readyzgives the first start up to ten minutes; - the liveness check uses
/livez, which answers while the process runs and never reads the database, so a database outage never restarts a pod; - the readiness check uses
/readyz, which answers200once the node knows which node owns which partition and while PostgreSQL answers within 2 seconds. From then on it answers every request, itself or by forwarding it; while the database does not answer, the pod leaves the Service and comes back by itself.
To check a configuration before you roll it out, run the image’s
config check with the same settings, for example as a one-off pod or a CI
step: tinyconductor config check --mode cluster reads every setting as a
start would, without connecting anywhere, and exits 1 with the reason when
one is wrong.
The engine reads its settings once, at startup. A change to the chart’s values rolls the nodes. After you change a Secret the chart only refers to, restart the nodes yourself:
kubectl -n tinyconductor rollout restart statefulset/tinyconductorThe script worker
Section titled “The script worker”Point the worker at the engine’s Service and give it an API client of its
own: create the client in the console (Access → Clients) or with
POST /v2/clients, give it the job permissions it needs, and let the worker
fetch its tokens from the engine’s token endpoint (or from your identity
provider).
kubectl -n tinyconductor create secret generic script-worker-client \ --from-literal=client-secret='<the client secret>'helm install script-worker ./tinyconductor-script-worker-0.1.0.tgz --namespace tinyconductor \ --set engine.url=http://tinyconductor:8080 \ --set auth.oauth.tokenUrl=http://tinyconductor:8080/oauth/token \ --set auth.oauth.clientId=script-worker \ --set auth.oauth.clientSecret.existingSecret=script-worker-clientIn a test environment a static token from a Secret works too
(auth.token.existingSecret).
| Value | Default | Meaning |
|---|---|---|
engine.url | (required) | The engine’s base URL |
engine.tenantId | The tenant to work for, when the credential may act in several | |
auth.token.existingSecret | A Secret with a static token, mounted as a file; for tests | |
auth.oauth.tokenUrl, clientId, clientSecret.existingSecret | The client-credentials grant: the worker’s API client | |
workerName | the worker’s own default | Worker name reported on activation |
replicaCount | 2 | Workers share nothing; add replicas for more job throughput |
concurrency | the CPU count | Jobs in flight per replica |
jobTimeoutMs, requestTimeoutMs, pollIntervalMs | 60000, 20000, 250 | The job lease and long-poll duration requested on activation, and the pause after an empty or failed poll |
retryBackoff | PT5S | Default failure backoff, ISO-8601 or milliseconds |
script.backend | quickjs | quickjs or boa |
script.cacheSize | 256 | Prepared-script handles cached, by SHA-256 of the source |
script.timeoutMs, script.instructionBudget, script.memoryLimitBytes, script.maxStackBytes, script.recursionLimit | 5000, 100000000, 67108864, 524288, 400 | The sandbox limits (Sandbox) |
script.numberPolicy | string | How unsafe integers are passed: string, bigint, reject or lossy |
shutdownGraceMs | 30000 | How long a stopping worker waits for its jobs; the pod’s grace period follows it |
logging.format | json | json or text. logLevel (RUST_LOG) sets the level |
metrics.expose, metrics.token.existingSecret | off | Serve /metrics to your scraper; it then needs a token |
metrics.vmServiceScrape.enabled, metrics.serviceMonitor.enabled | off | Scrape objects for the VictoriaMetrics operator and the Prometheus Operator, with that token |
settings | none | Any other CONDUCTOR_SCRIPT_WORKER_* setting this chart does not model yet |
Installing the pieces together
Section titled “Installing the pieces together”There is no umbrella chart: the engine and its workers have their own lifecycles, and you usually upgrade them separately. Install them in this order, into one namespace:
- the database (Cluster mode only), and its Secret;
tinyconductor;tinyconductor-script-worker, withengine.urlset to the engine’s Service,http://<release>-tinyconductor:8080(orhttp://tinyconductor:8080when the release itself is calledtinyconductor).
Workers in the same namespace can reach the engine through the default
NetworkPolicy. For a worker in another namespace, add that namespace to
networkPolicy.ingress.from of the engine chart.
Exposure
Section titled “Exposure”The engine listens on port 8080 behind a ClusterIP Service. To reach it from outside the cluster, turn on one of:
ingress.enabled, with youringress.className, hosts and TLS;httpRoute.enabled, a Gateway API route, with theparentRefsof your Gateway.
Then allow the ingress controller or gateway to reach the pods, for example:
networkPolicy: ingress: from: - namespaceSelector: matchLabels: kubernetes.io/metadata.name: ingress-nginxServe the engine over HTTPS only. Sign-in cookies are marked secure, and
the OIDC redirect URI must be https://<host>/auth/callback.
By default the pods may connect out only to DNS, to HTTPS (the identity
provider), when export is on to the OpenTelemetry collector, and in Cluster
mode to PostgreSQL’s port (narrow it with networkPolicy.egress.databaseTo).
Add more with networkPolicy.egress.extra, or turn the egress rules off
with networkPolicy.egress.enabled: false. An S3 document store on AWS is
reached over HTTPS; a MinIO or other S3-compatible service on another port
needs an extra rule for that port. Health checks from Kubernetes
itself are not affected by these rules.
Scaling
Section titled “Scaling”Bundled mode runs one pod and cannot be scaled out. Give it faster cores and faster storage; when one team outgrows it, move to Cluster mode.
Cluster mode. The nodes share the cell’s partitions, and any node
answers any request of its cell. To add nodes, raise replicaCount: the new
nodes take over their share of the partitions, one partition per second.
Make sure the database accepts about 20 more connections per node first
(Connections).
PostgreSQL is the part every node shares; when it is busy, more nodes add
load instead of capacity, and the next step is another cell with its own
database. See Performance and scaling.
The chart has no autoscaler, on purpose:
- more nodes do not always mean more throughput, because every node works against the same database, and CPU use on the nodes says little about how busy the database is;
- one busy tenant runs on one partition, so more nodes do not make it faster; give it a partition of its own instead (see Moving a busy tenant).
Change replicaCount deliberately, with the database’s capacity in view. A
node that is removed hands its partitions over before it stops.
Script workers share nothing. Scale them freely with the job load.
Upgrades
Section titled “Upgrades”-
Back up the database first. Every upgrade may update the schema, and an older engine refuses a database a newer one has migrated.
-
Download the chart of the new release, and upgrade with it. In Cluster mode the upgrade first runs the new release’s migration Job; the nodes roll only when it succeeded:
Terminal window helm upgrade tinyconductor ./tinyconductor-<new version>.tgz --namespace tinyconductor -f values.yaml -
Watch the rollout and the readiness of every pod.
The StatefulSet replaces its pods one at a time, from the highest number down. In Bundled mode that means the one pod stops and starts again; the engine loads its saved state before it reports ready.
In Cluster mode a change of settings rolls through the nodes one at a time: each stopping node hands its partitions over before it stops. A rolling upgrade to a new engine version, with old and new versions side by side, is not supported yet. Upgrade to a new version with a short stop instead: scale the nodes to zero, then upgrade, which migrates the database and then brings the nodes back at the new version.
kubectl -n tinyconductor scale statefulset/tinyconductor --replicas=0helm upgrade tinyconductor ./tinyconductor-<new version>.tgz --namespace tinyconductor -f values.yamlObservability
Section titled “Observability”- Metrics. The engine serves
/metricson its metrics port (9090,metrics.port), on the Service<release>-tinyconductor-metrics, without a credential: the port is never routed by the ingress, and the NetworkPolicy admits onlynetworkPolicy.metrics.scraperNamespace(observabilityby default). The policy is what protects the port, so keepnetworkPolicy.enabledon, on a cluster whose network plugin enforces NetworkPolicies. With the VictoriaMetrics operator setmetrics.vmServiceScrape.enabled: true; with the Prometheus Operator,metrics.serviceMonitor.enabled: true. Add the labels your scraper selects with.labelsof either. - Alerts and service-level objectives.
metrics.prometheusRule.enabled: trueinstalls the alert rules and the service-level-objective rules as a PrometheusRule. The same rules are in theobservability/folder of the examples download of the release (tinyconductor-0.1.0-examples.tar.gz). - OpenTelemetry. Set
otel.endpointto your collector, for examplehttp://otel-collector.observability:4318withotel.insecure: trueon a private network, or anhttps://address. Put an authentication header in a Secret and name it inotel.headers.existingSecret. Every pod adds its pod and namespace name to its telemetry. The collector gateway example in the examples download (observability/otel-collector-gateway.yaml) forwards traces to Tempo, metrics to a Prometheus-compatible store and logs to Loki. - Dashboards. Import
observability/grafana-dashboard.jsonfrom the examples download.
The script worker chart has its own metrics.vmServiceScrape and
metrics.serviceMonitor, which scrape with the worker’s metrics token, and
the same otel settings.
Known limits
Section titled “Known limits”- A crashed node’s tenants wait for its leases. When a node stops without
warning, the tenants on its partitions answer
503withRetry-Afteruntil the leases run out and another node takes the partitions over: about 10 to 12 seconds with the defaultCONDUCTOR_PARTITION_LEASE_MS. A clean stop and a restart in place hand over in seconds. - Rolling upgrades across engine versions are not supported yet (see Upgrades).
See also How the engine runs your processes.