Observability
This page describes what TinyConductor exposes for monitoring: the Prometheus metrics, the health endpoints, the logs, OpenTelemetry (OTLP) export of traces, metrics and logs, how one trace follows a process instance from request to worker, service-level objectives, and the audit and reporting sources. It ends with the known limitations.
Two terms appear often below:
- A partition is one engine thread with the tenants placed on it; see How the engine runs your processes.
- The exporter copies saved records into the query tables that searches, history and reports read. Export lag is how far those tables are behind the engine.
Ready-to-use files are in the observability/ folder of the examples
download of the release, tinyconductor-0.1.0-examples.tar.gz:
| File | What it is |
|---|---|
prometheus-scrape.yml | Scrape jobs for the engine and the script worker, each with its bearer token |
prometheus-alerts.yml | Alert rules for partitions and leases, export lag, forwarding, acknowledgement latency, incidents and the script worker |
prometheus-slo-rules.yml | Recording rules and burn-rate alerts for the service-level objectives |
grafana-dashboard.json | A Grafana dashboard over the metrics below, with data source and tenant variables, exemplars on the latency panels, and an API server and SLO row |
kubernetes-servicemonitor.yaml | Services, ServiceMonitors and a PrometheusRule for the Prometheus Operator on Kubernetes |
otel-collector.yaml | An OpenTelemetry Collector configuration that receives OTLP over gRPC and HTTP and writes it to files, for trying things out |
otel-collector-gateway.yaml | A production collector that forwards traces to Tempo, metrics by remote write and logs to Loki, with tail sampling |
docker-compose.otel.yml | A Compose service that runs the first collector next to TinyConductor |
Every metric and label these files use is checked by the test suite against what the engine and the script worker really expose, so a dashboard panel or an alert cannot silently point at a metric that does not exist.
At a glance
Section titled “At a glance”| Embedded | Bundled | Cluster | |
|---|---|---|---|
GET /metrics on the metrics port (9090) | No metrics | The full metric set | The full metric set |
GET /livez, GET /healthz, GET /readyz | Yes | Yes | Yes |
GET /api/v1/console/health | Yes | Yes | Yes |
GET /api/v1/console/health/background | kept: false | Yes | Yes, for the node that answers |
OpenMetrics with exemplars on /metrics | No metrics | Yes | Yes |
| JSON logs on stdout, with trace ids | Yes | Yes | Yes |
OTLP export (OTEL_* variables) | Traces and logs | Traces, metrics and logs | Traces, metrics and logs |
| One trace from the request to the job worker | Yes | Yes | Yes |
| … across message correlation and call activities | Yes | Yes | Yes |
| … across timers | A timer firing starts a new trace | The same | The same |
| Spans for applying, saving and exporting in a partition | No | Yes | Yes, and for forwarding between nodes |
| Where the trace context of an instance is kept | In memory, lost on restart | In the saved frames, kept across a restart | In the saved frames, kept across a restart, a failover and a tenant move |
| Audit log API | Yes | Yes | Yes |
Metrics
Section titled “Metrics”Scraping /metrics
Section titled “Scraping /metrics”Bundled and Cluster mode serve GET /metrics on a port of their own,
9090 (CONDUCTOR_SERVER_METRICS_PORT), on the same address as the API
(CONDUCTOR_SERVER_LISTEN_HOST). The port serves nothing else and needs no
credential; the API port serves no /metrics. Metric labels carry tenant
ids, so the port is for your metrics scraper only: do not publish it
outside the network the scraper runs in, never route it through an ingress,
and limit who reaches it with a firewall rule or, on Kubernetes, the chart’s
NetworkPolicy, which admits only the scraper’s namespace
(Kubernetes). CONDUCTOR_SERVER_METRICS_ENABLED=false
turns the port off. Embedded mode keeps no metrics.
/metrics serves the Prometheus text format (text/plain; version=0.0.4)
by default. A scraper that lists application/openmetrics-text in its
Accept header gets the OpenMetrics 1.0 format instead, which also carries
exemplars. Prometheus asks for OpenMetrics on its own when
exemplar storage is on (--enable-feature=exemplar-storage).
scrape_configs: - job_name: tinyconductor scrape_interval: 15s static_configs: - targets: ["tinyconductor.internal:9090"]Check it by hand on the engine’s host:
curl -s localhost:9090/metrics | grep '^# TYPE'A series appears once its value has been recorded at least once. For example, the timer-skew histogram appears after the first timer fires, and the command histogram appears after the first command. A gauge series disappears again when what it describes goes away, for example a job type with no available or activated jobs.
The same metrics can also be pushed over OTLP; see OpenTelemetry (OTLP).
Naming
Section titled “Naming”Inside the engine, metric and attribute names are dot-separated
OpenTelemetry names (bpm.engine.loop.iteration.duration, unit s). On
/metrics they follow the OpenTelemetry-to-Prometheus compatibility rules:
- dots become underscores (
bpm.tenant.idbecomesbpm_tenant_id); - the unit becomes a suffix: seconds
_seconds, bytes_bytes, and a unit-less ratio gauge_ratio; - monotonic counters end in
_total; - labels are written in name order;
- the process’s OpenTelemetry resource appears once as
target_info{service_name="tinyconductor",service_version=…,tinyconductor_mode=…} 1.
Durations are in seconds. Over OTLP the instruments keep their dotted names and units.
Labels
Section titled “Labels”| Label | Values |
|---|---|
bpm_tenant_id | Tenant id, or other beyond the tenant label limit (see Label cardinality) |
tenant | The same on the tinyconductor_* metrics |
command | On tinyconductor_command_seconds: create, complete-job, publish-message, deploy or other |
result | On tinyconductor_job_activation_wait_seconds: jobs (the activation handed out at least one job) or empty (it waited and found none) |
outcome | On tinyconductor_command_seconds and tinyconductor_job_activation_wait_seconds: ok, rejected (refused by the engine or a quota), unavailable (not applied; retry) or in_doubt (the answer was lost; retry with the same idempotency key). On tinyconductor_forwarded_requests_total and tinyconductor_tenant_moves_total: see the table below |
partition | A partition number |
cause | How a node took a partition: initial, reclaim, takeover or handover |
bpm_incident_kind | model (raised by the process, resolvable in the console) or engine |
bpm_element_intent | ELEMENT_ACTIVATING, ELEMENT_ACTIVATED, ELEMENT_COMPLETING, ELEMENT_COMPLETED, ELEMENT_TERMINATING, ELEMENT_TERMINATED |
bpm_job_type | The job type, or other beyond 256 distinct job types |
bpm_metric_label | On the overflow counter: which label was folded, bpm.tenant.id or bpm.job.type |
http_request_method | The request method: GET, POST, PUT, PATCH, DELETE, HEAD, OPTIONS, CONNECT, TRACE, or _OTHER for anything else |
http_route | The route template the request matched, for example /v2/process-instances/{processInstanceKey}, never the path with its keys. Absent when no route matched |
http_response_status_code | The response status code, for example 200 or 404 |
No metric carries instance, element, job or other entity keys, message names, variable values, user subjects or error text.
Engine metrics
Section titled “Engine metrics”Bundled and Cluster mode publish the same metrics.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tinyconductor_command_seconds | histogram | command, outcome | From the arrival of a request that changes something to its answer, after the change is saved. Job activations are not counted here |
tinyconductor_job_activation_wait_seconds | histogram | result, outcome | From the arrival of a job activation to its answer. A worker’s activation waits for work up to its request timeout, so a long time here with result="empty" is an idle worker, not a slow engine |
bpm_engine_element_transitions_total | counter | tenant, element intent | Element lifecycle transitions |
bpm_engine_jobs_available | gauge | tenant, job type | Jobs ready to be activated |
bpm_engine_jobs_activated | gauge | tenant, job type | Jobs handed to workers and not yet finished |
bpm_engine_timer_firing_skew_seconds | histogram | tenant | How late a timer fired, from its due instant to the firing |
bpm_engine_incidents_open | gauge | tenant, incident kind | Open incidents |
bpm_metrics_cardinality_overflow_total | counter | metric label | Measurements whose tenant or job-type label was folded into other |
http_server_request_duration_seconds | histogram | method, route, status code | Time the API server took to answer a request, per route template |
target_info | gauge | resource attributes | Always 1; the labels identify the process |
The partition, lease and exporter metrics are listed below.
The engine does not know how many jobs a worker fleet can run at once, so it
publishes the job backlog (available and activated jobs) rather than a
saturation ratio. Workers publish their own in-flight gauges, for example
bpm_script_worker_jobs_in_flight.
Histogram buckets, in seconds:
- command time: 0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10;
- timer skew: 0.01, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30, 60;
- HTTP request duration: 0.005, 0.01, 0.025, 0.05, 0.075, 0.1, 0.25, 0.5, 0.75, 1, 2.5, 5, 7.5, 10.
Useful queries:
# p95 time from request to acknowledgement, per command (job activations,# which wait for work, are measured apart and do not raise it)histogram_quantile(0.95, sum by (le, command) (rate(tinyconductor_command_seconds_bucket{outcome="ok"}[5m])))
# Job activations per second that waited and found no worksum(rate(tinyconductor_job_activation_wait_seconds_count{result="empty"}[5m]))
# Commands that could not be applied, per secondsum by (command, outcome) (rate(tinyconductor_command_seconds_count{outcome=~"unavailable|in_doubt"}[5m]))
# Element transitions per second, per tenantsum by (bpm_tenant_id) (rate(bpm_engine_element_transitions_total[5m]))
# Job backlog per job typesum by (bpm_job_type) (bpm_engine_jobs_available)
# Slowest API routes (p95)topk(5, histogram_quantile(0.95, sum by (le, http_route) (rate(http_server_request_duration_seconds_bucket[5m]))))
# Share of API requests answered with a server errorsum(rate(http_server_request_duration_seconds_count{http_response_status_code=~"5.."}[5m])) / sum(rate(http_server_request_duration_seconds_count[5m]))Partitions, leases and the exporter
Section titled “Partitions, leases and the exporter”These metrics describe the partition engine, explained in How the engine runs your processes. Bundled and Cluster mode publish them. In Cluster mode every node reports its own values; sum or compare them across nodes.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
tinyconductor_partition_owned | gauge | partition | 1 on the engine that owns the partition, 0 on every other engine. Every engine reports every partition, so a partition that nobody owns, or that two engines own, shows up in a sum |
tinyconductor_partition_lease_epoch | gauge | partition | The number of the lease this engine holds on the partition. It goes up by one each time the partition changes owner, also when an engine restarts and takes its partitions back |
tinyconductor_partition_lease_renew_failures_total | counter | Lease renewals that failed: PostgreSQL could not be reached, or a lease had been lost | |
tinyconductor_clock_offset_seconds | gauge | This node’s clock minus the database’s clock. When its size exceeds CONDUCTOR_MAX_CLOCK_SKEW_MS, the node gives up its partitions and reports not ready until the clock is back in range | |
tinyconductor_partition_takeovers_total | counter | cause | Partitions this engine took: initial (never owned before), reclaim (its own, after a restart), takeover (after another owner stopped renewing) or handover (released by another owner) |
tinyconductor_partition_commit_seconds | histogram | partition | One save to PostgreSQL, from start to commit |
tinyconductor_partition_commit_batch_frames | histogram | partition | Frames saved together in one commit |
tinyconductor_tenant_queue_wait_seconds | histogram | tenant | How long a command waited for its tenant’s turn on its partition |
tinyconductor_tenant_engine_seconds_total | counter | tenant | Processing time the tenant’s commands used |
tinyconductor_tenant_quarantined | gauge | tenant | 1 while the tenant is quarantined |
tinyconductor_exporter_lag_frames | gauge | tenant | Saved changes of the tenant that the exporter has not read yet |
tinyconductor_exporter_lag_records | gauge | tenant | Saved records of the tenant that are not in the query tables yet, including those in changes the exporter has not read yet: the whole backlog |
tinyconductor_exporter_lag_seconds | gauge | tenant | Age of the oldest part of that backlog; 0 when caught up |
tinyconductor_pruned_frames_total | counter | partition | Saved changes removed from the change log because every tenant’s saved state and the query tables already cover them |
tinyconductor_expired_claims_total | counter | partition | Idempotency keys forgotten after CONDUCTOR_DEDUP_RETENTION (removed with the hour in which they expired) |
tinyconductor_retention_purged_records_total | counter | tenant | Records of finished instances removed after the tenant’s history retention, one by one or with their day (the instances’ searchable history goes with them) |
tinyconductor_retention_moved_rows_total | counter | tenant | Rows of history set aside when their day was removed because they are still needed (running instances, recently finished ones, definitions, live messages); each row at most once |
tinyconductor_exporter_quarantined | gauge | tenant | 1 while the exporter has stopped for the tenant in a partition, because one of the tenant’s saved changes cannot be copied into the query tables. The tenant’s other changes stay saved, and its new instances are refused. See When the export of a tenant stops |
tinyconductor_exporter_backpressure_total | counter | tenant | The tenant’s commands that were slowed or refused because its query tables were too far behind |
tinyconductor_forwarded_requests_total | counter | outcome | Requests this node forwarded to the owner of a partition: ok, partition_moved, unavailable, unreachable, unauthorized (the owner refused the forwarded request: the nodes’ secrets differ, the clocks are more than two minutes apart, or the caller may not act in the tenant), answer_lost, or unencodable (this node could not encode the request, so nothing was sent; a fault in the engine) |
tinyconductor_tenant_moves_total | counter | outcome | Tenant moves that ended: done or failed. A move that waits for a partition does not count as failed: it goes on by itself |
tinyconductor_tenant_move_retries_total | counter | — | Attempts of a tenant move’s next step that had to wait, for example because the partition the tenant moves into was not serving yet. The node tries again after 1, 2, 4, then at most 8 seconds |
tinyconductor_tenant_move_blocked_seconds | gauge | — | How long the longest-waiting tenant move this node carries out has been waiting at its current step; 0 when none waits. While a move waits after its first step, the tenant’s requests are held and then answered 503 |
The tenant label follows the same limit as bpm_tenant_id (see
Label cardinality).
Buckets:
- commit time: 0.0005, 0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5 seconds;
- frames per commit: 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024;
- queue wait: 0.0001, 0.00025, 0.0005, 0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 1 seconds.
Useful queries:
# Partitions without an owner (should be empty)sum by (partition) (tinyconductor_partition_owned) == 0
# Partitions per enginesum by (instance) (tinyconductor_partition_owned)
# p95 save time, and frames per savehistogram_quantile(0.95, sum by (le) (rate(tinyconductor_partition_commit_seconds_bucket[5m])))histogram_quantile(0.5, sum by (le) (rate(tinyconductor_partition_commit_batch_frames_bucket[5m])))
# p99 wait for a tenant's turnhistogram_quantile(0.99, sum by (le, tenant) (rate(tinyconductor_tenant_queue_wait_seconds_bucket[5m])))
# Share of processing time per tenantsum by (tenant) (rate(tinyconductor_tenant_engine_seconds_total[5m])) / scalar(sum(rate(tinyconductor_tenant_engine_seconds_total[5m])))
# Tenants whose searches are more than 10 seconds behindtinyconductor_exporter_lag_seconds > 10Documents
Section titled “Documents”| Metric | Type | Labels | Meaning |
|---|---|---|---|
tinyconductor_documents_store_errors_total | counter | operation, store | Failures of the document store: operation is upload, download, delete, link or sweep, store is memory, local or s3. The log line of each failure names its cause. See Documents |
Bundled and Cluster mode publish it once the first failure happens.
Exemplars
Section titled “Exemplars”An exemplar is one sample observation attached to a histogram bucket, together with the trace it came from. In a Grafana latency panel it appears as a dot you can click to open that request’s trace.
Every latency histogram keeps the latest traced observation of each bucket:
acknowledgement and pipeline stages, loop iteration and checkpoint commit,
timer skew, and HTTP request duration. An observation counts only when it
was made inside a sampled trace (see OTEL_TRACES_SAMPLER below). The
exemplar carries trace_id and span_id:
http_server_request_duration_seconds_bucket{http_request_method="POST",http_response_status_code="200",http_route="/v2/process-instances",le="0.025"} 41 # {trace_id="4bf92f3577b34da6a3ce929d0e0e4736",span_id="00f067aa0ba902b7"} 0.0183Exemplars appear only in the OpenMetrics format; the Prometheus text format
has no place for them. To use them, scrape with OpenMetrics (Prometheus with
exemplar-storage does), and in the Grafana Prometheus data source add an
exemplar link to your trace back end on the trace_id label.
The engine does not measure the PostgreSQL server itself. Run a PostgreSQL exporter next to it for that.
Label cardinality
Section titled “Label cardinality”Every label value is a new time series in your metrics back end. Tenant ids are bounded by configuration, but a large Cluster can still have many:
| Variable | Default | Effect |
|---|---|---|
CONDUCTOR_METRICS_TENANT_LABEL_LIMIT | 100 | The first tenants seen keep their own bpm_tenant_id value, up to this many; later tenants are reported as other. 0 reports every tenant as other, which aggregates the whole cell. A value that is not a non-negative integer stops the engine at startup |
Job types are capped the same way at 256 distinct values. Every folded
measurement increments bpm_metrics_cardinality_overflow_total, so a rising
counter tells you the cap is in use. Counters and histograms of folded
tenants add up under other; a gauge under other shows the value most
recently recorded for any folded tenant. The limit does not apply to trace
attributes: spans are sampled and not aggregated.
Export lag alarms
Section titled “Export lag alarms”The exporter fills the query tables that searches, history and the analytics
views read. It keeps track of each tenant separately. When a tenant’s lag
reaches half of either limit, the engine slows that tenant’s new instances
and logs, once per change, a WARN line export lag alarm is firing; new instances of this tenant are slowed, then refused. At the limit it refuses
them. When the lag drops back, it logs an INFO line export lag alarm cleared; new instances of this tenant are admitted again. Both lines carry
the fields tenant, lag_records and lag_seconds.
| Variable | Default | Limit |
|---|---|---|
CONDUCTOR_EXPORTER_MAX_LAG | 1000000 | Records a tenant’s searches may be behind |
CONDUCTOR_EXPORTER_MAX_LAG_AGE_MS | 60000 | Age of the oldest record not yet exported, in milliseconds |
The alert rules in
prometheus-alerts.yml
warn at half of the default limits, where the slowing starts:
- alert: TinyConductorExporterLagRecords expr: max by (tenant) (tinyconductor_exporter_lag_records) > 500000 for: 5m labels: severity: warningWhen you change a limit, change the matching threshold in your rules.
Grafana dashboard
Section titled “Grafana dashboard”Import grafana-dashboard.json
(Dashboards → New → Import) and pick your Prometheus data source. It has
five rows: commands and the engine, partitions, leases and the exporter,
forwarding and tenant moves, the script worker, and the API server with the
service-level indicators. The tenant variable filters every tenant-labelled panel. The
latency panels show exemplars when Prometheus stores them.
Kubernetes
Section titled “Kubernetes”The Helm charts create the scrape objects
(a VictoriaMetrics VMServiceScrape with metrics.vmServiceScrape, a
Prometheus Operator ServiceMonitor with metrics.serviceMonitor) and the
PrometheusRule (metrics.prometheusRule) for you, with the same rules as the files in the examples download. Without
the charts, apply
kubernetes-servicemonitor.yaml
after adjusting the namespace and labels. It defines a Service and a
ServiceMonitor for the engine’s metrics port (no credential) and for the
script worker (with its metrics token read from a Secret), and a
PrometheusRule to hold the alert and SLO rules of the observability/
folder. Create the Secrets from the scrape
credentials first:
kubectl -n tinyconductor create secret generic tinyconductor-scrape --from-literal=token=<scrape-secret>kubectl -n tinyconductor create secret generic tinyconductor-script-worker-scrape --from-literal=token=<worker-token>For exemplars, turn on exemplar storage in the Prometheus resource
(enableFeatures: [exemplar-storage]).
Service-level objectives
Section titled “Service-level objectives”A service-level objective (SLO) is a target for the share of good events,
for example “99 % of accepted requests are acknowledged within 250 ms over
28 days”. The share that may be bad (here 1 %) is the error budget. The
prometheus-slo-rules.yml
file defines four indicators as Prometheus recording rules, and alerts on
how fast each budget is being spent. The targets are examples; choose the
ones your users need.
| Objective | Good event | Target | Modes | Recorded as |
|---|---|---|---|---|
| Acknowledgement latency | A command that changes something is answered within 250 ms | 99 % | Bundled, Cluster | tinyconductor:ack_latency_good:ratio_rate5m (also 30m, 1h, 6h) |
| API availability | An API request is answered without a 5xx | 99.9 % | Bundled, Cluster | tinyconductor:api_availability:ratio_rate5m (also 30m, 1h, 6h) |
| Timer punctuality | A timer fires within 1 s of its due time | 99 % | Bundled, Cluster | tinyconductor:timer_on_time:ratio_rate1h |
| Search freshness | A tenant’s searches are less than 60 s behind | 99 % of the time | Bundled, Cluster | tinyconductor:projection_fresh:ratio_avg1h |
Each indicator is good events divided by all events over a window, for example:
- record: tinyconductor:ack_latency_good:ratio_rate5m expr: | sum by (command) (rate(tinyconductor_command_seconds_bucket{outcome="ok",le="0.25"}[5m])) / sum by (command) (rate(tinyconductor_command_seconds_count{outcome="ok"}[5m]))The alerts follow the usual multi-window pattern. The burn rate is the bad share divided by the budget: at 1 the budget lasts exactly the 28 days. A page fires when the budget burns more than 14.4 times too fast over both the last hour and the last 5 minutes (2 % of the month’s budget gone in one hour), and a ticket at 6 times over 6 hours and 30 minutes:
- alert: TinyConductorAckLatencyBudgetBurnFast expr: | (1 - tinyconductor:ack_latency_good:ratio_rate1h) / (1 - 0.99) > 14.4 and (1 - tinyconductor:ack_latency_good:ratio_rate5m) / (1 - 0.99) > 14.4To change an objective, change the bucket (le="0.25") and the target
(0.99) together. A latency threshold must be one of the histogram’s bucket
boundaries listed under Engine metrics.
Script worker metrics
Section titled “Script worker metrics”The script worker serves /metrics, /healthz and /readyz on its own
listener, CONDUCTOR_SCRIPT_WORKER_METRICS_BIND (default 127.0.0.1:9464, reachable
only from the same host). Job types appear as labels and can name business
functions, so the listener is protected:
- set
CONDUCTOR_SCRIPT_WORKER_METRICS_TOKEN(orCONDUCTOR_SCRIPT_WORKER_METRICS_TOKEN_FILE) and/metricsanswers401unless the request carriesAuthorization: Bearer <token>; - a listener other hosts can reach (for example
0.0.0.0:9464in a container) is refused at startup unless a token is set; /healthzand/readyznever need the token, so container and Kubernetes probes keep working.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
bpm_script_worker_executions_total | counter | job_type, backend, outcome | Finished script jobs; outcome is completed, failed or bpmn_error |
bpm_script_worker_failures_total | counter | job_type, backend, class | Failed script jobs by class: syntax, runtime, timeout, cpu_budget, memory, stack, value, configuration, job_error |
bpm_script_worker_execution_seconds | histogram | job_type, backend | Script execution time. Buckets 0.001 to 10 s |
bpm_script_worker_api_errors_total | counter | operation | Failed calls to the engine: token, activate, complete, fail, error, client |
bpm_script_worker_jobs_in_flight | gauge | Jobs executing now | |
bpm_script_worker_script_cache_hits_total | counter | Prepared-script cache hits | |
bpm_script_worker_script_cache_misses_total | counter | Prepared-script cache misses | |
bpm_script_worker_script_cache_entries | gauge | Prepared scripts in the cache |
With OTLP metric export on, the worker pushes the same metrics as
bpm.script_worker.executions, bpm.script_worker.failures,
bpm.script_worker.execution (seconds), bpm.script_worker.api_errors,
bpm.script_worker.jobs_in_flight and bpm.script_worker.script_cache.*.
/healthz answers 200 while the process runs, and /readyz answers 200
after the first successful activation call. See
Script tasks.
Health endpoints
Section titled “Health endpoints”| Route | Credential | Embedded, Bundled | Cluster |
|---|---|---|---|
GET /livez | none | {"status":"UP"} while the process runs. It reads neither the database nor the engine: use it as the liveness probe | The same |
GET /healthz | none | {"status":"UP"|"DOWN","version":…,"mode":…,"node":…,"partitions":…,"log":…,"exporter":…}, with a reason when DOWN. In Embedded mode node, log and exporter are null and there is one partition | The same body, for this node |
GET /readyz | none | 200 when the engine can serve, 503 while it cannot. Bundled also needs PostgreSQL to answer within 2 seconds; the 503 body is the health report with the reason | 200 once the node has loaded the partition map and while PostgreSQL answers within 2 seconds, 503 with the reason otherwise |
GET /api/v1/console/health | operations:health | Yes | Yes |
GET /api/v1/console/health/background | operations:health | Embedded: {"kept": false, …}; Bundled: yes | Yes, for the node that answers |
Health reports the node’s identity, its partitions, the observed saving and
the largest exporter lag of any tenant (never which tenant). In Cluster mode
each node reports only itself: owned and loaded count the partitions this
node holds.
{"status":"UP","version":"0.1.0","mode":"bundled", "node":{"cell":1,"owner":"tinyconductor-bundled"}, "partitions":{"count":24,"owned":24,"loaded":24}, "log":{"commits":5,"frames":5,"meanBatch":1.0,"maxBatch":1}, "exporter":{"lagRecords":0,"lagSeconds":0.0}}status is DOWN, with a short reason, while a partition has stopped and
has not reloaded yet.
GET /api/v1/console/health (permission operations:health) is the summary
shown in the health area of TinyConductor Console. Its figures are for the
caller’s own tenant (scope: "tenant"), except for the cell’s
administrators: a caller who also holds tenant administration, which only
the operator tenant (the starting tenant) can hold, gets the command figures
of every tenant on the node (scope: "cell"):
{"tenantId":"default","scope":"tenant","node":"tinyconductor-bundled", "readiness":{"ready":true,"status":"UP","probes":[]}, "acknowledgement":{"samples":13,"cumulativeSeconds":0.061, "operations":[ {"operation":"create","count":12,"meanSeconds":0.004,"p50Seconds":0.005,"p95Seconds":0.01}, {"operation":"deploy","count":1,"meanSeconds":0.013,"p50Seconds":0.025,"p95Seconds":0.025}]}, "lag":{"applicable":true,"answeredBy":"tinyconductor-bundled","exportedHere":true, "pendingRecords":0,"pendingFrames":0,"oldestAgeSeconds":0.0,"stopped":null}}nodeis the node that answered;nullin Embedded mode.acknowledgementcounts the tenant’s commands that were saved, from their arrival at this node to their answer, since the node started: how many (samples), their total time, and one row per kind of command with its count, mean, and the bounds within which half (p50Seconds) and 95 % (p95Seconds) of them were answered (nullabove 10 s). The bounds are those oftinyconductor_command_seconds. Job activations are not counted: a worker’s activation waits for work by design. In Cluster mode each node counts the commands that arrived at it.- With
scope: "cell",acknowledgementsums every tenant’s commands on this node and addstenants, one row per tenant (tenantId,samples,cumulativeSeconds,meanSeconds,p50Seconds,p95Secondsand itsoperations).lagstays the caller’s own tenant. lagis where the tenant’s export to the query tables stands. On the node that owns the tenant’s partition (exportedHere),pendingRecordsandpendingFramesare the saved records and commands not exported yet,oldestAgeSecondshow long the oldest has waited, andstoppednames the partition and the reason when the tenant’s export stopped. On another node,partitionandnodeUrlname the tenant’s partition and the node that owns it (nullwhile none does). In Embedded modelagis{"applicable": false}: its searches are brought up to date before each answer.
GET /api/v1/console/configuration carries the same export position as
reporting (without applicable); the console’s analytics areas show it as
how far the reports trail the engine.
GET /api/v1/console/health/background (permission operations:health) is
what the janitor, retention, back-pressure and the export did on the node
that answers since it started, as JSON: the figures /metrics reports as
text, for people and tools that read no metrics. A service client sees the
whole node ("scope": "node"); anyone else sees only their own tenant’s rows
and none of the per-partition figures other tenants share ("scope": "tenant", with janitor.partitions, committedSeq and floor null):
{"kept":true,"scope":"node","tenantId":"default","node":"tinyconductor-bundled", "janitor":{"passes":42,"lastPassAt":"2026-09-30T10:00:00.000Z", "partitions":[{"partition":1,"expiredClaims":12,"prunedFrames":400}]}, "retention":{"dayPartitions":{"created":30,"dropped":2}, "tenants":[{"tenantId":"default","purgedRecords":90,"movedRows":4}]}, "backpressure":[{"tenantId":"default","slowed":3,"refused":1}], "export":[{"partition":1,"committedSeq":1200,"floor":1100, "tenants":[{"tenantId":"default","exportedThrough":700,"pendingRecords":5, "pendingFrames":1,"oldestAgeSeconds":1.5,"stopped":null}]}]}janitor: how many passes the node’s janitor has run and when the last one ended; per partition, the expired idempotency claims it deleted and the log frames it pruned.retention: the day partitions of history created and dropped, and per tenant the records purged and the rows kept past their day (moved into the long-running partition because their instance still ran or ended too recently).backpressure: per tenant, the creates slowed and the writes refused because its export lagged or stopped.export: per partition this node serves, the partition sequence committed and the pruning floor, and per tenant the last tenant sequence exported, what waits and since when, andstoppedwith the reason when its export stopped.
In Embedded mode, which keeps everything in memory, the answer is
{"kept": false, "scope": …, "tenantId": …, "reason": …}.
See Console for the health area itself.
The engine and the script worker write one flat JSON object per line to stdout.
CONDUCTOR_LOGGING_LEVEL sets the level (error, warn, info, debug or
trace; default info); RUST_LOG, when set, wins and takes the usual
target=level filter syntax. CONDUCTOR_LOGGING_FORMAT=text switches to one
human-readable line per event for local work; json is the default.
{"timestamp":"2026-09-23T20:52:42.255122Z","level":"warn","message":"request refused by the authentication boundary","target":"bpm::audit","tenant":"acme","trace_id":"4bf92f3577b34da6a3ce929d0e0e4736","span_id":"b7ad6b7169203331","event":"auth.refused","status":401,"method":"POST","route":"/v1/process-instances"}| Key | Meaning |
|---|---|
timestamp | RFC 3339, UTC |
level | error, warn, info, debug or trace |
message | The event text |
target | The emitting module, for example bpm_runtime::observability |
tenant | The tenant the line is about: the event’s own tenant, else the tenant of the request it belongs to. Absent when no tenant applies |
trace_id | The W3C trace id of the span the line was written in. Present whenever the line belongs to an API request or a worker job, whether or not traces are exported; absent otherwise |
span_id | The id of that span |
| any other key | A structured field of the event, for example consumer_id, mode, status. A field named like one of the keys above is written as field.<name> |
Every line has timestamp, level and message at the top level, so one
query works for every TinyBlox component, for example
jq 'select(.level == "error" and .tenant == "acme")'.
With trace_id a log search finds every line of one request, and a trace
back end that links logs (Grafana Tempo with Loki, for example) jumps from a
span to its lines.
Useful filters:
RUST_LOG=info # defaultRUST_LOG=warn,bpm::audit=warn # warnings and access refusals onlyRUST_LOG=info,sqlx::postgres::notice=warn # hide PostgreSQL notices printed while migrations run at startupLog lines can also be exported over OTLP (see below). Stdout stays on
either way; RUST_LOG filters both.
Access refusals: bpm::audit
Section titled “Access refusals: bpm::audit”Every 401 and 403 the API answers, and every 400 for a missing tenant
selection, is logged at warn on the bpm::audit target:
{"timestamp":"…","level":"warn","message":"request refused by the authentication boundary", "target":"bpm::audit","tenant":"acme","trace_id":"…","span_id":"…","event":"auth.refused", "status":403,"subject":"alice","method":"GET","route":"/v2/process-instances/search","reason":"…"}Tokens and credentials are never logged. The same refusals are served as
ACCESS entries by the audit log API.
What logs contain
Section titled “What logs contain”Logs name tenants, subjects, routes, keys and error reasons. Tokens and request bodies are not logged. An error reason can quote an expression or a value from the process, so treat logs as holding process data and apply the same access rules as to the database.
OpenTelemetry (OTLP)
Section titled “OpenTelemetry (OTLP)”The engine (every mode) and the script worker export traces, metrics and
logs over OTLP to an OpenTelemetry Collector or any OTLP back end. They are
configured only through the standard OpenTelemetry environment variables;
there are no TinyConductor-specific names. With none of them set, nothing
leaves the process and /metrics works as described above.
The quickest start is a collector next to the engine:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:4318 # every signal, http/protobuftinyconductor --mode bundledVariables
Section titled “Variables”| Variable | Default | Values and effect |
|---|---|---|
OTEL_SDK_DISABLED | false | true turns every exporter off, whatever else is set |
OTEL_SERVICE_NAME | tinyconductor (engine), bpm-script-worker | service.name of the resource |
OTEL_RESOURCE_ATTRIBUTES | none | key=value,…, values percent-encoded, for example deployment.environment.name=prod,service.namespace=bpm. A service.name here is used when OTEL_SERVICE_NAME is unset |
OTEL_TRACES_EXPORTER | otlp when an endpoint is set | otlp or none |
OTEL_METRICS_EXPORTER | otlp when an endpoint is set | otlp, prometheus, none, or otlp,prometheus. /metrics is always served; prometheus and none only turn the push off |
OTEL_LOGS_EXPORTER | otlp when an endpoint is set | otlp or none. Stdout logging stays on either way |
OTEL_EXPORTER_OTLP_ENDPOINT | none | Base URL for every signal, for example http://collector:4318. With http/protobuf the engine appends /v1/traces, /v1/metrics and /v1/logs |
OTEL_EXPORTER_OTLP_{TRACES,METRICS,LOGS}_ENDPOINT | the general endpoint | Full URL for one signal, used as is |
OTEL_EXPORTER_OTLP_PROTOCOL, …_{TRACES,METRICS,LOGS}_PROTOCOL | http/protobuf | grpc or http/protobuf |
OTEL_EXPORTER_OTLP_HEADERS, …_{TRACES,METRICS,LOGS}_HEADERS | none | name=value,…, values percent-encoded (Authorization=Basic%20…). The per-signal list replaces the general one. Header values are never logged |
OTEL_EXPORTER_OTLP_TIMEOUT, …_{TRACES,METRICS,LOGS}_TIMEOUT | 10000 | Export timeout in milliseconds |
OTEL_EXPORTER_OTLP_COMPRESSION, …_{TRACES,METRICS,LOGS}_COMPRESSION | none | gzip or none |
OTEL_EXPORTER_OTLP_CERTIFICATE, …_{TRACES,METRICS,LOGS}_CERTIFICATE | system roots | PEM file of the CA that signs the collector’s certificate |
OTEL_EXPORTER_OTLP_CLIENT_CERTIFICATE, OTEL_EXPORTER_OTLP_CLIENT_KEY (and per signal) | none | PEM client certificate and key for mutual TLS; set both |
OTEL_EXPORTER_OTLP_INSECURE, …_{TRACES,METRICS,LOGS}_INSECURE | false | true allows plaintext http:// to a collector that is not on loopback (see below) |
OTEL_TRACES_SAMPLER | parentbased_traceidratio | always_on, always_off, traceidratio, parentbased_always_on, parentbased_always_off, parentbased_traceidratio |
OTEL_TRACES_SAMPLER_ARG | 1.0 | Sampling ratio between 0 and 1 for the ratio samplers |
OTEL_METRIC_EXPORT_INTERVAL | 60000 | Milliseconds between two metric pushes |
OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE | cumulative | cumulative or delta |
OTEL_BSP_SCHEDULE_DELAY | 5000 | Milliseconds between two span exports |
OTEL_BSP_MAX_QUEUE_SIZE | 2048 | Spans buffered; beyond it spans are dropped, never waited for |
OTEL_BSP_MAX_EXPORT_BATCH_SIZE | 512 | Spans per export, at most the queue size |
OTEL_BLRP_SCHEDULE_DELAY, OTEL_BLRP_MAX_QUEUE_SIZE, OTEL_BLRP_MAX_EXPORT_BATCH_SIZE | 1000, 2048, 512 | The same bounds for log records |
OTEL_PROPAGATORS | tracecontext,baggage | tracecontext, baggage or none: which incoming headers are read |
A signal is exported when its exporter variable says otlp, or when the
variable is unset and an endpoint (general or for that signal) is set.
otlp without any endpoint sends to http://localhost:4318 (or
http://localhost:4317 with grpc).
Validation. Every variable above is checked before the process starts
anything else. An unsupported or malformed value (an exporter such as
zipkin, the protocol http/json, a sampler such as jaeger_remote, a
propagator such as b3, a negative timeout, an unreadable certificate file)
stops the engine or the worker with exit code 2 and a message that names the
variable and lists what is accepted, for example:
tinyconductor: OTEL_EXPORTER_OTLP_PROTOCOL: unsupported protocol; accepted: grpc, http/protobufOther OTEL_* variables are not supported by this build; each one present
is reported once as a WARN line at startup and otherwise ignored.
Plaintext. Spans and logs name tenants, keys and routes. Plaintext
http:// is therefore accepted only to a loopback collector
(localhost, 127.0.0.1, ::1); for any other host use https://, or set
OTEL_EXPORTER_OTLP_INSECURE=true to accept an unencrypted link on purpose
(a collector on the same private container network, for example).
Overhead. Exports run on background threads through bounded queues. A slow or unreachable collector drops telemetry; it never slows an API request or an engine turn. In Cluster mode each engine turn reads the stored trace context of its instance in its own transaction (one small indexed query), and writes one row when a timer, a message or a child instance hands the trace on. The latest measurements are in Overhead.
Resource
Section titled “Resource”Every signal carries service.name, service.version,
service.instance.id (the engine’s CONDUCTOR_CLUSTER_NODE_NAME in Cluster mode, a random
id per process otherwise, the worker’s CONDUCTOR_SCRIPT_WORKER_WORKER_NAME or
host name for the worker), tinyconductor.mode (embedded, bundled or
cluster), tinyconductor.cell.id in Cluster mode, the SDK attributes, and
whatever OTEL_RESOURCE_ATTRIBUTES adds. On /metrics the same resource is
the target_info series.
| Span | Kind | Where | Attributes |
|---|---|---|---|
{method} {route}, for example POST /v1/process-instances | server | every API request, every mode | http.request.method, http.route (the template, never the path), http.response.status_code, bpm.tenant.id; error.type on 5xx |
bpm.command.admit | internal | every command, every mode | tenant, command, whether an idempotency key was sent, outcome |
bpm.command.forward | client | Cluster: on the node that forwards a request to the partition’s owner; the owner continues the same trace | tenant, partition, owner host, outcome |
bpm.partition.apply | internal | applying the command in its partition, every mode | tenant, partition, tenant sequence, queue wait in milliseconds |
bpm.partition.commit | client | Bundled and Cluster: one save of a batch of frames. It carries no tenant and links to the apply spans of the frames it saved, rather than being their parent | partition, lease epoch, first and last sequence, frames, number of tenants |
bpm.export.batch | consumer | Bundled and Cluster: the exporter writing one tenant’s records into the query tables | tenant, partition, first and last tenant sequence, records |
bpm.timer.fire | internal | Bundled and Cluster: a timer firing; it starts a new trace | tenant, partition |
bpm.partition.claim | internal | Bundled and Cluster: a node taking a partition. It carries no tenant | partition, lease epoch, cause |
bpm.message.correlate | consumer | a published message reaching a waiting instance or starting one, a child of the publishing request, with a link to the instance’s previous trace | bpm.message.key; bpm.process.instance.key in Cluster |
bpm.job.activation | producer | each activated job of a traced instance | bpm.job.key, bpm.job.type |
bpm.job.complete, bpm.job.fail, bpm.job.throw-error | internal | job completion, failure and BPMN error requests, v1 and v2 | bpm.job.key |
script.execute | consumer | script worker, one per job | bpm.job.key, bpm.job.type, bpm.script.backend; error.type is the failure class |
A link connects two traces without making one the parent of the other. Trace back ends show it as “linked span” and let you jump across.
Spans never carry variables, request or response bodies, headers other than
the W3C ones, tokens or error messages. When nothing is exported, spans are
not recorded, but they still issue valid W3C ids, so log lines keep their
trace_id and job headers keep carrying context.
Metrics over OTLP
Section titled “Metrics over OTLP”With metric export on, the engine pushes the instruments of
Engine metrics under their OpenTelemetry names (for
example bpm.engine.loop.iteration.duration, unit s) every
OTEL_METRIC_EXPORT_INTERVAL, from the same meter that renders /metrics,
so both show the same values. Cumulative temporality is the default, which
Prometheus-compatible back ends expect. Embedded mode keeps no metrics.
Logs over OTLP
Section titled “Logs over OTLP”With log export on, every log line that passes RUST_LOG is also sent as an
OTLP log record, with its severity, target, fields and, inside a span, the
trace and span id.
Examples
Section titled “Examples”A generic collector (the one in
otel-collector.yaml, for
example), on another host of a private network:
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318OTEL_EXPORTER_OTLP_INSECURE=trueOTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=stagingStart that collector with the Compose file, from the unpacked examples download:
cd tinyconductor-0.1.0-examplesOTEL_OUT=$(mktemp -d) docker compose -f observability/docker-compose.otel.yml up -dGrafana Alloy, forwarding to Tempo, Mimir and Loki (Alloy’s
otelcol.receiver.otlp listens on 4317 and 4318), over gRPC with TLS:
OTEL_EXPORTER_OTLP_ENDPOINT=https://alloy.observability:4317OTEL_EXPORTER_OTLP_PROTOCOL=grpcOTEL_EXPORTER_OTLP_CERTIFICATE=/etc/tinyconductor/otel-ca.pemOTEL_TRACES_SAMPLER_ARG=0.1 # keep 10 % of new traces; callers' decisions winGrafana Tempo directly, traces only, metrics by Prometheus scrape:
OTEL_TRACES_EXPORTER=otlpOTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://tempo:4317OTEL_EXPORTER_OTLP_TRACES_PROTOCOL=grpcOTEL_EXPORTER_OTLP_TRACES_INSECURE=trueOTEL_METRICS_EXPORTER=prometheusOTEL_LOGS_EXPORTER=noneJaeger (v2 accepts OTLP on 4317 and 4318), traces only:
OTEL_TRACES_EXPORTER=otlpOTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://jaeger:4318/v1/tracesOTEL_EXPORTER_OTLP_TRACES_INSECURE=trueOTEL_METRICS_EXPORTER=noneOTEL_LOGS_EXPORTER=noneA hosted OTLP gateway that needs a credential:
OTEL_EXPORTER_OTLP_ENDPOINT=https://otlp-gateway.example.net/otlpOTEL_EXPORTER_OTLP_HEADERS=Authorization=Basic%20<base64 of instance:token>OTEL_EXPORTER_OTLP_COMPRESSION=gzipTrace context
Section titled “Trace context”In every mode the engine carries a W3C trace context from the request that creates or wakes an instance to the workers of that instance’s jobs:
-
Every API request continues the caller’s
traceparentandtracestateheaders (or starts a new trace) in its server span. Instance creation, throughPOST /v1/process-instancesorPOST /v2/process-instances, records its admission span under it and makes that context the instance’s context. -
Every job of that instance, activated through
/v1/jobs/activationor/v2/jobs/activation, carries the instance’s current context incustomHeaders, with the span id of its activation span:"customHeaders": {"traceparent": "00-4bf92f3577b34da6a3ce929d0e0e4736-735de69e2a1e06a1-01","bpm.traceId": "4bf92f3577b34da6a3ce929d0e0e4736","bpm.correlationId": "01a0d00b-11fc-760a-a79b-fbd789741e86"}The engine sets these keys itself. Header values of the same name in the BPMN model are overwritten, so a model cannot forge a context. The context never enters process variables, the record stream or exports.
-
The worker continues the trace, and its completion, failure or error request carries
traceparentback, so the engine’s completion span joins the same trace.
Across timers, messages and call activities
Section titled “Across timers, messages and call activities”- Timers. A timer firing starts a new trace, with a
bpm.timer.firespan; the work that follows the timer belongs to that trace. The same holds for job timeouts and message expiry. - Messages. A message publication (
POST /v1/messages/publicationor the v2 route) opensbpm.message.correlateunder the publishing request. Every instance the message reaches, a waiting catch event or a message start event, moves onto the publisher’s trace: its next jobs continue the publishing request, and the correlation span links back to the trace the instance was in before. A message that waited in the buffer continues the trace of the request that published it, not of the request that later matched it. Publish with the caller’straceparentto keep one end-to-end trace. - Call activities. A child instance continues the trace of its parent.
Where the context of an instance is kept depends on the mode:
| Mode | Where | Survives a restart |
|---|---|---|
| Bundled, Cluster | In the saved frames, next to the commands that set it | Yes, and a failover or a tenant move too: the context moves with the frames |
| Embedded | In memory, for the latest 10 000 instances | No |
The context never enters process variables, the records, the query tables or the audit log, and the records are the same with tracing on or off. The test suite checks this on the Embedded corpus.
Trace context for job workers
Section titled “Trace context for job workers”Any worker that reads custom headers as a map can continue the trace with a
stock OpenTelemetry propagator: extract traceparent and tracestate from
customHeaders, start the job’s span as their child, and send the span’s
traceparent (optionally with bpm-correlation-id, the bpm.correlationId
value) as HTTP headers on POST /v1/jobs/{jobKey}/completion,
/failure, /error or POST /v2/jobs/{jobKey}/completion. In Java, for
example:
Context parent = W3CTraceContextPropagator.getInstance() .extract(Context.root(), job.getCustomHeaders(), MAP_GETTER);Span span = tracer.spanBuilder("handle " + job.getType()) .setParent(parent).setSpanKind(SpanKind.CONSUMER).startSpan();The TinyConductor script worker does exactly this: each job runs in a
script.execute span that is a child of the job’s activation span, and its
terminal command carries the span’s context. One trace then shows the
creating request, the engine’s admission and job activation, the script
execution and the completion request.
Audit and reporting as sources
Section titled “Audit and reporting as sources”- The audit log API (
GET /v1/audit-log, permissionaudit:read) answers who did what, and lists refused requests. It is served in every mode. - The
analytics.*SQL views (Bundled and Cluster) are stable views for BI tools and for SQL-based dashboards on process throughput, durations and incidents. See Analytics. GET /api/v1/history/process-definitions/{key}/dashboardreturns the flow-node statistics and open incidents of one process definition. It also includes the conformance test results that ship with this engine release.
Overhead of tracing
Section titled “Overhead of tracing”The overhead of tracing is measured with the latency and throughput harnesses, once with nothing exported and once exporting traces, metrics and logs of every request (100 % sampling) to a local collector. The budget: tracing on may raise the 95th-percentile acknowledgement latency by at most 2 % and lower throughput by at most 3 %.
Measured on an otherwise idle test machine (2026-09-24; Apple M5 Max
host; release build; PostgreSQL 16 on tmpfs; OpenTelemetry Collector 0.161.0
on the same machine, OTLP over HTTP, OTEL_TRACES_SAMPLER=always_on). The
runs alternated: off, on, off, on.
| Measurement | Nothing exported (two runs) | Everything exported, 100 % sampling (two runs) |
|---|---|---|
| Cluster, create acknowledgement p50 / p95 (300 requests, 8 in parallel) | 121.4 / 220.2 ms; 122.5 / 259.5 ms | 112.7 / 214.7 ms; 113.1 / 217.7 ms |
| Cluster, one node, completed instances per second (30 s, 64 clients, jobs and messages) | 45.8; 43.5 | 44.9; 44.6 |
| Cluster, engine CPU per completed instance | 11.5; 11.9 ms | 11.7; 11.7 ms |
Throughput with everything exported was within 1 % of the runs without export, inside the 3 % budget, and CPU per instance did not change measurably. Latency with tracing on was never higher than without it, but two runs without tracing differed by 18 % at the 95th percentile, so 300 requests cannot resolve a 2 % latency budget: the latency budget is not violated as far as this measurement can tell, rather than proven.
To check the overhead in your own installation, compare the engine’s CPU
per instance and the acknowledgement latency (tinyconductor_command_seconds)
at the same load, once without export and once with it. Alternate the runs
and repeat each at least twice: the difference between two identical runs is
the floor of what the comparison can show.
Known limitations
Section titled “Known limitations”- Exemplars over OTLP. Metrics pushed over OTLP carry no exemplars: the
OpenTelemetry library the engine uses does not fill them in. Exemplars are
available on
/metricsin the OpenMetrics format. - Embedded mode has only the in-process spans: admitting and applying a command, and the job and message spans.
- Where a trace starts again:
- timers, job timeouts and message expiry start new traces;
- a buffered message continues the trace of the request that published it, not of the request that later matched it;
- signals carry no trace context: a signal does not move a waiting instance onto the broadcaster’s trace, and an instance started by a signal starts without one;
- in Embedded mode the context is lost on restart, and only the latest 10 000 instances keep one.
- A gauge of tenants folded into
otherby the label limit shows the value most recently recorded for any of them, not their sum.