The deployment types
In an on-prem (self-hosted) deployment, Cribl Stream runs on your own infrastructure. Per the On-Prem Deployment page, two key factors determine the deployment type: the amount of incoming data (planned ingest per unit of time, e.g. MB/s or GB/day) and the amount of data processing (just routing through, versus a lot of transformations, regex extractions, field encryptions, or heavy re-serialization). The page lists five on-prem deployment types:
- Single-instance deployment: when volume is low and/or processing is light.
- Distributed deployment: to accommodate increased load, scaling up and perhaps out with multiple instances.
- Splunk App deployment: an existing Splunk Heavy Forwarder infrastructure, via Cribl App for Splunk.
- Orchestrated deployment: Cribl's Helm charts.
- Docker deployment: images from Cribl's public Docker Hub.
When distributed stops being optional
The Distributed Deployment page defines a Distributed deployment as a multi-instance deployment that is fit for large data volumes and increased processing. A single Leader Node manages all Worker Nodes: tracking and monitoring their activity metrics, keeping configurations in sync, and handling version control.
The page's basic concepts:
- Single instance: a standalone (not distributed) installation on one server.
- Leader Node: a Cribl Stream instance that runs in Leader mode, used to centrally monitor and author configurations for the Worker Nodes.
- Worker Node: an instance that runs as a managed Worker; the Leader fully manages its configuration, and by default Workers poll the Leader for configuration changes every 10 seconds.
- Worker Group: a collection of Worker Nodes that share the same configuration, mapped with a Mapping Ruleset, an ordered list of filters.
A Worker Process is a Linux process that handles data inputs, processing, and output, constrained in count by the number of physical or virtual CPUs available.
The page's scale guidance is qualitative. Per our recommendation (an operational judgment, not a doc figure), a constant need for more Worker Processes, or running out of CPU, is the signal to plan a Distributed deployment.
Leader vs. worker responsibilities
The division of labor, per the Distributed Deployment page: the Leader Node runs an API Process that handles all the API interactions, plus Config Helpers, one process per Worker Group, that help maintain configurations and previews. Each Worker Node runs an API Process for communication with the Leader and other API requests, plus Worker Processes that handle all the data processing.
A single-instance deployment is the same stack condensed into one: the API Process handles all API interactions and the Worker Processes handle all data processing.
At runtime the Leader has two primary roles: it is the central location for the operational metrics of the Worker Nodes, and it authors, validates, deploys, and synchronizes configurations across Worker Groups. Worker Nodes send a heartbeat to the Leader every 10 seconds, with system metrics and tracked facts such as hostname, IP address, GUID, and tags. On check-in, the Leader maps the Worker to a Worker Group from its configuration, tags, and Mapping Rules, and the Worker gets an updated configuration bundle if necessary; a Worker missing two consecutive heartbeats is removed from the Workers page.
Two failure modes: a Worker Node that goes down loses any data it is processing (streaming senders, or Push Sources, are handled only in memory), so configure persistent queues on Sources and/or Destinations that support it. If the Leader goes down, Workers keep receiving and processing incoming data from most Sources and Collection tasks keep running, but future scheduled Collection jobs fail, as does collection on the Amazon Kinesis Data Streams, Prometheus Scraper, and all Office 365 Sources (which function as Collectors).
Choosing the mode: a summary of the setup steps
The Set Up Leader and Worker Nodes page covers pointing each instance at its role: Leader Node or managed Worker Node, via environment variables, UI settings, teleporting, the CLI, or instance.yml. Internally the Leader mode is called master, an older name.
Via environment variables, CRIBL_DIST_MODE defines the mode (worker or leader) and CRIBL_DIST_LEADER_URL points at the Leader Node; a Leader uses 0.0.0.0 as the address, and a set CRIBL_DIST_LEADER_URL defaults the mode to worker. Environment variables take priority, any CRIBL_DIST_* variable disables Distributed mode configuration via the UI, and switching to Leader mode requires an auth token.
Via the UI: Settings > Global, then Distributed Settings > General Settings; choose Leader or Stream: Managed Worker (managed by Leader) for Mode, and set the Leader's Address and Port and the Worker's Address under Leader Settings. The CLI equivalents, mode-master and mode-worker, require -H, -p, and -u, and Cribl Stream must restart afterward. The page's example worker command:
./cribl mode-worker -H 192.0.2.1 -p 4200 -u <your-token>
The Workers tab tracks each Worker with Auto-refresh on by default; Workers missing 5 heartbeats, or with connections closed for more than 30 seconds, are removed from the list. Behind a load balancer, register all instances using the health endpoint each Node exposes:
curl http://<host>:<port>/api/v1/health
{"status":"healthy"}
Worker groups
The Manage Worker Groups page defines Worker Groups as collections of Worker Nodes that share the same configuration. You configure data flow per Worker Group, including Routes, Pipelines, Sources, and Destinations, and each Worker Group has separate version control.
Why more than one group? The Distributed Deployment page's example is organizational or geographic: US, APAC, and EMEA Worker Groups, each with their own distinct certifications and settings. Per our recommendation, align group boundaries with real constraints, because routing, pipelines, deployment, and version control all operate at group granularity.
To create one: in the sidebar, Worker Groups, then Add Worker Group. On Cribl.Cloud, if a Worker Group type field appears, select Hybrid. Enter a unique Worker Group name, and optionally enable teleporting to Workers, a Description, and Tags.
Cloning a Worker Group copies everything configured on the original: Sources, Pipelines, Packs, Routes, and Destinations. Deleting one removes access to its Version Control menu, but the page offers two recovery methods: create a new Group with exactly the same name (restoring its configuration; select a previous version if the deletion was committed), or stop the Cribl server, find the deletion commit, and revert to the last commit before the deletion.
Configuring multiple Worker Groups, or more than 10 Worker Processes, requires a Cribl Stream on-prem deployment with a certain license tier; on Cribl.Cloud, a certain plan also allows Cribl-managed Worker Groups.
High availability
The Distributed Deployment page points to the Leader High Availability/Failover page for setting up the second Leader. The Manage High Availability Deployment page is the operating manual for a pair already in place: monitoring Leaders, performing an upgrade, and disabling standby Leader Nodes. Cribl recommends maintaining a standby Leader for continuity in your on-prem Distributed deployment; to disable one, contact support.
Monitoring is at Products > Stream on the top bar, then Monitoring > System > Leaders in the sidebar. The Leaders table shows the status of all your Leaders and their roles (primary or standby), and is only accessible when a standby Leader is configured.
Upgrades follow a fixed order, and upgrading the Leader Node through the UI is not supported for Distributed deployments with a standby Leader: use the command line (CLI).
- Stop the standby Leader first.
- Upgrade the standby Leader.
- Start the standby Leader.
- Stop the primary Leader; the upgraded standby takes over automatically through the HA failover mechanism.
- Upgrade the now-stopped former primary Leader.
- Start the former primary Leader; it rejoins and assumes the standby role.
- Upgrade the Worker Nodes in rolling batches; do not upgrade them to a version newer than the Leader.
Frequently asked questions
What on-prem deployment types does Cribl Stream offer?
The docs list five: single-instance, distributed, Splunk App, orchestrated, and Docker. Single-instance fits low volume and light processing, distributed accommodates increased load, and the other three cover reusing a Splunk Heavy Forwarder, Helm-chart deployment, and Docker Hub images.
How do the Leader and Worker Node roles differ in Cribl Stream?
The Leader Node handles all API interactions and runs one Config Helper process per Worker Group, while Worker Nodes handle all data processing through their Worker Processes. Worker Nodes send a heartbeat to the Leader every 10 seconds and pick up their Worker Group's configuration bundle.
What happens to data processing if a Cribl Stream Leader Node goes down?
Worker Groups continue receiving and processing incoming data from most Sources, and Collection tasks already in process keep running. However, future scheduled Collection jobs fail, as does data collection on the Amazon Kinesis Data Streams, Prometheus Scraper, and all Office 365 Sources.
How do you upgrade a Cribl Stream high-availability deployment?
You upgrade through the command line, because UI upgrades are not supported with a standby Leader configured. The order is standby first, then the primary (with the standby taking over through the HA failover mechanism), then the Worker Nodes in rolling batches, never newer than the Leader.
Verified against Cribl Stream 4.20 documentation on October 3, 2026.