Where the Money Actually Goes in a Splunk Bill
The line item that surprises most Splunk customers is ingestion. Splunk's cost scales with processing capacity, and what that capacity has to chew through each month is volume: every GB of events that reaches the indexers lands on the bill, whether or not anyone searches it. In practice, a large share of what gets indexed is verbose and low-value: success requests logged at debug detail, flow logs from networks nobody monitors, events from sourcetypes tied to applications that no longer exist. Cribl Stream is the layer in the middle: it sits between your log sources and the Splunk indexers, takes the full firehose, and decides what actually gets forwarded. Because its routing, filtering, and sampling all happen in flight, before indexing, the volume it drops or thins simply never reaches the Splunk bill.
The Cribl Stream Path to Cheaper Indexing
A Stream deployment is four moving parts. Events enter at a Source, a Route decides where they go, a Pipeline of functions processes what makes it through, and a Destination delivers the result. For a Splunk shop, that destination is the HEC receiver. Everything between the source and that receiver is where cost decisions live.
The levers are JavaScript. In the Cribl documentation on building custom logic to route and process data, the key mechanism is the Filter field, which appears on both Routes and Functions. A Route's filter accepts a JavaScript expression that must evaluate to true or false: true means the route processes the event, false means it passes through unchanged. The default is true, so by default every event goes everywhere until you say otherwise. Functions add value expressions for shaping fields in place, and Cribl Expressions (pre-built helpers that begin with C., such as C.Lookup or C.Time) cover the common operations.
A route-scoping filter looks exactly as simple as that sounds:
sourcetype=='access_combined'
sourcetype=='access_combined' && host.endsWith('dnto.ca')
Those two lines are copied from Cribl's regex-filtering use case: the first route matches all events of one web access log sourcetype; the second narrows the same route to a single host's domain. The docs also point out the safety net: the Advanced mode button on any Filter expression field opens a modal where you paste the expression and validate it against sample JSON input before the expression ever sees production data.
Drop the Noise with Regex Filtering
For the "nobody needs these events" class of noise, Cribl ships an out-of-the-box Regex Filter Function that drops events in real time (what the Regex Filtering use case calls filtering data in motion). The docs describe it as the in-flight equivalent of classic Splunk nullqueueing with TRANSFORMS: the end state is the same (the events never reach the indexers), but the matching condition is a regex against a field, scoped by a JavaScript Filter expression, which the docs note is far more flexible than the old TRANSFORMS approach.
The worked example: filter out any sourcetype=='access_combined' events whose _raw field contains the pattern Opera. Once deployed, you verify in Splunk that the search for those events comes back empty. The Filter input is what keeps the blast radius small. The docs show adding host.endsWith('dnto.ca') so the drop applies to one host's domain rather than every event named access_combined. The pattern generalizes to your own verbose debug noise: identify the pattern in _raw worth dropping, put a regex on it, scope the drop with field conditions in the Filter, and confirm in search before trusting it at full volume.
Sample What You Can Sample
Not everything you want to keep has to be kept at full volume. The Sampling Function filters out events based on an expression and a sampling rate: each sampling rule has a Filter and a Sampling Rate (an integer N, defaulting to 1), and the function keeps 1 in every N matching events. A rate of 30 keeps one in thirty; stack rules to sample one class of event at 5:1 while everything else flows unchanged at 1:1. The reference page also notes each Worker Process executes the function independently on its own share of events, a consequence of Stream's shared-nothing architecture. If you want sampled data to stop there, the Final toggle prevents it feeding downstream Functions.
Cribl's ingest-time sampling use case targets highly verbose, voluminous data (CDN logs, ELB access logs, VPC flow logs) with the goal of keeping enough sample that analysis stays statistically significant. A Regex Extract function pulls the HTTP status out of _raw into a __status field (fields starting with __ are special fields usable anywhere in a Pipeline), then:
Filter: __status == 200
Sampling Rate: 5
That single rule (scoped to sourcetype=='access_combined') drops four of every five success responses while every 400 and 500 still arrives 1:1. That is the line the docs draw for what should not be sampled: the events you troubleshoot with. The use case explicitly samples the verbose successes so that "all potentially erroneous events" keep flowing for investigation. Every event that passes through the Sampling Function gets an index-time sampled::<rate> field added, so your statistical queries can account for the thinning instead of assuming the volume is full.
Ship the Result via HEC
What survives filtering and sampling still needs to reach Splunk. The Splunk HEC Destination is the documented path: Cribl streams to an HEC receiver through the event, raw, or S2S endpoints, and the data arrives to Splunk "cooked and parsed," entering the data pipeline's indexing segment. The destination page calls it the recommended destination for Splunk Cloud and lists TLS support as available. The configuration knobs that matter when reviewing a Splunk target:
- Endpoints and load balancing: one or more Splunk HEC endpoint URLs, with per-endpoint load weights when Load balancing is on. Supported paths:
/services/collector/event(default),/services/collector/raw, and/services/collector/s2s. - Authentication: Manual, which exposes the HEC Auth token field (where your Splunk HEC token goes), or Secret, which selects a stored, reusable secret that references that token.
- TLS: the Use TLS toggle (off by default), server validation via the SNI server name and a Trusted server CA path, optional mutual authentication with certificate, private key path, and passphrase, and minimum/maximum TLS version thresholds.
- Backpressure behavior: block, drop, or queue events when all receivers are exerting backpressure; the queue option engages the persistent queue, which buffers events on disk through outages.
- Advanced: Compress is on by default to shrink payload bodies, Body size limit defaults to
4096KB, with a Cribl-recommended maximum of 1 MB when sending to Splunk Cloud, and the Retries section applies exponential backoff to5xxresponses while429responses need to be added explicitly to be retried.
One misconfiguration the docs call out by name: do not toggle Enable Indexer Acknowledgement on for the Splunk token. With it on, the Splunk receiver expects a Channel GUID to be passed in and the request fails. The page shows the resulting 400 error with "Data channel is missing".
What we would review with you
Nothing above is theoretical: these are the three levers we walk through in a hands-on demo against your own data.
- Filter: the sourcetypes, hosts, and
_rawpatterns with no search demand, and the route or Regex Filter expressions that drop them in flight. - Sample: the voluminous streams (access logs, flow logs, CDN) that can tolerate 5:1 or 30:1 without losing troubleshooting value, and how the
sampled::fields keep your statistics honest. - Break out: the data that belongs on a cheaper destination, or that shouldn't go to Splunk at all.
Use the demo form below to schedule a session, and we'll bring this setup (routes, regex filters, sampling rules, and the HEC destination) pointed at your own sources.
Frequently asked questions
What is the main lever a Cribl Stream deployment uses to cut Splunk ingestion costs?
Cribl Stream filters, samples, and shapes events in flight, so the volume it drops or thins never reaches the Splunk indexers. Its routing, filtering, and sampling all happen before indexing.
How does the Cribl Stream Regex Filter Function drop noise without touching other events?
It drops events in real time when a regex matches a field, with the Filter input scoping the drop to the events you name. That keeps the drop off every other event of the sourcetype.
How do you keep statistical queries honest after sampling in Cribl Stream?
Every event that passes a Sampling Function gets an index-time sampled field that records the rate at which it was sampled. Statistical queries can use that field to account for the thinning instead of assuming full volume.
What setting on the Cribl Stream HEC destination must stay off for a Splunk token?
Do not turn on Enable Indexer Acknowledgement for the Splunk token. With it on, the Splunk receiver expects a Channel GUID to be passed in and the request fails with a 400 error.
Verified against Cribl Stream 4.20 documentation on September 29, 2026.