Cold-Tier Storage with Cribl Lake and S3, and Replaying It Back

Route low-value data to cheap object storage now, search it later, or replay it into Splunk on demand.

Warm tier vs. cold tier

In a routing setup, "warm" and "cold" describe how often data is queried, not how important it is. The warm tier is whatever you search on a daily basis, and it stays indexed in Splunk. The cold tier is everything else: long-tail log sources, data that is audited only occasionally. Cold-tier storage addresses the cost of keeping that second group hot.

Cribl Stream offers two kinds of Destination for the cold tier (Cribl Lake, or Amazon S3 in a bucket you control) and a way to get the data back. Both are non-streaming, support TLS, and partition the stored objects so the data works with Cribl Search. And when you need archived data back in a live pipeline, Replay re-ingests it: "an easy way to selectively ingest, and re-ingest, data into systems of analysis", in the words of Cribl's S3 storage and replay use case.

Option A: Cribl Lake

The Cribl Lake Destination is the managed option: it "delivers data to Cribl Lake and automatically selects a partitioning scheme that works well with Cribl Search". It is available only in Cribl.Cloud, from both Cribl-managed Cloud and customer-managed hybrid Worker Groups, with the hybrid Groups needing Cribl Stream version 4.8 or later and outbound HTTP/S access to port 443.

Configuration is short: the Destination's settings, rearranged from that page:

General Settings
  Output ID             unique name to identify this Cribl Lake Destination
  Lake dataset          Cribl Lake Dataset to send data to

Optional Settings
  Backpressure behavior Block (default) or Drop
  Tags                  filter/group Destinations; not added to events

The data lands in a Cribl Lake Dataset, in a bucket whose name follows the pattern on the page's hybrid-access troubleshooting section: lake-main-<organizationId>.s3.<region>.amazonaws.com. One more detail: an incoming _time that is null or a string is converted to Date.now() / 1000, producing time-based partitioning.

Option B: S3

The other option is your own bucket. The Data Lake Amazon S3 Destination "sends data to Amazon Simple Storage Service (Amazon S3) and uses a partitioning scheme that is designed to work with Cribl Search". Its configuration, rearranged from the page:

General Settings
  Output ID             unique name to identify this S3 definition
  S3 bucket name        constant or JS expression, evaluated only at init time

Optional Settings
  Region                drop-down or custom Region; empty = global S3 endpoint
  Partition by fields   fields to partition the path by; time included automatically
  Backpressure behavior Block (default) or Drop
  Tags

Authentication          Auto (default) | Manual | Secret

One field deserves attention: leave Region empty and the global S3 endpoint is used, which "does not work for buckets that require a Region-specific endpoint". And the partitioning is what makes the archive searchable: the effective partition is <4-letters>-YYYY/<2-letters>-MM/<2-letters>-DD/<2-letters>-HH/<list/of/fields>: the page's example is hjhh-2023/ag-07/ai-25/aj-14/<list/of/fields>. The alphabetic prefixes are deliberate: S3 returns bucket content in ascending alphanumeric order, "oldest-first" by default, and the prefixes make the sort "newest-first", "allowing Search to find files that will likely have the requested objects soonest".

Authentication: Auto (the default) uses the AWS SDK for JavaScript, trying environment variables, IAM Identity Center (SSO), the shared credentials file, IAM Roles for EC2/ECS, a JSON file on disk, then other provider classes, in that order. Manual takes an Access key and Secret key, which the page recommends keeping as stored secrets referenced with C.Secret(); Secret selects a stored key pair from Cribl Stream's secrets manager. Assume Role covers accessing resources across Regions. In Advanced Settings, the cold-storage dial is Storage class: Standard by default, with Reduced Redundancy Storage, Standard, Infrequent Access, One Zone, Infrequent Access, Intelligent Tiering, Glacier, and Deep Archive among the options; file-closing limits default to a 32 MB size, 300 seconds open, 30 seconds idle, and 100 open files.

Replaying archived data

The Using S3 Storage and Replay use case walks the full round trip: write to S3, point a Collector at the store, and send selected events back through Cribl Stream.

Format matters: JSON captures the parsed event "with all metadata and modifications it contains at the time it reaches the Destination step" (one event per line, NDJSON), while raw captures the content of the _raw field. Cribl recommends the default JSON, since "Reusing that information makes sense, and will make your replay simpler".

In the example Destination, the store path is built from time plus index, host, and sourcetype, with a key prefix from a Global Variable, so replay can find files without opening them. Copied from the page:

Partitioning (one line):
`${C.Time.strftime(_time ? _time : Date.now() / 1000, '%Y/%m/%d')}/${index ? index : 'no_index'}/${host ? host : 'no_host'}/${sourcetype ? sourcetype : 'no_sourcetype'}`

Filename:
`${C.Time.strftime(_time ? _time : Date.now() / 1000, '%H%M')}`

The read side creates an S3 Collector with ID Replay, using the Auto-populate from option to pull the configuration from the S3 Destination. Its Path field extracts tokens from the store's paths, and it "is vital that this scheme match your partitioning scheme in the Destination definition", in this example:

/${MYENV}/${_time:%Y}/${_time:%m}/${_time:%d}/${index}/${host}/${sourcetype}/${_time:%H}${_time:%M}-${extended}

With the tokens defined, Stream can exclude files "that have no chance of matching our target events" without downloading them, then inspects only what is left. Under Result Settings > Event Breakers, the Cribl ruleset parses the packaged JSON. Add a __replayed field set to true; the example archival route matches !__replayed with Final unset, so replayed events are not re-archived.

One caveat: the page warns that Collectors and Replay "do not support the S3 Glacier, S3 Deep Glacier, or Azure archive tier, due to long retrieval times for those storage classes", while supporting S3 Glacier Instant Retrieval when using the S3 Intelligent-Tiering storage class. Once defined, the Collector "can be controlled via scheduling, manual runs, or API calls".

Reading Cribl Lake back

Cold-tier data in Cribl Lake is read back with the Cribl Lake Collector. It "gathers data from Cribl Lake". Cribl.Cloud only, like the Destination, and hybrid Worker Groups need version 4.8 or higher.

Collector Settings: a Collector ID (the page's example: myLakeCollector), a Storage Location (Cribl Lake is the default, plus each external Storage Location your user can access), and a Lake Dataset, which lists only Datasets that use the selected Storage Location. Result Settings covers event-breaking rulesets (default System Default Rule), fields defined via JavaScript expressions, Result Routing (Send to Routes on by default, or off to pin the Collector to a specific Pipeline and Destination), and Throttling; Advanced Settings add a Time to live for job artifacts (default 4h).

For a run or schedule you can set a Filter expression; Cribl Stream "evaluates that expression against both the Lake object path and the collected event fields". The page's warning is aimed at source: filtering on it "restricts which Lake files Cribl Stream pulls, not only which events it keeps", so source.endsWith('.log') "can return no results even when events contain a source field ending in .log". The page's advice: if you must, match both the file extension and the event values (its example: source.endsWith('log') || source.endsWith('gz')), or filter on source later. Verify in Preview mode first.

Choosing between the two

Limited to what the four pages state, the decision comes down to who operates the store, which storage classes remain, and how data comes back:

On Cribl.Cloud, the Lake path removes the bucket, IAM, and endpoint work; in a bucket you control, with the storage class of your choosing, the S3 Destination with its documented replay path is the counterpart.

Frequently asked questions

Which Cribl Stream destinations can hold cold-tier data outside the Splunk index?

Cribl Stream offers two cold-tier destinations: the managed Cribl Lake destination, and an Amazon S3 bucket you control. Both are non-streaming, support TLS, and partition objects so the data works with Cribl Search.

What is the availability difference between the Cribl Lake destination and the S3 destination?

The Cribl Lake destination and Lake Collector are Cribl.Cloud only, and hybrid Worker Groups need Cribl Stream version 4.8 or later with outbound port 443. The S3 destination works with your own bucket under the Auto, Manual, Secret, and Assume Role options.

How does Cribl Stream partition S3 objects so Cribl Search finds them quickly?

The S3 partition is built from time plus the fields you pick, and alphabetic prefixes flip the default oldest-first sort into newest-first. That ordering lets Cribl Search find likely files soonest.

Why can't Replay in Cribl Stream read back Glacier or Deep Archive data from S3?

Collectors and Replay do not support S3 Glacier, S3 Deep Glacier, or Azure archive tiers, due to long retrieval times. S3 Glacier Instant Retrieval is supported under the S3 Intelligent-Tiering storage class.

Verified against Cribl Stream 4.20 documentation on September 29, 2026.

Ready to Reduce Your SIEM Costs?

Get a personalized demo and see how Cribl can save you 50-80% on data costs.

Schedule Free Demo