Why extraction comes first
Every field you can filter on, route on, or search for has to exist in the event first. In Cribl Stream that happens in two stages: a stream of raw data is broken into discrete events, and then Functions inside a Pipeline pull named fields out of those events. The Event Breakers documentation describes the payoff of the first stage: "Once you have converted unstructured raw data into structured events, you can then send them through Pipelines for further routing and data processing."
The order is laid out on the Event Processing Order page: Sources come first, Event Breakers "can, optionally, break up incoming bytestreams into discrete events", and after that Routes "map incoming events to Processing Pipelines and Destinations." If an event reaches a route without fields, there is nothing to match on. Extraction is the prerequisite, and Cribl Stream's three main tools are event breakers, the Grok Function, and the Regex Extract Function.
Event breakers: from message to events
The docs describe Event Breaker Rulesets as "ordered collections of event-breaking rules that help you define the boundaries and structure of raw log data." Each is a reusable knowledge object in the Knowledge Library: apply one directly on a Source, or inside a Pipeline with the Event Breaker Function. Cribl compares incoming raw data against each rule's filter condition in order, and applies the breaking logic of the first matching rule for that stream.
The docs warn that a broad rule placed high in the list can capture events intended for a more specific rule below it. "To avoid this risk, order your rules from specific to general."
If no custom rule matches, the built-in System Default Rule takes over: a filter condition of true, the breaker pattern [\n\r]+(?!\s), an event byte limit of 51200, and a 150-byte Auto Timestamp scan depth. Its purpose is "to ensure that streams of unstructured logs are broken into structured data". The page also notes that Event Breakers "always" add the cribl_breaker field to output events.
The complete list of available types, from the Event Breaker Types page:
- Azure Virtual Network (VNet) Flow, for Azure VNet flow logs
- CSV, for data that adheres to the CSV standard, with row delineation, field extraction, and quoting rules
- File Header, for logs with a standard file header structure, such as Bro, IIS, or Apache access logs
- JSON Array, for single, large JSON objects containing a nested array of records
- JSON Newline Delimited, for data where each event is a complete JSON object followed by a newline
- Regex, for log data that does not fit the CSV, JSON Array, or File Header formats
- Timestamp, for streams with non-standard or highly varied timestamp formats
The same page's advice: "If you're unsure which Event Breaker to use", the Regex Event Breaker is "a good default Event Breaker because of its flexibility".
The regex breaker, configured
The Regex Event Breaker is "the default and most flexible Event Breaker type," allowing you to define event boundaries using regular expressions. Patterns use the JavaScript (ECMAScript) regex engine. Three considerations from the page:
- Breaks are continuous: the pattern applies continuously, any content before the match is part of the current event, and the break occurs at the start of the match.
- Matches are consumed by default: to set the break point without discarding the content that signifies the new event, such as a timestamp, use a positive lookahead such as
(?=pattern). "This is the standard practice for multiline logs." - Avoid capturing groups: parentheses like
(pattern)"will cause further, often unintended, splitting of the stream."
The configuration example below is copied from the regex event breaker page:
- Event Breaker:
[\n\r]+(?=\d+-\d+-\d+\s\d+:\d+:\d+). "This setting breaks after a newline or carriage return, but only if followed by a timestamp pattern." - Max Event Bytes:
51200. "The default setting."
The input is multi-line log data with a non-standard format:
2020-05-19 16:32:12 moen3628 ipsum[5213]: Use the mobile TCP feed, then you can program the auxiliary card!
Try to connect the FTP sensor, maybe it will connect the digital bus!
Try to navigate the AGP panel, maybe it will quantify the mobile alarm!
2020-05-19 16:32:12 moen3628 ipsum[5213]: Use the mobile TCP feed, then you can program the auxiliary card!
Try to connect the FTP sensor, maybe it will connect the digital bus!
Try to navigate the AGP panel, maybe it will quantify the mobile alarm!
The Regex Event Breaker generates two output events:
{
"_raw": "2020-05-19 16:32:12 moen3628 ipsum[5213]: Use the mobile TCP feed, then you can program the auxiliary card! \n Try to connect the FTP sensor, maybe it will connect the digital bus!\n Try to navigate the AGP panel, maybe it will quantify the mobile alarm!",
"_time": 1589920332
}
{
"_raw": "2020-05-19 16:32:12 moen3628 ipsum[5213]: Use the mobile TCP feed, then you can program the auxiliary card!\n Try to connect the FTP sensor, maybe it will connect the digital bus!\n Try to navigate the AGP panel, maybe it will quantify the mobile alarm!",
"_time": 1589920332
}
Both events carry the same _time, because the lookahead sets the break point at the start of a timestamp without consuming the matched text.
Grok
The Grok Function "extracts structured fields from unstructured log data, using modular regex patterns" instead of a raw regular expression: you compose a pattern from named tokens with the syntax %{PATTERN_NAME:FIELD_NAME}. The Source field setting defaults to _raw. Cribl Stream "ships common Grok Patterns for basic scenarios" as pattern files you add and edit under Knowledge > Grok Patterns.
The example below is copied from the Grok page. Example event:
{"_raw": "2020-09-16T04:20:42.45+01:00 DEBUG This is a sample debug log message"}
- Pattern:
%{TIMESTAMP_ISO8601:event_time} %{LOGLEVEL:log_level} %{GREEDYDATA:log_message} - Source Field:
_raw
Event after extraction:
{"_raw": "2020-09-16T04:20:42.45+01:00 DEBUG This is a sample debug log message",
"_time": 1600226442.045,
"event_time": "2020-09-16T04:20:42.45+01:00",
"log_level": "DEBUG",
"log_message": "This is a sample debug log message",
}
The page notes the new fields added to the event: event_time, log_level, and log_message. It also lists grokconstructor.appspot.com as a useful site for creating and testing Grok patterns.
Regex Extract
The Regex Extract Function "uses regex capture groups to pull structured data out of raw event strings." The page gives it three jobs: structure raw logs, extract names and values from strings (including dynamic key-value discovery), and clean and reduce data.
Settings in the Regex Extract modal:
- Regex: a "Regex literal." that must contain named capturing groups:
(?<foo>bar), wherefoois the field name andbaris the pattern to match. Special_NAME_Nand_VALUE_Ngroups extract both the name and value of a field; Add Regex chains extra conditions. - Source field: defaults to
_raw, so the Function applies the regex to the string in the_rawfield of every event. If the value is blank (null), "the Function will silently fail". - Overwrite existing fields: toggled on, the Function "replaces the field's value"; toggled off (the default), it "retains both values as an array containing the original and new values."
- Field name format expression: a JavaScript expression to transform or prefix field names generated by
_NAME_Ngroups; left blank, names are sanitized to valid JavaScript identifiers. - Max exec: "The maximum number of times the regex pattern should loop through the source field to find matches." Defaults to
100.
The page's "Single Field From Simple Event" example isolates one value from a noisy key-value line. Sample event:
{
"_raw": "metric1=23, metric2=42, dc=23, abc=xyz"
}
To extract only metric1, the page gives the regex metric1=(?<metric1>\d+). Resulting output:
{
"_raw": "metric1=23, metric2=42, dc=23, abc=xyz",
"metric1": "23"
}
For key-value discovery across varying fields, the page's example pairs the regex (?<_NAME_0>[\w-]+)="?(?<_VALUE_0>(?<=")[^"]*|\S*) with the field name format expression ${name}_XX: _NAME_N supplies the field name, _VALUE_N the value.
Order of operations
Breaking comes before extraction: on the Event Processing Order page, Event Breakers sit right after the Source, and the processing Pipelines (where your Functions live) come after Routes. The docs are explicit: "Do not use regex capturing groups in this pattern, as the Event Breaker is designed for structural discovery, not field extraction". For that job, the page points to the Regex Extract Function later in your Pipeline for key-value pair discovery.
So the Pipeline shape is: breaker rules decide where events start and end; Grok or Regex Extract Functions decide which fields each event carries; and routes plus everything downstream work on those fields. Keep specific rules ahead of general ones, and validate with the Data Preview window, which shows "how incoming events are structured after breaking."
Frequently asked questions
What is the order of breaking and extraction in Cribl Stream?
Breaking comes before extraction. Event Breakers sit right after the Source, and the functions that pull fields out run in the Pipelines that come after the Routes.
How does the Cribl Stream Grok Function name the fields it extracts?
You compose a pattern from named tokens, and the name in the token becomes the event's field. The source field defaults to _raw, and Stream ships common Grok Pattern files you can add and edit.
Why does Cribl Stream warn against capturing groups in Event Breaker patterns?
Capturing groups cause further, often unintended, splitting, because the breaker exists for structural discovery. The docs point to the Regex Extract Function later in the Pipeline for field extraction.
What does the Cribl Stream Regex Extract Function do?
The Function uses regex capture groups to pull structured data out of raw event strings. The docs give it three jobs: structure raw logs, extract names and values from strings, and clean and reduce data. The source field defaults to _raw, and a blank value makes it fail silently.
Verified against Cribl Stream 4.20 documentation on October 1, 2026.