Where masking has to happen to actually work
PII in a log stream is a timing problem as much as a content problem: once an unredacted value reaches a destination, it has been copied somewhere you no longer control. The Masking and Obfuscation use case in the Cribl docs opens with a section titled "Masking and Anonymization of Data in Motion", with the redaction happening on the processing path before any destination receives the event.
The page is quick to the point: "To mask patterns in real time, we use the out-of-the-box Mask Function", and it is "similar to sed, but with much more powerful functionality". The function acts on events as they flow through a Pipeline, and per the Mask Function documentation, marking a function Final makes it "stop feeding data to the downstream Functions", so everything downstream of the Mask Function sees the redacted value. The pages state why the work matters: "Masking is especially useful for redacting PII (personally identifiable information) and other sensitive data", and what happens to the original: "Cribl does not retain any content that matches the Match Regex that is not inserted into the Replace Expression". When a route needs the unredacted value for an archival store, the escape hatch: "If you need to send the unmodified values to a Destination", "you can use the Destination field for the associated Route to narrow the scope of the Mask Function".
The Mask function
The settings on the Mask Function page. Filter selects which events feed the function; it "Defaults to true, meaning it evaluates all events". Masking rules "defines pairs of patterns to match and replace". Each rule row holds a Match Regex ("Pattern to replace. Supports capture groups.", with /g added "to replace all matches") and a Replace Expression ("A JavaScript expression or literal to replace all matching content"). Apply to fields is "Fields on which to apply the masking rules. Defaults to _raw", and it supports wildcards, nested addressing, and order-sensitive negations. Advanced settings add Evaluate fields, which stamps extra key-value fields onto events where a rule matched, and Depth, which sets "The depth to which the Function will traverse nested properties in an event. Defaults to 5".
One housekeeping note before the examples: "This Function uses the JavaScript (ECMAScript) regex engine", so when testing patterns in an external tool, the page says to "select the ECMAScript (JavaScript) flavor".
The page's sensitive-data example hashes a Social Security number arriving in the Key=Value pattern social=#########. The rules, copied from the Mask Function page, are:
Match Regex: (social=)(\d+)
Replace Expression: `${g1}${C.Mask.md5(g2)}`
The capture groups keep the key in the output while hashing only the number: social= is group g1, the number is group g2. The page warns against folding the key into the match: "Without the key name preceding the hashed value, it isn't clear what value is being hashed".
The credit card example uses the same shape with a pattern that tolerates separators: "The cardNumber values may be digits alone or may include hyphens or spaces as separators".
Match Regex: (cardNumber=)(\b(?:\d[ -]*?){13,16}\b)
Replace Expression: `${g1}${C.Mask.md5(g2)}`
A third example on the page skips hashing and substitutes from the event itself, referencing an existing field by "prepending event to it" in the Replace Expression.
Lookup-based replacement
The using-lookups page states the concept in one line: "Use lookup files to enrich events in your Pipelines". Its guide list points to two companions that pair lookups with masking, "Lookups as Filters for Masks" and "Lookups and Regex Magic". On the pages covered here, a lookup works as a file-based reference looked up at processing time.
The Configure Lookups page walks the setup. First, where the file lives: "either the global Knowledge Library or to a specific Pack". Second, how it gets there: "Upload a New File" to bring an existing file in, or "Create with Text Editor" to "Enter comma-separated values (CSV) directly using the built-in editor". Storage Mode is then "Store in Memory" or "Store on Disk".
Third, the key. The page is precise about what it is: "an index is a fast-access data structure built from one or more header fields". For disk-based files, "Cribl Stream automatically indexes the left-most column of each lookup file by default. To take advantage of this, place your primary key fields in the first column". You can define "a maximum of four indexes" per file, and each can match case-sensitively or not.
Fourth, reference. Once deployed, the page offers "the Lookup Function or the C.Lookup() expression to enrich, route, or filter data based on lookup results", the expression form being callable in "any field that accepts JavaScript expressions, such as in the Filter field". Fifth, lifecycle. Files update "manually through the UI or programmatically using the API, which is the ideal method for automated or frequent updates"; in-memory files can be edited in the UI, disk-based ones must be deleted and re-uploaded, and "If the re-uploaded file uses the same filename, any references to the lookup file by name will continue to work without changes". And deleting a file "actively referenced in Pipelines or Routes can result in runtime errors or failed deployments", so "verify that it is no longer in use or update all references as needed".
The pages do not walk a complete recipe for replacing sensitive values from a lookup table, so this is our recommendation, not doc text: keep a lookup file mapping known sensitive values to redaction placeholders, and drive replacement at mask time with a C.Lookup() call. The composition rests on the pages' own parts: the Mask function's replace field "accepts a full JS expression that evaluates to a value", and C.Lookup() can be called in any expression-capable field. Rotating a placeholder is then a one-row edit to the lookup file, updated through the API the pages call "the ideal method for automated or frequent updates".
The obfuscation patterns the docs describe
Underneath the configuration, the use case page lists "several masking methods that are available under C.Mask" for use in a Replace Expression:
C.Mask.random: Generates a random alphanumeric string
C.Mask.repeat: Generates a repeating char/string pattern, e.g., XXXX
C.Mask.REDACTED: The literal 'REDACTED'
C.Mask.md5: Generates a MD5 hash of given value
C.Mask.sha1: Generates a SHA1 hash of given value
C.Mask.sha256: Generates a SHA256 hash of given value
"Almost all methods have an optional len parameter which can be used to control the length of the replacement". The parameter "can be either a number or string". Pass the captured group itself and the output comes out the same length as the original, which is how "length preserving replacement" works.
Every result row on the page derives from one source string, cardNumber=214992458870391, matched against /(cardNumber=)(\d+)/g:
C.Mask.random(), default length:cardNumber=HRhcC.Mask.random(g2), length preserving:cardNumber=DroJ73qmyaro51u3C.Mask.repeat():cardNumber=XXXXC.Mask.REDACTED:cardNumber=REDACTEDC.Mask.md5(g2):cardNumber=f5952ec7e6da54579e6d76feb7b0d01fC.Mask.md5(g2, 12), left 12 characters of the hash:cardNumber=d65a3ddb2749
Two constraints round out the list: the page notes "Replacement length will not exceed that of the hash algorithm output; MD5: 32 chars, SHA1: 40 chars, SHA256: 64 chars", and the left or right N-character substring is its way of shortening a hash. The Replace Expression is not limited to these methods; the simplest example on the Mask Function page is a plain literal, to "replace that value (if found) with Trans AM".
What masking does not fix
The pages are candid about the tools' sharp edges, and none of them claims masking is a complete answer. The limits they do document:
- First match only, by default: a Match Regex "will stop after the first match" unless it carries /g, which "will make the Function replace all matches". A second occurrence in the same field sails through unmasked.
- Bounded depth: the depth setting "Defaults to 5", so deeply nested values need it adjusted on purpose.
- Type changes: "The output of the Mask Function is a string even when the Replace expression produces a number", so a field that must stay numeric downstream needs follow-up; the page suggests a Numerify Function.
Where the pages are silent, they are silent, and that is where we draw the line between documentation and our judgment. Everything fetched covers data in motion, and none of the pages says a copy already delivered somewhere gets recalled, so we treat masking as a gate at the edge of your network, not a scrubber for downstream stores. Hashing carries its own judgment call. The pages list md5, sha1, and sha256 as masking methods without discussing whether a hash stays identifiable, so acceptability is a data policy decision. Our recommendation is to treat hashes of low-cardinality values, like SSNs, as linkable across events, and to keep any lookup file that could map a placeholder back to an original out of destinations that receive masked data.
Frequently asked questions
Where does Cribl Stream masking happen in the data path?
Masking happens in motion, on the processing path, before any destination receives the event. Marking the Mask Function Final stops it from feeding downstream, so everything after it sees the redacted value.
How does a Cribl Stream Mask Function rule replace a sensitive value?
Each rule pairs a Match Regex, which supports capture groups, with a Replace Expression that can be a JavaScript expression or a literal. Capture groups let you keep the key and hash only the value.
What replacement methods does Cribl Stream offer under C.Mask?
The use case page lists C.Mask.random, C.Mask.repeat, C.Mask.REDACTED, and the hash methods md5, sha1, and sha256. Almost all take an optional length parameter, and passing the captured group preserves the original length.
What does the Cribl Stream Mask Function do by default when a field has multiple matching values?
A Match Regex stops after the first match unless it carries the /g flag, which makes the Function replace all matches. Without it, a second occurrence in the same field sails through unmasked.
Verified against Cribl Stream 4.20 documentation on October 1, 2026.