[high] Keep wildcards after backslashes in replace_string transformation - #566
Conversation
ReplaceStringTransformation (skip_special=False) converted the value to its plain representation, applied the regular expression and parsed the complete result again. The plain representation of a literal backslash followed by a wildcard (C:\Windows\ + *) is identical to an escaped literal asterisk, so every such value lost its wildcard, even when the regular expression did not match at all: contains '\Temp\' became endswith '\Temp\*' and startswith 'C:\Windows\' became an exact match. The regular expression still operates on the same plain representation, but only the replaced text is parsed again. Text not matched by the regular expression is taken over from the original SigmaString parts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
One of the new tests aliases the same mutable value list into both input and expected values, which can mask regressions if the transformation ever mutates in-place.
Review effort: Lite
Findings: 1
Open (1)
What changed in this PR
This PR fixes a semantic corruption in ReplaceStringTransformation where a literal backslash immediately before a wildcard (*/?) could be re-parsed as an escape sequence after replacement, silently turning intended wildcards into literal characters (or dropping the backslash). The new implementation preserves unmatched SigmaString parts verbatim and only re-parses the replacement text, keeping existing pipeline regex matching behavior intact while preventing wildcard/backslash loss.
Changes:
- Reworks
ReplaceStringTransformation(non-skip_specialpath) to do span-awarefinditerreplacement and only parse replaced fragments back intoSigmaStringparts. - Adds targeted parametrized regression tests plus an end-to-end conversion test for common Windows-path cases ending in backslash.
| File | Description |
|---|---|
sigma/processing/transformations/values.py |
Implements span-aware plain-string replacement that preserves unmatched original SigmaString parts (fixing backslash + wildcard ambiguity). |
tests/test_processing_transformations.py |
Adds regression tests covering no-match, in-path replacements, and end-to-end conversion behavior for trailing backslashes before wildcards. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Convert string values to lists in test for transformations. Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

BLUF
skip_special=False,ReplaceStringTransformationrenders the value to plain text, applies the regex, then parses the whole result again. A literal backslash followed by a wildcard (C:\Windows\+*) renders to the same text as an escaped literal asterisk (\*). The workaround that doubles backslashes intentionally skips\*, so the wildcard becomes a literal*(or?) and the backslash is lost.Image|contains: '\Temp\'converts toImage endswith "\Temp\*", andCommandLine|startswith: 'C:\Windows\'converts to an exact match onC:\Windows*. A^C:→%SystemDrive%replacement gives the literal%SystemDrive%\Windows*. Windows directory paths ending in a backslash withcontains/startswithare very common in SigmaHQ rules, and these rules silently stop matching.SigmaStringparts (plain strings,SpecialChars,Placeholders).mainand pass with the fix; the other 2 guard existing behaviour. The full suite, black and mypy pass.Priority: high
Details
SigmaString.to_plain()escapes literal*/?but not backslashes. That makes the plain form ambiguous:to_plain()['C:\Windows\', WILDCARD_MULTI]C:\Windows\*['C:\Windows*'](escaped literal*)C:\Windows\*apply_string_valuethen ranre.sub(r"\\(?![*?])", r"\\\\", replaced)on the whole result and parsed it withSigmaString(). The first form therefore always came back as the second.The new
_replace_plainbuilds the same plain string and records which plain-text span each original part (character, special character, placeholder) produced. It then walksre.finditer():\+ wildcard intact.match.expand()), plus any part that a match cuts through, is collected and parsed with the previous rules: the backslash-doubling post-processing, andinsert_placeholders()when the value contained placeholders. So replacement strings keep the meaning they had before (e.g.*in a replacement is still a wildcard,\*is still a literal asterisk).finditer+expandproduce the same matches and the same replacement text asre.sub(including empty matches on Python ≥ 3.7).Behaviour differences vs
main, checked with a differential fuzz of 200k random values/regexes/replacements against the previous implementation:%in a plain part: the literal text is no longer re-tokenized into a bogus placeholder by the whole-stringinsert_placeholders().Related: #562 (open, mine) adds
SigmaString.to_plain(escape_backslash=True)forto_dict/from_dictround-trips. This PR deliberately does not use it here. Escaping backslashes beforere.subwould change the text that users' regexes run against (e.g. a UNC\\serverwould become\\\server), which would break existing pipelines. The two PRs are independent.Testing
test_replace_string_backslash_before_wildcard(parametrized): no-match regex over['C:\Windows\', *],[*, '\Temp\', *]and['C:\Windows\', ?, 'x']leaves the value unchanged.^C:→%SystemDrive%keeps the trailing wildcard. A replacement inside the path keeps\+?. A replacement ending in a backslash keeps the following wildcard. An escaped literal*is still unchanged, and\\→/still works; these two pass onmainas well and are regression guards.test_replace_string_backslash_before_wildcard_conversion: a full rule converted withTextQueryTestBackendgives the same query with and without a no-opreplace_stringpipeline.test_replace_string_*tests (specials, placeholders, backslashes, numbers) are unchanged and pass..github/workflows/test.ymlwere run locally on Python 3.12:black --check .is clean,pytestgives 1588 passed and 1 skipped, andmypyreports no issues.🤖 Generated with Claude Code