Skip to content

[high] Keep wildcards after backslashes in replace_string transformation - #566

Merged
thomaspatzke merged 2 commits into
SigmaHQ:mainfrom
elhoim:fix/replace-string-backslash-wildcard
Sep 27, 2026
Merged

thomaspatzke merged 2 commits into
SigmaHQ:mainfrom
elhoim:fix/replace-string-backslash-wildcard

Conversation

@elhoim

@elhoim elhoim commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

BLUF

  • Problem: with the default skip_special=False, ReplaceStringTransformation renders the value to plain text, applies the regex, then parses the whole result again. A literal backslash followed by a wildcard (C:\Windows\ + *) renders to the same text as an escaped literal asterisk (\*). The workaround that doubles backslashes intentionally skips \*, so the wildcard becomes a literal * (or ?) and the backslash is lost.
  • Detection impact: this happens on every value, even when the regex matches nothing. Image|contains: '\Temp\' converts to Image endswith "\Temp\*", and CommandLine|startswith: 'C:\Windows\' converts to an exact match on C:\Windows*. A ^C: → %SystemDrive% replacement gives the literal %SystemDrive%\Windows*. Windows directory paths ending in a backslash with contains/startswith are very common in SigmaHQ rules, and these rules silently stop matching.
  • Fix: the regex still runs on the unchanged plain representation, so existing pipelines' regexes see exactly what they saw before. Only the replaced text is parsed again. Text the regex did not match is copied over from the original SigmaString parts (plain strings, SpecialChars, Placeholders).
  • Tests: 8 parametrized cases plus 1 end-to-end conversion test. 7 of them fail on main and pass with the fix; the other 2 guard existing behaviour. The full suite, black and mypy pass.

Priority: high

Details

SigmaString.to_plain() escapes literal */? but not backslashes. That makes the plain form ambiguous:

parts to_plain()
['C:\Windows\', WILDCARD_MULTI] C:\Windows\*
['C:\Windows*'] (escaped literal *) C:\Windows\*

apply_string_value then ran re.sub(r"\\(?![*?])", r"\\\\", replaced) on the whole result and parsed it with SigmaString(). The first form therefore always came back as the second.

The new _replace_plain builds the same plain string and records which plain-text span each original part (character, special character, placeholder) produced. It then walks re.finditer():

  • Unmatched original parts are appended unchanged. This is what keeps \ + wildcard intact.
  • The replacement text (match.expand()), plus any part that a match cuts through, is collected and parsed with the previous rules: the backslash-doubling post-processing, and insert_placeholders() when the value contained placeholders. So replacement strings keep the meaning they had before (e.g. * in a replacement is still a wildcard, \* is still a literal asterisk).

finditer + expand produce the same matches and the same replacement text as re.sub (including empty matches on Python ≥ 3.7).

Behaviour differences vs main, checked with a differential fuzz of 200k random values/regexes/replacements against the previous implementation:

  • Inputs with a backslash before a wildcard: the wildcard and the backslash are now kept. This is the fix.
  • Replacement text ending in a literal backslash directly before an original wildcard: the wildcard is now kept instead of being turned into a literal. This is the same bug class.
  • Values that contain both a placeholder and a literal % in a plain part: the literal text is no longer re-tokenized into a bogus placeholder by the whole-string insert_placeholders().
  • No other differences were found.

Related: #562 (open, mine) adds SigmaString.to_plain(escape_backslash=True) for to_dict/from_dict round-trips. This PR deliberately does not use it here. Escaping backslashes before re.sub would change the text that users' regexes run against (e.g. a UNC \\server would become \\\server), which would break existing pipelines. The two PRs are independent.

Testing

  • test_replace_string_backslash_before_wildcard (parametrized): no-match regex over ['C:\Windows\', *], [*, '\Temp\', *] and ['C:\Windows\', ?, 'x'] leaves the value unchanged. ^C: → %SystemDrive% keeps the trailing wildcard. A replacement inside the path keeps \ + ?. A replacement ending in a backslash keeps the following wildcard. An escaped literal * is still unchanged, and \\ → / still works; these two pass on main as well and are regression guards.
  • test_replace_string_backslash_before_wildcard_conversion: a full rule converted with TextQueryTestBackend gives the same query with and without a no-op replace_string pipeline.
  • The existing test_replace_string_* tests (specials, placeholders, backslashes, numbers) are unchanged and pass.
  • The steps from .github/workflows/test.yml were run locally on Python 3.12: black --check . is clean, pytest gives 1588 passed and 1 skipped, and mypy reports no issues.

🤖 Generated with Claude Code

ReplaceStringTransformation (skip_special=False) converted the value to
its plain representation, applied the regular expression and parsed the
complete result again. The plain representation of a literal backslash
followed by a wildcard (C:\Windows\ + *) is identical to an escaped
literal asterisk, so every such value lost its wildcard, even when the
regular expression did not match at all: contains '\Temp\' became
endswith '\Temp\*' and startswith 'C:\Windows\' became an exact match.

The regular expression still operates on the same plain representation,
but only the replaced text is parsed again. Text not matched by the
regular expression is taken over from the original SigmaString parts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

One of the new tests aliases the same mutable value list into both input and expected values, which can mask regressions if the transformation ever mutates in-place.

Review effort: Lite
Findings: 1 Low severity

Open (1)
What changed in this PR

This PR fixes a semantic corruption in ReplaceStringTransformation where a literal backslash immediately before a wildcard (*/?) could be re-parsed as an escape sequence after replacement, silently turning intended wildcards into literal characters (or dropping the backslash). The new implementation preserves unmatched SigmaString parts verbatim and only re-parses the replacement text, keeping existing pipeline regex matching behavior intact while preventing wildcard/backslash loss.

Changes:

  • Reworks ReplaceStringTransformation (non-skip_special path) to do span-aware finditer replacement and only parse replaced fragments back into SigmaString parts.
  • Adds targeted parametrized regression tests plus an end-to-end conversion test for common Windows-path cases ending in backslash.
File Description
sigma/​processing/​transformations/​values.py Implements span-aware plain-string replacement that preserves unmatched original SigmaString parts (fixing backslash + wildcard ambiguity).
tests/​test_processing_transformations.py Adds regression tests covering no-match, in-path replacements, and end-to-end conversion behavior for trailing backslashes before wildcards.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tests/test_processing_transformations.py Outdated
Convert string values to lists in test for transformations.

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
@thomaspatzke
thomaspatzke merged commit dd96800 into SigmaHQ:main Sep 27, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants