feat(security): add MemoryGuard and sensitive_field_guard for OWASP ASI06 prompt injection defense - #2321
Open
adithyaspillai wants to merge 1 commit into
Open
Conversation
…SI06 defense Closes 567-labs#2316
jordandunmire97-ai
approved these changes
Jun 28, 2026
This was referenced Jul 29, 2026
jxnl
added a commit
that referenced
this pull request
Aug 3, 2026
## Summary - accumulate declared, nested, and unknown numeric OpenAI/Anthropic usage fields across retries while preserving non-numeric metadata - add corrective feedback when a Responses API retry receives no tool call - preserve raw iterable type hints through sync and async v2 parallel-tool wrappers - strengthen API-key-free coverage for current and future SDK usage counters ## Consolidated and superseded items - closes #2493 - consolidates contributor work from #2498, #2500, and #2501 with original commit authorship preserved - supersedes #2497 because it drops unknown `model_extra` counters - supersedes #2499 because its hand-maintained provider field lists would drift as SDKs evolve ## Validation - focused changed-surface suite: `110 passed` - broad offline v2/coverage suite: `2368 passed, 91 skipped, 73 deselected` - Ruff check and format check: passed - scoped `ty check`: passed - `uv lock --check`: passed - pre-commit hooks and `git diff --check`: passed The 73 deselected tests require live provider credentials. An unfiltered local run confirmed its 22 failures were provider network connections in the restricted environment; GitHub provider jobs remain the authoritative validation for those paths. ## Intentionally skipped - provider additions or expansions: #2436, #2435, #2423, #2409, #2384, #2322, #2306, #2298, #2283, #2168, #2086; issues #2408, #2383, #2365, #2260, #2084, #2076 - broad architecture, product, security, or streaming decisions: #2394, #2392, #2357, #2356, #2355, #2351, #2321, #2307, #2287, #2263; issues #2479, #2403, #2393, #2391, #2316, #2272, #2056 - dependency batch: #2433 - nontrivial examples and editorial/resource additions: #2468, #2405, #2401, #2354, #2346, #2311, #2305; issue #2404 These remain open because they need dedicated product, architecture, provider, security, dependency, or editorial review and are not required for the `1.15.5` patch release. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes retry usage totals and reask message content on failure paths; scope is limited and heavily covered by tests, with no auth or data-store changes. > > **Overview** > Bundles three v2 retry and wrapper fixes for a patch release. > > **Retry usage accounting** replaces hand-maintained token field sums with generic `_accumulate_models` on Pydantic usage objects. Numeric fields (including nested models and `model_extra` counters) add across retries; booleans and other non-numeric metadata are not treated as billable. OpenAI and Anthropic paths share this logic. > > **OpenAI Responses reask** appends a user correction when `RESPONSES_TOOLS` validation fails but the output has no tool calls (e.g. reasoning-only), so retries include feedback instead of repeating the same request. > > **Parallel tools** in `patch_v2` skips `prepare_response_model` and does not replace `response_model` with the handler’s prepared wrapper for parallel modes, keeping raw `Iterable[...]` hints so schemas and parsed results include every member type. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit bbddca1. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->
Collaborator
|
1.16 triage: keeping this open with issue #2316 for separate security-product review. It adds new public |
This was referenced Aug 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
feat(security): OWASP ASI06 Memory Poisoning Defense
Closes #2316
What is this?
Adds two new security utilities to
instructor/security.pyto protect structured output agents from OWASP ASI06 — Memory Poisoning attacks, where malicious content embedded in parsed documents (emails, web pages, PDFs) can poison structured outputs before they are written to agent memory or passed to downstream tools.What was added
MemoryGuard
A Pydantic model mixin that scans all string fields (including nested dicts and lists) for prompt injection patterns after the LLM response is parsed. If an injection pattern is detected, it raises a
ValueErrorwhich triggers Instructor's built-in retry loop automatically.sensitive_field_guard
A field-level
BeforeValidatorfor guarding specific high-value fields likerole,permissions, oruser_idagainst known bad values.Files changed
instructor/security.py— new file withMemoryGuardandsensitive_field_guardinstructor/__init__.py— exports both utilities from the top-level packagetests/test_security.py— 15 unit tests, no API key or external dependencies requiredWhy this approach
pyproject.tomlmodel_validatorandBeforeValidatorpatterns already documented in Instructor's validation conceptsValueError, Instructor's retry loop handles it without any extra wiringTesting
All 15 tests pass with no API calls required.