Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -123,6 +123,7 @@ ENV/
env.bak/
venv.bak/
docs/*
!docs/ai_configuration.md
!docs/extract_ai_user_guide.md
!docs/github-app-deployment.md
!docs/search_ai_mode.md
Expand Down
215 changes: 215 additions & 0 deletions docs/ai_configuration.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,215 @@
# AI model configuration

`wrangles/ai_defaults.yml` is the packaged catalog for AI model selection and
defaults. Set `WRANGLES_AI_CONFIG` to a versioned replacement YAML file to use
your own catalog. Start from a copy of the packaged file. Configuration is
loaded locally and cached; call `wrangles.ai_config.clear_cache()` after changing
an already-loaded file.

## Providers, models, and operations

Version 2 groups models under their providers. Model entries describe lifecycle,
default roles, model-specific defaults, and known supported parameter values.
Operation settings hold concurrency, timeouts, caching, and extraction prompts.
Optional `applications` metadata is a list of non-empty strings,
such as `[embeddings]` or `[data_extraction, description_writing]`. These labels help readers find
models; `default_for` selects defaults. Application labels are not sent to APIs.
Provider `documentation` links, including `model_cards`, are reference material
and are kept separate from request `endpoints`.

This is an excerpt; a replacement file should contain the complete catalog:

```yaml
version: 2
providers:
openai:
endpoints:
responses: https://api.openai.com/v1/responses
models:
gpt-6-luna:
status: active
applications: [data_extraction, description_writing]
default_for: [global, test, extract.ai, generate.ai]
defaults:
reasoning:
effort: none
text:
verbosity: low
supported_values:
reasoning.effort: [none, low, medium, high, xhigh, max]
text.verbosity: [low, medium, high]
operations:
extract.ai:
provider: openai
protocol: responses
defaults:
default_concurrency: 32
request_timeout_seconds: 12
retries: 1
```

`status` is `active`, `deprecated`, or `retired`. It describes the model's status
in this catalog, not a live availability check against the provider. A model
can serve several roles through `default_for`; each role has one default within
a provider. Retired entries cannot hold default roles. Deprecated entries remain
usable, including through default roles. Explicit model selection remains available
for compatibility, including unlisted custom
models and dated snapshots. Provider errors still determine actual availability.
Selecting a deprecated model logs a warning naming its provider and model, then
continues normally. This also applies to configured defaults, saved definitions,
and tests. The warning is issued once per operation, outside row and retry loops;
reading or resolving the catalog does not emit warnings.

The `global` role is a fallback for text extraction and generation. Embeddings
and URL retrieval use their own roles and never inherit a text model. The `test`
role is selected explicitly by a test or trial runner; importing pytest does
not change production model selection.

`supported_values` contains lists of enum values, not flags for individual enum
members. It records capabilities for the models we configure; it is not an
automatically discovered inventory of every provider model. A provider entry
does not implement a new adapter. For example, the reserved Anthropic section
does not enable Anthropic extraction.

Use `protocol_defaults` for tuning that applies only to a specific protocol:

```yaml
gpt-4o-mini:
status: active
default_for: []
protocol_defaults:
chat_completions:
temperature: 0.2
```

## Resolution and overrides

An operation selects its configured provider and protocol, then its model by
default role. Model defaults and protocol-specific defaults are combined with
operation defaults; operation defaults take precedence. Explicit caller
arguments take precedence over resolved defaults, including `False` and `0`.
Endpoint overrides remain available in APIs that already expose them.

All packaged operations default to one additional retry after the first attempt.
Set `retries: 0` to disable retries. Temperature is model-specific: modern OpenAI
models leave it unset, legacy GPT-4o Chat Completions keeps `0.2`, and the configured
Gemini URL-retrieval model keeps `0.1`. Unset temperature uses the provider's default.

For extraction, a saved definition's model retains its existing precedence over
the caller's `model`. Both tuning and runtime defaults are resolved for that
selected model before applying explicit caller overrides.
Explicit reasoning takes precedence over saved reasoning, which takes precedence
over configured reasoning. Declared reasoning and verbosity enums are checked
against the requested value; unsupported values are warned about and omitted.
Recipe reasoning includes `max`. Saved `ReasoningEffort` remains `none|low` to
preserve compatibility with existing editors that use that narrower contract.

The config loader rejects malformed declarations, conflicting default roles,
and configured enum defaults outside declared supported values. It does not
contact providers or validate credentials. Provider compatibility and extraction
quality still require live validation when adopting a new model.

Python callers can inspect the effective policy without making an AI request:

```python
from wrangles import ai_config

extraction = ai_config.resolve("extract.ai")
embeddings = ai_config.resolve("embeddings")
trial = ai_config.resolve("extract.ai", role="test")

# Existing helper remains available.
assert ai_config.extract_ai() == extraction
```

`load()` returns a defensive copy of the active YAML structure. `resolve()` and
`extract_ai()` return defensive copies of flat effective policies. Changing a
returned dictionary does not alter cached configuration.

## Integrated callers

| Operation | Callers | Configured settings |
| --- | --- | --- |
| `extract.ai` | Python and recipe extraction | Model, endpoints, model tuning, concurrency, timeout, retries, strictness, storage, cache, prompt |
| `embeddings` | `openai.embeddings` and recipe `create.embeddings` | Provider, model, endpoint, batch size, concurrency, timeout, retries, precision, dimensions, Jina task/normalization/truncation |
| `search.retrieve_link_content` | Python, recipe, and Gemini URL-context client | Model, endpoint/API version, concurrency, timeout, retries, temperature/top-p/top-k/token limits/stop sequences |
| `generate.ai` | Python and recipe generation | Model, endpoint, reasoning/text tuning, concurrency, timeout, retries, strictness |
| `huggingface` | Generic recipe task wrangle | Explicit model, endpoint, timeout, retries, task parameters |

OpenAI embeddings retain `text-embedding-3-small` and their existing dimensions
unless explicitly configured otherwise. Jina requires an explicit model or a
Jina catalog model assigned the `embeddings` role; the package does not invent a
Jina model default. Explicit Jina URLs retain their existing provider inference.

The catalog records Jina v5's `task` enum on the model: `retrieval.query`,
`retrieval.passage`, `text-matching`, `clustering`, and `classification`. Its
default is `text-matching`, matching the provider's documented default. Use
`retrieval.passage` for indexed documents and `retrieval.query` for search queries;
explicit caller `task` overrides the model default. Validation uses the selected
model's catalog enum. Uncataloged older Jina models retain their existing task
validation, including v3's `separation` value. See the
[Jina API schema](https://api.jina.ai/openapi.json) for model-specific values.

Gemini URL retrieval uses the configured model and Google's URL-context tools.
`search.ai_mode` delegates its underlying model to SerpAPI/Google and has no
selectable LLM model in this API.

For Google, `endpoints.base_url` is the SDK service root
`https://generativelanguage.googleapis.com`. The retrieval operation sets
`api_version: v1beta`; the SDK appends the model and method. With the configured
`gemini-3.8-flash`, the complete request URL is
`https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent`.
Both the base URL and version are configurable. See the
[Google API reference](https://ai.google.dev/api/generate-content).
Google model names with or without the SDK's `models/` prefix share the same
catalog defaults. Declare only one spelling for each model in the catalog.

OpenAI and Jina use complete request URLs in `endpoints`: OpenAI
`/v1/responses`, `/v1/chat/completions`, and `/v1/embeddings` on `api.openai.com`,
and Jina `/v1/embeddings` on `api.jina.ai`. Provider documentation links are not
request endpoints.

Hugging Face's generic task wrangle retains its required explicit `model`:
different tasks cannot share one model default. Its operation declares
`requires_model: true`, and model entries can supply task-specific `parameters`
and lifecycle status. Explicit parameters override configured values. Requests
use the configured HF Inference base plus the model ID, currently
`https://router.huggingface.co/hf-inference/models/{model}`. The wrangle preserves
raw JSON results and retries transient failures only. See the
[HF Inference reference](https://huggingface.co/docs/inference-providers/en/providers/hf-inference).

The public `openai.chatGPT` wrapper has been removed. Legacy Chat Completions
remains available through `extract.ai(protocol="chat_completions")`. Extraction
resolves configuration once per operation and uses a private transport for its
individual rows.

Generation remains unreleased. Its operation keeps `low` reasoning; extraction
keeps `none` where supported. Its existing direct-Python and recipe strictness
defaults are represented by `strict` and `recipe_strict`, respectively.

Extraction uses the operation's `defaults` directly. There is no profile registry
or caller-selectable preset behavior; the unused `profile` label has been removed.
A named preset system is outside the current configuration work.

Provider request options use explicit allowlists so runtime settings and catalog
metadata cannot leak into API payloads. Explicit request arguments still override
configured options. Extraction retains `messages`/`examples` aliases and recipe
output-shape controls because existing callers use them. Private transport
arguments and unused generation scaffolding have been removed where redundant.

This catalog governs WranglesPY callers. WranglesXL saved-model authoring and
WranglesJS note-generation calls still select models outside Python. Their model
defaults require a separate client integration. SerpAPI AI Mode and WrangleWorks
saved-model service endpoints own their server-side model selection.

## Version-1 overrides

Existing `version: 1` files with an `extract_ai` section remain supported with
their original replacement semantics. Existing `model_capabilities` flags
remain accepted on that compatibility path and retain packaged capability
inheritance. Generation still follows a version-1 extraction model override;
other operations use packaged defaults because version 1 did not configure them.

For new files, copy version 2 and edit the provider/model catalog and operation
settings. The compatibility API `model_capabilities()` remains available for
existing extraction internals; new configuration should use `supported_values`.
13 changes: 9 additions & 4 deletions docs/extract_ai_configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,12 +6,15 @@ in recipe YAML, see [`extract_ai_user_guide.md`](extract_ai_user_guide.md).
The packaged defaults and base prompt live in
`wrangles/ai_defaults.yml`. Set `WRANGLES_AI_CONFIG` to the path of a
versioned replacement YAML file to override the complete configuration.
Version 2 separates the provider/model catalog from operation settings. See
[AI model configuration](ai_configuration.md) for structure, model-default roles,
and compatibility with existing version-1 override files.

## Runtime defaults

- Provider: `openai`
- Protocol: `responses`
- Model: `gpt-5.4-mini`
- Model: `gpt-6-luna` (also the default for `generate.ai`)
- Default worker concurrency (`default_concurrency`): 32 per `extract.ai` call
- Network timeout per HTTP attempt: 12 seconds
- Retries: 1 additional attempt per row after a retryable failure
Expand All @@ -21,9 +24,10 @@ versioned replacement YAML file to override the complete configuration.
Recipes and Python calls can override these settings individually. Saved XL
models and recipe outputs are compiled through the same definition compiler.

When `threads` is omitted, the call uses `extract_ai.default_concurrency`.
When `threads` is omitted, the call uses
`operations.extract.ai.defaults.default_concurrency` in version 2.
An explicit `threads` value can raise or lower concurrency for that call.
Custom `WRANGLES_AI_CONFIG` files should use `extract_ai.default_concurrency`.
Version-1 files continue to use `extract_ai.default_concurrency`.

Each retry receives the full configured timeout. Queued rows and retry delays
do not consume that timeout, so a complete batch can take much longer than one
Expand Down Expand Up @@ -56,7 +60,8 @@ wrangles:

Direct Python calls can likewise pass `store=False`. A per-call value overrides
the configuration. A replacement `WRANGLES_AI_CONFIG` file should set
`extract_ai.store` explicitly; if it omits that key, the runtime's fallback
`operations.extract.ai.defaults.store` explicitly (`extract_ai.store` in version
1); if it omits that key, the runtime's fallback
is `true`.

Response storage is separate from the local result cache. Cache hits do not
Expand Down
29 changes: 27 additions & 2 deletions docs/extract_ai_user_guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,29 @@ Use `extract.ai` when each input row should produce one or more consistently
named attributes. You can define the attributes in an Excel saved model or
directly in a recipe. Both routes compile to the same output contract.

## Model defaults and capabilities

`wrangles/ai_defaults.yml` is the packaged source of model configuration.
The `extract.ai` default role selects `gpt-6-luna`. Omit `model` in ordinary
recipes to follow the configured default. Saved extraction model selection
retains its existing precedence; saved model names are not automatically migrated.

Version 2 groups models by provider, separates lifecycle status from default
roles, and records model defaults and supported enum values together. Extraction
settings such as concurrency, cache, and the base prompt belong to the operation.
See [AI model configuration](ai_configuration.md) for the schema, caller coverage,
override precedence, test-role selection, and version-1 compatibility.

The public parameters remain `reasoning: {effort: ...}` and
`verbosity: low | medium | high`. Supported values in the catalog describe the
model; they are separate from the value requested by a recipe. Existing saved
model validation and legacy model-family compatibility behavior are preserved.

For each model upgrade, verify capabilities against
[OpenAI's model documentation](https://developers.openai.com/api/docs/models)
and run credentialed extraction checks. Mocked tests verify request construction,
not provider availability or extraction quality.

## Start with the output

Define the result you want before writing general instructions or examples.
Expand Down Expand Up @@ -163,7 +186,10 @@ wrangles.connectors.train.extract.write(
definition,
name="Voltage schema",
variant="ai",
settings={"GPTModel": "gpt-5.4-mini", "ReasoningEffort": "none"},
settings={
"GPTModel": wrangles.ai_config.extract_ai()["model"],
"ReasoningEffort": "none",
},
)
```

Expand Down Expand Up @@ -310,7 +336,6 @@ wrangles:
- Title
- Technical Data
api_key: ${OPENAI_API_KEY}
model: gpt-5.6-luna
reasoning:
effort: low
instructions:
Expand Down
8 changes: 8 additions & 0 deletions pytest-local.ini
Original file line number Diff line number Diff line change
@@ -1,11 +1,18 @@
[pytest]
testpaths =
tests/test_ai_cache.py
tests/test_ai_config.py
tests/test_ai_caller_config.py
tests/test_ai_definition.py
tests/test_container_smoke.py
tests/test_data.py
tests/test_dataframe.py
tests/test_extract_ai_metadata.py
tests/test_extract_ai_metadata_context.py
tests/test_openai_extract_ai.py
tests/test_search-ai_mode.py
tests/test_search_ai_extraction.py
tests/test_standardize_clean_markup.py
"tests/test_back_ processes.py"
tests/recipes
tests/connectors/test_access.py
Expand All @@ -28,6 +35,7 @@ testpaths =
tests/connectors/test_sqlite.py
tests/connectors/test_ssh.py
tests/connectors/test_test.py
tests/connectors/test_train_extract_ai.py
addopts =
-p no:cacheprovider
--ignore=tests/recipes/wrangles/test_extract.py
Expand Down
2 changes: 1 addition & 1 deletion requirements.txt
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ pymysql

# AI / LLM
openai
google-genai
google-genai>=1.24.0

# Search
serpapi
8 changes: 6 additions & 2 deletions scripts/test-local.ps1
Original file line number Diff line number Diff line change
Expand Up @@ -21,13 +21,17 @@ foreach ($name in @(
"AWS_SECRET_ACCESS_KEY",
"AWS_SESSION_TOKEN",
"GEMINI_API_KEY",
"GOOGLE_API_KEY",
"HUGGINGFACE_TOKEN",
"OPENAI_API_KEY",
"SERPAPI_API_KEY",
"WRANGLES_PASSWORD",
"WRANGLES_USER"
"WRANGLES_USER",
"WRANGLES_AI_CONFIG"
)) {
[Environment]::SetEnvironmentVariable($name, $null, "Process")
# Newer .NET versions can preserve an empty environment entry when passed
# $null. Wrangles treats a missing credential differently from an empty one.
Remove-Item -LiteralPath ("Env:" + $name) -ErrorAction SilentlyContinue
}

$localConfig = Join-Path $repoRoot "pytest-local.ini"
Expand Down
Loading
Loading