Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
182 changes: 176 additions & 6 deletions docs/ai_answers.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,11 @@ The existing [AI model catalog](ai_configuration.md) supplies
`provider: typesafe`, `protocol: systemone`, and the pinned model `jev-1.13.0`
as defaults. `extract.ai` keeps its current API and behavior.

All four wrangles support [question templates](#questions-from-row-values):
insert row values with `{{ column_name }}`, or use `for_each` to ask a question
about each candidate in a row's list or dictionary. Typesafe caching is
[off by default](#runtime-settings-and-compatibility).

## Input and output schema

Every wrangle takes a nonempty `questions` mapping. Each key is a nonblank
Expand Down Expand Up @@ -286,7 +291,168 @@ complete ordered list inside that question, as shown for `tone` in the
`ai.choose` and `ai.score` require exactly three output names; `ai.true_false`
requires exactly two, in the order shown in the schema table. Every name must
be a nonblank string and unique across all questions in the wrangle. Partial
lists, an empty list, and a single nonempty string are invalid.
lists, an empty list, and a single nonempty string are invalid for ordinary
questions. A question with `for_each` instead creates one column containing a
list or dictionary of answers matching its source collection, as described below.

## Questions from row values

### Insert column values into a question

All four AI wrangles support `{{ column_name }}` placeholders in instructions
and criterion description string values, including those inside nested objects
and arrays. Replace each character outside `A–Z`, `a–z`, `0–9`, and `_` in a
column name with `_`: `Manufacturer Name` becomes `{{ Manufacturer_Name }}`.
The resulting alias must start with a letter or `_`. These are literal
substitutions, not Python expressions or Jinja filters. Missing or ambiguous
references raise an error before provider calls. For example, `Product Name`
and `Product-Name` both produce `Product_Name` and cannot be referenced
unambiguously in the same row. `${...}` remains the syntax for recipe variables.

Question labels, option labels, dictionary keys, and output names stay literal.
Substitution runs once, so braces inside inserted data remain part of that data.
String values are inserted as text; other finite JSON values are serialized as
JSON text. For `ai.score`, rendered criterion descriptions also become the
probability keys and must remain nonblank and unique.

Templates also work without `for_each`. Given columns `Description`,
`Proposed Category`, and `Manufacturer Name`, this recipe asks one question per
row and creates `category_fit`, `category_fit_confidence`, and
`category_fit_probabilities`:

```yaml
wrangles:
- ai.score:
input: Description # Templates may reference other columns in the row.
questions:
category_fit:
instructions: >-
How strongly does this description support "{{ Proposed_Category }}"?
The manufacturer is "{{ Manufacturer_Name }}".
criteria: [Contradicted, Weakly supported, Supported but incomplete, Clearly supported]
```

### Repeat a question with `for_each`

To ask the same question about several candidates, add `for_each`. Its `values`
must be an exact column name; the cell must contain a list or a dictionary.
It does not accept literal candidate collections or JSON-encoded strings.
`variable` names the placeholder for each candidate value and takes precedence
over a column with the same name within that question.

| Field | Meaning |
| --- | --- |
| `for_each.values` | Required exact source column name, including its spaces and punctuation; a record field name in Python. Each cell contains a list or a dictionary with string keys. |
| `for_each.variable` | Required local placeholder name, such as `category`. Use ASCII letters, digits, or underscores, starting with a letter or underscore. |
| `output` | Optional single output column name or one-item list. Omitted or blank uses the question label. |

**Example input**

```yaml
- Description: CONTAINER 24X15X11 CHARCOAL
Manufacturer Name: Example Supply
CandidateCategories:
category_1: Containers
category_2: Shelves
category_3: "" # Reserve a named slot without asking a question.
```

**Template**

```yaml
wrangles:
- ai.score:
input: Description # Shared input sent to the provider.
# cache: true # Optional opt-in; Typesafe caching defaults to off.
questions:
category_fit:
for_each:
values: CandidateCategories # Exact column name, resolved per row.
variable: category # Current list item or dictionary value.
instructions: >-
How strongly does this description support "{{ category }}"?
The manufacturer is "{{ Manufacturer_Name }}".
Distinguish missing information from contradictory information.
criteria:
- Contradicted
- Weakly supported
- Supported but incomplete
- Clearly supported
# output: Category Fits # Optional single name; defaults to category_fit.
```

The full row supplies template references and the candidate collection even
when `input` selects only `Description`. Referenced values become part of the
rendered questions sent to the provider. Every generated question, including
ordinary questions in the same wrangle, is sent together in one request per
row. Candidate lists can differ between rows.

**Illustrative `category_fit` cell**

```yaml
category_1:
value: Containers
score: 2.77
confidence: 0.77
probabilities:
Contradicted: 0.01
Weakly supported: 0.03
Supported but incomplete: 0.14
Clearly supported: 0.82
category_2:
value: Shelves
score: 0.23
confidence: 0.77
probabilities:
Contradicted: 0.82
Weakly supported: 0.14
Supported but incomplete: 0.03
Clearly supported: 0.01
category_3: {}
```

A dictionary-valued cell produces a dictionary of answers using the input's
original keys. Its values supply `{{ category }}`. A list-valued cell such as
`[Containers, Shelves]` produces a list of answer dictionaries. Each answer
retains its `value` in either shape. Both input shapes preserve their order.

An empty or whitespace-only candidate string reserves a slot without generating
a provider question. That slot returns `{}` at its original dictionary key or
list position. For example, `{category_1: Containers, category_2: ""}` returns
an answer under `category_1` and `{}` under `category_2`; `[Containers, ""]`
returns `[<answer>, {}]`. Supply these placeholders in the input when you need
consistent named slots. Missing keys and list positions are not added
automatically. Other JSON values keep their usual meaning and are not skipped.

An empty source dictionary returns `{}` and an empty source list returns `[]`.
When every question expands to an empty collection or only blank-string
placeholders, no provider request is needed.

Repeated `choose`, `score`, and `true_false` questions retain their usual
answer fields inside each item. For example, a repeated `true_false` answer
contains `value`, `probability_true`, and `true_criteria`. An omitted or blank
`output` uses the question label; a string or one-item list renames this single
column. Ordinary questions retain their existing positional output names.

Use `split.dictionary` directly on a dictionary result to create columns such
as `category_1`, each containing that candidate's answer dictionary. A further
`split.dictionary` step can expose its answer fields with a prefix such as
`category_1_score`. For example:

```yaml
- split.dictionary:
input: category_fit # Creates category_1, category_2, and category_3.
- split.dictionary:
input: category_1
output:
- "*": category_1_* # Exposes category_1_value, category_1_score, etc.
```

No `split.list` step is needed for dictionary sources. The runnable
[category scoring fixture](../tests/fixtures/ai_question_templates/README.md)
includes `ai_category_judge.recipe`, a detailed four-level rubric, a blank
candidate slot, empty collections, and a small Python runner. Its two steps
score the candidates and split the keyed answers into columns.

## Python

Expand All @@ -310,6 +476,8 @@ print(answers["product_class"]["choice"])

A string or dictionary input returns named answers; a list returns one set of
named answers per input, in order. Question definitions match the YAML schema.
For templates, pass dictionary records containing the referenced fields;
`for_each.values` remains a field name in Python calls as well.

## Runtime settings and compatibility

Expand All @@ -320,11 +488,13 @@ are 10 threads, 30 seconds, and one retry. Use `retries: 0` to disable retries.
Only transient failures are retried; invalid credentials, question definitions,
and malformed results fail. Errors do not become fabricated answers.

Successful results are cached in memory by request identity. The packaged
defaults enable a one-hour TTL, up to 512 entries, and up to 65,536 bytes per
cached value. Concurrent duplicate requests share one in-flight request.
Use `cache: false` or `cache_ttl` in seconds to override the corresponding
catalog values. Process-level environment controls take precedence:
Caching is off by default for all four Typesafe answer wrangles. Set
`cache: true` in a recipe or `cache=True` in Python to enable it. When enabled,
successful results are cached in memory by request identity, and concurrent
duplicate requests share one in-flight request. The configured limits are a
one-hour TTL, up to 512 entries, and up to 65,536 bytes per cached value.
Use `cache_ttl` in seconds to override the lifetime. Process-level environment
controls take precedence:

- `WRANGLES_AI_CACHE_ENABLED`
- `WRANGLES_AI_CACHE_TTL_SECONDS`
Expand Down
127 changes: 74 additions & 53 deletions docs/ai_configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -137,37 +137,47 @@ returned dictionary does not alter cached configuration.
| `generate.ai` | Python and recipe generation | Model, endpoint, reasoning/text tuning, concurrency, timeout, retries, strictness |
| `huggingface` | Generic recipe task wrangle | Explicit model, endpoint, timeout, retries, task parameters |

The four answer wrangles provide structured answers to common types of
questions through [Typesafe](https://docs.typesafe.ai/introduction).
`ai.choose` selects the best-fitting option, `ai.score` locates the input along
ordered criteria, and `ai.true_false` returns the probability that a statement
is true. `ai.answers` answers any combination of these question types together.
They use Typesafe's `systemone` protocol at
`https://api.typesafe.ai/v1/systemone`, with `jev-1.13.0` as their pinned default
model. Each operation has its own default role; changing an extraction or global
model does not change these operations. Explicit unlisted Typesafe model names
remain available. This adapter supports only `provider: typesafe` and
`protocol: systemone`; adding another catalog provider alone does not implement
an adapter for it.
Provider request options use explicit allowlists so runtime settings and catalog
metadata cannot leak into API payloads. Explicit request arguments still override
configured options. Extraction retains `messages`/`examples` aliases and recipe
output-shape controls because existing callers use them. Private transport
arguments and unused generation scaffolding have been removed where redundant.

Their packaged runtime defaults are 10 concurrent requests, a 30-second timeout,
and one additional attempt after a transient failure. Successful results use a
bounded in-memory cache with a one-hour TTL, at most 512 entries, and a maximum
value size of 65,536 bytes. Duplicate in-flight requests share their result, and
periodic cache logging is disabled. `WRANGLES_AI_CACHE_*` environment controls
apply to these operations independently of the existing
`WRANGLES_EXTRACT_AI_CACHE_*` controls. See [AI answers](ai_answers.md) for
question schemas, examples, output columns, and cache overrides.
This catalog governs WranglesPY callers. WranglesXL saved-model authoring and
WranglesJS note-generation calls still select models outside Python. Their model
defaults require a separate client integration. SerpAPI AI Mode and WrangleWorks
saved-model service endpoints own their server-side model selection.

Keep credentials outside the catalog. Supply `api_key` explicitly, use the local
`TYPESAFE_API_KEY` environment variable, or use `api_key: ${TYPESAFE_API_KEY}` in
a hosted recipe with that managed secret. Hosted secrets are supplied as recipe
variables; they are not placed in the worker's environment.
## Provider settings

### OpenAI

OpenAI embeddings retain `text-embedding-3-small` and their existing dimensions
unless explicitly configured otherwise. Jina requires an explicit model or a
Jina catalog model assigned the `embeddings` role; the package does not invent a
Jina model default. Explicit Jina URLs retain their existing provider inference.
unless explicitly configured otherwise.

OpenAI uses complete request URLs in `endpoints`: `/v1/responses`,
`/v1/chat/completions`, and `/v1/embeddings` on `api.openai.com`.
Provider documentation links are not request endpoints.

The public `openai.chatGPT` wrapper has been removed. Legacy Chat Completions
remains available through `extract.ai(protocol="chat_completions")`. Extraction
resolves configuration once per operation and uses a private transport for its
individual rows.

Generation remains unreleased. Its operation keeps `low` reasoning; extraction
keeps `none` where supported. Its existing direct-Python and recipe strictness
defaults are represented by `strict` and `recipe_strict`, respectively.

Extraction uses the operation's `defaults` directly. There is no profile registry
or caller-selectable preset behavior; the unused `profile` label has been removed.
A named preset system is outside the current configuration work.

### Jina

Jina requires an explicit model or a Jina catalog model assigned the `embeddings`
role; the package does not invent a Jina model default. Explicit Jina URLs retain
their existing provider inference. Jina uses the complete `/v1/embeddings`
request URL on `api.jina.ai` in `endpoints`.

The catalog records Jina v5's `task` enum on the model: `retrieval.query`,
`retrieval.passage`, `text-matching`, `clustering`, and `classification`. Its
Expand All @@ -178,6 +188,8 @@ model's catalog enum. Uncataloged older Jina models retain their existing task
validation, including v3's `separation` value. See the
[Jina API schema](https://api.jina.ai/openapi.json) for model-specific values.

### Google Gemini

Gemini URL retrieval uses the configured model and Google's URL-context tools.
`search.ai_mode` delegates its underlying model to SerpAPI/Google and has no
selectable LLM model in this API.
Expand All @@ -192,10 +204,38 @@ Both the base URL and version are configurable. See the
Google model names with or without the SDK's `models/` prefix share the same
catalog defaults. Declare only one spelling for each model in the catalog.

OpenAI and Jina use complete request URLs in `endpoints`: OpenAI
`/v1/responses`, `/v1/chat/completions`, and `/v1/embeddings` on `api.openai.com`,
and Jina `/v1/embeddings` on `api.jina.ai`. Provider documentation links are not
request endpoints.
### Typesafe

The four answer wrangles provide structured answers to common types of
questions through [Typesafe](https://docs.typesafe.ai/introduction).
`ai.choose` selects the best-fitting option, `ai.score` locates the input along
ordered criteria, and `ai.true_false` returns the probability that a statement
is true. `ai.answers` answers any combination of these question types together.
They use Typesafe's `systemone` protocol at
`https://api.typesafe.ai/v1/systemone`, with `jev-1.13.0` as their pinned default
model. Each operation has its own default role; changing an extraction or global
model does not change these operations. Explicit unlisted Typesafe model names
remain available. This adapter supports only `provider: typesafe` and
`protocol: systemone`; adding another catalog provider alone does not implement
an adapter for it.

Their packaged runtime defaults are 10 concurrent requests, a 30-second timeout,
and one additional attempt after a transient failure. Caching is off by default
(`defaults.cache.enabled: false`). Set `cache: true` in a recipe or `cache=True`
in Python to enable a bounded in-memory cache with a one-hour TTL, at most 512
entries, and a maximum value size of 65,536 bytes. When enabled, duplicate
in-flight requests share their result. Periodic cache logging is disabled.
`WRANGLES_AI_CACHE_*` environment controls take precedence over caller and
catalog settings, independently of the existing `WRANGLES_EXTRACT_AI_CACHE_*`
controls. See [AI answers](ai_answers.md) for question schemas, examples, output
columns, and cache overrides.

Keep credentials outside the catalog. Supply `api_key` explicitly, use the local
`TYPESAFE_API_KEY` environment variable, or use `api_key: ${TYPESAFE_API_KEY}` in
a hosted recipe with that managed secret. Hosted secrets are supplied as recipe
variables; they are not placed in the worker's environment.

### Hugging Face

Hugging Face's generic task wrangle retains its required explicit `model`:
different tasks cannot share one model default. Its operation declares
Expand All @@ -206,29 +246,10 @@ use the configured HF Inference base plus the model ID, currently
raw JSON results and retries transient failures only. See the
[HF Inference reference](https://huggingface.co/docs/inference-providers/en/providers/hf-inference).

The public `openai.chatGPT` wrapper has been removed. Legacy Chat Completions
remains available through `extract.ai(protocol="chat_completions")`. Extraction
resolves configuration once per operation and uses a private transport for its
individual rows.

Generation remains unreleased. Its operation keeps `low` reasoning; extraction
keeps `none` where supported. Its existing direct-Python and recipe strictness
defaults are represented by `strict` and `recipe_strict`, respectively.

Extraction uses the operation's `defaults` directly. There is no profile registry
or caller-selectable preset behavior; the unused `profile` label has been removed.
A named preset system is outside the current configuration work.
### Anthropic

Provider request options use explicit allowlists so runtime settings and catalog
metadata cannot leak into API payloads. Explicit request arguments still override
configured options. Extraction retains `messages`/`examples` aliases and recipe
output-shape controls because existing callers use them. Private transport
arguments and unused generation scaffolding have been removed where redundant.

This catalog governs WranglesPY callers. WranglesXL saved-model authoring and
WranglesJS note-generation calls still select models outside Python. Their model
defaults require a separate client integration. SerpAPI AI Mode and WrangleWorks
saved-model service endpoints own their server-side model selection.
The Anthropic provider entry is reserved for a future adapter. It does not
enable Anthropic extraction or other runtime support.

## Version-1 overrides

Expand Down
Loading
Loading