Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 34 additions & 10 deletions docs/ai_configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,8 +90,9 @@ operation defaults; operation defaults take precedence. Explicit caller
arguments take precedence over resolved defaults, including `False` and `0`.
Endpoint overrides remain available in APIs that already expose them.

All packaged operations default to one additional retry after the first attempt.
Set `retries: 0` to disable retries. Temperature is model-specific: modern OpenAI
Packaged operations default to one additional retry after the first attempt,
except Gemini URL retrieval, which defaults to no retries. Set `retries: 0`
to disable retries. Temperature is model-specific: modern OpenAI
models and Gemini URL retrieval leave it unset, while legacy GPT-4o Chat Completions
keeps `0.2`. Unset temperature uses the provider's default (`1.0` for Gemini 3).

Expand Down Expand Up @@ -133,13 +134,15 @@ returned dictionary does not alter cached configuration.
| `ai.choose`, `ai.score`, `ai.true_false`, `ai.answers` | Python and recipe structured answers | Provider, model, endpoint, concurrency, timeout, retries, cache |
| `extract.ai` | Python and recipe extraction | Model, endpoints, model tuning, concurrency, timeout, retries, strictness, storage, cache, prompt |
| `embeddings` | `openai.embeddings` and recipe `create.embeddings` | Provider, model, endpoint, batch size, concurrency, timeout, retries, precision, dimensions, Jina task/normalization/truncation |
| `search.retrieve_link_content` | Python, recipe, and Gemini URL-context client | Model, endpoint/API version, concurrency, timeout, retries, temperature/top-p/top-k/token limits/stop sequences |
| `search.retrieve_link_content` | Python, recipe, and Gemini URL-context client | Model, endpoint/API version, concurrency, per-URL deadline, retries, thinking level, temperature/top-p/top-k/token limits/stop sequences |
| `generate.ai` | Python and recipe generation | Model, endpoint, reasoning/text tuning, concurrency, timeout, retries, strictness |
| `huggingface` | Generic recipe task wrangle | Explicit model, endpoint, timeout, retries, task parameters |

Provider request options use explicit allowlists so runtime settings and catalog
metadata cannot leak into API payloads. Explicit request arguments still override
configured options. Extraction retains `messages`/`examples` aliases and recipe
Configured provider request options use explicit allowlists so runtime settings
and catalog metadata cannot leak into API payloads. Explicit request arguments
still override configured options. Gemini retrieval also accepts caller-supplied
`GenerateContentConfig` options through keyword arguments. Extraction retains
`messages`/`examples` aliases and recipe
output-shape controls because existing callers use them. Private transport
arguments and unused generation scaffolding have been removed where redundant.

Expand Down Expand Up @@ -190,15 +193,36 @@ validation, including v3's `separation` value. See the

### Google Gemini

Gemini URL retrieval uses the configured model and Google's URL-context tools.
Gemini URL retrieval defaults to `gemini-3.5-flash` and uses Google's
URL-context tools. Its packaged settings explicitly select
`thinking_level: minimal`, `request_timeout_seconds: 10`, and `retries: 0`.
The Google SDK dependency requires version `1.64.0` or later for these controls.
Use `model_id` to select another model, such as `gemini-3.6-flash`, and
`thinking_level` to override the configured thinking level. The wrapper accepts
`minimal`, `low`, `medium`, and `high`; actual support depends on the model.
Explicit `request_timeout_seconds` overrides the configured deadline.

The deadline covers one URL's provider request, including all configured retries.
A timeout produces the existing per-URL `Failure` result and preserves output
ordering. Cancellation and client cleanup may take a little longer. Queued URLs
receive their own deadline when their worker starts, so a batch can exceed 10
seconds when it needs multiple waves of requests.

Additional caller keyword arguments are forwarded to the SDK's
`GenerateContentConfig` and override configured generation defaults, including
options such as `max_output_tokens`, `top_p`, and `temperature`. Retrieval's
prompt, URL-context tool, response format, and HTTP settings remain managed by
the wrapper. An explicit `thinking_config` replaces the thinking configuration
as a whole, including any selected `thinking_level`. See the
[retrieval guide](search_retrieve_link_content.md) for recipe and Python usage.
`search.ai_mode` delegates its underlying model to SerpAPI/Google and has no
selectable LLM model in this API.

For Google, `endpoints.base_url` is the SDK service root
`https://generativelanguage.googleapis.com`. The retrieval operation sets
`api_version: v1beta`; the SDK appends the model and method. With the configured
`gemini-3.8-flash`, the complete request URL is
`https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent`.
`api_version: v1beta`; the SDK appends the model and method. With the default
`gemini-3.5-flash`, the complete request URL is
`https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent`.
Both the base URL and version are configurable. See the
[Google API reference](https://ai.google.dev/api/generate-content).
Google model names with or without the SDK's `models/` prefix share the same
Expand Down
46 changes: 45 additions & 1 deletion docs/search_retrieve_link_content.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,8 +90,52 @@ dual output mode: `output: [page_data, page_text]` stores the result list in
their existing meanings. See [AI configuration](ai_configuration.md) for model
and runtime settings.

## Model, thinking, and request options

The packaged defaults are `gemini-3.5-flash`, `thinking_level: minimal`, a
10-second deadline per URL, and no retries. Omitted options use the active AI
configuration, so a replacement catalog can change these defaults. Override
`model_id`, `thinking_level`, or `request_timeout_seconds` in the recipe when
needed. For example:

```yaml
- search.retrieve_link_content:
input: URL
output: Page Content
api_key: ${GEMINI_API_KEY}
model_id: gemini-3.6-flash
thinking_level: minimal
request_timeout_seconds: 10
max_output_tokens: 2048
seed: 42
```

`thinking_level` accepts `minimal`, `low`, `medium`, or `high`; support depends
on the selected model. Additional options such as `max_output_tokens`, `top_p`,
and `temperature` are passed as top-level Gemini `GenerateContentConfig`
settings. Explicit options override configured generation defaults. The
installed Google SDK validates these options; their availability depends on
the selected model and SDK version.

Retrieval manages `system_instruction`, `tools`, `response_mime_type`,
`response_modalities`, and `http_options`; those names and their SDK aliases
cannot be passed as additional options. Use `prompt`, `output_format`, and
`request_timeout_seconds` for the corresponding controls. For advanced thinking
settings, an explicit `thinking_config` replaces the entire thinking
configuration, including any `thinking_level`, rather than merging with it.

The positive, finite `request_timeout_seconds` value limits the total provider
request time for each URL, including any retries enabled in the AI catalog.
When the deadline expires, that URL returns a `Failure` result with an error
and no extracted content; its position in the output is preserved. Cancellation
and client cleanup may add a little time after the deadline. The deadline
applies separately to each URL after its worker starts, so batches that need
several waves of concurrent requests can take longer than 10 seconds.

## Direct Python calls

Column substitution applies to recipe execution, where a full row is available.
Direct calls to `wrangles.search.retrieve_link_content(...)` continue to use
the supplied prompt literally and do not resolve column placeholders.
the supplied prompt literally and do not resolve column placeholders. The
`thinking_level`, `request_timeout_seconds`, and additional Gemini generation
options are also available as Python keyword arguments.
1 change: 1 addition & 0 deletions pytest-local.ini
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ testpaths =
tests/test_dataframe.py
tests/test_extract_ai_metadata.py
tests/test_extract_ai_metadata_context.py
tests/test_lookup_variants.py
tests/test_openai_extract_ai.py
tests/test_search-ai_mode.py
tests/test_search_ai_extraction.py
Expand Down
2 changes: 1 addition & 1 deletion requirements.txt
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ pymysql

# AI / LLM
openai
google-genai>=1.24.0
google-genai>=1.64.0

# Search
serpapi
Loading
Loading