Repository navigation
Centralize AI model configuration and integrate WranglesPY callers - #1200
Conversation
|
@copilot - OpenAI recently recommended moving to newer models. https://developers.openai.com/api/docs/models Please check our code to ensure we are up-to-date wherever we name specific models, as well as whether a model supports reasoning (eg in openai_responses.py) and test configuration. I made one change in this branch already (5.5-mini to gpt-6-luna). This new gpt-6-luna model should be our new default. Also, identify opportunities to improve how we manage the model names - they change often. perhaps a single YAML config file that we point to from the various code / tests that need a model. |
Co-authored-by: ebhills <53243273+ebhills@users.noreply.github.com>
Implemented in
Validation: 1,736 local tests passed (6 skipped, 140 deselected); package/wheel checks passed; secret scan clean; CodeQL: 0 alerts. A fallback read-only review found no significant issues after the automated review service was unavailable. The local-config guard still reports six pre-existing uncovered files, unrelated to this change. I could not access current OpenAI documentation or verify live API behavior. Can you confirm that Recommended disposition: Needs decision Next steps
|
|
Initial description is more current than the Copilot summary because it was revised upon completion of all changes. |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The declared Google SDK minimum is incompatible with client cleanup, and Hugging Face retries omit transient response-body failures.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Centralizes AI provider/model configuration and applies it across extraction, generation, embeddings, URL retrieval, and Hugging Face callers.
Changes:
- Adds the version-2 AI model catalog and resolution APIs.
- Integrates configured defaults, retries, capabilities, and endpoints.
- Expands offline tests and documentation.
Recommended disposition: Request changes
Next steps
- PR assignee: Raise the Google SDK minimum and expand Hugging Face transient-error retries.
- AI agent:
@codex address both review threads and add focused regression tests. - Reviewer: Verify fixes, resolve threads, and approve after review is re-requested.
| File | Description |
|---|---|
wrangles/search.py |
Resolves retrieval configuration. |
wrangles/recipe_wrangles/search.py |
Exposes configured retrieval defaults. |
wrangles/recipe_wrangles/main.py |
Configures Hugging Face requests and retries. |
wrangles/recipe_wrangles/generate.py |
Integrates recipe generation defaults. |
wrangles/recipe_wrangles/extract.py |
Adds maximum reasoning effort. |
wrangles/recipe_wrangles/create.py |
Exposes embedding configuration defaults. |
wrangles/openai.py |
Configures embeddings and private chat transport. |
wrangles/openai_responses.py |
Adds capability and request-option handling. |
wrangles/generate.py |
Applies generation policy and model capabilities. |
wrangles/extract.py |
Resolves effective extraction settings. |
wrangles/clients/gemini.py |
Configures Gemini requests and cleanup. |
wrangles/ai_defaults.yml |
Defines the provider/model catalog. |
wrangles/ai_config.py |
Implements catalog validation and resolution. |
tests/test_openai_extract_ai.py |
Expands extraction configuration coverage. |
tests/test_ai_config.py |
Tests catalog contracts. |
tests/test_ai_caller_config.py |
Tests embedding and retrieval integration. |
tests/samples/extract ai judge example.wrgl.yml |
Uses configured extraction model. |
tests/recipes/wrangles/test_huggingface_config.py |
Tests Hugging Face configuration. |
tests/recipes/wrangles/test_generate.py |
Uses configured generation model. |
tests/recipes/wrangles/test_generate_config.py |
Tests generation policy integration. |
tests/recipes/wrangles/test_extract.py |
Migrates extraction tests to defaults. |
tests/fixtures/search_ai_mode/run_search_ai_mode.py |
Selects the configured test role. |
scripts/test-local.ps1 |
Clears additional environment overrides. |
requirements.txt |
Sets the Google SDK minimum. |
pytest-local.ini |
Expands credential-free test coverage. |
docs/extract_ai_user_guide.md |
Documents model defaults. |
docs/extract_ai_configuration.md |
Updates extraction configuration guidance. |
docs/ai_configuration.md |
Documents the new catalog. |
.gitignore |
Includes the new documentation file. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Add handling for ChunkedEncodingError and ContentDecodingError in exception block. Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

AI model selection and runtime defaults were split between the YAML configuration and individual callers. This PR introduces a provider/model catalog and makes existing WranglesPY AI callers resolve their settings consistently, with explicit caller options taking precedence.
Scope and behavior
Group models under providers with lifecycle status,
applications, default roles, model/protocol defaults, and supported enum values. Keep operation settings for concurrency, timeouts, retries, caching, and prompts. Validate the catalog locally without a compiler or provider discovery service.Change the
extract.aidefault fromgpt-5.4-minitogpt-6-luna; use the configured model roles for extraction, generation, embeddings, retrieval, and opt-in tests. Keeptext-embedding-3-smallas the OpenAI embedding default and large as an option. Google URL retrieval defaults togemini-3.8-flash.Integrate Python and recipe extraction/generation, OpenAI/Jina embeddings, Gemini URL retrieval, and the existing generic Hugging Face wrangle. Configure provider endpoints, model-specific Jina tasks, and Google SDK base URL/API version. Update Hugging Face to its HF Inference router endpoint while retaining explicit model selection and raw JSON output; require
google-genai>=1.24.0.Default to one additional retry; preserve explicit
retries: 0and other explicit overrides. Resolve both tuning and runtime settings after a saved extraction definition selects its model. Forward supported configured request options, fix legacy Chat Completions transport retries, and close Gemini clients.Translate configured, saved, and explicit reasoning/verbosity to the legacy Chat Completions parameter names, preserving caller precedence and capability checks. This sends the configured
reasoning_effort: nonerequired for the default model's function-tool requests.Check reasoning and verbosity against model enums, preserve generated output schemas when text options are supplied, remove redundant internal arguments and the unused profile label, and log deprecated-model selections once per operation without blocking execution—including deprecated configured defaults.
Clean up ignored
seedarguments in Responses integration recipes while retaining legacy Chat Completions and warning-compatibility coverage. Clarify the live dimension-extraction smoke test, require all three correct values, allow only whitespace around/between number and unit, and include actual results in failures.Remove pandas downcasting warnings from Excel reads by filling only columns containing blanks, preserving inferred values/types. Use a sufficiently long synthetic JWT signing key in its recipe-variable test and correct an invalid regex escape.
Compatibility and API impact
wrangles.openai.chatGPTwrapper and its wrapper-specific tests. External code calling that helper must migrate towrangles.extract.ai; the internal transport and explicit legacychat_completionsprotocol remain supported.ReasoningEffortlimited tonone|lowfor existing editor compatibility; recipe reasoning includesmax.Validation
9e21d41e: complete credential-free local suite viascripts/test-local.ps1— 2,315 passed, 6 skipped, 140 deselected.7429256e: 252 focused offline tests passed for configuration, extraction, caller integration, and Jina validation. Remote CI subsequently reported 3,017 passed, 6 skipped, 1 failed; the earlier Jina and legacy Chat Completions failures no longer appeared.TestExtractAI.test_ai: none of its three exact strings matched, with no provider error logged. The log did not include returned values, so the precise mismatch is unconfirmed. Atb5d07d9b, its dimension instructions and diagnostics are clearer; numeric and unit correctness are still required for all three rows. A remote rerun is needed to establish the live result.b5d07d9b: 151 focused offline tests passed — 136 extraction tests and 15 Excel/JWT tests. The Excel/JWT tests ran with the reported FutureWarning and JWT key-length warning treated as errors. Six additional offline assertion cases confirmed that only harmless unit spacing is tolerated; incorrect values, units, blanks, and provider errors still fail.Remaining work and rollback
WRANGLES_AI_CONFIG. Reverting this PR restores the previous configuration/caller behavior and removed public helper; there is no persisted-data migration.