Skip to content

Centralize AI model configuration and integrate WranglesPY callers - #1200

Merged
ebhills merged 9 commits into
mainfrom
update-Open-AI-model-defaults
Sep 26, 2026
Merged

ebhills merged 9 commits into
mainfrom
update-Open-AI-model-defaults

Conversation

@ebhills

@ebhills ebhills commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

AI model selection and runtime defaults were split between the YAML configuration and individual callers. This PR introduces a provider/model catalog and makes existing WranglesPY AI callers resolve their settings consistently, with explicit caller options taking precedence.

Scope and behavior

  • Group models under providers with lifecycle status, applications, default roles, model/protocol defaults, and supported enum values. Keep operation settings for concurrency, timeouts, retries, caching, and prompts. Validate the catalog locally without a compiler or provider discovery service.

  • Change the extract.ai default from gpt-5.4-mini to gpt-6-luna; use the configured model roles for extraction, generation, embeddings, retrieval, and opt-in tests. Keep text-embedding-3-small as the OpenAI embedding default and large as an option. Google URL retrieval defaults to gemini-3.8-flash.

  • Integrate Python and recipe extraction/generation, OpenAI/Jina embeddings, Gemini URL retrieval, and the existing generic Hugging Face wrangle. Configure provider endpoints, model-specific Jina tasks, and Google SDK base URL/API version. Update Hugging Face to its HF Inference router endpoint while retaining explicit model selection and raw JSON output; require google-genai>=1.24.0.

  • Default to one additional retry; preserve explicit retries: 0 and other explicit overrides. Resolve both tuning and runtime settings after a saved extraction definition selects its model. Forward supported configured request options, fix legacy Chat Completions transport retries, and close Gemini clients.

  • Translate configured, saved, and explicit reasoning/verbosity to the legacy Chat Completions parameter names, preserving caller precedence and capability checks. This sends the configured reasoning_effort: none required for the default model's function-tool requests.

  • Check reasoning and verbosity against model enums, preserve generated output schemas when text options are supplied, remove redundant internal arguments and the unused profile label, and log deprecated-model selections once per operation without blocking execution—including deprecated configured defaults.

  • Clean up ignored seed arguments in Responses integration recipes while retaining legacy Chat Completions and warning-compatibility coverage. Clarify the live dimension-extraction smoke test, require all three correct values, allow only whitespace around/between number and unit, and include actual results in failures.

  • Remove pandas downcasting warnings from Excel reads by filling only columns containing blanks, preserving inferred values/types. Use a sufficiently long synthetic JWT signing key in its recipe-variable test and correct an invalid regex escape.

Compatibility and API impact

  • Remove the public wrangles.openai.chatGPT wrapper and its wrapper-specific tests. External code calling that helper must migrate to wrangles.extract.ai; the internal transport and explicit legacy chat_completions protocol remain supported.
  • Preserve version-1 configuration replacement/capability behavior, explicit custom models, saved-model selection precedence, extraction input/output contracts, and used public aliases and overrides.
  • Keep saved ReasoningEffort limited to none|low for existing editor compatibility; recipe reasoning includes max.
  • Generation remains unreleased, with its existing Python/recipe strictness distinction preserved.

Validation

  • At 9e21d41e: complete credential-free local suite via scripts/test-local.ps1 — 2,315 passed, 6 skipped, 140 deselected.
  • At 7429256e: 252 focused offline tests passed for configuration, extraction, caller integration, and Jina validation. Remote CI subsequently reported 3,017 passed, 6 skipped, 1 failed; the earlier Jina and legacy Chat Completions failures no longer appeared.
  • The remaining remote failure was TestExtractAI.test_ai: none of its three exact strings matched, with no provider error logged. The log did not include returned values, so the precise mismatch is unconfirmed. At b5d07d9b, its dimension instructions and diagnostics are clearer; numeric and unit correctness are still required for all three rows. A remote rerun is needed to establish the live result.
  • At b5d07d9b: 151 focused offline tests passed — 136 extraction tests and 15 Excel/JWT tests. The Excel/JWT tests ran with the reported FutureWarning and JWT key-length warning treated as errors. Six additional offline assertion cases confirmed that only harmless unit spacing is tolerated; incorrect values, units, blanks, and provider errors still fail.
  • Configuration coverage and diff checks passed. The full offline suite was not rerun. Local provider requests are mocked; no local live-provider validation, deployment, or production smoke test was performed. Fresh remote CI results remain pending.

Remaining work and rollback

  • XL Note Generators and the saved-extract-model dropdown/catalog delivery will be next follow-up.
  • LiteLLM/proxy integration, automatic capability discovery, broader extraction API redesign, and legacy protocol retirement are future enhancements.
  • Provider availability, account permissions, and extraction quality for the selected defaults and HF router still need live validation.
  • Model/default choices can be changed through WRANGLES_AI_CONFIG. Reverting this PR restores the previous configuration/caller behavior and removed public helper; there is no persisted-data migration.

@ebhills

ebhills commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator Author

@copilot - OpenAI recently recommended moving to newer models. https://developers.openai.com/api/docs/models

Please check our code to ensure we are up-to-date wherever we name specific models, as well as whether a model supports reasoning (eg in openai_responses.py) and test configuration.

I made one change in this branch already (5.5-mini to gpt-6-luna). This new gpt-6-luna model should be our new default.

Also, identify opportunities to improve how we manage the model names - they change often. perhaps a single YAML config file that we point to from the various code / tests that need a model.

Co-authored-by: ebhills <53243273+ebhills@users.noreply.github.com>

Copilot AI commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

@copilot - OpenAI recently recommended moving to newer models. https://developers.openai.com/api/docs/models

Please check our code to ensure we are up-to-date wherever we name s...

Implemented in 7915e85:

  • Reused wrangles/ai_defaults.yml: extraction and generation now share the gpt-6-luna default; capability flags are configurable, including dated snapshots.
  • Ordinary live tests/examples inherit the default. Explicit overrides, saved-model fixtures, legacy compatibility tests, and separate embedding defaults remain intact.
  • Added offline regression coverage and configuration documentation.

Validation: 1,736 local tests passed (6 skipped, 140 deselected); package/wheel checks passed; secret scan clean; CodeQL: 0 alerts. A fallback read-only review found no significant issues after the automated review service was unavailable. The local-config guard still reports six pre-existing uncovered files, unrelated to this change.

I could not access current OpenAI documentation or verify live API behavior. Can you confirm that gpt-6-luna supports reasoning effort none/low and low text verbosity? The configured flags remain unverified pending that check; I cannot certify the model audit as fully current.

Recommended disposition: Needs decision

Next steps

  1. PR assignee: Verify those parameters against current documentation and run the extraction/generation live checks, then re-request review.
  2. Reviewer: Verify 7915e85, required checks, and provider compatibility; resolve the conversation and submit a fresh approval when satisfied.

@ebhills ebhills changed the title Update OpenAI model defaults Centralize AI model configuration and integrate WranglesPY callers Sep 26, 2026
@ebhills

ebhills commented Sep 26, 2026

Copy link
Copy Markdown
Collaborator Author

Initial description is more current than the Copilot summary because it was revised upon completion of all changes.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The declared Google SDK minimum is incompatible with client cleanup, and Hugging Face retries omit transient response-body failures.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Centralizes AI provider/model configuration and applies it across extraction, generation, embeddings, URL retrieval, and Hugging Face callers.

Changes:

  • Adds the version-2 AI model catalog and resolution APIs.
  • Integrates configured defaults, retries, capabilities, and endpoints.
  • Expands offline tests and documentation.

Recommended disposition: Request changes

Next steps

  1. PR assignee: Raise the Google SDK minimum and expand Hugging Face transient-error retries.
  2. AI agent: @codex address both review threads and add focused regression tests.
  3. Reviewer: Verify fixes, resolve threads, and approve after review is re-requested.
File Description
wrangles/​search.py Resolves retrieval configuration.
wrangles/​recipe_wrangles/​search.py Exposes configured retrieval defaults.
wrangles/​recipe_wrangles/​main.py Configures Hugging Face requests and retries.
wrangles/​recipe_wrangles/​generate.py Integrates recipe generation defaults.
wrangles/​recipe_wrangles/​extract.py Adds maximum reasoning effort.
wrangles/​recipe_wrangles/​create.py Exposes embedding configuration defaults.
wrangles/​openai.py Configures embeddings and private chat transport.
wrangles/​openai_responses.py Adds capability and request-option handling.
wrangles/​generate.py Applies generation policy and model capabilities.
wrangles/​extract.py Resolves effective extraction settings.
wrangles/​clients/​gemini.py Configures Gemini requests and cleanup.
wrangles/​ai_defaults.yml Defines the provider/model catalog.
wrangles/​ai_config.py Implements catalog validation and resolution.
tests/​test_openai_extract_ai.py Expands extraction configuration coverage.
tests/​test_ai_config.py Tests catalog contracts.
tests/​test_ai_caller_config.py Tests embedding and retrieval integration.
tests/​samples/​extract ai judge example.wrgl.yml Uses configured extraction model.
tests/​recipes/​wrangles/​test_huggingface_config.py Tests Hugging Face configuration.
tests/​recipes/​wrangles/​test_generate.py Uses configured generation model.
tests/​recipes/​wrangles/​test_generate_config.py Tests generation policy integration.
tests/​recipes/​wrangles/​test_extract.py Migrates extraction tests to defaults.
tests/​fixtures/​search_ai_mode/​run_search_ai_mode.py Selects the configured test role.
scripts/​test-local.ps1 Clears additional environment overrides.
requirements.txt Sets the Google SDK minimum.
pytest-local.ini Expands credential-free test coverage.
docs/​extract_ai_user_guide.md Documents model defaults.
docs/​extract_ai_configuration.md Updates extraction configuration guidance.
docs/​ai_configuration.md Documents the new catalog.
.gitignore Includes the new documentation file.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread wrangles/recipe_wrangles/main.py Outdated
Add handling for ChunkedEncodingError and ContentDecodingError in exception block.

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
@ebhills
ebhills marked this pull request as draft September 26, 2026 16:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants