Skip to content

feat(sarvam): add Sarvam provider support - #2423

Open
EshwarCVS wants to merge 1 commit into
567-labs:mainfrom
EshwarCVS:feat/sarvam-provider-and-mode-eval
Open

feat(sarvam): add Sarvam provider support#2423
EshwarCVS wants to merge 1 commit into
567-labs:mainfrom
EshwarCVS:feat/sarvam-provider-and-mode-eval

Conversation

@EshwarCVS

Copy link
Copy Markdown

Description

Adds Sarvam AI as a first-class Instructor provider for structured outputs on
Indic chat models (sarvam-30b, sarvam-105b) via Sarvam's OpenAI-compatible
chat completions API.

Users can initialize clients with from_sarvam() or from_provider("sarvam/...").
The integration supports Mode.TOOLS, Mode.JSON_SCHEMA, and Mode.MD_JSON,
with Mode.MD_JSON as the default.

Changes

  • Register Provider.SARVAM in v2 provider specs, modes, and auto_client
  • Add from_sarvam() factory and _build_sarvam() routing
  • Wire Sarvam through OpenAI-compatible handlers
  • Add sarvam-30b and sarvam-105b to supported model lists
  • Add integration docs at docs/integrations/sarvam.md
  • Add v2 unit tests and live smoke tests (simple, stream, retries)

Testing

uv run ruff check instructor tests/llm/test_sarvam tests/v2/test_sarvam_client.py
uv run ruff format instructor tests/llm/test_sarvam tests/v2/test_sarvam_client.py
uv run pytest tests/v2/test_sarvam_client.py tests/v2/test_client_unified.py -k SARVAM -q
uv run pytest tests/llm/test_sarvam/test_simple.py tests/llm/test_sarvam/test_stream.py tests/llm/test_sarvam/test_retries.py -q

Live smoke tests require `SARVAM_API_KEY`.

…ovider

Add first-class Sarvam AI integration for the OpenAI-compatible Indic chat API (sarvam-30b, sarvam-105b) through from_sarvam() and from_provider("sarvam/...").

- Register Sarvam in v2 provider specs, modes, and auto-client routing
- Reuse OpenAI-compatible handlers for TOOLS, JSON_SCHEMA, and MD_JSON
- Default mode to MD_JSON for structured extraction
- Add integration docs, smoke tests, and v2 client/handler coverage
@EshwarCVS

Copy link
Copy Markdown
Author

@jxnl Please review whenever you can.

Open for suggestions. Thanks.

@EshwarCVS

Copy link
Copy Markdown
Author

@jxnl @jxnl-oai @vm

Could you please review
Thanks.

@sujay119

Copy link
Copy Markdown

Does this act as a bridge between both

jxnl added a commit that referenced this pull request Aug 3, 2026
## Summary

- accumulate declared, nested, and unknown numeric OpenAI/Anthropic
usage fields across retries while preserving non-numeric metadata
- add corrective feedback when a Responses API retry receives no tool
call
- preserve raw iterable type hints through sync and async v2
parallel-tool wrappers
- strengthen API-key-free coverage for current and future SDK usage
counters

## Consolidated and superseded items

- closes #2493
- consolidates contributor work from #2498, #2500, and #2501 with
original commit authorship preserved
- supersedes #2497 because it drops unknown `model_extra` counters
- supersedes #2499 because its hand-maintained provider field lists
would drift as SDKs evolve

## Validation

- focused changed-surface suite: `110 passed`
- broad offline v2/coverage suite: `2368 passed, 91 skipped, 73
deselected`
- Ruff check and format check: passed
- scoped `ty check`: passed
- `uv lock --check`: passed
- pre-commit hooks and `git diff --check`: passed

The 73 deselected tests require live provider credentials. An unfiltered
local run confirmed its 22 failures were provider network connections in
the restricted environment; GitHub provider jobs remain the
authoritative validation for those paths.

## Intentionally skipped

- provider additions or expansions: #2436, #2435, #2423, #2409, #2384,
#2322, #2306, #2298, #2283, #2168, #2086; issues #2408, #2383, #2365,
#2260, #2084, #2076
- broad architecture, product, security, or streaming decisions: #2394,
#2392, #2357, #2356, #2355, #2351, #2321, #2307, #2287, #2263; issues
#2479, #2403, #2393, #2391, #2316, #2272, #2056
- dependency batch: #2433
- nontrivial examples and editorial/resource additions: #2468, #2405,
#2401, #2354, #2346, #2311, #2305; issue #2404

These remain open because they need dedicated product, architecture,
provider, security, dependency, or editorial review and are not required
for the `1.15.5` patch release.

<!-- CURSOR_SUMMARY -->
---

> [!NOTE]
> **Medium Risk**
> Changes retry usage totals and reask message content on failure paths;
scope is limited and heavily covered by tests, with no auth or
data-store changes.
> 
> **Overview**
> Bundles three v2 retry and wrapper fixes for a patch release.
> 
> **Retry usage accounting** replaces hand-maintained token field sums
with generic `_accumulate_models` on Pydantic usage objects. Numeric
fields (including nested models and `model_extra` counters) add across
retries; booleans and other non-numeric metadata are not treated as
billable. OpenAI and Anthropic paths share this logic.
> 
> **OpenAI Responses reask** appends a user correction when
`RESPONSES_TOOLS` validation fails but the output has no tool calls
(e.g. reasoning-only), so retries include feedback instead of repeating
the same request.
> 
> **Parallel tools** in `patch_v2` skips `prepare_response_model` and
does not replace `response_model` with the handler’s prepared wrapper
for parallel modes, keeping raw `Iterable[...]` hints so schemas and
parsed results include every member type.
> 
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit
bbddca1. Configure
[here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants