Skip to content

Release v1.1.0: fix payer eligibility and published tool paths - #53

Merged
haolin-chen-actava merged 14 commits into
mainfrom
codex/bug-fixes
Oct 7, 2026
Merged

haolin-chen-actava merged 14 commits into
mainfrom
codex/bug-fixes

Conversation

@haolin-chen-actava

@haolin-chen-actava haolin-chen-actava commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Published task instructions point to a tool-reference directory that the runtime image did not provide. The pa_t035 payer task also hides its coverage termination date and expects a final state that payer tools cannot produce. This PR restores the published path and makes the eligibility failure visible in the existing request form. The payer can end the case in the existing denied state and deliver a denial letter.

Changes

  • Mark the Python package, module version, and lock file as 1.1.0; add the changelog and upgrade commands.
  • Alias /opt/healthverse-task-assets to /opt/chi-bench-task-assets. Per-task tool-reference selection and tool-name rewriting remain visible through the published path.
  • Allow completed intake with failed eligibility to supply the denial determination and reason ELIG_TERMED. Block approval and clinical routing after failed eligibility. Create a pending letter request; the agent must generate and deliver the letter.
  • Verify the intake record as the determination source, including matching case, completed intake, failed eligibility, and denial without override.
  • Add coverage status and known coverage dates to the existing request-form template. Keep get_member_coverage outside payer MCP.
  • Provide repair_pa_t035.py to restore termination date 2025-11-30 from the original synthesis facts, correct the unreachable requirement_checked expectation to denial with a denial letter, and repair task and marathon copies.
  • Provide regenerate_payer_request_forms.py for the 25 payer request forms and their marathon copies. Remove rendering timestamps so unchanged content retains the same PDF hash.
  • Retain the competition Rules and Overview references requested for this branch, align them with Kaggle, and link the organizer's submission and API-cost replies. Correct the submission example's stale package command to prepare.
  • Document the existing P2P behavior. No P2P runtime or sign-off rule changes are included.

Data upgrade

The downloaded fixtures are ignored and are not part of this PR. After downloading the public dataset and approved handbook, run:

uv sync --extra dev
uv run playwright install chromium
uv run python scripts/repair_pa_t035.py --data-dir data
uv run python scripts/regenerate_payer_request_forms.py --data-dir data
uv run cb data verify
uv run cb docker build

Use playwright install --with-deps chromium on Linux if needed. Runtime 1.1.0 is separate from dataset revision chi-bench-v1.0.0. The scripts apply local corrections; this PR does not publish a new Hugging Face revision or update the Harbor hub listing. Preserve old trial evidence and record the runtime commit and image digest for new runs.

Validation

  • Full default suite: 258 passed, 3 skipped, 1 deselected.
  • Ruff lint and format checks, lock consistency check, and git diff --check: passed.
  • Actual downloaded data: both repair scripts repeat without changing files; all 50 request PDFs retain text and page counts; all 25 marathon copies match their task copies. Data verification passes.
  • Docker test image built from the existing local runtime with current source, package metadata, and repaired task data: version 1.1.0, all 50 PDF hashes match, pa_t035 has the termination date and denial expectation, published tool-reference alias resolves to the selected payer reference.
  • Live MCP test in that image: 41 payer tools, coverage lookup absent, eligibility failure recorded, case denied, denial letter generated and delivered. All 18 intake/outcome checks pass; no clinical review records are created.
  • The six actual public P2P task contracts were checked against runtime finalization and stage checks during the preceding investigation. No conflicting signed_off expectation was found; P2P code is unchanged.

Build limit: The normal Docker build stalled at Docker Hub frontend resolution. The syntax-free retry stalled at the python:3.12-slim metadata lookup. Both owned builds were stopped. The successful test image uses the already available local runtime base, with an offline package install. This does not verify a fresh base-image download or dependency installation. No agent-model trial or LLM judge call was made for these checks.

Policy follow-up: Live inspection confirms CHI-Bench is a Community Hackathon with Writeup submissions, while Rules Section 9 describes CSV upload and Kaggle Accuracy scoring. Section 11 requires a Kaggle entry in addition to private materials. The email reply does not remove that entry requirement. The cost thread also says all 75 public results are not required for private evaluation. The communication drafts recommend correcting the Rules to require a Writeup plus the private email package; this is not adopted policy. This PR preserves the published CSV requirements in the local reference.

Announcements and FAQ drafts are outside the repository. No competition message has been posted.

@haolin-chen-actava
haolin-chen-actava marked this pull request as ready for review October 7, 2026 06:12
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T06:18:18.523382Z 926e2fa Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 926e2fa498

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docker/Dockerfile
COPY data/care_management/shared/tool_reference.md /opt/chi-bench-task-assets/tool_reference.md
# Published task instructions still use the healthverse asset path. Alias the
# directory so per-task selection and tool-name rewriting remain visible there.
RUN ln -s /opt/chi-bench-task-assets /opt/healthverse-task-assets

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Mirror the tool-reference alias in the Harbor image

This only fixes images built from docker/Dockerfile. I checked the published-task export path: scripts/export_harbor_tasks.py copies docker/Dockerfile.harbor, which creates /opt/chi-bench-task-assets but never creates /opt/healthverse-task-assets. Consequently, newly exported Harbor tasks still cannot read the path advertised in their instructions; add the same directory alias to that image definition.

Useful? React with 👍 / 👎.

Comment thread README.md
Comment on lines +98 to +99
uv run python scripts/repair_pa_t035.py
uv run python scripts/regenerate_payer_request_forms.py

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Repair pa_t035 in the Harbor build path

These commands repair only a user's local checkout. The published-task image built by docker/Dockerfile.harbor still downloads the unchanged chi-bench-v1.0.0 dataset and copies it directly without running either script, so a fresh Harbor export retains pa_t035's missing termination evidence and unreachable requirement_checked expectation. Invoke the repair during that build or publish and pin a repaired dataset revision before releasing this behavior.

Useful? React with 👍 / 👎.

source_record_id=intake.id,
recommendation="deny",
original_recommendation="deny",
rationale="ELIG_TERMED: Member coverage is inactive on the requested date of service.",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Derive the eligibility denial reason from coverage facts

Any completed intake with member_eligible=False is now labeled ELIG_TERMED, even when eligibility failed because no coverage was found or coverage starts after the requested date—states already represented by the coverage API/model. In those cases this rationale also propagates to the case and denial letter, producing a materially incorrect reason code and appeal record; restrict this source to terminated coverage or derive the reason from the actual coverage status and dates.

Useful? React with 👍 / 👎.

@haolin-chen-actava
haolin-chen-actava merged commit c05b225 into main Oct 7, 2026
9 checks passed
@haolin-chen-actava
haolin-chen-actava deleted the codex/bug-fixes branch October 7, 2026 06:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants