Skip to content

About

GitHub-native ML System Cards: evidence-backed, PR-first documentation for ML systems

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

ML System Card (GitHub-native, PR-first)

Stakeholder-adaptable documentation for ML prototypes, stored in-repo as YAML and kept current through a PR-first workflow powered by GitHub. This README explains how the system works, how to integrate it into a project, and how contributors can use it day-to-day.

Live demo: ml-system-card.vercel.app

If this project helps your ML documentation, governance, or research workflow, a GitHub star helps others discover the idea.

Project docs

  • Contribution guide: CONTRIBUTING.md
  • Security policy: SECURITY.md
  • Code of conduct: CODE_OF_CONDUCT.md
  • Release checklist: docs/RELEASE_CHECKLIST.md
  • Known limitations: docs/KNOWN_LIMITATIONS.md
  • Environment template: .env.example
  • License: LICENSE with rationale in docs/LICENSE_RECOMMENDATION.md

Why this exists

ML prototypes evolve quickly and involve many audiences (data scientists, engineers, PMs, governance). Documentation often lags behind. The ML System Card solves this by:

  • living with the code in the repo (auditable, reviewable, versioned),
  • using LLMs to analyze and understand the entire codebase (plus configs, API specs, metrics files) to draft structured, cited updates,
  • guiding humans through a pre-merge review (accept/edit/reject) with clear signals for low-confidence content,
  • landing changes via pull requests that auto-merge once checks pass.

This produces a single source of truth that’s always close to the implementation.


Scope (and what’s intentionally out of scope for now)

In scope (phase 1):

  • Canonical ML System Card as YAML in docs/ml_system_card.yaml.
  • Individual stakeholder views and notes in docs/stakeholders.yaml (extendable).
  • PR-first workflow with a draft PR used as the proposal carrier (no time-based CI artifacts).
  • Auto-merge with safety guards (schema, path allow-list, diff caps, evidence coverage).
  • Micro-receipt (tiny JSON) committed after merge for audit.
  • Review Dashboard: A Next.js web application for reviewing, editing, and approving proposals.
  • PDF Export: Built-in print styles for exporting the card to PDF (supports Light/Dark mode).

Out of scope (phase 1 – future work):

  • Integrations with external experiment trackers (MLflow / Weights & Biases).
  • Organization-wide registry or portal.
  • On-prem/self-hosted LLM serving.
  • Advanced analytics and dashboards.

Key decisions at a glance

  • Canonical card file: docs/ml_system_card.yaml (YAML).
  • Validation schema: lib/ml_system_card.schema.json (JSON Schema, versioned via $id).
  • Stakeholders (individuals) + prompts: docs/stakeholders.yaml.
  • Language: default matches the repository language; per-stakeholder override supported.
  • Generation mode: PR-first — create a draft PR immediately and let it carry proposal JSON; promote to a final PR after review.
  • Analysis Trigger: Code changes trigger the CI Diff Inspector (checks for drift), but they do not automatically trigger the Generator (LLM). Analysis is currently demand-driven (User must click "Update").
  • Pipeline Architecture:
    1. Extractor: Analyzes code and extracts candidate facts.
    2. Reasoner: Harmonizes facts with the existing card and applies policy rules.
    3. Verifier: Independent QA pass that cross-references anchors against actual code to prevent hallucinations.
    4. Notes: Synthesizes stakeholder-specific summaries based on verified facts.
  • Confidence policy (inferred fields only):
    • OK: ≥ 0.80 (Verified)
    • Warn: 0.65–0.79 OR Invalid Anchors → must explicitly accept or edit
    • Require: < 0.65 → blocks apply until accepted or edited (reject allowed)
  • Auto-merge: gated by schema validation, writer determinism, evidence coverage, a strict path allow-list, and diff-size caps.
  • Noise Control: An Automated Diff Inspector runs in CI to reject PRs that contain only formatting changes (semantic noise), ensuring a clean git history.
  • Labels: ml-system-card, auto-merge-ok; optional confidence-low, stakeholder-<id>, superseded.

Configuration

The system is configured via GitHub Actions Secrets and Variables.

Required Secrets

  • LLM_API_KEY: Your OpenAI or Azure API key.

Optional Variables

  • LLM_MODEL: The model ID to use (default: gpt-5.1). Examples: gpt-4-turbo, gpt-3.5-turbo.
  • LLM_PROVIDER: The provider to use (default: openai). Options: openai, azure.

Repository layout

app/                        # Next.js Review Dashboard and UI
docs/
  ml_system_card.yaml         # Canonical ML System Card (YAML)
  ml_system_card.anchors.json  # Persistent anchor index (JSON; written on Apply)
  stakeholders.yaml           # Default + custom stakeholder definitions
  .proposals/                 # Temporary proposal JSON (only on PR branches)
  .card_runs/                 # Tiny permanent receipts written after merge
lib/
  ml_system_card.schema.json  # JSON Schema (versioned via $id)
scripts/                    # Node.js scripts for Analysis, Generation, and Apply actions

Local development quickstart

Prerequisites

  • Node.js 20.9.0 or newer (the Next.js app and workflow scripts require modern Node 20 APIs)
  • npm (ships with Node 20)
  • Git (for invoking the scripts’ incremental analysis helpers)

Install dependencies once per workspace:

# Scripts workspace (workflow utilities, deterministic writers, validators)
cd scripts
npm install

# Next.js app (control plane UI)
cd ../app
npm install

Everyday commands

Workspace Purpose Command
scripts/ TypeScript lint (ESLint with type-awareness) npm run lint
scripts/ Type-check (no emit) npm run typecheck
scripts/ Compile utilities for GitHub Actions npm run build
scripts/ Vitest suite (determinism, anchors, incremental analysis) npm test
app/ Lint Next.js app npm run lint
app/ Production build for UI npm run build
app/ Local dev server npm run dev

Validation sweep

To mirror CI locally, run the following (from the repo root):

cd scripts && npm run lint && npm run typecheck && npm run build && npm test
cd ../app && npm run lint && npm run build

GitHub integration prerequisites

  • A GitHub OAuth application for the control-plane UI. Grant it the scopes read:user, read:org, and repo so signed-in operators can browse and triage private repositories.
  • GitHub Actions secrets for LLM providers and any other downstream tooling referenced in the workflows.
  • (Optional) a fallback GITHUB_TOKEN/ML_SYSTEM_CARD_GITHUB_TOKEN for non-interactive scripts; the web UI now prefers per-user OAuth tokens.

Place a .env.local inside app/ if running the UI locally:

GITHUB_CLIENT_ID=replace-with-github-oauth-client-id
GITHUB_CLIENT_SECRET=replace-with-github-oauth-client-secret
NEXTAUTH_SECRET=generate-a-random-string
NEXTAUTH_URL=http://localhost:3000
NEXT_PUBLIC_DEFAULT_REPO_OWNER=
NEXT_PUBLIC_DEFAULT_REPO_NAME=

If you still rely on a service token (for example, when deploying in environments without GitHub login), set GITHUB_TOKEN=replace-with-github-token as well.

LLM-powered generation uses the following environment variables (see .env.example for defaults):

  • LLM_API_KEY, LLM_BASE_URL, LLM_PROVIDER, LLM_MODEL
  • feature toggles: LLM_ENABLED (default 0) and LLM_DRY_RUN (default 1)

Sampler defaults (temperature, top-p, max tokens) and retry limits live in scripts/src/constants.ts. The runtime loader in scripts/src/config/env.ts centralizes environment parsing so GitHub Actions can override values per workflow.

Prompt templates for extractor, reasoner, and stakeholder notes are versioned in scripts/prompts/*.md. Update the version front-matter field and bump the semantic ID (e.g., extractor.v2) when editing prompts; tests in scripts/src/__tests__/prompts.test.ts ensure required placeholders remain intact. Each generator run snapshots repository context to docs/.analysis/<runId>.json (with a cache at docs/.analysis/cache.json) and emits intermediate pass artifacts beside the proposal (docs/.proposals/<runId>{.extractor,.reasoner,.notes}.json).

Before persisting a proposal, the generator enforces redaction and safety guardrails (scripts/src/safety/). Secrets matching common token patterns are replaced with <redacted>, deny-listed files (e.g., .env, config.*) are excluded from analysis, and evidence coverage plus diff-size thresholds are enforced (COVERAGE_THRESHOLD, MAX_DIFF_LINES, MAX_CARD_GROWTH_RATIO).

Dispatching workflows manually

  • Generator: POST /repos/:owner/:repo/actions/workflows/generator.yml/dispatches with { "ref": "card/proposal/<runId>", "inputs": { "runId": "...", "baseSha": "..." } }
  • Apply: POST /repos/:owner/:repo/actions/workflows/apply.yml/dispatches with { "ref": "card/proposal/<runId>", "inputs": { "runId": "..." } }

The Next.js /runs/new and /review/[pr] pages wrap these APIs, so the manual calls above are mainly for automation or debugging.

Persistent Anchor Index

A permanent, deterministic evidence map that keeps the card YAML clean while making every fact traceable.

  • File: docs/ml_system_card.anchors.json (committed on the same Apply that updates docs/ml_system_card.yaml).
  • Purpose: map card JSONPaths → repo anchors (e.g., src/api.py#L42-L73@<commit>), so viewers can click any field and see the exact evidence lines.
  • Schema (concise):
    • cardSha, runId, generatedAt
    • anchorsByPath { <jsonPath>: [ { path, startLine, endLine, commit, fingerprint, kind } ] }
  • Determinism: sorted keys; arrays sorted by path → startLine → endLine → commit; LF newline.
  • CI: evidence-coverage gate reads this file (e.g., ≥95% of non-null facts must have ≥1 anchor). Link validation ensures anchors point to real files at the referenced commit.
  • Privacy: never anchor to .env / secrets.*. If sensitive, anchor to ADR/README instead.

You do not need to create these files manually. The web app’s “Initialize ML System Card” button opens a draft PR that seeds them for you.


The ML System Card (what’s inside)

Top-level keys (alphabetical for clean diffs and deterministic writing):

  • ai — provenance for AI-generated content (engine, runId, fieldMeta[] for inferred items).

  • business — use case, non-goals, KPIs/targets, pilot scope.

  • devInsight — concise code intelligence for developers (see example below).

  • governance — policies/regs, risk register (likelihood/impact/mitigation), assessments (fairness/privacy/safety), sign-offs.

  • integration — API I/O schema, latency/throughput, cost envelope, fallbacks, observability, UX notes, plus ART extensions:

    • security {auth ∈ [apiKey, oauth2, mtls], scopes?[]}
    • errorModel[] {httpStatus, code?, message?}
    • idempotency {keys?[], policy?}
    • versioningPolicy {scheme ∈ [semver], notes?}
    • operationalQualities[] {name, desiredDirection ∈ [maximize, minimize, target_range], targets[], measures[]}
    • driftSignals[], feedbackChannels[], incidentReporting?
  • meta — title, owners, maturity (Ideation/PoC/Pilot/Production-Shadow/Production), links, tags, timestamps.

  • mlCore — problem, datasets (lineage/license), features, baselines, qualities, failure modes, plus ART extensions:

    • qualities[] {category ∈ [predictive, calibration, fairness, robustness, reliability, safety, explainability, data_quality], evaluationProtocol, targets[], measures[]}
    • training {data, compute, schedule?}, artifactURIs[]
  • provenance — branch, commit, lastGeneratedAt, changelog.

  • stakeholderNotes — short Markdown notes per stakeholder ID.

devInsight (developer-focused slice)

devInsight:
  codeOverview:
    languages: ["python","typescript"]
    entrypoints: ["src/service.py","api/server.ts"]
    components: [{name, summary, keyFiles:[...]}]
  architecture:
    publicApis: [{name, path, method?, signature?, file}]
    dataFlow: ["request -> api/server.ts -> model/infer.py -> response", "..."]
    depsSummary: ["fastapi","pydantic","numpy","react"]
  qualitySignals:
    testsPresent: true
    coverageHint: "approx 62% (from coverage.xml)"
    complexityHints: ["src/feature.py: high branching"]
    todoHotspots: ["trainer.py (5 TODOs)"]
  runtimePerf:
    latencyMsP50: 120
    latencyMsP95: 280
    notes: "GPU path faster by ~30%"

Deterministic ordering for tiny diffs:

  • Arrays:
    • mlCore.metrics → sort by (slice ASC, metric ASC)
    • mlCore.failureModes → A–Z (case-insensitive)
    • business.kpis → by name ASC
    • governance.riskRegister → (likelihood DESC, impact DESC, item ASC)
    • governance.signOffs → by timestamp ASC
    • provenance.changelog → by date ASC
  • YAML writer: 2-space indent, LF endings, no trailing spaces, stable key order.

AI provenance & guards (only where useful, typically for inferred items and notes):

Each inferred field or note can carry a small metadata entry (stored under ai.fieldMeta[] or inside the stakeholder note) with:

  • path (JSONPath of the field),
  • source.kind (inferred|extracted|manual) and confidence,
  • repoSources[] (file:lines that informed the value),
  • needs_review (true until accepted or edited),
  • guard.locked (true if manually edited; AI won’t overwrite),
  • guard.skip_generation (true if a suggestion was rejected; won’t reappear unless watched files change),
  • guard.watch_paths[] and guard.last_accept_commit (re-enable suggestions only when relevant files change).

Stakeholders (individuals, not groups)

Default entries live in docs/stakeholders.yaml and can be extended:

  • Data Scientist
  • ML Engineer
  • Data Engineer
  • Software Developer
  • UX Researcher
  • Product Manager
  • Project Manager
  • Domain Expert
  • Governance, Compliance & Ethics Officer

Each stakeholder entry has:

- id: product-manager
  title: Product Manager
  description: Align ML work with business goals and delivery
  roles: ["Product Manager"]
  language: auto           # or "de", "en", ...
  field_paths:
    - $.business.*
    - $.provenance.*
  prompt_template: |
    Draft a brief, non-technical summary covering use case, KPIs, pilot scope,
    costs, and key risks. Keep it under 180 words.
  watch_paths:
    - docs/roadmap.md
    - docs/**

The web app renders stakeholder-specific tabs using field_paths and shows the stakeholder’s short textMd note.


End-to-end workflow (PR-first; the PR is the proposal carrier)

A) Initialize (zero seed work)

  1. In the web app, click Initialize ML System Card for your repository/branch. This opens a draft PR that seeds:
    • docs/ml_system_card.yaml (minimal card),
    • docs/stakeholders.yaml (defaults, language auto-detected),
    • lib/ml_system_card.schema.json (validation schema).

B) Select & Generate

  1. Still in the web app, select branch, optional commit, and a specific stakeholder.
  2. Click Generate (AI).

C) Draft PR (proposal carrier)

  1. The app writes docs/.proposals/<runId>.json into the existing draft PR branch. The JSON bundles:
    • card_patch (RFC-6902 changes the AI proposes),
    • confidence_report (per-path: kind, confidence, sources),
    • sources (deep links like file#L10-L40),
      • notes (per stakeholder: textMd, language, citations, confidence),
  2. Draft PR labels: ml-system-card, proposal. The PR body lists base SHA, engine, runId, and a link back to the review UI.

D) Review Workspace (pre-merge, in the app)

Preview mode addition: The UI uses field-level diffs only by default.

  1. The app loads the proposal JSON and shows field-level diffs with badges:
    • extracted (from parsing/code intelligence),
    • inferred (LLM understanding),
    • manual (human-authored from earlier versions).
  2. Confidence gates (for inferred items only):
    • OK (≥ 0.80): normal.
    • Warn (0.65–0.79): must explicitly Accept or Edit (no bulk-accept).
    • Require (< 0.65): blocks Apply until Accepted or Edited (Reject allowed).
  3. Manual edits set guard.locked: true; rejects set guard.skip_generation: true. Shared fields updates reflect in every stakeholder view.

E) Promote draft PR → final PR

  1. Click Apply to promote:
    • Replace the proposal JSON with a deterministically written docs/ml_system_card.yaml (and docs/stakeholders.yaml if changed).
    • Remove docs/.proposals/<runId>.json.
    • Update PR:
      • remove proposal, add auto-merge-ok,
      • expand body with a summary, Low-confidence table (Warn + Require accepted/edited), “Stakeholder notes updated”, <details> for All inferred and Extracted, base SHA, engine, runId.
    • Auto-close older open card PRs to the same base branch with label superseded.

F) Auto-merge (squash)

  1. The PR auto-merges once all gates pass:
    • AJV schema validation passes against lib/ml_system_card.schema.json,
    • Path whitelist is satisfied: only docs/ml_system_card.yaml, docs/stakeholders.yaml (and optionally lib/ml_system_card.schema.json) changed,
    • Diff safety (e.g., ≤ 500 changed lines and ≤ 25% size growth),
    • Labels include ml-system-card and auto-merge-ok.
  2. The branch is deleted after merge.

G) Micro-receipt (tiny, permanent)

  1. A small JSON receipt is committed to the base branch:
docs/.card_runs/<yyyy-mm-dd>-<shortSha>.json

Contents (example):

{
  "base_sha": "<sha analyzed>",
  "engine": "system",
  "run_id": "<runId>",
  "changed_paths": ["$.mlCore.problem", "$.integration.latencyMs"],
  "low_confidence_rows_only": [
    {"path":"$.mlCore.problem","confidence":0.62,"sources":["README.md#L1-L40"],"disposition":"accepted"}
  ],
  "yaml_hash": "<sha256 of final YAML>",
  "schema_id": "lib/ml_system_card.schema.json#1.0.0"
}

LLM strategy (dual-model, tool-augmented, two-pass)

This section defines how the system analyzes a repository, produces evidence-backed facts, and drafts stakeholder notes while preserving determinism, safety, and research-grade reproducibility.

Goals

  • High precision facts with file-line citations and per-field confidence.
  • Deterministic outputs (byte-stable YAML and proposal JSON).
  • Human-in-the-loop via confidence gates (OK/Warn/Require).
  • Reproducible evaluations for generator behavior (micro-receipts, seeds, metrics).

Roles

  • Extractor Agent — extractor & evidence gatherer. Builds static signals (imports/routes/OpenAPI/config/metrics), performs AST/CFG parsing, and emits typed evidence tables with repo anchors and confidences.

  • GPT-5.1 Reasoner — synthesizer & verifier. Fuses evidence across files, resolves contradictions, writes Pass-1 facts (strict JSON) and Pass-2 stakeholder notes (short Markdown) deterministically.

Sampling & defaults

  • Pass-1 (facts): temperature=0.1, top_p=0.9, max_tokens≈2k.
  • Verifier: temperature=0.0, max_tokens≈512.
  • Pass-2 (notes): temperature=0.2–0.3, max_tokens≈512.
  • One retry only when confidence < 0.65 or citations are incomplete.

Tools (function-calling, minimal interface)

  • repo.search(globs|regex, max_hits, max_lines)
  • code.ast_summary(file) · code.symbols(file) · code.cfg(file)
  • openapi.parse(file) · openapi.contracts() (paths, methods, schemas)
  • metrics.load(paths) (e.g., coverage.xml, latency.json)
  • docs.extract(headings?) (README/ADR snippets)
  • schema.validate(json, lib/ml_system_card.schema.json)
  • yaml.write_deterministic(card) · yaml.diff(old,new)->RFC6902
  • citations.anchor(file, line_start, line_end) (stable anchors; re-anchors on drift)

Secret/PII scrubbing (pre-prompt): ENV patterns, API keys, tokens, emails, secrets in .env, secrets.*, config.*.

File ranking, chunking & caching

Incremental analysis (clarified):

  • Base commit = last merge that updated docs/ml_system_card.yaml.

  • Analyze changed files since base + 1–2 hop neighbors in the import graph.

  • Always include critical paths: OpenAPI, configs, CI definitions, metrics artifacts (bench/*.json, latency.json, coverage.xml), README/ADRs.

  • Re-anchor using content fingerprints when lines move; mark deleted anchors stale.

  • Fall back to a full scan on schema bump, major lockfile/dependency jumps, or mass renames.

  • Ranker prioritizes: entrypoints, API/router files, train/infer modules, data loaders, configuration, metrics, and docs.

  • Chunking at function/class granularity with local imports/context; hard caps on tokens per chunk.

  • Caching keyed by (path, content_hash); invalidates on hash change.

  • Incremental mode: analyze changed files + 1–2 hop neighbors in the import graph.


Pipeline (expanded)

  1. Static ingest (Extractor) Outputs:

    • static_graph.json (imports, entrypoints, route map)
    • openapi_summary.json (if any OpenAPI present; else empty with low_confidence=true)
    • metrics_table.json (coverage, latency histograms, throughput, costs)
    • docs_extract.json (README/ADR sections)
  2. File-level extraction (Extractor) For each ranked file, produce:

    {
      "path": "src/service.py",
      "purpose": "REST entrypoint and orchestration",
      "keyFunctions": [{"name":"handle","signature":"(req)","summary":"..."}],
      "inputs": ["HTTP JSON"],
      "outputs": ["JSON 200/400"],
      "risks": ["unbounded payload size"],
      "citations": ["src/service.py#L12-L58"],
      "confidence": 0.82
    }

    If confidence < 0.65 or citations empty → retry once; otherwise mark Require. Cache results.

  3. System synthesis (GPT-5.1) Combine file summaries + static signals →

    • components.json {name, responsibilities[], contracts[], keyFiles[], risks[], citations[], confidence}
    • system_narrative.json {dataFlowBullets[], keyEntryPoints[], failurePaths[], citations[], confidence} Sanity: each referenced file must exist; unresolved refs → Require.
  4. Pass-1: Facts (GPT-5.1) Emit card_facts.json (strict JSON) that fills structured fields only. Every non-null field includes:

    • source.kind ∈ {extracted|inferred|manual}
    • repoSources[] (at least one file#Lstart-Lend)
    • confidence ∈ [0,1]

    Normalization before write:

    • Units: latency in ms, throughput in req/s, costs in per-1k.
    • Numbers: ≤3 significant digits; timestamps in ISO-8601.
    • Enums validated (e.g., maturity, security.auth, versioningPolicy.scheme).
    • Sorting rules applied (metrics, risks, sign-offs, etc.). Failing range/enum/shape checks → Require.
  5. Verifier mini-pass (GPT-5.1, deterministic) Cross-checks for:

    • Contradictions (e.g., latencyP50 > latencyP95)
    • Unit mismatches (s/ms, ms/s) and impossible values (<0, NaN)
    • KPI target/current consistency
    • Duplicate or overlapping JSONPaths Violations → Require with a terse reason.
  6. Pass-2: Stakeholder notes (GPT-5.1) For each stakeholder (docs/stakeholders.yaml), project accepted facts via field_paths and draft ~160-word textMd in repo language (or override).

    • No new facts; references by JSONPath only (optional inline [@$.path]).
    • Style guardrails: concise, actionable, audience-appropriate.
    • No tool calls in Pass-2 to avoid drift; only the accepted Pass-1 payload is visible.

Confidence & retry policy

Band Threshold Handling in Review Workspace
OK confidence ≥ .80 Eligible for bulk-accept
Warn .65 ≤ c < .80 Must Accept/Edit explicitly (no bulk)
Require c < .65 or any validation/verifier failure Blocks Apply until Accepted/Edited (Reject allowed)

Retry rule: Only the Extractor extraction step and Pass-1 fact emission may retry once for better evidence; otherwise mark Require.

Evidence coverage & citation rules

  • Coverage gate: ≥ 95% of non-null Pass-1 fields must have ≥1 repoSources anchor.
  • Anchors must be stable (file#Lstart-Lend); the system re-anchors after rebases using a fuzzy matcher.
  • If evidence is missing → leave the field null, cite nothing, set low confidence.

Conflict resolution & guards

  • If a suggested value conflicts with a locked field (guard.locked:true) → do not overwrite; surface the conflict as Require with rationale.
  • If a field was rejected (guard.skip_generation:true) → do not re-suggest unless a file in guard.watch_paths[] changed since guard.last_accept_commit.

Multilingual & localization

  • Detect repo language from docs; per-stakeholder language overrides allowed.
  • Notes inherit repo language unless overridden; numbers and dates follow ISO conventions regardless of language.

Performance & cost controls

  • Rank top ≤200 important files per run; rotate cohorts across runs.
  • Cap tokens per step; truncate long comments while preserving anchors.
  • Incremental analysis on diffs + dependency neighbors for repeat runs.

Reproducibility & telemetry

  • Record in micro-receipt: engine (both models), run_id, base_sha, yaml_hash, schema_id, changed_paths[], and all accepted Warn/Require rows.
  • Store per-step timings and counts in the proposal’s dev block (e.g., file_summaries_count, components_count).

Validation & CI

  • Required check: schema validation with AJV against lib/ml_system_card.schema.json.
  • Optional check: writer determinism (re-serialize and compare).
  • Path whitelist: block auto-merge if files outside the allowed set changed.
  • Evidence coverage: computed from docs/ml_system_card.anchors.json (e.g., ≥95% of non-null facts must have ≥1 anchor; anchors must resolve at the referenced commit).

PR labels & automation

  • Always add: ml-system-card
  • Ready to merge: auto-merge-ok
  • If any Require (< 0.65) items were accepted: confidence-low
  • Stakeholder note updated: stakeholder-<id> (e.g., stakeholder-product-manager)
  • Superseded PRs: superseded

Before promoting a draft PR, auto-close older open card PRs to the same base branch with superseded.


Quickstart

  1. Open the web app and click Initialize ML System Card for your repo/branch. This creates a draft PR with the seed files.
  2. Click Generate (AI). The app writes a proposal JSON to the draft PR.
  3. Review & apply in the app. Accept/edit/reject suggestions; low-confidence gates will guide you.
  4. Promote the PR. The proposal JSON is replaced by deterministic YAML. Auto-merge follows once checks pass.
  5. A micro-receipt is written to docs/.card_runs/.

Contributing

  • Edit shared fields in docs/ml_system_card.yaml when you know the source of truth—updates benefit every stakeholder view.
  • Add or tweak stakeholders in docs/stakeholders.yaml (keep definitions focused; add field_paths and a clear prompt_template).
  • If an AI suggestion is unhelpful:
    • Edit it (locks the field against future overwrites), or
    • Reject it (suppresses repeats until watched files change).

FAQ

Why a PR for proposals instead of CI artifacts? The draft PR branch carries the proposal JSON. It’s discoverable, auditable, and short-lived—then replaced by the final YAML before merge. No expiring artifacts to chase.

Can anyone trigger generation? Anyone with repo write access can generate and create PRs. Auto-merge is gated by checks and strict file whitelisting.

How does language work? The system detects the repository language and writes in that language by default. A stakeholder entry can override with language: "de" (or similar) for their note.

What prevents noisy re-suggestions? Manual edits lock fields; rejections suppress suggestions until watched files actually change.



Editing policy (evidence rules) — addition

Category Evidence (anchors) required for non-null? Notes
Structured facts (APIs, SLOs, metrics, protocols, risks) Yes Needed to satisfy the CI anchor coverage threshold.
Narrative/meta (stakeholderNotes.*, descriptions, owners) No (encouraged if linkable) Keep concise; date notes if useful.
Business KPI targets Yes Treat numeric targets as structured facts.

Manual edits happen via schema-driven forms (no raw file edits). Edited fields become source.kind: "manual" and can be locked to prevent future overwrite by generation.

Roadmap (post-phase-1 ideas)

  • Export to PDF/HTML
  • Org-wide registry/search portal
  • Tracker integrations (MLflow / W&B)
  • Schema migrations and richer analytics
  • Optional on-prem LLM serving

License

Add a license appropriate for your organization (e.g., MIT/Apache-2.0) at the repository root as LICENSE.


Appendix: Example docs/stakeholders.yaml (snippet)

- id: data-scientist
  title: Data Scientist
  description: Builds models & experiments; cares about datasets, metrics, failure modes
  roles: ["Data Scientist"]
  language: auto
  field_paths:
    - $.mlCore.*
    - $.provenance.*
  prompt_template: |
    Summarize datasets, splits, baselines, current metrics (with important slices),
    known failure modes, and open technical risks. Keep it concise.
  watch_paths:
    - README.md
    - reports/**
    - models/**

Appendix: Example micro-receipt

{
  "base_sha": "a1b2c3d",
  "engine": "system",
  "run_id": "gh-run-12345",
  "changed_paths": [
    "$.mlCore.problem",
    "$.integration.latencyMs",
    "$.stakeholderNotes.product-manager"
  ],
  "low_confidence_rows_only": [
    {
      "path": "$.mlCore.problem",
      "confidence": 0.62,
      "sources": ["README.md#L1-L40"],
      "disposition": "accepted"
    }
  ],
  "yaml_hash": "sha256:…",
  "schema_id": "lib/ml_system_card.schema.json#1.0.0"
}

About

GitHub-native ML System Cards: evidence-backed, PR-first documentation for ML systems

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages