refactor(agents)!: consolidate the Data Science workstream into skill-primary coach jobs - #2612
refactor(agents)!: consolidate the Data Science workstream into skill-primary coach jobs#2612Bill Berry (WilliamBerryiii) wants to merge 39 commits into
Conversation
…ounded in the CSE playbook - author two MIT-grounded reference packs with provenance and derivation labels - migrate MVE methodology from the experiment-designer instruction into the skill - add permissive-license and mixed-content classes to the licensing posture - declare both packs MIT AND CC-BY-4.0 and correct reproduction claims - register in data-science and experimental collections with 9 eval stimuli 📚 - Generated by Copilot
- normalize CRLF to LF before drift, ms.date, and orphan-scaffold comparisons - stop false drift that rewrote every reference page on a CRLF checkout - restore orphaned-scaffold removal that a CRLF template tail suppressed - add Test-DocContentEqual and CRLF regression tests 🔧 - Generated by Copilot
- add EDA section sequence, plot selection, and scale thresholds - add dashboard components, caching rules, and validation budgets - add evaluation interview, difficulty balance, and dataset contract - record external evaluator vocabulary as cite-only provenance 📊 - Generated by Copilot
- add DS_CATALOG_V1 entity, relationship, and profile contracts - add evidence-led feasibility study and interchange traceability - split ML reproducibility and readiness from experiment framing - extend ds-dataops with derived-dataset persistence conventions 📚 - Generated by Copilot
…tion - relocate experiment-design out of the data-science collection - add an experiment-readiness reference for scoping and vetting - add a feasibility-to-PRD handoff to requirements-author 🧪 - Generated by Copilot
- remove gen-data-spec, gen-jupyter-notebook, and gen-streamlit-dashboard - remove test-streamlit-dashboard and eval-dataset-creator - route analysis and a new evaluation job to owning skills - expand the coach skill-boundary table from five to seven ♻️ - Generated by Copilot
…rtions - remove five per-agent stimulus, expectation, and signature files - assert seven-skill boundaries and skill-primary job routing - add behavior-conformance stimuli for both new skills - regenerate agent inventory and behavior spec from source ✅ - Generated by Copilot
- render a declared catalog model as an entity relationship diagram - validate catalog input and reject unsafe rendering payloads - commit a uv-managed dependency set with tests 📐 - Generated by Copilot
- update the canonical deck reference and coaching state handling - align the dt-coach agent and canonical-deck prompt 🎨 - Generated by Copilot
- harden the adr-author sensitive-content scanner and its tests - update licensing-posture and disclaimer-language guidance 🔒 - Generated by Copilot
- remove five retired specialist agents from the themed manifest - register ds-analysis-authoring and ds-evaluation-design - regenerate the canonical manifest and collection regions 📦 - Generated by Copilot
…ents - rewrite the data-scientist role and three lifecycle guides - drop retired agents from the agent catalog and installer sample - regenerate reference pages and author the new skill page tails 📝 - Generated by Copilot
- regenerate plugin trees for the new and relocated skills - drop per-agent outputs for the retired specialists - record third-party attribution for the new reference packs 🔧 - Generated by Copilot
…nsolidation # Conflicts: # .github/plugin/marketplace.json # .github/skills/installer/hve-core-installer/SKILL.md # collections/data-science.collection.md # collections/data-science.collection.yml # collections/experimental.collection.md # collections/experimental.collection.yml # collections/hve-core-all.collection.md # collections/hve-core-all.collection.yml # collections/project-planning.collection.md # docs/reference/README.md # docs/reference/instructions/README.md # docs/reference/skills/README.md # evals/behavior-conformance/skill-behavior.eval.yaml # plugins/data-science/.github/plugin/plugin.json # plugins/data-science/README.md # plugins/experimental/.github/plugin/plugin.json # plugins/experimental/README.md # plugins/hve-core-all/.github/plugin/plugin.json # plugins/hve-core-all/README.md # plugins/project-planning/README.md # scripts/tests/docs/Generate-AssetDocs.Tests.ps1
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.OpenSSF ScorecardScorecard details
Scanned Files
|
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #2612 +/- ##
==========================================
+ Coverage 82.85% 86.96% +4.11%
==========================================
Files 166 105 -61
Lines 22508 12077 -10431
Branches 29 29
==========================================
- Hits 18648 10503 -8145
+ Misses 3857 1571 -2286
Partials 3 3
Flags with carried forward coverage won't be shown. Click here to find out more.
🚀 New features to boost your workflow:
|
Eval Execution✅ Status: Passed — no merge-blocking failures (96 advisory assertion failure(s) present)
|
- reject YAML aliases, anchors, explicit tags, and merge keys - bound input size and validate RFC 3339 timestamps - detect lineage cycles with a three-colour depth-first search - pin jsonschema format-nongpl extras for format assertions * - Generated by Copilot
- seed ds-catalog, ds-feasibility, and architecture-diagrams harnesses - cover valid, malformed, duplicate-key, and missing-block parser paths - unblock the Fuzz Tests lane that requires tests/corpus to exist * - Generated by Copilot
- move budgets to budgets.json with target, ceiling, and rationale - warn above target and fail only above ceiling - stop shared instruction growth from failing unrelated changes - require a rationale whenever a ceiling exceeds its target * - Generated by Copilot
- resolve broken Docusaurus links to numbered design-session pages - replace two terms the inclusive-language gate rejects - add data-science vocabulary and stems to the cspell word list * - Generated by Copilot
…nsolidation # Conflicts: # docs/reference/README.md # docs/reference/skills/README.md # evals/behavior-conformance/skill-behavior.eval.yaml
Jamie Kim (jkim323)
left a comment
There was a problem hiding this comment.
Thanks for the exceptionally thorough work on this PR! I especially appreciate the effort to make the new skill-first model explicit rather than simply moving content around. Once the remaining contract and package-closure issues are addressed, this should provide a much stronger and more maintainable Data Science workstream foundation.
|
I noticed a licensing inconsistency that I hope to have a better understanding of... The new Is the intended policy that these skills are CC-BY-4.0 because they derive from Playbook documentation, with |
Jamie Kim (jkim323)
left a comment
There was a problem hiding this comment.
One more...! Thank you for working on it :)
…dback - ship privacy-standards, telemetry-foundations, architecture-diagrams, and adr-author in the data-science package - route privacy-standards from the catalog and pipeline job-registry rows - add experiment-design to the experimental, hve-core-all, and project-planning recipes - reduce reproduced upstream checklist text and refresh Data Science skill contracts - regenerate package documentation and reference pages 🔒 - Generated by Copilot
The Data Workstream Coach requires the shared disclaimer instructions during initialization, but the standalone data-science recipe omitted them, so the reference resolved in the monorepo and dangled after package installation. - add rules/shared/disclaimer-language.instructions.md to the data-science recipe - regenerate package documentation 🔒 - Generated by Copilot
…nderer Adjacent string literals inside list displays are deliberate line wraps, but CodeQL cannot distinguish them from a missing comma. Explicit parentheses express the grouping without changing behavior. - wrap eight multi-line string expressions in render_catalog_erd.py 🔒 - Generated by Copilot
…ections The handoff examples cited section names that do not exist in the shipped feasibility-study template, so a producer following them could not satisfy the validation rule requiring every entry to resolve to a named study section. - replace invented section names with template sections and FS-### anchors 🔒 - Generated by Copilot
…' into 2610-ds-workstream-consolidation
…rose - extract the feasibility profile block by deterministic scanning so many unterminated begin markers cannot force quadratic backtracking - convert recursive parser exhaustion into the documented validation error in both the catalog and feasibility validators - evaluate narrative anchors and requirement headings against prose only, blanking the authoritative block and fenced code - add regression tests for bounded extraction, nesting depth, and fenced content 🔒 - Generated by Copilot
…workstream-consolidation # Conflicts: # .github/agents/data-science/eval-dataset-creator.agent.md # evals/agent-behavior/expectations/eval-dataset-creator.expectations.yml
Port population coverage, validation provenance, detecting-metric guidance, and responsibility/safety metric vocabulary from the retired evaluation agent into ds-evaluation-design. Keep assessment ownership with rai-planner and extend the existing knowledge stimulus. 🔒 - Generated by Copilot
Replace three supporting Markdown outputs with one sectioned evaluation guide while preserving curation notes, metric selection, tooling recommendations, review checkboxes, and the risks-and-detecting-metrics boundary. BREAKING CHANGE: ds-evaluation-design now emits one evaluation guide instead of three separate supporting Markdown documents. 🔒 - Generated by Copilot
Replace the remaining plural supporting-document references after ds-evaluation-design consolidated its curation, metric, and tooling material into one evaluation guide. 🔒 - Generated by Copilot
- drop all ten design-session documents from the repository - resolve the Playbook licensing contradiction at its source rather than in place 🗑️ - Generated by Copilot
- replace the extra with base jsonschema in ds-catalog and ds-feasibility - re-lock both skills, removing 14 extra-only packages including MPL-2.0 and GPL-compound licenses - allowlist pkg:pypi/pyyaml, whose PyPI metadata declares no SPDX expression 🔒 - Generated by Copilot
- rename the top-level evidence_refs field to evidence_sections with a closed section set - scope candidate evidence_refs to FS-### display references and keep study_item_id as the durable UUID - replace the single global rule with field-specific validation and align both examples and the PRD consumer 🔗 - Generated by Copilot
- replace 33 fundamentals items with one applicability summary per section heading - regroup 11 production items into five repository-original readiness domains - correct the skill and provenance claims that asserted item labels were preserved 📄 - Generated by Copilot
|
Thanks for catching this, and sorry for the delayed reply. Responding to your licensing question above. Your reading of the policy is the correct one. The design-session documents contradicted that, and when I inventoried them the problem was wider than the two files first identified: the MIT claim appeared across seven of the ten documents, including two copyable Rather than correct them in place, we removed the design-session content from the repository entirely ( Separately, the Dependency Review license gate now passes. The six incompatible-license findings came from |
Close the stimulus-presence gap that failed all four Eval Execute lanes after this branch modified the licensing-posture instruction. 🧪 - Generated by Copilot
- PyYAML timestamp construction raises a bare ValueError for out-of-range dates, bypassing the yaml.YAMLError conversion and the CatalogRenderError contract - re-raise CatalogRenderError first, since it subclasses ValueError, then convert remaining scalar errors - add a regression test for out-of-range day, month, and day-of-month values 🐛 - Generated by Copilot
- PyYAML timestamp construction raises a bare ValueError for out-of-range dates, bypassing the documented validation-error contract - completes the A9 defect class across ds-catalog, ds-feasibility, and architecture-diagrams - add regression tests for out-of-range day, month, and day-of-month values 🐛 - Generated by Copilot
- stage the agent as workspace instructions with its two declared skills - replace two order-coupled graders with six semantic invariants - add generator environment and grader regression coverage - clarify smoke, functional, and A/B coverage ownership 🔬 - Generated by Copilot
Pull Request
Description
Consolidates the Data Science workstream from an agent-primary surface into a skill-primary one governed by
data-workstream-coach.The five Data Science specialist agents delegated work that either belongs to a skill (durable, reusable conventions) or is already served by built-in VS Code Copilot tooling (notebook and dashboard scaffolding). Research confirmed the durable capability in those agents was authoring convention, not orchestration, so it is absorbed into skills and the retired agent surfaces are dropped rather than rehomed.
Retired agents
Removes five specialist agents under
.github/agents/data-science/:eval-dataset-creatorgen-data-specgen-jupyter-notebookgen-streamlit-dashboardtest-streamlit-dashboardNew skills
ds-analysis-authoring— EDA notebook sequencing and analytical dashboard authoring conventions, including a plot-selection table, axis-scale thresholds,cache_dataversuscache_resourceguidance, and interaction latency budgets.ds-evaluation-design— evaluation dataset design for AI systems, including an interview and review protocol, metric selection, provenance requirements, and a dataset contract template. It sets a 30-pair floor and a category distribution with a per-category floor, and deliberately does not freeze an evaluator catalog.Extended skills
ds-cataloggains dataset profile contract coverage.ds-dataopsgains persistence and versioning coverage.Coach and registry
data-workstream-coachreduces itsagents:frontmatter from six entries to one (Experiment Designer), and now states that coaching governs decision ownership rather than abstention from producing work.analysisjob target withds-analysis-authoring, adds anevaluationjob routed tods-evaluation-design, and moves from five-skill to seven-skill boundaries.Catalog and evals
marketplace.jsonupdates thedata-scienceandhve-core-allpackages to the new membership, with matching component maturity across both packages.Related Issue(s)
Closes #2610
Type of Change
Select all that apply:
Code & Documentation:
Infrastructure & Configuration:
AI Artifacts:
hve-builderand addressed all actionable findings.github/instructions/*.instructions.md).github/prompts/*.prompt.md).github/agents/*.agent.md).github/skills/*/SKILL.md).github/hooks/*/*.json)evals/)Other:
.ps1,.sh,.py)Sample Prompts (for AI Artifact Contributions)
User Request:
"Help me build an evaluation dataset for our retrieval-augmented support assistant."
Execution Flow:
data-workstream-coachmatches the request to theevaluationjob in the job registry.ds-evaluation-design.Output Artifacts:
An evaluation dataset contract from
templates/evaluation-dataset-contract.md, covering system boundary, category distribution, metric selection with rationale, provenance, and review disposition.Success Indicators:
The contract meets the 30-pair floor, every category satisfies its floor, each selected metric ties to a declared failure mode, and every pair carries provenance.
Testing
Local validation on the merged branch:
npm run lint:marketplace— passnpm run lint:plugin-output— passnpm run validate:skills— passnpm run lint:frontmatter— passnpm run lint:md— passnpm run lint:yaml— passnpm run lint:json— passnpm run lint:asset-docs— passDiff hygiene:
main(35 modified Markdown files) each report zero date-only diffs.docs/reference/pages are regenerated throughnpm run docs:generate; generated regions are not hand-edited.CI-owned lanes are deferred to this PR's checks rather than run locally: the eval lint lanes (
vally, schema, text, safety), stimulus presence, changed-artifact execution, and content moderation.Checklist
Required Checks
AI Artifact Contributions
hve-builderreview mode to review contributionhve-builderreviewRequired Local Checks
The following local-safe validation commands must pass before merging:
npm run validate:local(Targeted checks were run individually; the aggregate was not accepted as final evidence)npm run validate:docs(Docusaurus installation not current in this working tree)npm run spell-check(Introduced findings were fixed; the repository-wide run has pre-existing failures from other in-flight work)npm run lint:md-links(Not accepted as final aggregate evidence)Security Considerations
Additional Notes
origin/main, including the immutable-snapshot packaging migration from feat(plugins): migrate packages to immutable snapshots #2577. Thecollections/andplugins/trees deleted by that migration were accepted as deleted; package membership for this change lives solely in.github/plugin/marketplace.json.eval-dataset-creator), refactor(collections)!: consolidate platform backlog collections into project-planning #2602 (collection topology), and refactor(skills): make cross-artifact references portable #2599 (artifact portability).