Skip to content

Latest commit

 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SpecKit

Autonomous spec-quality gate — refuses to let a weak specification get built.

Built with Google ADK 2.1 for the Google for Startups AI Agents Challenge (Refactor track).

Architecture pattern: The Proof-Carrying Gate — a model that can rewrite its own evidence cannot be its own gate. The retriever, grader, and gate have disjoint capabilities, and the ship/escalate decision runs in LLM-free Python.

Status (2026-06-11): version 1.14.0 · 480 tests passing with uv run pytest eval/ -q. Live demo: speckit-ui.run.app — the SpecKit UI on Cloud Run (public, scale-to-zero; replay + live, the agent runs in-process). Vertex AI Agent Engine deployment was verified live 2026-06-08 (reasoningEngines/2215848532236042240, smoke-tested) and then torn down to avoid a 24/7 standing cost — reproducible via deploy/deploy_agent_engine.py. The build carries the 1.14.0 changes (cross-run learning, hardened grounding floor, laundering benchmark) and the empty/ack-PRD guard (#1); see submission/DEPLOYMENT_PROOF.md.

What It Does

SpecKit sits between a product idea and the code an AI builds from it. It interviews you until it has a complete brief (with a measurable success target), drafts a production-grade PRD, then grades it against 12 decision-quality dimensions — citing real evidence from your project corpus. It returns Ship, Revise, or Kill.

  • Interview: asks one question at a time until it has a success target (metric + baseline + target + time window), weakest assumption, and target users. Will not finish the brief with the target left blank.
  • Draft: produces a PRD grounded in your codebase, past PRDs, postmortems, and analytics baselines.
  • Gate: grades against 12 dimensions (4 load-bearing, 8 writing-quality). Every verdict cites retrieved evidence. Never passes a claim the evidence contradicts.
  • Revise: auto-fixes writing-quality problems and re-grades, looping up to 4 times. Decision-quality problems escalate to the human.
  • Kill: ≤6 pass, unfalsifiable weakest assumption, or landfill strategic fit.

Architecture

root_agent (LlmAgent)
├── grill_interviewer (LlmAgent, tools=[save_brief])
└── refine_loop (LoopAgent, max_iterations=4)
    ├── prd_author (LlmAgent, output_key="prd")
    ├── evidence_retriever (LlmAgent, tools=[search], output_key="evidence")
    ├── prd_grader (LlmAgent, output_schema=PrdVerdict, output_key="verdict")
    └── gate (GateChecker, custom BaseAgent — escalate on Ship/Kill)
Node Role Model
root_agent Route: raw idea → interviewer, existing draft → refine_loop gemini-2.5-flash
grill_interviewer Multi-turn interview, collects Brief via save_brief gemini-2.5-flash
prd_author Draft/revise PRD with corpus grounding gemini-2.5-flash
evidence_retriever Search corpus for evidence across 12 dimensions gemini-2.5-flash
prd_grader 12-dimension quality gate — structured PrdVerdict gemini-2.5-pro
gate Deterministic: escalate on Ship/Kill, loop on Revise Python (no LLM)

Module Layout

prd_agent/              # New architecture (LlmAgent + LoopAgent)
├── __init__.py         # Exposes root_agent
├── agent.py            # Component tree
├── schemas.py          # PrdVerdict, DimensionVerdict, Brief
├── gate.py             # GateChecker(BaseAgent)
├── grounding.py        # VertexAiSearchTool + corpus ingestion docs
├── instructions.py     # Fat agent prompts
├── tools.py            # save_brief, save_sliced_issues
├── a2a.py              # A2A Agent Card + endpoint
├── memory.py           # Learning: persist outcomes, rework signals
├── handler.py          # Entry point for async issue processing
└── publisher.py        # VerdictPublisher (GitHub integration)

planner_agent/          # Peer agent that consults SpecKit via A2A
├── __init__.py         # Exposes product_planner
├── agent.py            # Single LlmAgent
├── a2a_client.py       # A2A discovery and invocation logic
└── tools.py            # consult_speckit tool

eval/                   # Tests for prd_agent and planner_agent
agent_engine.yaml       # Vertex AI Agent Engine deploy config
corpus/                 # Seed corpus

Quick Start

Prerequisites

  • Python 3.11+, uv, Google Cloud project with Vertex AI APIs
  • gcloud auth application-default login

Setup

git clone https://github.com/zaeem-rafiq/AgenticPRD.git && cd AgenticPRD
uv sync
cp .env.example .env   # Edit with your GCP project ID

# Provision corpus (one-time, requires billing)
uv run python scripts/provision_datastore.py

# Run locally
adk web prd_agent

Run Tests

uv run pytest eval/ -v    # 480 tests, all passing (verified 2026-06-11)
uv run ruff check .

Deploy

# Vertex AI Agent Engine
adk deploy agent_engine \
  --project=$GOOGLE_CLOUD_PROJECT \
  --region=$GOOGLE_CLOUD_LOCATION \
  --agent_engine_config_file=agent_engine.yaml \
  ./prd_agent

# A2A endpoint
uvicorn prd_agent.a2a:a2a_app --port 8080
curl http://localhost:8080/.well-known/agent.json

# Cloud Run (alternative)
docker build -t speckit -f deploy/Dockerfile .

Quality Gate — 12 Dimensions

Load-Bearing (decision-quality)

# Dimension Fail →
D1 Problem Existence Revise/Kill
D2 Success Target Revise/Kill
D3 Weakest Assumption Revise/Kill
D4 Strategic Coherence Revise/Kill

Writing-Quality (auto-revise)

# Dimension Fail →
W1 User Definition Revise
W2 Requirements Coherence Revise
W3 Technical Feasibility Revise
W4 Evidence Grounding Revise
W5 Scope Boundaries Revise
W6 Risk Awareness Revise
W7 Edge Cases Revise
W8 Execution Clarity Revise

Verdict Calibration

Verdict Condition
Ship 11-12 pass, no load-bearing fail
Revise 7-10 pass, or any load-bearing fail
Kill ≤6 pass, unfalsifiable assumption, or landfill fit

Spec-Laundering Benchmark

A reproducible, LLM-free benchmark measures the deterministic gate against a labeled taxonomy of spec-laundering attacks (baseline-swap, annotation-bypass, calibration-spoof, ungrounded-pass) plus honest controls. In each laundered fixture the grader is a clean Ship — so a catch is the structural gate alone.

uv run python eval/run_laundering_benchmark.py

Deterministic gate: 100% catch on structural vectors (12/12), 0 false positives on clean specs, 80% overall (12/15). The remaining 3 are semantic laundering vectors (a swapped metric name, an ungrounded claim outside grounded mode, an inverted meaning) — out of structural scope by design and deferred to the grounded grader. The honest misses are reported, not hidden; an undisclosed 100% would be the red flag. Fixtures: eval/benchmark/laundering/.

Constitution

8 non-negotiable principles — see GEMINI.md.

License

Apache 2.0

About

SpecKit — autonomous spec-quality gate (Google ADK 2.1 + Vertex AI Agent Engine)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages