Autonomous spec-quality gate — refuses to let a weak specification get built.
Built with Google ADK 2.1 for the Google for Startups AI Agents Challenge (Refactor track).
Architecture pattern: The Proof-Carrying Gate — a model that can rewrite its own evidence cannot be its own gate. The retriever, grader, and gate have disjoint capabilities, and the ship/escalate decision runs in LLM-free Python.
Status (2026-06-11): version 1.14.0 · 480 tests passing with
uv run pytest eval/ -q. Live demo: speckit-ui.run.app — the SpecKit UI on Cloud Run (public, scale-to-zero; replay + live, the agent runs in-process). Vertex AI Agent Engine deployment was verified live 2026-06-08 (reasoningEngines/2215848532236042240, smoke-tested) and then torn down to avoid a 24/7 standing cost — reproducible viadeploy/deploy_agent_engine.py. The build carries the 1.14.0 changes (cross-run learning, hardened grounding floor, laundering benchmark) and the empty/ack-PRD guard (#1); see submission/DEPLOYMENT_PROOF.md.
SpecKit sits between a product idea and the code an AI builds from it. It interviews you until it has a complete brief (with a measurable success target), drafts a production-grade PRD, then grades it against 12 decision-quality dimensions — citing real evidence from your project corpus. It returns Ship, Revise, or Kill.
- Interview: asks one question at a time until it has a success target (metric + baseline + target + time window), weakest assumption, and target users. Will not finish the brief with the target left blank.
- Draft: produces a PRD grounded in your codebase, past PRDs, postmortems, and analytics baselines.
- Gate: grades against 12 dimensions (4 load-bearing, 8 writing-quality). Every verdict cites retrieved evidence. Never passes a claim the evidence contradicts.
- Revise: auto-fixes writing-quality problems and re-grades, looping up to 4 times. Decision-quality problems escalate to the human.
- Kill: ≤6 pass, unfalsifiable weakest assumption, or landfill strategic fit.
root_agent (LlmAgent)
├── grill_interviewer (LlmAgent, tools=[save_brief])
└── refine_loop (LoopAgent, max_iterations=4)
├── prd_author (LlmAgent, output_key="prd")
├── evidence_retriever (LlmAgent, tools=[search], output_key="evidence")
├── prd_grader (LlmAgent, output_schema=PrdVerdict, output_key="verdict")
└── gate (GateChecker, custom BaseAgent — escalate on Ship/Kill)
| Node | Role | Model |
|---|---|---|
| root_agent | Route: raw idea → interviewer, existing draft → refine_loop | gemini-2.5-flash |
| grill_interviewer | Multi-turn interview, collects Brief via save_brief | gemini-2.5-flash |
| prd_author | Draft/revise PRD with corpus grounding | gemini-2.5-flash |
| evidence_retriever | Search corpus for evidence across 12 dimensions | gemini-2.5-flash |
| prd_grader | 12-dimension quality gate — structured PrdVerdict | gemini-2.5-pro |
| gate | Deterministic: escalate on Ship/Kill, loop on Revise | Python (no LLM) |
prd_agent/ # New architecture (LlmAgent + LoopAgent)
├── __init__.py # Exposes root_agent
├── agent.py # Component tree
├── schemas.py # PrdVerdict, DimensionVerdict, Brief
├── gate.py # GateChecker(BaseAgent)
├── grounding.py # VertexAiSearchTool + corpus ingestion docs
├── instructions.py # Fat agent prompts
├── tools.py # save_brief, save_sliced_issues
├── a2a.py # A2A Agent Card + endpoint
├── memory.py # Learning: persist outcomes, rework signals
├── handler.py # Entry point for async issue processing
└── publisher.py # VerdictPublisher (GitHub integration)
planner_agent/ # Peer agent that consults SpecKit via A2A
├── __init__.py # Exposes product_planner
├── agent.py # Single LlmAgent
├── a2a_client.py # A2A discovery and invocation logic
└── tools.py # consult_speckit tool
eval/ # Tests for prd_agent and planner_agent
agent_engine.yaml # Vertex AI Agent Engine deploy config
corpus/ # Seed corpus
- Python 3.11+, uv, Google Cloud project with Vertex AI APIs
gcloud auth application-default login
git clone https://github.com/zaeem-rafiq/AgenticPRD.git && cd AgenticPRD
uv sync
cp .env.example .env # Edit with your GCP project ID
# Provision corpus (one-time, requires billing)
uv run python scripts/provision_datastore.py
# Run locally
adk web prd_agentuv run pytest eval/ -v # 480 tests, all passing (verified 2026-06-11)
uv run ruff check .# Vertex AI Agent Engine
adk deploy agent_engine \
--project=$GOOGLE_CLOUD_PROJECT \
--region=$GOOGLE_CLOUD_LOCATION \
--agent_engine_config_file=agent_engine.yaml \
./prd_agent
# A2A endpoint
uvicorn prd_agent.a2a:a2a_app --port 8080
curl http://localhost:8080/.well-known/agent.json
# Cloud Run (alternative)
docker build -t speckit -f deploy/Dockerfile .| # | Dimension | Fail → |
|---|---|---|
| D1 | Problem Existence | Revise/Kill |
| D2 | Success Target | Revise/Kill |
| D3 | Weakest Assumption | Revise/Kill |
| D4 | Strategic Coherence | Revise/Kill |
| # | Dimension | Fail → |
|---|---|---|
| W1 | User Definition | Revise |
| W2 | Requirements Coherence | Revise |
| W3 | Technical Feasibility | Revise |
| W4 | Evidence Grounding | Revise |
| W5 | Scope Boundaries | Revise |
| W6 | Risk Awareness | Revise |
| W7 | Edge Cases | Revise |
| W8 | Execution Clarity | Revise |
| Verdict | Condition |
|---|---|
| Ship | 11-12 pass, no load-bearing fail |
| Revise | 7-10 pass, or any load-bearing fail |
| Kill | ≤6 pass, unfalsifiable assumption, or landfill fit |
A reproducible, LLM-free benchmark measures the deterministic gate against a labeled taxonomy of spec-laundering attacks (baseline-swap, annotation-bypass, calibration-spoof, ungrounded-pass) plus honest controls. In each laundered fixture the grader is a clean Ship — so a catch is the structural gate alone.
uv run python eval/run_laundering_benchmark.pyDeterministic gate: 100% catch on structural vectors (12/12), 0 false positives on clean specs, 80% overall (12/15). The remaining 3 are semantic laundering vectors (a swapped metric name, an ungrounded claim outside grounded mode, an inverted meaning) — out of structural scope by design and deferred to the grounded grader. The honest misses are reported, not hidden; an undisclosed 100% would be the red flag. Fixtures:
eval/benchmark/laundering/.
8 non-negotiable principles — see GEMINI.md.
Apache 2.0