World model first. Agent second.
Open research on structured latent world models, planning, and self-improving agents for StarCraft: Brood War.
Important
Current status: M0 complete, M1 in progress. This repository contains a reproducible research scaffold, real replay acquisition, model smoke tests, and CI research infrastructure. It does not yet contain a trained competitive StarCraft agent.
Delivery target: an agent that plays humans on the current StarCraft: Remastered Battle.net public matchmaking and ladder. Background window capture and a targeted keyboard input have been verified; mouse gameplay control and the full live loop have not. See the runtime probe and runtime gate. The active implementation checklist is TODO.md; the public Remastered adapter starts with a version-locked, read-only process probe and does not yet control a match.
Most game agents are trained primarily to answer:
What action should I take in this state?
This project asks a complementary question:
If I take this action, what happens next?
The long-term agent should learn a compact internal model of Brood War, imagine candidate futures, and use those imagined futures for planning.
structured observation
↓
entity encoder + memory
↓
latent world model
↓
imagined futures
↓
planner / policy
↓
structured game action
- Structured research data. BWAPI/replay tooling exposes units, resources, orders, map data, and actions for offline research. The current Battle.net client still needs a verified player-observable live interface.
- World-model-first research. We compare mature model-based RL baselines with JEPA-style predictive representation learning.
- Current data pipeline. A weekly GitHub Actions job fetches bounded high-MMR replay snapshots on current competitive maps.
- Falsifiable milestones. Complex methods do not advance until simpler baselines pass measurable gates.
- CI-first research. CPU baselines, multi-seed checks, data validation, package builds, security scanning, and dependency maintenance run on GitHub Actions.
The current-client runtime gate must be passed before this project can claim human public-match or ladder play.
| Milestone | Goal | Status |
|---|---|---|
| M0 | PyTorch project, reproducibility, synthetic latent-dynamics tests | ✅ complete |
| M1 | Versioned replay/state/action dataset and headless extraction | 🚧 current |
| M2 | Real-data world-model baselines and predictive representation | planned |
| M3 | Latent planning and negative controls | planned |
| M4 | Latent skills and discrete diffusion action proposals | research |
| M5 | Self-play population and fixed evaluation league | research |
See ROADMAP.md and the open GitHub Issues.
StarData / public replays / current high-MMR snapshots
↓
replay parser
↓
headless game-state reconstruction
↓
observation_t + action_t + observation_t+1
↓
versioned TransitionV1
↓
Parquet / dataset loader
↓
PyTorch tensors
↓
world-model training
The repository already contains:
- a provenance-traceable real replay fixture;
- a weekly current high-MMR CWAL replay snapshot workflow;
- a pinned first 2026 snapshot manifest with match IDs and SHA-256 hashes;
- current competitive map metadata;
- historical StarData metadata.
Read docs/CURRENT_REPLAYS.md and DATA_CARD.md.
The first real comparison will not assume that the newest method wins.
simple supervised dynamics
↓
Dreamer/RSSM-style baseline
vs
structured JEPA-style latent dynamics
↓
short-horizon latent planning
↓
only then: latent actions / diffusion proposals / league self-play
See MODEL_CARD.md, docs/RESEARCH_PLAN.md, and docs/REFERENCES.md.
Target development environment:
Windows 11 host WSL2 Ubuntu
------------------------------- ------------------------------
StarCraft / game data Python 3.12
optional BWAPI runtime ---> PyTorch + CUDA
local replay extraction datasets / training / eval
Clone and install:
git clone https://github.com/Mrbaeksang/starcraft-ai.git
cd starcraft-ai
curl -LsSf https://astral.sh/uv/install.sh | sh
uv python install 3.12
uv sync --locked --extra cu130 --group devVerify an RTX machine:
uv run --extra cu130 scai doctor
uv run --extra cu130 scai smoke-train --steps 100 --device cuda
make EXTRA=cu130 checkCPU-only:
uv sync --locked --extra cpu --group dev
make EXTRA=cpu checkEvery push / pull request runs:
- Ruff lint + formatting;
- Python compilation;
- pytest;
- package build;
- model smoke training with multiple seeds.
Scheduled automation also runs:
- nightly model-size × seed CPU research matrix;
- weekly current high-MMR replay acquisition;
- CodeQL security analysis;
- Dependabot updates.
GitHub-hosted CPU runners are intentionally used for everything that does not require a licensed StarCraft installation or long GPU training. See docs/CI_STRATEGY.md.
.
├── src/starcraft_ai/ # model, training, CLI
├── collector/ # replay/runtime integration boundary
├── data/sources/ # provenance + source manifests
├── data/snapshots/ # pinned replay identities/checksums
├── tests/ # deterministic CI tests + real replay fixture
├── scripts/ # acquisition / repository tooling
├── docs/ # architecture and research protocol
├── site/ # GitHub Pages landing page
├── MODEL_CARD.md
├── DATA_CARD.md
├── CHANGELOG.md
└── AGENTS.md # Codex/agent operating rules
- No benchmark without a reproducible command, seed, commit, and data manifest.
- No “SOTA” or “human-level” claim without a directly comparable public protocol.
- Non-cheating evaluation must preserve player-visible information boundaries.
- Game binaries and commercial assets are never committed.
- Every complex method must have a simpler baseline and a negative control.
- A lower training loss alone is not evidence of a better game model.
Read docs/REPRODUCIBILITY.md.
Research, negative results, replay tooling, benchmarks, and infrastructure contributions are welcome.
AI coding agents must read AGENTS.md first.
StarCraft, StarCraft: Brood War, StarCraft: Remastered, Battle.net, and Blizzard Entertainment are trademarks of Blizzard Entertainment. This project is independent and is not affiliated with or endorsed by Blizzard Entertainment.