Pilot listens for a wake word ("Pilot"/"Hey Pilot") or a push-to-talk key, transcribes speech locally, and runs actions natively. Deterministic requests are executed locally with command-mapping and execute almost instantaneously, whereas requests that require reasoning or visual context are routed to an LLM with tool calling!
Local commands
- Media: play/pause, next/prev, volume, mute
- Apps: open, focus, minimize by name or process
- Windows: snap left/right, fullscreen, minimize all, focus only one app
- Text: type anything via clipboard injection, dictation mode
- Shortcuts: copy, paste, undo, new tab, close tab
- System: lock screen, set exact volume level
- Reminders: "remind me in 5 minutes to check the build"
LLM delegation
- "What's the weather in Toronto?" (spoken answer)
- "Describe what's on my screen" (takes screenshot and speaks a description)
- "Explain this error" (screenshot + spoken explanation)
- Anything the local dictionary doesn't cover
Voice I/O
- Wake word: say "Pilot" or "Hey Pilot"
- Push-to-talk: hold F8 (configurable) to bypass wake detection entirely
- TTS responses via Kokoro (local, runs without an API key)
...and MUCH more planed :)!
Requirements: Windows 10/11, MSVC 2022 Build Tools, CMake 3.21+, Ninja, Git, Python 3.x.
# Clone
git clone https://github.com/Gurvirr/Hey-Pilot.git
cd Hey-Pilot
# Download models (~580 MB total)
./scripts/fetch_deps.ps1
# Optional: add a free Gemini API key for LLM routing
cp pilot.cfg.example pilot.cfg
# Edit pilot.cfg, set gemini_api_key (get one free at https://aistudio.google.com)
# Build
./build.ps1
# Run
./build/pilot_listen.exeSay "Pilot" or hold F8 for push-to-talk.
Check ROADMAP.md for upcoming features and pilot.cfg.example for options.
Mic -> WASAPI -> lock-free ring -> Wake/VAD -> STT -> Router -> OS Driver
|
(no match)
|
LLM Orchestrator -> TTS
(+ screenshot if visual)
Every stage runs on its own thread, meaning no stage will ever be blocked by another.
Stack
| Layer | Choice |
|---|---|
| Audio capture | WASAPI event-driven, shared mode |
| Wake word | sherpa-onnx KWS (GigaSpeech Zipformer, BPE) |
| VAD | Silero VAD via sherpa-onnx |
| STT | NVIDIA Parakeet-TDT-0.6B-v2 via sherpa-onnx (CPU int8 or CUDA fp32) |
| Command routing | Local FST/trie with Levenshtein fuzzy matching |
| LLM | Gemini 2.5 Flash (vision) + Groq Llama (text fallback) |
| TTS | Kokoro-82M int8 via sherpa-onnx |
| OS automation | Win32 / WinRT / Core Audio / WinHTTP |
| Screen capture | GDI + WIC JPEG encode |
WASAPI event-driven capture, resampled to 16 kHz mono f32. It packs audio into 30ms frames and pushes them through a lock-free ring buffer with zero heap allocations in the hot path.
./build/pilot_audio.exe # capture until Ctrl+C
./build/pilot_audio.exe 3 # capture for 3 s then exitStreaming keyword spotting via sherpa-onnx (Zipformer KWS). Detects "Pilot" / "Hey Pilot" using a PIMPL wrapper. Silero VAD handles speech boundaries, while a low-level keyboard hook handles push-to-talk.
Local speech-to-text via NVIDIA Parakeet (RNN-Transducer). Runs about 20x real-time on CPU, with a CUDA path available for a 4.4x speedup on GPU.
$p = "models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8"
./build/pilot_transcribe.exe $p "$p/test_wavs/0.wav"Transcribed text is matched against a command-map (exact + fuzzy verbs). Commands are executed natively.
./build/pilot_act.exe open notepad
./build/pilot_act.exe --dry "next track" # route only, no executionUnmatched requests are sent to Gemini 2.5 Flash (with desktop screen captures for visual context) or Groq Llama. The agent speaks answers back using Kokoro TTS locally.
STT runs locally on CPU by default (~235ms for a 3s clip). With a discrete NVIDIA GPU, you can get ~53ms using the fp32 model.
# Install CUDA 12 + cuDNN 9 runtime DLLs via pip wheels (no full Toolkit needed)
./scripts/setup_cuda.ps1
# Build against the CUDA sherpa package
$cuda = "third_party/sherpa-onnx-v1.13.2-cuda-12.x-cudnn-9.x-win-x64-cuda"
cmake -S . -B build-cuda -G Ninja -DCMAKE_BUILD_TYPE=Release -DSHERPA_DIR=$cuda
cmake --build build-cuda
# Run with GPU
./build-cuda/pilot_listen.exe --provider=cudaNote: the int8 model is actually slower on GPU than CPU (it's a CPU optimization). Use the fp32 model for GPU. See scripts/fp16_to_fp32.py to convert.
Copy pilot.cfg.example to pilot.cfg. Key options:
[stt]
model_dir = models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8
provider = cpu # or cuda
[llm]
gemini_api_key = # free at https://aistudio.google.com
groq_api_key = # free at https://console.groq.com
[tts]
speaker = 0 # 0=female, 3=male (adam), 4=male (michael)
[input]
ptt_key = 0x77 # F8src/
audio/ WASAPI capture, ring buffer
wake/ Wake word interface + sherpa-onnx KWS backend
vad/ Silero VAD wrapper
stt/ STT interface + Parakeet backend
action/ Action types, command router, Win32 OS driver
llm/ Gemini + Groq backends, orchestrator, screen capture
tts/ Kokoro TTS wrapper
input/ Global PTT keyboard hook
scripts/
fetch_deps.ps1 Download all models (run once after cloning)
setup_cuda.ps1 Install CUDA runtime DLLs for GPU acceleration
fp16_to_fp32.py Convert Parakeet fp16 model to fp32 for GPU use
I have a lot of ideas to make Pilot even better! You can check out ROADMAP.md for some of the planned features. Stay tuned!
Source code: Apache-2.0
Third-party models have their own licenses. Parakeet-TDT-0.6B-v2 is CC-BY-4.0 (attribution required). Kokoro-82M is Apache-2.0. sherpa-onnx is Apache-2.0. Silero VAD is MIT.
