Skip to content
GurvirrPublic

About

A low-latency voice-to-action computer-use agent.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

82 Commits

Folders and files

Repository files navigation

Hey Pilot

A bare-metal, low-latency voice-to-action computer-use agent, written in C++.

GitHub License
Hey Pilot

Pilot listens for a wake word ("Pilot"/"Hey Pilot") or a push-to-talk key, transcribes speech locally, and runs actions natively. Deterministic requests are executed locally with command-mapping and execute almost instantaneously, whereas requests that require reasoning or visual context are routed to an LLM with tool calling!

What it can do

Local commands

  • Media: play/pause, next/prev, volume, mute
  • Apps: open, focus, minimize by name or process
  • Windows: snap left/right, fullscreen, minimize all, focus only one app
  • Text: type anything via clipboard injection, dictation mode
  • Shortcuts: copy, paste, undo, new tab, close tab
  • System: lock screen, set exact volume level
  • Reminders: "remind me in 5 minutes to check the build"

LLM delegation

  • "What's the weather in Toronto?" (spoken answer)
  • "Describe what's on my screen" (takes screenshot and speaks a description)
  • "Explain this error" (screenshot + spoken explanation)
  • Anything the local dictionary doesn't cover

Voice I/O

  • Wake word: say "Pilot" or "Hey Pilot"
  • Push-to-talk: hold F8 (configurable) to bypass wake detection entirely
  • TTS responses via Kokoro (local, runs without an API key)

...and MUCH more planed :)!

Getting Started

Requirements: Windows 10/11, MSVC 2022 Build Tools, CMake 3.21+, Ninja, Git, Python 3.x.

# Clone
git clone https://github.com/Gurvirr/Hey-Pilot.git
cd Hey-Pilot

# Download models (~580 MB total)
./scripts/fetch_deps.ps1

# Optional: add a free Gemini API key for LLM routing
cp pilot.cfg.example pilot.cfg
# Edit pilot.cfg, set gemini_api_key (get one free at https://aistudio.google.com)

# Build
./build.ps1

# Run
./build/pilot_listen.exe

Say "Pilot" or hold F8 for push-to-talk. Check ROADMAP.md for upcoming features and pilot.cfg.example for options.

How it works

Mic -> WASAPI -> lock-free ring -> Wake/VAD -> STT -> Router -> OS Driver
                                                          |
                                                   (no match)
                                                          |
                                                   LLM Orchestrator -> TTS
                                                   (+ screenshot if visual)

Every stage runs on its own thread, meaning no stage will ever be blocked by another.

Stack

Layer Choice
Audio capture WASAPI event-driven, shared mode
Wake word sherpa-onnx KWS (GigaSpeech Zipformer, BPE)
VAD Silero VAD via sherpa-onnx
STT NVIDIA Parakeet-TDT-0.6B-v2 via sherpa-onnx (CPU int8 or CUDA fp32)
Command routing Local FST/trie with Levenshtein fuzzy matching
LLM Gemini 2.5 Flash (vision) + Groq Llama (text fallback)
TTS Kokoro-82M int8 via sherpa-onnx
OS automation Win32 / WinRT / Core Audio / WinHTTP
Screen capture GDI + WIC JPEG encode

Milestones

1. Audio loop

WASAPI event-driven capture, resampled to 16 kHz mono f32. It packs audio into 30ms frames and pushes them through a lock-free ring buffer with zero heap allocations in the hot path.

./build/pilot_audio.exe        # capture until Ctrl+C
./build/pilot_audio.exe 3      # capture for 3 s then exit

2. Wake word & VAD

Streaming keyword spotting via sherpa-onnx (Zipformer KWS). Detects "Pilot" / "Hey Pilot" using a PIMPL wrapper. Silero VAD handles speech boundaries, while a low-level keyboard hook handles push-to-talk.

3. Speech-to-Text

Local speech-to-text via NVIDIA Parakeet (RNN-Transducer). Runs about 20x real-time on CPU, with a CUDA path available for a 4.4x speedup on GPU.

$p = "models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8"
./build/pilot_transcribe.exe $p "$p/test_wavs/0.wav"

4. Router & execution

Transcribed text is matched against a command-map (exact + fuzzy verbs). Commands are executed natively.

./build/pilot_act.exe open notepad
./build/pilot_act.exe --dry "next track"   # route only, no execution

5. LLM escalation

Unmatched requests are sent to Gemini 2.5 Flash (with desktop screen captures for visual context) or Groq Llama. The agent speaks answers back using Kokoro TTS locally.

GPU acceleration (optional)

STT runs locally on CPU by default (~235ms for a 3s clip). With a discrete NVIDIA GPU, you can get ~53ms using the fp32 model.

# Install CUDA 12 + cuDNN 9 runtime DLLs via pip wheels (no full Toolkit needed)
./scripts/setup_cuda.ps1

# Build against the CUDA sherpa package
$cuda = "third_party/sherpa-onnx-v1.13.2-cuda-12.x-cudnn-9.x-win-x64-cuda"
cmake -S . -B build-cuda -G Ninja -DCMAKE_BUILD_TYPE=Release -DSHERPA_DIR=$cuda
cmake --build build-cuda

# Run with GPU
./build-cuda/pilot_listen.exe --provider=cuda

Note: the int8 model is actually slower on GPU than CPU (it's a CPU optimization). Use the fp32 model for GPU. See scripts/fp16_to_fp32.py to convert.

Configuration

Copy pilot.cfg.example to pilot.cfg. Key options:

[stt]
model_dir = models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8
provider  = cpu   # or cuda

[llm]
gemini_api_key =  # free at https://aistudio.google.com
groq_api_key   =  # free at https://console.groq.com

[tts]
speaker = 0  # 0=female, 3=male (adam), 4=male (michael)

[input]
ptt_key = 0x77  # F8

Project structure

src/
  audio/      WASAPI capture, ring buffer
  wake/       Wake word interface + sherpa-onnx KWS backend
  vad/        Silero VAD wrapper
  stt/        STT interface + Parakeet backend
  action/     Action types, command router, Win32 OS driver
  llm/        Gemini + Groq backends, orchestrator, screen capture
  tts/        Kokoro TTS wrapper
  input/      Global PTT keyboard hook
scripts/
  fetch_deps.ps1    Download all models (run once after cloning)
  setup_cuda.ps1    Install CUDA runtime DLLs for GPU acceleration
  fp16_to_fp32.py   Convert Parakeet fp16 model to fp32 for GPU use

Roadmap

I have a lot of ideas to make Pilot even better! You can check out ROADMAP.md for some of the planned features. Stay tuned!

License

Source code: Apache-2.0

Third-party models have their own licenses. Parakeet-TDT-0.6B-v2 is CC-BY-4.0 (attribution required). Kokoro-82M is Apache-2.0. sherpa-onnx is Apache-2.0. Silero VAD is MIT.

About

A low-latency voice-to-action computer-use agent.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Contributors

Languages