Machine Learning / LLM Engineer. I build LLM systems end to end, from architecture and application logic through to production: self-hosted inference (vLLM), agentic workflows (LangGraph), RAG, and high-volume document AI, with the MLOps to run it reliably.
- Self-hosted, privacy-first LLM systems (vLLM, on-prem GPU inference)
- RAG and agentic pipelines (LangGraph, MCP, vector databases)
- Document AI at scale (OCR, structured extraction) and voice/ASR
- MLOps: Docker, GitLab CI/CD, Airflow, monitoring
- Local Multimodal AI Chat: self-hostable multimodal chat with local LLMs (Ollama/OpenAI), PDF RAG, image chat, Whisper voice. 200+ stars, tested, CI on Ubuntu/Windows.
- Telegram PDF-RAG Bot: question answering over your PDFs with RAG (LangChain + FAISS).
- MCP Server: YouTube to Notion: a Model Context Protocol server that summarizes YouTube videos into Notion.
LinkedIn 路 leonsander.com 路 YouTube 路 leonsander.consulting@gmail.com
Stack: Python 路 vLLM 路 LangGraph 路 RAG 路 FastAPI 路 PostgreSQL/pgvector 路 Chroma 路 PyTorch 路 Hugging Face 路 Docker 路 Airflow

