CTO & Co-founder of Nexa AI (acquired by Qualcomm)
personal website → zackli.ai
Website · GitHub · LinkedIn · Twitter/X · zackli4ai@gmail.com
I'm Zack Li, an entrepreneur, builder, and lifelong learner. I build scalable AI systems end to end — from frontier research to inference runtimes that ship on real devices. My interests lie in scalable software architecture design, Gen AI inference and Physical AI. For more information, visit zackli.ai.
I co-founded Nexa AI (acquired by Qualcomm) as Chief Technology Officer.
- Engineering — we built Nexa SDK (now GenieX), a single runtime that runs frontier LLMs and VLMs across NPU, GPU, and CPU in one line of code. It hit #1 on GitHub Trending and #1 Product of the Day, earned 8k+ stars, and was adopted or featured by Qualcomm, AMD, NVIDIA, Intel, and IBM.
- Research — I co-authored the Octopus and Omni model series: on-device agents built on functional tokens, and sub-billion multimodal models for vision and audio. They reached the top 2 trending on Hugging Face, were spotlighted at Google I/O, and OmniAudio was featured in Google DeepMind's Gemmaverse. We also shipped Hyperlink, a fully local file-search agent that reached 30k users in two months.
- Business — I owned enterprise delivery for our customers: HP, Lenovo, İşbank; initiated partnerships with AMD, Intel, Hugging Face and more; and drove technical showcases at CES 2024–2026, Snapdragon Summit, Microsoft Ignite and more.
After the acquisition, I'm a Senior Staff Machine Learning Engineer at Qualcomm, where our team is building GenieX, the generative AI inference runtime for Qualcomm platforms. I also work on llama.cpp for the Snapdragon Hexagon NPU.
Before Nexa, I built on-device AI and software systems at Google, Amazon Lab126, and Cadence. I hold an M.S. from Stanford University (2019), where I was a teaching assistant for CS229T Machine Learning Theory, and a B.E. from Tongji University (2017).
Selected publications (Google Scholar)
- Octo-planner: On-Device Language Model for Planner-Action Agents — EMAS 2025, pp. 141–156
- Octopus: On-Device Language Model for Function Calling of Software APIs — NAACL 2025, pp. 329–339
- DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models — ICDM 2025 — Best Paper Runner-Up
- AutoNeural: Co-Designing Vision-Language Models for NPU Inference — arXiv:2512.02924, 2025
- Octopus v2: On-Device Language Model for Super Agent — arXiv:2404.01744, 2024
- OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference — arXiv:2412.11475, 2024
- On-Device Language Models: A Comprehensive Review — arXiv:2409.00088, 2024
- GenieX — Qualcomm's generative AI inference runtime, bringing frontier LLMs and VLMs to local NPU, GPU, and CPU in a few lines of code. Grown out of Nexa SDK.
- llama.cpp — the C/C++ LLM inference engine behind much of local AI; I work on the Snapdragon Hexagon NPU path.
- ai-hub-models — Qualcomm AI Hub's model zoo of optimized models ready to deploy on-device.
Silicon
- Qualcomm — developer blogs on Snapdragon Android, the Hexagon NPU, Granite 4.0, and IoT & robotics; featured on Qualcomm's X and LinkedIn, plus seven "This Week in AI" posts (1, 2, 3, 4, 5, 6, 7).
- NVIDIA — RTX blogs on our local agent and CES booth, plus a 1.3M-view video.
- AMD — blogs on NPU image generation and DeepSeek R1 speedups with NexaQuant; posts from AMD, AMD Enterprise, and AMD Developer, plus an interview on its developer channel.
- Intel — spotlighted NexaSDK on its developer channel for on-device LLMs on Intel's NPU.
Models
- Google — DeepMind featured OmniAudio in the Gemmaverse, a showcase of models built on Gemma; highlighted in the 2024 Gemma keynote, with posts from Google AI Developers and Google Devs.
- OpenAI — its Developer Relations team amplified our work running open models on consumer hardware.
- Hugging Face — ranked our models among the most-downloaded from 2022 to today; CEO Clément Delangue called to collaborate.
- IBM — listed NexaSDK as a supported framework for Granite 4.0, next to vLLM, llama.cpp, and MLX.
- Alibaba — the Qwen team called out our day-0 Qwen3-VL support and Qwen3 on the Qualcomm NPU.
- Liquid AI — partnered with us on the LFM2.5 launch and LFM2.5 Thinking.
- pyannoteAI — ran their diarization model on Qualcomm NPUs through Nexa SDK for CES 2026.
Platforms
- Microsoft — demoed our product on stage at the Ignite 2025 keynote and listed us as a software partner for Windows.
- Apple — the MLX team reposted our work running multimodal models on Apple silicon.
Industry talks — Microsoft Ignite 2025 · CES 2026 (NVIDIA & AMD booths) · Qualcomm Snapdragon Summit · GenAI Summit San Francisco · PyTorch Conference · AMD Advancing AI · IBM TechXchange · Open Data Science Conference





