vla.cpp
The open-source project I'm P.I.C of at VinRobotics - bringing Vision-Language-Action models to local hardware, in the spirit of llama.cpp for robotics.
View on GitHub ↗AI Software Engineer
I'm Khanh D. Nguyen. I work on inference engineering, model optimization, and edge AI for robotics - building lean C/C++ inference engines that bring VLMs, VLAs, and LLMs to CPUs, accelerators, and robots. Inspired by llama.cpp: a few lines of simple code, big results on your own hardware.
I build and ship software where deep learning meets real hardware. My work focuses on designing SDKs for computer-vision models, optimizing neural networks for AI accelerators, and decomposing operations so they run fast and in parallel - getting capable models onto edge devices and robots.
A few things I'm proud of - small in code, big in scope.
The open-source project I'm P.I.C of at VinRobotics - bringing Vision-Language-Action models to local hardware, in the spirit of llama.cpp for robotics.
View on GitHub ↗Efficient CPU inference for language-conditioned manipulation - shared SIMD micro-kernels (AVX2 / NEON) plus IMPACT, a language-conditioned policy that supplies 30+ actions per second on a Raspberry Pi 5.
View project page ↗Neural network Representation EXchange - a hands-on study of how different frameworks represent and serialize neural networks.
View on GitHub ↗Multiple Appropriate Facial Reaction Generation - synthesizing diverse listener reactions from a speaker's audio-visual cues. 3rd place at REACT 2024.
View on GitHub ↗Always happy to talk edge AI, optimization, and robotics.