AI Software Engineer

Making capable AI models run fast on real hardware.

I'm Khanh D. Nguyen. I work on inference engineering, model optimization, and edge AI for robotics - building lean C/C++ inference engines that bring VLMs, VLAs, and LLMs to CPUs, accelerators, and robots. Inspired by llama.cpp: a few lines of simple code, big results on your own hardware.

khanh.cpp
// about

About me

I build and ship software where deep learning meets real hardware. My work focuses on designing SDKs for computer-vision models, optimizing neural networks for AI accelerators, and decomposing operations so they run fast and in parallel - getting capable models onto edge devices and robots.

Edge AI Model optimization Robotics VLM VLA LLM C / C++ Python · PyTorch

Mentor

An Thai Le

Assistant Professor at VinUniversity and Director of Foundation AI at VinRobotics, working on robots that plan, learn, and act reliably with limited compute and data.

Now

AI Software Engineer

Inference engineering · Model optimization · Edge AI for robotics

Research

M.S. in AI Convergence

Multimodal learning & affective behavior analysis - 6 peer-reviewed papers (CVPRW, ECCVW, RSSW)

// projects

Pinned projects

A few things I'm proud of - small in code, big in scope.

vla.cpp

The open-source project I'm P.I.C of at VinRobotics - bringing Vision-Language-Action models to local hardware, in the spirit of llama.cpp for robotics.

C++ggmlInferenceVLA

K.D. Nguyen, H.T. Ho, C.T. Nguyen, T.Q. Duong, L.D. Le, D.M.H. Nguyen, V.A. Ngo, A.T. Le

View on GitHub ↗

vla.simd

Efficient CPU inference for language-conditioned manipulation - shared SIMD micro-kernels (AVX2 / NEON) plus IMPACT, a language-conditioned policy that supplies 30+ actions per second on a Raspberry Pi 5.

C++SIMDCPU inferenceVLA

K.D. Nguyen, H.M. Truong, A.T. Le

View project page ↗

REX

Neural network Representation EXchange - a hands-on study of how different frameworks represent and serialize neural networks.

IRFrameworksCompiler
View on GitHub ↗

FaceGen

Multiple Appropriate Facial Reaction Generation - synthesizing diverse listener reactions from a speaker's audio-visual cues. 3rd place at REACT 2024.

GenerativeGaussian Mixture ModelsVAEMultimodal
View on GitHub ↗

ReadItDown

Native markdown viewer and editor in Linux/Windows/macOS

MarkdownEditorTerminal
View on GitHub ↗
// contact

Say hello

Always happy to talk edge AI, optimization, and robotics.