I'm Hari — a production LLM/ML systems engineer with 10+ years shipping models at scale. Today I lead AI engineering at Best Buy India: LLM fine-tuning (LoRA/QLoRA via Unsloth/PEFT, multi-GPU) and serving (vLLM, GKE, hot-swappable adapters), plus real-time fraud detection at 40k events/sec. Before that, GPU-accelerated cybersecurity ML on 8×A100 DGX with NVIDIA Morpheus, supply-chain optimization, and computer vision for semiconductor QA. My roots run all the way back to enterprise mainframes — so I know how real systems stay up.
Selected outcomes from production work. Numbers reflect real systems I've built, trained, and shipped.
A rare engineer who has worked both ends of the stack — the COBOL/Db2 systems that quietly run enterprises, and the LLM systems reinventing them.
I'm a production LLM/ML systems engineer. My focus is the unglamorous, decisive part: taking a model from "works in a notebook" to "serves real traffic" — reliably, observably, and at scale.
I started on enterprise mainframes — COBOL, JCL, Db2 — where reliability isn't optional. That foundation became my edge: I moved through data science, GPU-accelerated ML (8×A100 DGX, NVIDIA Morpheus), and into LLM engineering — fine-tuning with LoRA/QLoRA via Unsloth/PEFT and serving with vLLM and hot-swappable adapters.
Today I lead AI engineering at Best Buy India: real-time fraud detection at 40k events/sec, self-hosted LLM initiatives, and the MLOps that keeps it all alive across GCP and Azure.
I'm now going deeper — toward inference internals and frontier-scale AI: turning "framework user" into "framework builder." I build in public, write what I learn, and chase the hard problems.
Leading with the LLM systems stack — then the full lifecycle that gets a model into production and keeps it there.
Open-source work, from local-first AI tooling to production-minded utilities — plus the deeper systems I'm building now.
Interactive tool to find the optimal vLLM-compatible LLM for any GPU — 108 models filtered by VRAM, quantization & KV-cache fit across A100 → B200.
Open the live app →Code ↗Daily AI news brief from free sources (RSS, arXiv, Hacker News), synthesized with Gemini and exported as polished PDF + Markdown.
Explore repo →Run any LLM fully locally on your Mac with a single command — private, offline, and zero cloud dependency.
Explore repo →Remote command executor & project launcher — single binary, zero dependencies, real-time SSE streaming.
Explore repo →A self-contained personal & family finance tracker that runs entirely on your machine — no accounts, no cloud, no dependencies. Just Python and a browser.
Explore repo →Config-driven LoRA → eval → vLLM serve. Plus inference internals from scratch — KV-cache, paged attention, speculative decoding.
Shipping in public, 2026Sharing what I learn, and keeping the fundamentals sharp.
Open to senior AI / ML / research-engineering roles and ambitious collaborations. If you're pushing the frontier — let's talk.