Breaking LLM inference’s autoregressive bottleneck 🛠️
We've teamed up with @haozhangml, @YimingBob, and @aaronzhfeng, among others from UCSD to achieve a massive 3.13X speedup for LLM inference on Google Cloud TPUs using Diffusion-Style Speculative Decoding (DFlash).
Read the
Reasoning VLAs can think. They just can't think fast. Until now.
Introducing FlashDrive⚡
🚀 716 ms → 159 ms on RTX PRO 6000 (up to 5.7×)
✅ Zero accuracy loss
FlashDrive = streaming inference + DFlash speculative reasoning + ParoQuant W4A8
Real-time reasoning for autonomous
“the experts will always tell you it can't be done. build it anyway!”
It is true in more ways than one, but it is a “feature” more than a bug: experts are experts because they can see ways things can fail, and are trained to be careful and precise to ensure correctness. Risk