AI inference, speculative decoding, open source. Built novel decoding algorithms – default in Hugging Face Transformers (160+ ⭐). Making AI faster + cheaper
reality has no reset button. that's why the next training paradigm might be dreaming.
@dwarkesh_sp's new episode makes the case for training in imagination. think of teaching a robot to hold a glass. the outcome is verifiable (did it break?) but not replayable: once it shatters,
1/ On Training in Imagination -
Dwarkesh's episode has a segment on dreaming as one of the next training paradigms. The idea is that a model learns mostly inside its own, by imagining what would happen, instead of trying out for real.
We have a recent paper on exactly this
Speculative decoding has shown a lot of promise, though broader adoption has taken time due to the complexity of building production-ready tooling and high-quality draft models.
We’re releasing SpecBundle, a collection of large-scale EAGLE-3 draft models trained with SpecForge
Even w/o training, you can still use speculative decoding. No need to train a speculator per model.
Our spec decoding algos for heterogeneous vocabs (open-sourced in HF Transformers; not yet in vLLM) let any off-the-shelf model serve as the speculator. ♻️
That means day-0
Speculative decoding is a powerful way to improve inference performance, but in practice it has been hard to adopt.
Training a unique draft model per LLM is time-consuming, and production-ready training utilities that work cleanly with vLLM have been limited.
Speculators
inference is perhaps the most valuable emerging software category.
as models get smarter and more economically valuable, compute will increasingly be spent drawing samples from the models.
if you'd like to work on inference at openai, reach out — [email protected]. include a