1. X
  2. Joe Ton
Log inSign up
Joe Ton
757 posts
user avatar
Joe Ton
@joeton
Learning to optimize AI on hardware.
Seattle, WA
joeton.ai
Joined April 2022
27
Following
50
Followers
RepliesRepliesMediaMedia
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

  • user avatar
    Joe Ton
    @joeton
    May 15
    TTFT It's basically how long you have to wait until AI first responds. For example, when you first talk to an AI, there's a short delay before first message appears. That delay is called TTFT. Important for user experience. Shorter the better. More technical: 👇
    47
  • user avatar
    Joe Ton
    @joeton
    May 4
    Is HBM with 1 TB enough for KV caching?
    36
  • user avatar
    Joe Ton
    @joeton
    May 2
    I wish the AMD dev program allows more than $100 of credit.
    33
  • user avatar
    Joe Ton
    @joeton
    May 1
    Attention Matching it is.
    29
  • user avatar
    Joe Ton
    @joeton
    Feb 16
    To solve the data movement bottleneck in AI scaling, we need to move beyond the traditional von Neumann architecture -- toward tighter coupling of compute and memory. My concern: will this lead to even stronger vendor lock-in from hardware-specific optimizations/models?
    63