TTFT
It's basically how long you have to wait until AI first responds.
For example, when you first talk to an AI, there's a short delay before first message appears. That delay is called TTFT.
Important for user experience. Shorter the better.
More technical: 👇
To solve the data movement bottleneck in AI scaling, we need to move beyond the traditional von Neumann architecture -- toward tighter coupling of compute and memory.
My concern: will this lead to even stronger vendor lock-in from hardware-specific optimizations/models?