I should not have stopped that codec training early 🙃
generated frames vs codec reconstructions - 0.28 DINO-FDD
codec's own reconstructions vs real (the floor) - 1.86 DINO-FDD
well can't do shit now, trying to optimize the inference like I do get 24fps but I'm not happy with
in a LTX2.3 rabbit hole these days and while reading the arch I saw it has a fat linear (per block) that does (4096 -> 36864) which then feeds to self/mlp scale, shift, gate, some 4 AV cross and some more cross modalities, around a third of the model all driven by ONE timestep