On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken. We've seen video models as being built on next frame prediction.
Why listen
It goes beyond the title with direct discussion of like, kind, model, including: We know LLMs are built on next token prediction.
Key takeaways
01On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken
02We've seen video models as being built on next frame prediction
03Each time we made the model bigger and each time we trained it for longer, it got significantly better
Best for
research-minded practitioners comparing model behavior