Every time a new trillion-dollar industry emerges, there's a need for this independent testing group. When Meta released Lama 4 on our held-out private benchmarks, the model was actually underperforming, but on all of the major public benchmarks, it was showing incredible capabilities.
Why listen
It goes beyond the title with direct discussion of think, like, actually, including: When Meta released Lama 4 on our held-out private benchmarks, the model was actually underperforming, but on all of the major public benchmarks, it was showing incredible capabilit.
Key takeaways
01Every time a new trillion-dollar industry emerges, there's a need for this independent testing group
02When Meta released Lama 4 on our held-out private benchmarks, the model was actually underperforming, but on all of the major public benchmarks, it was showing incredible capabilit
03In an ideal world, take a frontier model and have it train the next version of itself, but obviously that's very expensive and slow, and so what we're doing is forming a set of pro
Best for
research-minded practitioners comparing model behavior