GPUs are today, of course, the workhorse of AI inference in production, but they aren't actually optimized for that. My guest today is Anton McGonnell, VP of product at SambaNova, a Bay Area company that has raised over $2 billion to develop a chip optimized for AI inference.
Why listen
It goes beyond the title with direct discussion of inference, really, data, including: Today's guest built a new chip that is pushing the frontier of real-time AI speed and bandwidth.
Key takeaways
01GPUs are today, of course, the workhorse of AI inference in production, but they aren't actually optimized for that
02My guest today is Anton McGonnell, VP of product at SambaNova, a Bay Area company that has raised over $2 billion to develop a chip optimized for AI inference
03In this episode, Anton explains why Agentic AI has changed the shape of inference workloads and how SambaNova's chip sidesteps the memory bottleneck that slows GPUs when generating
Best for
listeners looking for a practical AI episode debrief