
Practical AI — YouTube · 2026-07-18
YouTubeCorey Sanders on AI's Unique Infrastructure Needs
Hosts: Unknown
Guests: Corey Sanders
Why it matters
CoreWeave's Corey Sanders explains why AI training needs specialized infrastructure beyond standard public cloud setups.
Key claims
- AI training requires large tranches of deeply interconnected GPU deployments unlike standard cloud workloads
- GPUs are so expensive that any slowdown, including failures or storage loading issues, becomes economically significant
- Sanders first realized these infrastructure needs while working on AI platforms during his final years at Microsoft
- Traditional public cloud fungibility (deploy more compute or storage as needed) doesn't work for AI workloads with specialized networking
Radar summary
Summary
Corey Sanders, a leader at CoreWeave, explains why AI training workloads demand fundamentally different infrastructure than traditional cloud computing. He contrasts the commoditized, on-demand model of public clouds—where compute and storage can be deployed incrementally—with the tightly interconnected, pre-planned deployments required for AI training at scale, dependent on specialized networking like InfiniBand and purpose-built storage.
Sanders traces his own realization about these differences back to his later years at Microsoft, where he was pulled into discussions on optimizing AI infrastructure deployment despite industry solutions being his official role. That experience highlighted how AI workloads break the standard model of fungible cloud resources, requiring upfront planning rather than just-in-time provisioning.
Since moving to CoreWeave, Sanders says he came to appreciate even more deeply the breadth of customization required across the stack—from GPU-level observability and caching to bare-metal Kubernetes orchestration tuned specifically for AI. He frames this depth of specialization as economically justified given the cost of GPUs, arguing customers will pay a premium for infrastructure that extracts maximum value from every expensive accelerator.
- AI training requires large tranches of deeply interconnected GPU deployments unlike standard cloud workloads
- GPUs are so expensive that any slowdown, including failures or storage loading issues, becomes economically significant
- Sanders first realized these infrastructure needs while working on AI platforms during his final years at Microsoft
- Traditional public cloud fungibility (deploy more compute or storage as needed) doesn't work for AI workloads with specialized networking
- CoreWeave differentiates via observability, storage, and orchestration tuned for massive-scale AI training
- Specialized caching, bare-metal Kubernetes, and GPU-aware orchestration are key to squeezing maximum value from accelerators
Source material
Full source text
AI training and technology.
What is AI training and how to use AI training?
Look, I mean, at a very basic level, there are kind of two big aspects of sort of what of how AI is used or leveraged.
One is on the training side, right?
One is the actual creation of these models, right?
Creation of these weights, right?
The making, sort of doing the learning.
In a very similar way to the way the human brain works, but doing the learning to actually create the output that we then engage with, whether it be open AI, whether it be anthropic, whether it be, you know, meta, sort of all of those have that sort of background in training.
And that requires very specific infrastructure to be able to deliver.
And the infrastructure is very expensive, right?
As we all know, and it's very, you know, deployed in sort of large tranches and very interconnected, right?
One of the biggest requirements for many of these training workloads is there sort of these large tranches of deployments that are all deeply interconnected.
So they're all working together.
And, you know, as part of that training area, some of the things that, you know, core we've identified and, you know, certainly even the broader market identified is the potential places where that work slows down because the GPUs are so expensive, right?
Any sort of slowdown and being able to get the job done is impactful.
And so that includes things like failures in the GPUs that includes things like storage being loaded into the GPUs, right?
And so all the way down at the infrastructure level, what's going wrong with the GPUs, are there failures into all the way up to the orchestration over which job is running on which infrastructure to be able to optimize the output, right?
And so this is where I think AI focused direction, that's a very different set of requirements than just your standard public cloud with a whole bunch of computer racks, right?
Right?
It requires a bunch of both design and and how the hardware is being laid out and specifically how it's interconnected.
It's very unique to AI workloads all the way up to then how the jobs are being orchestrated with knowledge over what's being run down at the core level of the platform.
And this is a bunch of places that I think, you know, core we've has a bunch of differentiation on things like our observability platform, our storage platform, right?
There's a bunch of places where we offer unique services targeting that type of AI massive scale training workload that then is just very different than your again, your classic cloud workload.
So yeah, go ahead pause there because the second half of the story I'll get to never go ahead, Chris, I'm ready you had you had a follow up and I'll talk for I'll talk 45 minutes uninterrupted if you let me so let me be let me pause and let you kind of jump in.
Go ahead.
Yeah, no, it's fine.
And don't lose the second half of that story because you have which is why I had the follow up on there.
I think the I am curious as you're talking about that transition.
Like if you bring like for you personally, when you were before you came to core weave, when did you have that realization like is this is as whatever this epiphany is is occurring?
You know, was it all at once?
Did it happen over time?
Like when did you realize I want to do something different?
And don't lose that second half of the story.
Yeah, totally.
Yeah, no, no, no, right.
And the second half of the story is very similar to the first half, just a different a different sort of outcome, actually.
But yeah, I mean, for me, look, I got the I got the honor of my final years at Microsoft getting to work on sort of a lot of the AI infrastructure aspects in the platform.
And it was it was kind of a funny thing because it was sort of not my day job.
Like my day job, I was working on industry solutions.
So like financial services, you know, you know, retail services, etc, trying to think through how these large verticals could leverage AI in their workloads, which actually is the second part of the answer.
But the first part of the answer, I ended up getting pulled into a lot of the discussion over how at Microsoft, were we optimizing the deployment of that infrastructure, right?
And that that realization really struck me there, which was like, you know, you know, the prior to that, a lot of the a lot of the big cloud strategy was, you end up having a bunch of space and power.
And as the needs come up, you deploy what you need to, right?
So you need more compute in a given location, you deploy it.
And within, you know, days, you have a huge amount of additional compute, you need a bit more storage in a location, you can go deploy a bit more storage, and you can kind of add, and that sort of commoditization, that sort of fungibility, is a big value add for the big clouds at the scale that they're running.
But it turns out, you can't really do that, when you're dealing with sort of AI workloads, again, all interconnected, all wired together with, let's say, InfiniBand or Rocky, very specialized network behind it, very specialized storage underneath it.
Suddenly, it's like, you've got to pre plan things, right?
You've got to, you got to be thinking about it ahead of time.
And so it definitely was, it definitely for me was an experience of learning, right?
It started there.
And then as as I moved over to core weave, I think that the realization of how many places from, again, caching, specific caching for AI workloads to specific approaches to Kubernetes, which is, of course, a standard orchestration solution, but can deliver focused AI enablement for bare metal based deployment to really squeeze out all the juice of that GPU, right?
All the way up to again, the orchestrator being sort of aware, that realization of like, how many aspects of the platform needed to be could be and then needed to be customized for AI specific workloads.
That was really, I feel like I didn't fully get it until I came to core weave.
But I saw it and the beginnings of it when I went in Microsoft and I don't know whether that's a Microsoft or Corey thing or more just a point in time thing for me personally, like it's more probably that learning over time that you start seeing more and more of this customization realized wow, like, this is because of how much value and cost there is behind this.
It is worth it to do this customization because people will pay for it to get the most out of those very expensive GPUs.
Thank you.