
BG2Pod with Brad Gerstner and Bill Gurley · 2025-09-11
PodcastYouTubeOpenAI Enterprise, GPT-5, and Why Digital Autonomy Lags Physical
Hosts: Brad Gerstner, Bill Gurley
Guests: Sherwin Wu, Olivier Godement
Why it matters
OpenAI's Sherwin Wu and Olivier Godement on enterprise deployments, GPT-5 tradeoffs, real-time voice, and RFT.
Key claims
- OpenAI Platform framed as mission-critical for distributing AGI benefits through B2B channels including Fortune 500, startups, and government.
- Three flagship enterprise deployments: T-Mobile (AI voice support, GA real-time API), Amgen (GPT-5 for drug R&D and regulatory docs), Los Alamos (O3 deployed on-prem on air-gapped Venado supercomputer with physically transported model weights).
- Forward Deployed Engineering (FDE) team, a Palantir-style function, embeds deeply with customers to build scaffolding, integrations, and evals.
- Keys to successful deployments: top-down executive buy-in, bottom-up tiger team with institutional knowledge, evals-first, and a 'helpline' to push from ~46% to 99% quality.
Radar summary
Summary
Sherwin Wu and Olivier Godement, who lead engineering and product for the OpenAI Platform, walk through OpenAI's enterprise business and the work of its Forward Deployed Engineering (FDE) team. They detail three flagship deployments: T-Mobile, where OpenAI models power both text and natural-sounding voice support; Amgen, where GPT-5 accelerates drug R&D and the regulatory/document workflow; and Los Alamos National Labs, where the O3 reasoning model was deployed on-prem on an air-gapped supercomputer called Venado, with model weights physically transported into the facility. They frame the platform as essential to OpenAI's mission of distributing AGI's benefits, since many high-impact use cases in healthcare, telecom, and government only flow through B2B channels.
On the model side, they discuss GPT-5 as OpenAI's first launch framed as a full system rather than a single model, with deliberate focus on behavior, style, and tone alongside capability. Key tradeoffs include reasoning depth vs. latency, and the team describes customer-driven feedback that made GPT-5 very literal in instruction following, sometimes requiring prompt rewrites. They cover the recent GA of the real-time voice API (which unifies STT, reasoning, and TTS rather than stitching them), and Reinforcement Fine Tuning (RFT) as a more powerful successor to supervised fine-tuning, with early wins from Rogo (financial docs) and Accordance (tax).
On the broader question of why digital autonomy lags self-driving despite a lower safety bar, they point to two factors: decades of accumulated work and real-world scaffolding (roads, signs, laws) versus AI agents being dropped into unstructured environments with no standard interfaces. They also address the MIT report that 95% of AI deployments fail, attributing failures to missing scaffolding, lack of eval discipline, and absence of a cross-functional tiger team with top-down sponsorship. The episode closes with rapid-fire picks: Sherwin is long esports and short the AI tooling/evals category, Olivier is long healthcare/pharma AI and short memorization-based education, and both share AGI-pilled moments and career advice emphasizing critical thinking and AI-native fluency.
- OpenAI Platform framed as mission-critical for distributing AGI benefits through B2B channels including Fortune 500, startups, and government.
- Three flagship enterprise deployments: T-Mobile (AI voice support, GA real-time API), Amgen (GPT-5 for drug R&D and regulatory docs), Los Alamos (O3 deployed on-prem on air-gapped Venado supercomputer with physically transported model weights).
- Forward Deployed Engineering (FDE) team, a Palantir-style function, embeds deeply with customers to build scaffolding, integrations, and evals.
- Keys to successful deployments: top-down executive buy-in, bottom-up tiger team with institutional knowledge, evals-first, and a 'helpline' to push from ~46% to 99% quality.
- GPT-5 launched as a full system with focus on behavior and craft, not just benchmarks; major tradeoffs around reasoning tokens vs. latency, and very literal instruction following that broke some old prompts.
- Real-time API unifies STT, reasoning, and TTS into a single speech-to-speech model, with some customers still on stitched pipelines.
- Reinforcement Fine Tuning (RFT) is positioned as a step-change over supervised fine-tuning, with customers like Rogo and Accordance achieving best-in-class results on financial docs and tax.
- Digital autonomy lags physical autonomy because self-driving had ~15 years and standardized scaffolding (roads, signs), while AI agents are day-one and dropped into unstructured environments.
- Picks: Sherwin long esports, short AI tooling/evals startups; Olivier long AI in healthcare/pharma, short memorization-based education.
Source material
Full source text
We literally had to bring the weights of the model physically into their supercomputer.
In San Francisco, you could take a car from one part of SF to the other, fully autonomously.
As opposed to the digital world, I can't book a ticket online right now.
Physical autonomy is ahead of digital autonomy in 2025.
I think AI agents are like really in day one here.
Like chat GPT only came out in 2022.
And the slope I think is incredibly steep.
I actually do think self-driving cars have a good amount of scaffolding in the world.
You have roads, roads exist.
They're pretty standardized.
Yes, stoplights.
AI agents are just kind of dropped in the middle of nowhere.
We'll start with long, short game.
I'm short on the entire category of like tooling, EVALS products.
Healthcare is probably the industry that will benefit the most from AI.
I think I'm a JPL.
You're definitely a JPL.
The first one was the realization in 2023 that I would never need to code manually like ever again.
Hey folks, I'm Apoor Vaghrival and today at the OpenAI office, we had a wide ranging conversation about OpenAI's work in enterprise.
I have with me the head of engineering and head of product of the OpenAI platform, Sherwin Wu and Olivia Godinant.
OpenAI is well known as the creator of chat GPT, which is a product that billions across the world have come to love and enjoy.
But today we dive into the other side of the business, which is OpenAI's work in enterprise.
We go deep into their work with specific customers and how OpenAI is transforming large and important industries like healthcare, telecommunications and national security research.
We also talk about Sherwin and Olivia's outlook on what's next in AI, what's next in technology and their picks both on the long and short side.
This is a lot of fun to do.
I hope you really enjoy it.
Well, two world class builders, two people who make look building easy.
Sherwin, my Palantir 2013 classmate, tennis buddy with two stops at Cora and Opendoor through the IPO before joining OpenAI.
Before chat GPT, you've now been here for three years and lead engineering for all OpenAI platform.
Olivier, former entrepreneur, winner of the Golden Llama at Stripe where you were for just under a decade and now lead all of the product at OpenAI platform.
That's right.
Thanks for doing it.
Thank you.
Thanks for having us.
You as a shareholder, as a thought partner, kicking ideas back and forth.
I always learn a lot from you guys and so it's a treat.
It's a real treat to be do this for everybody.
You know, I'll open with people know OpenAI as the firm that build chat GPT, the product that they have in their pocket that comes with them every day to work, to personal lives.
But the focus for today is OpenAI for enterprise.
You guys lead OpenAI platform.
Tell us about it.
What's underneath the OpenAI platform for B2B for enterprise?
Yeah.
So this is actually a really interesting question too because when I joined OpenAI around three years ago to work on the API, it was actually the only product that we had.
So I think a lot of people actually forget this where the original product for from OpenAI actually was not chat GPT.
It was a B2B product.
It was the API we were catering towards developers.
And so I've actually seen the launch of chat GPT and everything downstream from that.
But at its core, I actually think the reason why we have a platform and why we started with an API is it kind of comes back to the OpenAI mission.
So our mission obviously is to build AGI, which is pretty hard in and of itself, but also to distribute the benefits of it to everyone in the world, to all of humanity.
And it's pretty clear right now to see chat GPT doing that because my mom, maybe even your parents are using chat GPT.
But we actually view our platform and especially our API and how we work with our customers, our enterprise customers as our way of getting the benefits of AGI, of AI, to as many people as possible to everyone in every corner of the world.
Chat GPT obviously is really, really, really big now.
It's I think like the fifth largest website in the world.
But we actually by working through developers using our API, we're actually able to reach even more people in every corner of the world and every different use case that you might have.
And especially with some of our enterprise customers, we're able to reach even use cases within businesses and end users of those businesses as well.
And so we actually view the platform as kind of our way of fully expressing our mission, of getting the benefits of AGI to everyone.
And so concretely though, what the platform actually includes today, the biggest product that we have is obviously our developer platform, which is our API.
Many developers, the majority of the startup ecosystem builds on top of this, as well as a lot of digital natives, Fortune 500 enterprises at this point.
We also have a product that we sell to governments as well in the public sector.
So that's all part of this as well.
And also an emerging product line for us in the platform is our enterprise products.
So we actually might sell directly to enterprises beyond just a core API offering.
Fascinating.
And maybe to double down, I think B2B is actually quite core to the open-air mission.
What we mean by distributing AGI benefits is, I want to live in a world where there are 10x more medicines going out every year.
I want to live in a world where education, public service, civil service, increasingly optimize to everyone.
And there are a large category of use cases that only go through B2B, frankly, unless you enable the enterprises.
And we talk to the parent here, I think that's probably the same fees that parent here.
It's like, hey, those are the businesses where actually making stuff happen in the real world.
And so if you do enable them, if you do accelerate them, like that's how essentially you benefit new to distribute AGI.
Yeah.
Well, maybe we can double click into that, Olivia.
You know, the reach for chat is obviously wide, billions of users.
But for enterprise, it's maybe tell us about it.
Maybe we go deep into a customer example or two.
And what is an organization that we have helped transform maybe?
And at what layers?
So if I were to step back, like we started our B2B efforts with the API like a few years ago.
Initially, the customers were startups, developers, indie hackers, extremely technically sophisticated people, like, you know, who are building like, you know, cool new stuff essentially, and taking massive like, you know, market that they can risk.
So we still have a bunch of customers in that category, and we love them.
And we keep building with them.
On top of that, you know, over the past couple of years, we've been working one more with traditional like enterprises.
And also like digital natives.
Essentially, I think basically everyone woke up like with chat GPT on like, those models are working.
There is a ton of value, and they could see essentially many use cases in enterprise.
Couple of examples which I like the most.
One which is very both fresh and you know, it's quite cool.
We'd be working a lot with T-Mobile.
So T-Mobile leading like US telco operator.
T-Mobile has like, you know, a massive customer spot load.
Like, you know, people asking like, you know, hey, I was charged like that amount of money was going on or you know, my cell phone like isn't working anymore.
A massive like, you know, share of that load is like, you know, voice calls.
People want to talk to someone.
And so for them, like, you know, to be able to essentially automate like more and more, and you know, to help like people like self-serve in a way, like, you know, debug their subscription was pretty big.
And so we've been working with T-Mobile pretty much for the past year.
At that point to basically automate like not only like text support, but also voice support.
And so today, like, you know, there are features like in the T-Mobile app that if you call actually handled by open IR models behind the scenes.
And you know, it does sound like supernatural, like, you know, human sounding latency quality wise.
So that one was really fun.
A second one, which is very just on that.
Can I ask you a follow up question?
So we've got text models.
We've got voice models, maybe even video models someday that are deployed at T-Mobile.
Yeah.
But what above the models or adjacent to the models might we have helped T-Mobile with, for example?
Yeah, there is a tone we're doing.
The first one is, you know, you have to put yourself in the shoes of an enterprise buyer.
Like their goal is to automate, you know, reduce, like, you know, optimize customer support.
And you're going from like a model, like tokens in tokens out to that case, it's hard.
Yeah.
And so, you know, first, like, there's a lot of design, like, you know, system design.
We do have a query now for what deployed engineers who are helping us quite a bit.
For deployed engineers.
Yeah, familiar to the far the term from Palantir.
Yeah, it's a great term.
Were you at these at Penantir?
I was not an F.D.
I was on I think they called it the dev side, right?
It's like software engineering.
I was also only an intern at Palantir.
But yeah, it's a great term.
I think it accurately describes what we're asking folks to do, which is like embed very deeply with customers and and honestly, like build things specific to their systems.
They're deployed onto these customers.
But yeah, we are obviously growing and hiring that team quite a bit because they've been very effective, like T-Mobile.
Four years of my life.
Yeah, yeah, yeah.
Forward deployed.
Yeah.
But go ahead.
So forward deployed engineering.
Forward deployed engineers and the sort of like systems and like integration is that doing is, you know, first, like, you know, you have to orchestrate those models.
Like those models are not just, you know, those models, like, know nothing about like, you know, the CRM, like, you know, and like what's going on.
And so you have to plug the model to like many, many different tools.
Many of those like tools, like in enterprise, do not even have like API or like clean interfaces, right?
It's the first time they're exposed to a third party system.
And so there is a lot of, you know, standing up like, you know, API gateways, like tools, connecting.
Then you have to essentially like define what good looks like, you know.
Again, like to put in your exercise for everyone, like, you know, defining like a golden set of evals is, you know, easier than sound.
Harder than it sounds.
And so we're spending like a bunch of time with them.
Evals are important.
Evals are super important.
Especially like audio evals.
I know audio evals are like extra hard to grade and get right.
But like the bulk of the use case here is actually audio.
And I have like, I don't know, five minute like call transfer.
How do you actually know that the right thing happened?
It's a pretty tough problem.
Yeah, it's pretty tough.
And then, you know, actually nailing down like the quality of the customer experience, like, you know, until it feels unnatural.
And here latency and interruptions, they're really like, you know, important part.
We shipped in GA, an API, real-time API.
I think it was last week.
A couple of weeks ago.
Yeah.
It was just last week, I think.
Which is like a beautiful work of engineering.
There was already a cracked team behind the scenes.
Which basically allows us to get like the most natural sounding, like, you know, voice experience without having like these weird interruptions on your leg where you can feel that essentially the thing is off.
So yeah, cobbling all that together, you know, and you get like, you know, a really good experience.
Yeah, that's a lot more than just models.
Yeah.
Yeah, I was going to say, one actually really great thing that I think we've gotten from the T-Mobile experience is actually working with them to improve our models themselves.
So for example, the real-time GA last week, we obviously released a new snapshot, the GA snapshot.
And a lot of the improvements that we actually got into the model came out of, you know, the learnings that we have from T-Mobile.
It brings in a lot of other changes from other customers, but because we were so deeply embedded into T-Mobile and we were able to understand what good looks like for them, we were able to bring that to some of our models.
That makes sense.
So this is a large customer with tens of millions of users, if not hundreds of millions.
And the before and after is on the support side, both tech support internally and then their customer support.
Yeah.
Makes sense.
Yeah.
Is there another one that you guys can share?
I like a lot Amgen.
Amgen, the healthcare business.
Amgen, yeah.
So we are working quite a bit with healthcare companies.
Amgen is one of the leading like healthcare companies.
We specialize into drugs for cancer or like, you know, inflammatory diseases.
They're based out of LA.
And we've been working essentially with Amgen to essentially speed up like the drug development and the conversation process.
So, you know, the sort of the the north star is like pretty bold.
And it's really interesting.
Like when you similarly, like, you know, we embedded like pretty deeply with Amgen to understand what other needs.
And it's really interesting.
Like when I look at those healthcare companies, I feel like there are two big buckets of needs.
One is like pure R&D.
It's like, you know, you're seeing like a massive amount of data and like you have super smart scientists who are trying to, you know, combine, test out things, you know.
So that's one bucket.
A second bucket is like, you know, much more like, you know, common across other industries.
It's like pure, like, you know, admin, document authoring, document scripting work, which is, you know, by the time like your R&D team has essentially locked the recipe of a medication, getting that medication to market is a ton of work.
Like you have to submit to like, value regulatory bodies, get a ton of reviews.
And, you know, when we looked at essentially those problems, what we knew what models were capable of, we saw like, you know, a ton of benefits, a ton of opportunities to automate and, you know, augment essentially the work of those teams.
And so, yeah, Amgen has been like a top customer of DPT-5, for instance.
Wow.
I mean, this could be hundreds of millions of lives if a new drug is developed faster.
Yeah, exactly.
Huge impact.
So that's, you know, that's I think one good example of like, a kind of impact on which you need to enable enterprises to do it.
Right.
You know, and so I think we're going to do more and more of those.
And yeah, frankly, like, you know, on a personal level, like, it's a delight, you know.
If I can play like, you know, a tiny role essentially, like doubling like, you know, the kind of medication that people, you know, get in the real world, that feels like, you know, a pretty good achievement.
Huge.
Huge, huge.
I know you had one as well.
So one of my favorite deployments that we've done more recently, actually, is with the Los Alamos National Labs.
So this is the like, government, national research lab that the US government is running in Los Alamos, New Mexico.
It's also where, you know, the Manhattan Project happened back in the 40s and 50s, back when it was a secret project.
So, you know, after that, they ended up formalizing it as a city and a program.
And then now it's a pretty sizable national laboratory.
This one is very interesting because one, just the depth of impact here is like unimaginable.
For me, it's like, on the scale of Amgen and some of these other larger companies.
But, you know, obviously, they're doing a lot of actual new research there.
So a lot of new science.
They're doing a lot of stuff with our defense department and defense use cases as well.
So very intense, you know, very intense stuff.
But the other thing that's actually very interesting about this one was that it's also a story of a very like, bespoke and like, new type of deployment that we've done.
So because they are so, they're government lab, they're so, you know, restrictive and high security and high clearance with a lot of their things, we couldn't just do a normal deployment with them.
They couldn't, you know, you can't have people doing national security research just hitting our APIs.
And so we actually did a custom on-prem deployment with them onto one of their supercomputers called Bonado.
And so this actually involves a bunch of, you know, very bespoke work with some FDEs, also with a lot of our developer team, to actually bring one of our reasoning models, O3, into their laboratory, into an air-gapped, you know, supercomputer Bonado, and actually deploy it and get it installed to work on their hardware, on their networking stack, and actually run it in this particular environment.
And so it was actually very interesting because we literally had to bring the weights of the model physically into their supercomputer.
In an environment, by the way, where you're not allowed to have, you know, it's very locked down for a good reason.
You're not allowed to have, like, cell phones or, like, any electronics with you as well.
And so I think that was a very unique challenge.
And then the other interesting thing about this deployment is just how it's being used, right?
So the interesting thing is because it's so locked down and on-prem, we actually do not have much visibility into exactly what they're doing with it.
But we do have, you know, they give us feedback.
Yeah, yeah.
They actually do have some telemetry, but it's, you know, within their own systems.
But we do know that it's, you know, being used for a bunch of different things is being used for aiding them in terms of speeding up their experiments.
They have a lot of data analysis use cases, a lot of notebooks that they're running with reams of data that they're trying to process.
They're actually using it as a thought partner, which is something that's pretty interesting to me.
O3 is, like, pretty smart as a model.
And a lot of these people are tackling really tough, you know, novel research problems.
And a lot of times they're kind of using O3 and going back and forth with it on their experiment design, on, like, what they actually should be using it for.
Which is, you know, something that we couldn't really say about our older models.
And so, yeah, it's just being used by, for a lot of different use cases for the National Lab.
And the other cool thing is it's actually being shared between Los Alamos and some of the other labs, Lawrence Livermore, Sandia as well, because it's the supercomputer setup where they can all kind of connect with it remotely.
Fascinating.
I mean, we've just gone through three pretty large scale enterprise deployments, right?
Which might touch tens, if not hundreds of millions of people.
But there was this, on the other side of this is the MIT report that came out a couple of weeks ago.
95% of AI deployments don't work.
A bunch of, you know, scary headlines that even shook the markets for a couple of days.
Like, you know, put this in perspective, like, for every deployment that works, there's presumably a bunch that don't work.
So maybe we can, you know, maybe talk about that.
Like, what does it take to build a successful enterprise deployment, a successful customer deployment, and the counterfactual, based on your experience serving all these large enterprises?
I think at that point, I may have worked with like a couple of hundreds.
I think enterprises.
So, okay, I'm going to pattern match.
What I've seen being like clear leading indicator of success.
Number one is like the interesting combination of like top-down, like buy-in and enabling like, you know, very clear group of like a tiger team, essentially, like, you know, at the enterprise, which is sometimes a mix of like open AI, like, you know, enterprise employee.
So, you know, typically, like, you know, you take like T-Mobile, like the top leadership was like, it's a priority.
Right, right.
But then letting the team like, you know, organize and be like, okay, if you want to start small, start small, you know, and then you can scale it up, essentially.
So that would be part number one.
So top-down buy-in and a bottom called a tiger team.
Tiger team, you know, people like, you know, a mix of like technical skills and like people who just have like the organizational knowledge, like institutional knowledge, you know.
It's really funny, like in the enterprise, like customer support, to give you an example, like what we found is that the vast majority of the knowledge is in people's heads, right?
Which is probably like a thing that, you know, with FDS, like in general, but like, you know, you take a customer support, you would think that, you know, everything is like perfectly documented, like in a GIR, et cetera.
The reality is like the standard like operating procedures, like the SOPs, are larger than people said.
And so unless you have that tiger team, like mix of like technical and like, you know, subject matter expert, really hard like to get something at the ground.
That would be one.
Two would be evals first.
Like, whatever we define as good evals, like that gives like a very clear common goal for people to hit.
Whenever like, you know, the customer like fails to come up with good evals, it's a moving target.
You never know essentially, you know, if you made it or not.
And, you know, evals are much harder than what it looks to get done.
And evals also oftentimes need to come up bottom up, right?
Because all of these things are kind of in people's heads, in the actual operator's heads.
Like, it's actually very hard to have a top-down mandate of like, you got like, this is how the evals should look.
A lot of it needs the bottoms up adoption.
Right.
Yeah.
Yeah.
Yeah.
And so we've been building quite a bit of tooling on evals.
We have like an evals product and, you know, we're working on more to essentially solve like, you know, that problem or, you know, make it as easy as we can.
The last thing is, you know, you want a helpline, essentially.
You have your evals, the goal is to get to 99%.
You start at like, you know, 46.
You know, how do you get there?
And here, frankly, I think oftentimes, like, you know, a mix of like, like, I will say like almost wisdom from people who've done it before.
Like, you know, a lot of that is like, you know, like art, sometimes more than science.
Or like, you know, knowing like the quirks of the model, the behavior, sometimes we even need to fine tune ourselves, the models, you know, when there are some clear limitation.
And, you know, being patient, getting your way, you know, up there, and then, you know, ship.
Can we go under the hood a little bit?
You know, one of the things that we think about a lot is autonomy more broadly, right?
What is the makeup of autonomy on one side, you know, in San Francisco, you could take a car from one part of SF to the other fully autonomously.
No humans involved.
No, you press about it.
Right.
They've done billions of rides.
I think it was like what, three and a half billion rides.
On the test, this is on the Tesla FSD.
I think we almost done like million, tens of millions of rides.
That's a lot of autonomy.
Yeah.
In the physical world, as opposed to the digital world, I can't book a ticket online right now.
There's all sorts of problems that happen if I have my operator try to book a ticket.
And it's very counterintuitive because the bar for physical safety is so much higher.
The bar for physical safety is higher than the human's capability because lives are at stake.
Yeah.
The bar for digital safety, not that high because all you're going to lose is money.
Nobody's life is at stake, but yet physical autonomy is ahead of digital autonomy in 2025, what seems counterintuitive.
Like, why is that the case at, you know, at a technical level?
Why is it that what should sound easier is actually a lot harder?
Yeah.
So I think there are kind of two things at play here.
And I really like the analogy with self-driving cars because they've actually been like one of the best applications of AI, I think, that I've used recently.
But I think there are two things in play.
One of them is honestly just the timelines.
Like we've been working on self-driving cars for so long.
I remember when I, you know, back in like 2014, it was kind of like the advent of this and everyone was like, oh, it's happening in like five years.
It turns out it took like, I don't know, 10, 15 years or so for this time.
So there's been a long time for the technology to really mature.
And I think there's probably like dark ages, you know, back in like 2015 or 2018 or something where it felt like it wasn't going to happen.
It's true of the solution.
Yes.
Yes.
Yeah.
And then now we're, you know, finally seeing it get deployed, which is really exciting.
But it has been like, I don't know, 10 years, maybe even 20 years from the very beginning of the research.
Whereas I think AI agents are like really in day one here.
Like, chat GPT only came out in 2022.
So like around three years, like less than three years ago.
But I actually think that what we think about with AI agents and all that really, I think started with the reasoning paradigm that when we released the, the own preview model back in late last year, I think.
And so I actually think this whole reasoning paradigm with AI agents and the robustness that those bring has only really unfolded for like a year, less than a year, really.
And so I know you had a chart in your blog post, which I really like, which, you know, the slope is very meaningfully different now.
Like self-driving started very, very early.
Slope seems to be a little bit slower, but now it's reaching the promised land.
But man, like we started super recently with AI agents and the slope, I think, is incredibly steep and we'll probably see a crossover at some point.
But we really have only had like a year really to explore these things.
Do you think we haven't crossed over already when you look at like the coding work in particular?
Yeah, it's a good point.
It's like, you know, your chart actually shows AI agents as below self-driving, but like, you know, it was like, what is the y-axis?
Like by some measures, like I would not be surprised actually if, you know, AI products or AI agents product is making more revenue than Waymo at this point.
Like Waymo is making a lot, but like, just look at all the startups coming up, look at, you know, chat GPT and how many subscriptions are happening there and all of that.
And so maybe we have actually crossed and, you know, a couple years from now it's going to look very, very different.
Yeah, the y-axis is tangible felt autonomy.
Perfectly objective.
How do I feel about it?
Yeah, actually vibes more than revenue, but revenue is a good one.
We should probably redo that with revenue.
There's a second thing I wanted to mention on this as well, which is the scaffolding and the environment in which these things operate in.
So I actually remember in the early days of self-driving, a lot of the like researchers around self-driving were saying that the roads themselves will have to like change to accommodate self-driving right.
There might be like sensors everywhere so that the self-driving cars can interact with it, which I think is like, you know, retrospect overkill.
But I actually do think self-driving cars have a good amount of scaffolding in the world for them to operate in.
It's like not completely like unlimited.
You have roads, roads exist, they're pretty standardized.
You have stoplights.
People generally operate in like pretty normal ways.
And there are all these traffic laws that you can learn.
Whereas AI agents are just kind of dropped in the middle of nowhere and they kind of have to feel around for them.
And I actually think, you know, going off of what Olivier just said too, my hunch is some of the enterprise deployments that don't actually work out likely don't have the scaffolding or infrastructure for these agents to interact with as well.
A lot of the like really successful deployments that we've made, a lot of what our FDEs end up doing with some of these customers is to create almost like a platform or some type of scaffolding connectors organizing the data so that the models have something that they can interact with in a more standardized way.
And so my sense of self-driving cars actually have had this in some degree with roads over the course of their deployment.
But I actually think it's still very early in the AI agents space.
And I would not be surprised if a lot of these, a lot of enterprises, a lot of companies just don't really have the scaffolding ready.
So if you drop an AI agent in there, it kind of doesn't really know what to do and its impact will be limited.
And so I think once the scaffolding gets built out across some of these companies, I think the deployment also speed up.
But again, to our point earlier, I think there's no slowdown.
There's no, you know, things are still moving very fast.
That's great.
Well, you know, I've thought about autonomy as a three part structure.
You've got perception, you've got the reasoning, the brain, and then you've got the scaffolding, the last mile of making things work.
Maybe we can dive into the second part, which is the reasoning, which is the juice that you guys are building with GPT-5 most recently.
Huge endeavor.
Congrats.
The first time you guys have launched a full system, not a model or a set of models, but a full system.
Talk about that.
I mean, the full arc of that development, what was your focus?
I mean, honestly, the benchmarks all seem so saturated.
Like clearly it was more than just benchmarks that you were focused on.
And so what was the North Star?
Like tell us about GPT-5 soup to nuts.
It's been the work of love of many people for a long time.
And to your point, I think GPT-5 is amazingly intelligent.
You look at the benchmark, like the sweet bench and the likes.
It is going pretty high.
But I think to me equally important and impactful was I would say the craft, like the style, the tone, the behavior of the model.
So capabilities intelligence and behavior of the model.
On the behavior of the model, I think it's the first model, like large model release, for which we have worked so closely with a bunch of customers for like month and month essentially to better understand what are the concrete blocks, like what are the concrete blockers of the model.
And often it's not about like having a model which is way more intelligent, a model which is faster, a model that better follows instruction, a model that is more likely to say no when he doesn't know about something.
And so that super close customer feedback loop on GPT-5 was pretty impressive to see.
And I think all the love that GPT-5 has been getting in the past couple of weeks, I think people are starting to show that essentially, the builders.
And once you see it, like it's really hard essentially to come back to a model which is like extremely intelligent, but you know an exclusive academic essentially way.
Are there trade-offs that you made as you were going through it?
Like maybe what are the hardest trade-offs you made as you were building GPT-5?
I actually think a very clear trade-off, which I honestly think we are still iterating on, is the trade-off between the reasoning tokens and how long it thinks versus performance.
And honestly this is something that I think we've been working on with our customers since the launch of the reasoning models, which is these models are so, so smart.
Especially if you give it all this like thinking time.
I think the feedback I've been seeing around GPT-5 Pro has been pretty crazy too.
It's just like these unsolved problems.
Andre had a great tweet last night.
Yeah, I saw that Sam retweeted it.
But like these unsolved problems that none of the other models could handle.
You throw to GPT-5 Pro and it just like one-shots it.
It's pretty crazy.
But the trade-off here is you're waiting for 10 minutes.
It's quite a long time.
And so these things just get like so smart with more inference time.
But on the like product builder on the API side for some of these like business use cases, I think it's pretty tough to like you know manage that trade-off.
And for us it's been difficult to figure out where we want to fall on that spectrum.
So we've had to make some trade-offs on like how much of the model think versus like how intelligent should it get.
Because as a product builder there's a latency, there's a real latency trade-off that you have to deal with where you know your user might not be happy waiting 10 minutes for like the best answer in the world.
It might be more okay with the substandard answer and like no wait at all.
Yeah, I mean even between GPT-5 and GPT-5 thinking, I have to toggle it now because sometimes I'm so impatient I just want it ASAP.
Yeah, I think there's an ability to skip right?
And yeah, I'm impatient I just want that more simple answer.
That's right, that's right.
Well four weeks in GPT-5, how's the feedback?
Yeah, I think feedback has been very positive especially on the platform side which has been really great to see.
I think a lot of the things that Olivier mentioned have been you know come up in feedback from customers.
The model is extremely good at coding, extremely good at kind of like reasoning through different tasks.
But especially for like coding use cases and especially at the you know at the when it thinks for a while it'll usually solve problems that no other models can solve.
So I think that's been a big positive point of feedback.
The kind of robustness and the reduction in hallucinations has been a really big positive feedback.
Yeah, I think there's an eval that showed that hallucinations basically went to zero for a lot of this.
It's not perfect, there's a lot of work to be done but that's a big one.
I think because of the reasoning in there too it just makes the model more likely to say no, less likely to hallucinate answer.
So that's been something that people have really liked as well.
Other bit of feedback has been around instruction following.
So it's really good at instruction following.
This almost bleeds into like the constructive feedback that we're working on where for that it's so good at construction following that instruction following that people need to tweak their prompts or it's almost like too literal.
That's why it's interesting to head off actually because you know when you ask people developers like what do you want like you want the model for instructions of course.
But once you have a model which who is like that is like extremely literal essentially, then essentially forces you to express extremely clearly what you want.
Otherwise the model may go sideways and so that was an interesting feedback.
It's almost like the monkey paw where it's like developers and platform customers ask for better instruction following.
So yes we'll give you really good instruction following but it's like you know follows it almost to a tee and so it's obviously something that the team is actually working through.
I think a good example of this by the way is some customers would have these prompts.
I remember when we were testing GPT-5 one of the negative feedback that we got was the model was too concise.
We were like what's going on why is the model so concise.
And then we realized it was because they were using their old prompts from other models and with the other models they have to like you have to like really beg the model to be concise.
So there are like 10 lines of like be concise, really be concise, also keep your answer short.
And it turns out when you give that to GPT-5 it's like oh my gosh this person really wants it to be concise and so the response would be like one sentence which is too terse.
And so just by removing the extra prompts around being concise the model behaved in a much better way and much closer to what they actually end up.
Yeah turns out writing the right prompt is still important.
Yes yes yeah prompt engineering is still very very important.
Yeah on constructive feedback for GPT-5 there's actually been a good amount as well which we're all we're all working through.
One of them that I think is I'm really excited for the next you know snapshot to come out to fix some of this is code quality and like small like code like paradigms or like idioms that they might use.
I think there are like feedback around the types of code and the patterns in which it was using which I think we're working through as well.
And then the other bit of feedback which I think we've already made good progress on internally is around the trade-off of the reasoning tokens and thinking and latency around intelligence.
I think especially for the simpler problems you don't usually need a lot of thinking the thinking should ideally be a little bit more dynamic.
And of course we're always trying to squeeze as much reasoning and performance into as little reasoning tokens as possible.
So I'd imagine that that could kind of going down as well.
Yeah well huge congrats.
I mean it's been I know it's a work in motion for a bunch of our companies they've had incredible outcomes with GPT-5 one of them's expo cyber security business.
Just like a huge huge upgrade from whatever they were using prior to that.
I think they're gonna need a new eval soon.
That's right.
They're gonna need a new eval.
It's all about evals.
On the multimodality side of it obviously you guys announced the real-time APA last week.
I saw T-Mobile was one of the featured customers on there.
Talk about that like how obviously the text models are core leading the pack.
But then we got audio and we got video.
Talk about the progress on the multimodal models.
When should we expect to have like the next big unlock and what would that look like?
It's a good question.
The teams have been making amazing progress on multimodality.
On voice, image, video, frankly the last generation models have been unlocking like quite a few cold use cases.
One of the feedback that we received is you know because like text was so much leading the pack on the intelligence.
Like people felt like in patreon voice that the model was somewhat a little less intelligent.
And you know until you actually see it like it does feel weird like you know to have like you know to have a better answer like on text versus voice.
And so that's pretty much a focus that we have at the moment.
I think we like filled like part of that gap but not the full gap for sure.
So I think you know catching up I would say you know with the text like you know would be one.
A second one you know which is absolutely fascinating is the model is like excellent at the moment on like you know easy like casual conversation like talk to your coach or therapist.
And we basically had to teach the model like to speak essentially better like in actual like you know work economically valuable setups.
Give an example like the model has to be able to understand what a necessity is and you know what it's meant to spell SSN.
And if one digit is actually like you know fuzzy it should actually has to repeat versus you know guess.
You know there are lots of like you know intuitions like that that someone you know of course has you know of our voice that we are currently teaching the model.
And that's like an ongoing work actually with our customers.
You know until we actually confront the model to like actual like customer support calls actual set score.
It's really hard like to get a feel you know for those gaps.
So that's a top like you know priority as well.
This is a completely off script but an interesting question that comes up in voice models particularly the real time API is you know previously people were doing they were taking a speech input convert that to text.
Yeah.
Then have some layer of intelligence.
Yeah.
Then you would have a text to speech model that would sort of play it back.
Yeah.
And this would be it would be a stitch of these three parts.
But the real time API you guys have integrated all of that.
Yes.
And you know how does it happen because a lot of the logic is written in text.
A lot of the Boolean logic or any call it any function calling is written in text.
How does it work with the real time API.
Is that.
That's an excellent question.
So the reason why we should the real time API is that we saw that for the stitch model the stitch model.
Yeah.
Oh yeah.
Like a stitch together.
Speech to text thinking.
Yeah.
Like we saw essentially a couple of issues.
One slowness like you know hops essentially to loss of signal like you know a cross each model like the speech text model is less intelligent.
Yeah.
You'd lose emotion.
You'd be like accent.
Exactly.
Right.
Buses.
Yeah.
And you know when you when you are doing like actual voice like phone calls essentially like the signals are like so important like you know for the century to the stem.
Yeah.
One of the challenges that we have is what you mentioned which is you know it means like a slightly different architecture essentially for text versus voice.
And so that's something that we are actively working on.
But I think it was the right call to start essentially with let's make like the voice experience like natural sounding to a point why essentially you're feeling comfortable like putting in production and then working backward like to unify like the sort of the orchestration logic essentially across modalities.
And then to be clear like a lot of customers still stitch these together.
It's like kind of what worked in the last generation.
But what we're interested in seeing is more and more customers moving towards the real time approach because of how natural it sounds how much lower latency it is especially as we up level the intelligence of the model.
But also even like taking a step back I will say it's like pretty mind blowing to me that it works like the fact that like I think it's mind blowing that these elements work at all where you just train it on a bunch of text and it's just you know autoregressively coming up with the next token and it sounds super intelligent.
That's like mind blowing in and of itself.
But I think it's actually even more mind blowing that this speech to speech setup actually works correctly because you're literally taking the audio bits from it from someone speaking streaming or putting it into the the model and then it's generating audio bits back.
And so to me it's actually crazy that works at all.
Let alone the fact that it can understand accents and tone and pauses and things like that and then also be intelligent enough to handle a support call or something.
I mean if you've gone from text in text out to voice in voice out that's pretty crazy.
Yeah.
We have a bunch of companies in our portfolio that are using these models you know Parloa on the customer support side live kit on the infra side and you know there's a bunch of use cases we were starting to see that that that a speech to speech model could address.
Obviously a lot of the harder ones still still running on what you're calling the stitch model.
Yeah.
But I hope the day is not far when it's all on real time API.
It's going to happen at some point.
Right.
Right.
Right.
Right.
And actually maybe that's a good segue into talking about model customization because I suspect that you have such a wide variety of enterprise customers.
I think you mentioned what hundreds of customers or maybe more each of them has a different use case a different problem set a different color envelope of parameters that they're working in maybe latency maybe maybe power maybe others.
How do you how do you handle that.
Talk about what open AI offers enterprises who need a customized version of a great model to make it great for them.
Yeah.
So model customization is actually been something that we've invested very deeply in on the API platform since the very beginning.
So even you know preach LGBT days we actually had a supervised fine tuning API available and people were actually using it to great effect.
The most exciting thing actually I'd say around model customization it obviously resonates quite well with customers because they want to be able to bring in your own custom data and create your own custom version of you know O3 or O4 mini or something or GBD5 even suited to their own needs.
It's very attractive but the most recent development of that thing is very exciting has been the introduction of reinforcement fine tuning.
There's something we announced late last year I think in the 12 days of Christmas we've GA'd it since and we continue to iterate on it.
What does it break it down for us?
Yeah so it's called it's actually funny I think we made up the term reinforcement fine tuning it's like not a real thing until we announced it.
It's stuck now.
I see it all the time.
I remember we were discussing it and I was like I don't know about RFT guys.
You're not kidding.
Reinforce and fine tuning.
So it really it's introducing reinforcement learning into the fine tuning process.
So the original fine tuning API does something called supervised fine tuning and call it SFT.
It is not using reinforcement learning it is you know it's using supervised learning.
And so what that usually means is you need a bunch of data a bunch of prompt completion pairs you need to really supervise and tell exactly the model how it should be acting and then when you train it on our fine tuning API it moves it closer in that direction.
Reinforcement fine tuning introduces like RL or reinforcement learning to this loop.
Way more complex way more finicky but in order of magnitude more powerful.
And so that's actually what's really resonated with a lot of our customers.
It allows you to if you use RFT the discussion is less of like creating a custom model that's specific to your own use case.
It is you can actually use your own data and actually crank the RL yeah turn the crank on RL to actually create a like best in class model for your own particular use case.
And so that's kind of the main difference here with RFT the data set looks a little bit different instead of you know prompt completion pairs you really need a set of tasks that are very gradable you need a grader that is very objective that you can use here as well.
And so that's actually been something that we've invested a lot in over the last year and we've actually seen a couple a good number of customers get really good results on this.
We've talked about a couple of them across different verticals so ROGO which is a startup in the financial services space.
They have a very sophisticated AI team.
I think they hire some folks in DeepMind to run their AI program and they've been using RFT to get best in class results on you know parsing through financial documents and answering questions around it and doing tasks around that as well.
There's another startup called Accordance that's doing this in the tax space.
I think they've been targeting an eval called TaxBench which looks at you know CPA style tasks as well and because they because you know they're able to turn it into a very gradable setup they're actually able to turn the RFT crate and also get I think like soda results on a tax bench just using our RFT product as well.
And so it has kind of shifted the the discussion away from just customizing something for your own use case to like really leveraging your own data to create a best in class maybe best in the world model for something that you care about for your business.
Yeah I feel like the base models are getting so good at instruction following that for you know behavior like steering like you know you don't need to find you at that point like you know you can describe what you want and the model is pretty good at it.
But pushing the frontier on like actual capabilities my hunch is that RFT will pretty much become the norm.
Like you know if you are actually pushing in your field like you know intelligence like you know to a pretty high point like at some point like you know you need to hire a hire essentially with custom environments.
Yeah fascinating.
And even going back to the point earlier around like top down versus bottoms up for some of these enterprises a lot of the data that you end up needing for RFT require like very intricate knowledge about the exact task that you're doing and understanding how to grade it.
And so a lot of that actually comes from bottoms up like a lot I know a lot of these stars will work with experts in their field to try and get the right tasks and get the right feedback to craft some of these data sets.
Without further ado we're going to jump to my favorite section which is a rapid fire question.
We had a lot of great friends of ours send in some questions for you guys.
We'll start with the Ultimators favorite game which is a long short game.
Pick a business an idea a startup that you're long and the same short that you would bet against that there's more hype than there's reality.
Whoever's ready to go first long short.
My long is actually not in the AI space so this is going to be slightly different.
Here we go.
My short is though in the AI space so I'm actually extremely long esports and so what I mean by esports is the entire like professional gaming industry that's emerging around video games.
Very near and dear to my heart I play a lot of video games and so I watch a lot of this so obviously I'm pretty in the weeds on this.
But I actually think there's incredible untapped potential in esports and incredible growth to be had in this area.
So concretely what I mean are like you know a really big one is League of Legends all of the games that Riot Games puts out.
They actually have their own professional leagues.
They actually have professional tournaments believe it or not.
They rent out stadiums actually now.
But I just think it's like if you look at kind of what the youth and like what younger kids are looking and where their time is going it's predominantly going towards these things.
They spend a lot of time on video games.
They watch more esports and next year they get balls.
Yeah yeah yeah yeah a growing number of these too.
I've actually been to some of these events and it's very interesting.
It's very committed.
Yeah yeah I'm extremely long stuff.
And so they're booking out stadiums for people to go watch electronic sports.
Yeah yeah yeah it's I literally went to Oracle Arena the old warrior stadium to watch one of these I think before Covid.
And then the so it's just wow that's five years ago.
It was a while ago.
So I actually I've been following this for a while and I think it had a really big moment in Covid like everyone is playing video games.
And I think it's kind of like come back down.
So I think it's like undervalued you know it's like I think no one's really appreciating it now.
But it has all the elements to like really really take off.
And so the youth are doing it.
The other thing I'd say is it is huge in Asia like absolutely massive in Asia.
It is absolutely big in Korea and China as well.
Like you know we rented out Oracle Arena I think or like the event I went to was an Oracle Arena.
My senses in Asia they rent out like the entire stadiums like the soccer stadiums and the players are ready like celebrities.
So anyways you know as like you know the I know like Korean culture is really making its way into the US as well.
I think that's another tailwind for this whole thing.
But anyways esports I think is something you should keep an eye out on because there's a lot of room for growth.
Very unexpected.
Yeah.
Good to hear.
Short.
My short.
My short's a little spicy which is I'm short on the entire category of like tooling around AI products.
And so this encapsulates a lot of different things kind of cheating because some of these you know I think are starting to play out already.
But I think like two years ago it was maybe like Evals products or like frameworks or vector stores.
I'm pretty short those.
I think nowadays there's a lot of additional excitement around other tooling around AI models.
So our all environments I think are really big right now as well.
Unfortunately I'm very short on those.
Not really I don't really see a lot of potential there.
See a lot of potential and reinforcement learning and applying it.
But I think the startup space around our environment I think is really tough.
Main main thing is is one it's just a very competitive space.
There's just a lot of people kind of operating and then two if the last two years have shown us anything the space is evolving so quickly and it's so difficult to try and like adapt and make and understand what the exact stack is that will really you know carry through to the next generation of models.
I think that just makes it very difficult when you're in the tooling space because you know today's really hot framework or really hot tool might just not get used in the next generation of models.
So I've been noticing like the same pattern which is the the teams that build like breakout like startups in AI are extremely pragmatic.
They are not super like you know intellectual but like the perfect world etc.
And it's funny because I feel like you know our generation has basically started in tech in a very like stable like you know moment where you know technology had been building up like you know for years and years with like SaaS like cloud etc.
And so we were in a way like raised like you know in that very stable moment where it makes sense at that point to you know design like very like you know good like abstractions and toolings because you know you have a sense why it's going but it's so different today like don't wait for it to know what's going to happen next year or two so it's almost impossible like to define like the perfect tooling platform.
Right right right right well that's uh there's a lot of that going around right now.
Yes.
Spicy a lot of homework there.
Olivier over to you sir.
Long short.
I've been thinking a lot about education for the best month in the context of kids.
I'm pretty shocked on any education which basically emphasizes human memorization at that point and I say that having mostly been through the dedication myself but you know like I learned so much on like you know history facts like you know legal things that are you know some of it like does shape your way of thinking a lot of it frankly is just like you know knowledge tokens essentially and those knowledge tokens you know it turns out like you know algorithms are pretty good at it.
So I'm I'm quite short on that.
That's right you all need memory when strategy is bionic you can just think about it straight into your head.
Exactly exactly what am I long at um frankly I think healthcare is probably the industry that will benefit the most from AI in the next like year or two.
I would say more.
I think like all the ingredients are here for a perfect storm um a huge amount of like structured and structured data you know it's basically the heart of you know like the pharma companies uh the models are excellent at digesting processing that kind of data uh a huge amount of like admin like you know heavy like documents heavy like you know culture but at the same time like companies which are very technical very R&D friendly like you know companies like you know who's like sort of technology in a way is at the heart of what they do and so yeah I'm pretty bullish on that.
This is like life sciences so you mean life sciences research organizations that are producing drugs gotcha exactly yeah yeah it's almost like you know over the last 20 30 years these these like pharma or like biotech companies have have basically if you look at the work that they're doing like only a small amount of it is is actual research and so much of it ends up being admin and like you know documents and things like that and that area is just so ripe for you know something to happen with AI and I think that's what we're seeing with Amgen and some of these other customers exactly and it's also like not what they want to do it's I think it's good that we have some regulations there obviously but like just means that they have like reams and reams of things that kind of go through and so you know like when you have a technology that's able to really helps like bring down the the cost of something like that I think it'll just you know tear right through it and I think once governments and like you know institutions are going to realize that like if you step back like it is probably one of the biggest bottleneck to like human progress right you step back in the past decade like you know how many like through like breakthrough drugs have there been like you know not that many like you know how life would be different like if you double that rate essentially so once you realize what is at stake yeah my hunch is that we're going to see quite a bit of momentum in that space wow all right lots of homework there as well yeah next one favorite underrated AI tool other than chat GPT maybe I love granola oh man you still mind you still mind answer like two votes for granola there is something like yeah hey what about chat GPT record I like to record as well but there are some features in granola which I think are really done well like the whole like you know integration with your gold calendar is excellent yeah um and just you know the quality of like the transcription and like the summary is pretty good do you just have it on because I know your calendar is back to back you just have granola horn so the funny thing is that I don't use granola internally uh I usually have one for my personal life mostly I see yeah I see on dates I'm joking uh I'll say yeah granola is actually gonna be mine so uh two votes for granola um I was gonna say the easy answer for me is codex uh as a software engineer it's just like so it's gotten so good recently codex kly especially with GPT 5 especially for me I tend to be less time sensitive about like uh you know the iteration loop with with coding and so leaning in as GPT 5 on codex I think has been interesting really what about codex has changed because you know codex has also been through a journey codex has been around for a bit I remember like it's been launched for like a even more than over a year ago it's like what's what's changed about codex yeah I'll actually say it's been around for a bit so I feel like it's been less than a year for codex I would say the time dilation is so crazy in this field it feels like it's been around for a year ago with GPT 4 oh like you know that demo like that feels like ages ago one hadn't even come out yet probably since it hadn't happened yet the voice demo it's a naming thing okay but anyway yeah there was a codex model that's what I'm thinking there was a codex model yeah we are yeah we are we are you're not too blamed for that confusion also I think the the github thing was called codex that's right yes yes that's right but I'm talking about our uh coding product within chat GPT which are the codex cloud offering and then also codex kly so actually maybe if I were to narrow my answer a little bit more codex kly which I've really really liked I like the local environment set up the thing that's actually made it really useful in the last I'd say like month or so is one I think the team has done a really good job of just like getting rid of all the paper cuts like the small product polish and like paper cut things it just it kind of feels like a joy to use now I feel more reactive and then the second thing honestly is is GPT 5 like I just think GPT 5 really allows the product to shine it's you know at the end of the day this is kind of a there's a product that really is dependent on the underlying model and when you have to you know like iterate and go back and forth with model like four or five times to get it right to get it to like you know do the change that you want versus having it think a little bit longer and it just like one shots and does exactly what you want to do you get this like weird like bionic feeling where you're like I feel so mind melded with the model right now and like perfectly understands what I'm doing and so getting like that kind of dopamine hit and like feedback loop constantly with codex has made it kind of like an indispensable thing that I really really like nice and the other thing I'd say codex is just really good for me is so I use it in my for like personal projects I also use it to like help me understand codebases like as a as a engineering manager now I'm not as in the weeds on on the actual code and so you're actually able to use codex to really understand what's happening with the code base have it like ask questions and have an answer about things and really catch up to speed on things as well so like even the non-coding use cases are really useful with codex clay fascinating sam had this tweet about codex usage ripping I think like like yesterday so I wonder what's going on there but you're you're not alone yeah I think I think I'm not alone just judging from the twitter feedback I think people are really realizing how great of a combination codex client gpt5 are yeah I know that team is undergoing a lot of scaling challenges but I mean it the system hasn't gone down for me so props to them but we are in a gpu crunch so we'll see how you know how long that goes awesome awesome all right the next one um will there be more software engineers in 10 years or less there's about 45 million professional software engineers that's what you mean like full-time like a cool job yeah yeah because it's a hard one because like I think without a doubt there's gonna be a lot more software engineering going on yes there's actually a really great post that was shared I think in our internal slack it was like a reddit post recently I actually think that highlights this um is a really touching story it was a reddit post about someone who has a brother who's nonverbal I actually don't know if you saw this it was just posted um it's a person on reddit posted they have a nonverbal brother who they have to take care of the brother like they tried all these types of things to help the brother interact with the world use computers but like vision tracking didn't work because I think his vision wasn't good um all the tools didn't work and then uh this brother um ended up using chat gbt I don't think he used codex but he used chat gbt and basically taught himself how to create a set of tools that were tailor made to his nonverbal brother basically a custom software application just for them and because of that he he now has like custom setup that was written by his brother uh and allows him to like browse the internet I think the video was like him watching the simpsons or something like that which is really touching but I think that's actually what we'll see a lot more of like this guy's not a professional software engineer his title is not software engineer but he like did a lot of software engineering probably pretty good um good enough definitely for his brother to use so the amount of code the amount of like building that'll happen I think is just going to go through an incredible transformation right I'm not sure what that means for software engineers like myself maybe there's you know equivalent or maybe there's of course more showing yeah more of me more of me specifically way more of you that's right but definitely a lot more software engineering a lot of yeah I buy that completely like I buy completely a physicist that there is a massive software shortage yeah like in the world like we've been sort of accepting it you know for the past 20 years but like the goal of software was never to be that super rigid super hard to build you know uh artifact it was to be like you know customized like malleable and so I expect that we'll see like way more a sort of a reconfiguration of like people's like in a job in skill set where way more people code like you know I expect the product managers are going to code like more and more for instance uh yeah you made you made your pm's code recently if you're yeah we did that that was really fun uh we started like essentially not doing like PRDs like product requirements documents you know classic pm thing you write like five pages like my product doesn't etc and you know pm's have been basically like coding prototypes and one is pretty fast with gpt vibes and like codecs yeah just a couple hours I think freaking fast yeah and second like it sort of conveys like so much more information than a document yeah like you get a feel essentially for the feature like is it to write or not so yeah I expect that sort of going to be heavier we're going to see more and more yeah instead of writing English you can actually now write the actual thing you want yeah yeah yeah yeah and yeah that's amazing advice for high school students who are just starting out their career my advice is I don't know maybe it's ever going like prioritize critical thinking above anything else if you go in the field which requires like extremely high critical thinking like you know skills I don't know math physics or you know maybe curiosofice in that bucket you will be fine regardless if you go in the field that sort of turns down that thing and again gets back to like memorization like you know pattern matching I think you will probably be less future proof yeah as a you know uh what's a good way to sharpen critical thinking um use trashy bt and have it test you that's true having like you know workplace shooter who essentially knows how to put the bar like 20% of what you can do all the time you know uh is actually probably a really good way to do it yeah nice anything from you sir um mine is I think it's just I think we're actually in such a interesting like unique time period where um the like younger so like maybe this is a more general advice for not just like high school students but just like the younger generation even like college students it's like uh I think the advice would be don't underestimate how much of an advantage you have relative to the rest of the world right now because of how ai native you might be or how interesting like you know in the in the ways of the tools you are my hunch is like high schoolers college students when they come into the workplace they're going to have actually a huge leg up on uh how to use ai tools how to actually transform the workplace and my push for like some of the younger I guess high school students is like one like just really immerse yourself in this thing and then two just like really take advantage of the fact that you're you're in a unique time where like no one else in the workforce really understands these tools as deeply probably as you do a good example this is actually we had our first intern class uh recently at open ai um a lot of software interns and some of them were just like the most incredible cursor power users of like ever seen they were so productive yeah I was I was shocked yeah yeah yeah I was like yeah I know I know we can get good interns but like I don't know they'd be like this good yeah and I think part of it's like they've grown up um using these tools for better or worse in college yeah um but I think the meta level point is they they're they're so like ai native and even like I don't know me and Olivier we're like kind of ai native we're good open ai but like we haven't like been steeped in this and kind of grown up in this and so uh the advice here would just be like yeah leverage that like you know don't be afraid to kind of like go in and spread this knowledge and take advantage of it in the workplace because it is a pretty big advantage for them yeah I can't remember who said this to us at Bountour but every intern class was just getting faster smarter like laptops like smarter every generation you sure didn't peak in 2013 you know when I was a netterner that's right that's right he's a wet spy that's summer 2013 yeah two guys like that's right that's right that's right that's right yeah well lots happen you know lots happen since you guys joined open ai right what three years and almost three years um in your open ai journey what has been the the rose moment your favorite moment the bud moment where you're like most excited about something but but still opportunity ahead and the thorn toughest moment of your of your of your uh three-year journey the thorn is easy for me what we call the blip which is you know the coup of the board like that was already tough moment yeah it's funny because you know after the fact it's actually reunited quite a bit the companion like there was a feeling open ai had a pretty strong culture before but you know there was a feeling of like camaraderie essentially that was even stronger yeah but you know sure like top on the day off it's very rare to see that anti-fragility most orgs after something like that break break apart but i feel like opening i got stronger opening i came back it's a good point i feel it may to open a stronger for real now essentially when they look after the fact when they look at you know other like you know news and departures or you know whatever like you know bad news essentially i feel the company has built like you know a thicker skin and you know an ability to like recover like yeah way quicker and she think part i think it's definitely right part of it too i think is also just the culture i also think this is why it was such a low point for a lot of people so many people just at opening i care so deeply about what we're doing which is why they work so hard um you just care a lot about the working almost feels like your life's work like it's a very audacious mission and thing that you're doing which is why i think the blip was like so tough on a lot of people but also is what i think helped bring people back together and we were able to hold together and get that that thick skin as well yeah yeah i have a separate uh worst moment uh which was uh the big outage that we had in uh december of last year if you remember yeah you remember i do it was like a multi-hour outage um really highlights to uh to us how uh essential of almost like a utility the api was so the background is i think we had like a like a three four hour outage some time in uh november december last year really brutal pure self zero no one could hit chat gbt no one could hit the apis it was it was really rough um that was just really tough just from a like you know customer trust perspective i remember we like talked to a lot of our customers to kind of like um post-mortem them on what happened and kind of our plan moving forward thankfully we haven't had anything close to that uh since then and i've been actually really happy with um all the investments we've made in reliability over the last uh six months um uh but uh in that moment i think it was really it was really tough yeah on the happy side like on the roses um i think i have two of them the first one would be gbt5 was really good like the sprint up to gbt5 i think really like you know showed like the best of openai like you know having like cutting edge like science research like you know extreme like customer focus like you know extreme like you know infrastructure and inference like you know talent uh and the fact that we were able like to ship like such a big model and scale it to like you know many many many many tokens you know permanent like almost like immediately i think speaks to it so that one i really with no with no outages and no outages yeah really good reliability yeah like i remember when we shipped like dt for turbo like a year ago a year and a half ago we were terrified by you know like the the in-scapes of traffic and i feel we've really gotten like much better at you know shipping like those you know massive updates um the second uh rose like you know happy moment for me would be the first death day was really fun yeah it felt like a coming of age like openai like you know we are embracing that you know we have like a huge community developers you know we are going to ship models new products and i remember basically seeing like you know all my favorite people openai or not you know like you know home essentially nerding out on you know what are you building like you know what's coming up next it felt really like you know a special moment in time no that was actually going to be mine as well so i'll just pay you back off of that which is the very first dev day 2023 november uh i actually i remember it so i mean obviously a lot of good things have happened since then there's just a very i don't know why for me it was a very memorable moment which was uh one it was actually quite a rush up to dev day we were we shipped a lot so our team was just really really sprinting so it was like this high you know uh high stress environment kind of going up to add to that um you know of course because we're openai we did a live demo on sam's uh in sam's keynote of all the stuff that we shipped i just remember being in the back of the audience sitting with like the team and like waiting for the demo to happen once it finished happening we all just like let out a huge sigh of relief like oh my god thank you um uh and so there's i then there's just like a lot of like you know build up to it for me the most memorable thing was i remember right after dev day all the demos worked well all the talks worked well we had the after party and then i was just in a way mode driving home at night with the music playing it was just like such a great end to the dev day that was that was what i remember that was my rose for the last uh love it yeah that's awesome i assume you guys are but please tell me if you're agi-filled yes or no and if so what was the moment that got you there what was your aha moment when did you feel the agi i think i made jp i think i made jp you're definitely agi-pilled i am okay uh i've had a couple of them the first one was the realization in 2023 that i would never need to code manually like ever ever again i'm not i'm not the best color frankly you know i chose my job like you know for a reason um but realizing that you know what i thought was a given that we humans would have like to write like basically machine language like forever is actually not a given and that you know the pay surprise is huge feeling the agi um the second like field the agi moment for me was uh maybe the progress on voice and multimodality like you know text like at some point you get used to it like okay you know the machine can write pretty good text yeah voice makes it real but once you start actually talking like you know to something that he understands your tone like you know understand my accent like in french um it felt like sort of um uh 20 moment like okay machines are going beyond like cold mechanical deterministic like you know like logic to something like much more like emotional and like you know tangible yeah that's a great one yeah mine mine are uh uh so i do think i'm i am agi-pilled i probably gradually became agi-pilled over the last couple years i think there are two and and for me um yeah i think i i i actually get more get more shocked from the text models i know the multimodal ones are really great as well for me i think they actually line up with with two like general breakthroughs so the first one was right when i joined the company in september 2022 we it was pre-chat gbt yeah two months ago but at the time gbt4 already existed internally and i think we were trying to figure out how to deploy i think nick torley's talked about this a lot um early it is chat gbt but it was the first time i talked to gbt4 and it was it was like going from nothing to gbt4 was just the most mind-blowing experience for me i think for the rest of the world maybe going from nothing to gbt3.5 in chat was maybe the big one and then going from 3.5 to 4 but for me and i think for a lot of maybe some other people who joined around that time going from nothing to or not nothing but like what was publicly available at the time going from that to gbt4 was just incredible like i just remember asking throwing so many things out i was like there's no way this thing is gonna be able to give an intelligible answer and it just like knocks it out of the park it was it was absolutely incredible gbt4 was insane i remember gbt4 came out when i was interviewing with openai and i was still looking at the phone should i just join so that thing i was like okay i mean i mean there is no way i can walk on f-e-l-c at the point yeah yeah yeah yeah so gbt yeah gbt4 was just crazy um and then uh the other one was uh like is the other breakthrough which is like the reasoning paradigm i actually think the the purest represent representation of that for for me was deep research and throwing like like asking it to like really look up things that i didn't think it would be able to know and seeing it like think through all of it like be really persistent with the search get really detailed with the write-up and all of that that was pretty pretty crazy i don't remember the exact query that i threw it but i just remember like i feel like the field age moments for me are like i'll throw something at the model that i was like there's no way this thing will get and then it just like knocks out of the park like that is kind of the field age moment i definitely had that with deep research with some of the things i was asking yeah well this has been great thank you so much folks you guys are building the future you guys are inspiring us every day and appreciate the conversation yeah thank you so much thank you thank you As a reminder to everybody, just our opinions, not investment advice.