Latent Space · 2026-07-28

PodcastYouTube

ChatGPT Work: Codex's Leap from Devs to Everyone

Hosts: Alessio (Latent Space), Vibu

Guests: Akshay Nathan

ChatGPT WorkCodexAgentic AIHarness engineeringArtifactsSub-agentsMemoryChronicleEnterprise vs productivityProduct development metrics

Why it matters

OpenAI's Akshay Nathan on merging Codex into ChatGPT Work, the shared agent harness, and 10M users in weeks.

Key claims

  • ChatGPT Work launched as a merge of Codex and ChatGPT after internal UXR showed non-developers adopting Codex with pride and treating it as a superpower.
  • The Codex harness and ChatGPT Work harness are the same underlying agent; the split is opinionated UX, sandboxing defaults, and how much git/diff state is exposed.
  • The team is called 'Productivity' rather than 'Enterprise' to reflect that personal and work tasks increasingly share the same primitives.
  • Adoption is described as faster than GPT-5.0, with 10M users reached; the next goal is bringing the experience to ChatGPT's full consumer base.

Radar summary

Summary

Akshay Nathan, who leads the Productivity team at OpenAI, walks through the launch of ChatGPT Work — a unification of the Codex harness and ChatGPT aimed at all knowledge workers, not just developers. The trigger was an internal surprise: non-engineers at OpenAI were adopting Codex eagerly, feeling like they had a superpower. That observation pushed the team to dissolve the separate Codex app and merge it into ChatGPT, framing the product as "productivity" rather than narrowly "enterprise," since personal tasks (meal planning, tracking a lost package) increasingly live on the same primitives as work tasks.

The underlying harness is shared between Codex and ChatGPT Work, but the UX opinionates differently by surface — Codex mode exposes git diffs and stricter sandboxing, while Work mode abstracts the computer environment and leans into artifacts, plugins, and computer use. Nathan positions the Codex app as a durable brand for developers that is not being deprecated, while Work is the broader on-ramp. Memory, including the experimental Chronicle feature, is framed as a major investment area, with ChatGPT Work inheriting ChatGPT's existing memory system so context carries across products.

On numbers, Nathan confirms the launch has reached 10 million users faster than 5.0 and calls it a "culmination," while stressing the goal is to bring the magic to ChatGPT's full user base. He draws inspiration from OpenClaw for persistent computer environments, sees programmatic tool calling and sub-agents as expanding the ceiling of what MCP-connected agents can do, and argues the next sequencing step is taking the knowledge-work learnings and applying them to everyone in everyday life. He also pushes back on using AI to auto-generate performance reviews, framing the technology as agentic search for context, not a substitute for human judgment.

  • ChatGPT Work launched as a merge of Codex and ChatGPT after internal UXR showed non-developers adopting Codex with pride and treating it as a superpower.
  • The Codex harness and ChatGPT Work harness are the same underlying agent; the split is opinionated UX, sandboxing defaults, and how much git/diff state is exposed.
  • The team is called 'Productivity' rather than 'Enterprise' to reflect that personal and work tasks increasingly share the same primitives.
  • Adoption is described as faster than GPT-5.0, with 10M users reached; the next goal is bringing the experience to ChatGPT's full consumer base.
  • Memory is a major investment area, and ChatGPT Work inherits ChatGPT's memory system by default; experimental Chronicle adds deeper context from computer activity.
  • Sub-agents and programmatic tool calling are positioned as raising the ceiling on what MCPs can accomplish, with ultra mode now opt-in.
  • Nathan pushes back on fully AI-generated performance reviews, advocating for AI as agentic search for context rather than a replacement for human judgment.
  • He frames 'motion vs. progress' as the core trap for AI-era teams, with at-bats — the full idea-to-validation loop — as the team's preferred metric.

Source material

Full source text

Okay, we're here in the studio with Akshay from OpenAI.

Welcome.

Thank you.

And with our trustee co-host, Vibu.

So you recently launched ChatGPT Work.

You lead Core Product Engineering.

You know, it's been a long journey into all this.

I find it very interesting that you started with no-code or low-code with Walrus and Airtable.

And to some extent, ChatGPT Work is kind of like the super app of super apps of, well, here is the ultimate no-code.

You just write a prompt.

Yeah, yeah.

It's funny how things come full circle.

I mean, I think for a long time in my career, I mean, I started my career working consumer fintech, but then after that, there's this hypothesis that the things that we were able to do with code as engineers, if we could bring that to many more people in a more accessible way, then that would be truly magical.

We were working on a startup.

It's actually funny, before LLMs, before Vision LLMs on how to do automated testing with AI.

And it was just kind of jank back then, but doing what we can.

And then worked at Airtable for a while on the same thesis that if we can bring a database or the primitives behind a database to people, that would be really useful to them.

But once I think LLMs came onto the scene, it became clear that this was the missing piece, the missing technology required to bring the magic of code to everyone without them having to know what's going on underneath the hood.

And so I think this launch and a lot of the stuff that we've been up to is the manifestation of that.

How was stuff when you joined?

So you joined OpenAI in 2023.

Now we've got so much more stuff.

So ChatGPT, CodexApp, ChatGPT for work.

Have things changed?

Actually, I think the more interesting thing is how things haven't changed.

I guess one, I remember when I joined, it was like 500 people.

One thing I was worried about was I was looking for something more early stage and was it going to feel startup enough?

And I joined and I was like, this feels even more startup-y than I could ever imagine.

And that really hasn't changed even until now.

I mean, I think the level of bottoms-up ambition and the ability of anyone to do anything or have an idea and ship it is really cool.

But on the mission side, I think what was really compelling to me is this mission of bringing Frontier intelligence to everyone, building AGI and then bringing it to everyone.

And I think technology even back then that that vision is going to not be a linear progression.

We're probably going to try different products and have different things that succeed and don't.

But the vision has stayed the same and the mission has stayed the same.

And we're starting to see the pieces fall together.

And that's really cool.

You worked on enterprise.

A lot of people never touch ChatGPT for enterprise, God.

What is something that you learned from there that you're bringing into your work now?

I think how there's no one-size-fits-all solution in enterprise.

I remember in the early days of ChatGPT Enterprise, we would talk to customers and everyone, that was when, I think it was a year after ChatGPT was released, and everyone was so excited to bring AI into their enterprise.

And there was all these teams that were being stood up as the AI deployment team, these enormous budgets.

And if you asked anyone, what were they excited about?

What were they excited about?

At first, you'd get the baseline answers of, we have all this context and data and all this stuff.

But then if you asked them, what was a discrete use case that they want AI to enable in their workplace?

You get such a different variance, explosion of different types of answers.

And it's interesting, using these models and these products, you have this box and you can say anything to it, which is the magic.

But on the flip side, it also means that you don't know what to do with it.

And in enterprise, I think a big part of that is actually meeting the users where they are, what use case are they trying to solve, and then actually teaching them how they can use AI to gain leverage there.

Do you meaningfully differentiate that from forward deployed engineering?

I think there's the go-to-market side of it, and then there's the product side of it, I think.

You need to see more than a product side.

And I think however good we get at FDE Motion, I think at the end of the day, if we have a user who's looking at their computer or looking at their phone, it's our job in the product to be enabling them and showing them where to go.

So we're really excited about that.

Do you think there's been changes over the past three years of adoption?

So there have been step function changes, you have reasoning models and whatnot.

Is there still the same problems of enterprise has black box, don't know what to do with it, or have things changed?

I mean, we're seeing now that there's this huge uptake, right?

Everyone's extremely excited about it.

It feels like millions, hundreds of millions of people are using ChatGPT.

They understand how generally to work with AI.

But then every time a new capability gets unlocked, so now we're seeing with agents, there is probably a contingent of early adopters still who truly get it, who are like, you can do anything.

You just have to make sure the right context is there, it's connected to the right tools, and then that you're supervising it, but anything is possible.

But then there's this 10x or 100x bigger market, or they don't yet get that, or they don't yet see that.

And so I think that's the next stage here.

So I guess to answer your question, I think the adoption is there and growing fast, but I think the opportunity is far, far bigger than that.

That's where we want to play, especially with ChatGPT work.

Yeah.

Well, let's skip ahead to ChatGPT work.

Only a month ago or so announced, what was the sort of decision process that led into it?

There was this overall merging of the super app.

Is that what we're officially calling it?

You deprecated the browser as well.

Just, I guess, summarize your last couple of months of working on this thing.

Yeah.

It feels like forever now, but I guess it's only been a few months.

I think maybe the one impetus that is most salient is when we release codecs or even internally add codecs.

It was really surprising to us.

I think we recently put out some stats on this, that there's this real inflection of adoption among non-developers at OpenAI.

And I, you know, through this product development process, like, would go to, like, these UXR sessions to talk to people internally.

And the thing that stuck out to me is, like, one, like, you know, you go talk to, like, strategic finance or marketing or whatever, and they're all using codecs for, you know, their use cases.

That part's cool.

But the thing that really stuck out to me is how proud people were that they were using codecs.

Like, how, like...

It's like, I'm not supposed to be using it, but I am.

It was that.

It was, like, that they were, you know, early to this, like, new thing.

But it was also this thing of, like, they felt like they had a superpower, right?

And what we recognized then is that, like, the power of codecs, the power of agents, like, we already had this massive distribution base of people who have, you know, come to know and love ChatGPT.

Like, how do we show that to them?

Like, how do we bring it to them?

Which is, like, a hard product problem.

And it's, like, a tricky thing, right?

There's many ways you can go about it.

And so that's what we called the merge and the super app over time and ultimately launched it in ChatGPT work is how do we do that?

But it came from that initial realization that, like, the power was not only for developers.

Like, much, much earlier than probably even we thought.

Like, it could be extended to everyone.

How do you see the products differently?

So, like, who is it for, right?

So Codex started out even CLI, then app.

Now there's a merge of ChatGPT, Codex, and ChatGPT work.

So is it the opening for the average user, for enterprise, for work?

How do you position it?

I think we want to get to position it for if you're doing worky-related things, for lack of a better word.

I think productivity is, like, it's actually what, like, the pillar that I support.

Like, that's the name of the team.

And the reason for that, the reason we call it productivity and not, like, you know, enterprise or, like, work or something like that is because there's also personal productivity, right?

And, like, I think ChatGPT work is, I've seen people do things in their personal lives that you wouldn't classify as, like, work technically, but, like, these agents are, you know, super capable for.

Like, one recent example that someone posted about on our Slack is, like, someone has, like, a missed package.

Like, they didn't receive it.

And then they got, like, the picture of it, you know, from Amazon or wherever the courier was.

And they, like, asked ChatGPT work to, like, find out where that package is.

And, like, the agent, you know, is extremely tenacious.

And, like, I, like, took the image.

I, like, looked at a bunch of, like, listings around their neighborhood.

I figured out exactly the apartment complex in which the package was.

I gave them some information.

And so, like, I think there's all these things that, like, you, you know, worky or productivity-related things.

I think that's what we want the product to be.

You asked about Codex.

I think we think Codex is, you know, a durable brand.

But we have a principle that, like, the user, you know, we don't want a user to get stuck in a tab or an experience where they don't get the power of the product.

And so, like, basically everything that you can do, you know, in the Codex portion of the product on desktop, you can do and chat with your work and vice versa.

But we made some opinionated product decisions on, like, you know, how much of the git state, if you're in a git repo, do we want to expose to the end user?

Or how much do we want to make the experience of seeing the agents thinking, like, diff forward so that you're getting exposed to the diff side of the back?

And then, like, on the safety side, like, how do we want to think about, like, sandboxing and making sure that we have the right defaults in one state versus the other?

So there's, like, some opinions that go behind that.

But we do want, we don't want the user to need to choose which experience they're in.

That is a good goal for AGI, right?

Like, people don't want, like, to choose what version of AGI they want.

They just want the AGI to decide for them.

Can I get an answer?

Or, like, it's not super clear to me, is the Codex harness and the ChatGPT work harness the same?

Is it just UI affordances?

Or are there actually prompt level or even deeper differences?

So the harness is the same.

The harness is shared.

In both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plugins or computer use or artifacts.

You get that power regardless of what your experience you're in.

On the UX side, there's opinionated takes that we have when you're in Codex mode, what the UX should be, how the UX should behave.

And some stuff around the sandbox like I mentioned.

But the underlying harness and capabilities should be the same.

I think I'm just kind of curious, maybe we can, is there a query that we can run that would look different in the two modes?

Yeah, I try to create, like ask it to create like a retirement calculator or spreadsheet or something in both modes.

And then in Codex mode, you might have to be in a repo for this, but you'll see like the diffs of like the sheet that it's creating and stuff like that and the file edits.

But in REC, you won't be able to see that.

I think that's super clear.

And then also the other thing I wanted to dive into was your, the productivity team.

What else is there?

First of all, you know, what are the top level teams other than productivity?

Isn't productivity everything?

So, you know, we have a team focused on, on ChatGPT, like the, the core chat experience for consumer, which is like, you know, not, I think, all productivity.

Like there's, people are using ChatGPT every day for search to, you know, figure out how to write messages to loved ones, to think about how to like learn a new topic, et cetera.

And so, there's so much more inside to create images.

There's so much more in chat that, you know, the hundreds of millions of users are using that, you know, obviously that warrants like a very dedicated effort.

And there's teams focused on enterprise and infrastructure and API and stuff like that as well.

I will bring it up.

Yeah, so I have them both running.

Yeah.

This is work.

There's a codex version here.

I picked 5.6 Sol, so this will take a while.

I think, I think we'll just keep it in the background and, you know, as, as they finish, we'll look into some of the differences.

Yeah, but immediately, I think if you flip back to the codex version, you'll see that, uh, that assumes, it assumes Git.

Yeah, exactly.

Like, dynamic island assumes that you're in a Git repo.

Um, and you might miss some stuff because some of it is, like, in the actual chain of thought with those changes and how we display that.

Is there an unintuitive, like, is there a thing that you wanted to ship and then you got feedback and you were like, no, let's not do it?

Like, what's the thinking behind that?

In, uh, try to be your work?

Yeah.

I think one direction we could have gone with this is, like, keeping the experiences, like, completely separate.

So it's like, why?

Different apps.

Exactly.

Like, different apps or even in the same app, like, different, completely different experiences.

Like, why merge it all?

Like, what is, you know, Codex, obviously, people love.

Like, why bring these products together?

And I think the intuition here is that, like, all of our jobs are, changing dramatically with AI.

Like, for every few months, like, I feel like I wake up and I'm, like, doing a completely different thing than I was doing a few months ago.

And my hypothesis here is that, or I should say, our hypothesis is that, like, part of what we're building in this technology is giving people leverage, like, you know, the things, maybe it's the more mundane parts of your job or parts that, like, if you were able to automate, you'd be able to share more ideas faster or whatever, like, you're able to do now.

And because of that, like, that might actually blur the lines between someone who's, like, only writing code or creating strategy docs or, you know, planning events or helping with marketing or doing podcasts or whatever, right?

And so, like, these things are going to get blurred over time.

And so, like, trying to draw a hard boundary based on, like, who you are is going to be tough and, like, we should enable users to choose but we shouldn't box them in.

And so, a lot of, the work that went in here, like, you know, keeping the primitives the same, like, for example, plugins are, like, unified across this product and ChatGPT in the cloud was because of that.

It's this thesis that, like, eventually things are going to come together and we don't want to be, like, we want to be prescriptive about when to be in either experience but we don't want to box anyone in.

I wonder if there's users who are very tuned to the old ChatGPT harness that is effectively now replaced by the codex harness.

I can't imagine what that was but maybe they're more, the more conversational side.

Can you compare and contrast the two harnesses because only you've seen it?

Yeah, I mean, I think ChatGPT, the existing harness, like, still exists today.

It, like, exists in this app.

The classic, right?

You just start a new chat and you don't go under work, right?

Yeah, if you start a new chat and go to chat then you're talking to ChatGPT with the instant model.

technically do another.

I guess, you know, instant.

Yeah, so this one's not going to code or it's going to be inline.

It's not in a sandbox.

Oh, actually, we try to push you to go to work if you're creating a spreadsheet.

This is a router decision.

Sorry?

It's a router decision.

This is the decision that, you know, the model is making and then, like, you know, it sees that you're able to or you're trying to do something that would be better served in work mode.

But I think your question was, like, what are the advantages of, like, the chat, like, ChatGPT, chat harness?

It's more broadly, like, I want to basically do an oral history of harness engineering, right?

You know, the ChatGPT harness lasted us from, let's call it the 01 era until now.

And now it's being replaced by the codex harness.

Effectively.

And they're overlapping somewhat.

But I'm curious what changed if, you know, if there is.

My perspective on this is, like, there's sort of, like, a constant process of, like, divergence, convergence, divergence, convergence.

and chat, like, many of the use cases I was talking about before, like, you know, search or learning, I think we're really optimizing for latency and optimizing for personality and, like, different things that over time, like, the product, the reason people love ChatGPT is because we've been optimizing for those things and working on them for so long.

codex, what we learned was that, like, if you give the agent access to this infinitely flexible environment as a computer, you can do really, really powerful things.

And so, when we think about, like, okay, well, for knowledge work, like, what is, which mode should we choose, it was, like, it felt more natural to us to bring that to this, like, computer environment and, you know, maybe abstract some of the details of this computer away from users who might not be used to that, but, like, give them that same power.

But ultimately, I think that we want the power in all places, right?

We want to meet people where they are.

So I'm sure there'll be work down the road in order to get things to be equivalently capable in all scenarios.

But it's just a question of, like, what we've been focusing on the product on, historically, and what we're focusing on now.

I think alongside that, outside of just harness and when to use codex that you put your work, there's also the new models you've released, right?

Any guidance there?

So people love to min-max what to use, like, only use Terra on high reasoning versus for this, you know, you want to use Sol here, ignore all these.

There's 32 options.

Yeah, yeah, yeah.

But that being said, you know, for people that are expanding, so productivity, trying stuff for work that don't have the breakdown of what all this is, what's the advice, right?

Well, I mean, I think before the advice, like, the first thing is, like, none of this would be possible without these models.

Like, I think you asked earlier, like, you know, what was, like, the inspiration for work?

And, like, you know, early on, like, I mentioned, like, what we were seeing with codex, but that was also because the models were getting infinitely more capable.

That's happening again.

I think it's, like, another step function jump now.

And to answer the question on advice, like, we want this default to be the best possible.

Like, we want to be opinionated about the default.

and so we've chosen a default that we think is going to be the best for everyone.

And, you know, we have, for power users, options under the hood.

One could argue that there might be too many right now, and we're working on simplifying it.

But you can extend, you know, the reasoning level, and you can change between the different model classes if you need to, but the default should be the best for most use cases.

So my advice to most people would be to stick to that.

And then, you know, if you reach a situation in which you think that you want to try a different configuration, if you're not seeing either the efficiency on the cost side or the quality on the intelligence side, then you can change the defaults and see if you can get something better.

But we think the default should be good enough.

I have, I'm just going to run something by you since you have way more experience than me.

I've recently been doing so light, but with goal, with the idea that the goal basically augments the reasoning effort, but with more terminations and turns.

Is that a good way to think about it?

As opposed to so ultra or so, you know, extra high.

Yeah.

It's hard to say because it's like an interaction effect.

Exactly.

It's like there's a preference on, you know, for you as an individual, like how do you like to collaborate with the models?

Like how many of those like terminations, as you call them, do you want where, you know, you can steer or make sure that it's doing the right thing?

I think generally people should try whatever works for them.

I think that like using ultra or the like multi-agent setups are best for like when you have like tasks that are either incredibly complicated, like open expirations or very parallelizable.

I think even for tasks, using goal I think is best for tasks that you know that you'll be able to make consistent progress in a way that's verifiable over time.

But I think for most tasks, they actually don't fall into either of those buckets.

And so like at least when they're starting and so that's why I think the best first step is like trying it with the default configuration and then seeing like where you want to go from there.

Right.

you guys worked on a slider which actually is super helpful for reducing the amount of panic.

It's nice on mobile at least.

There's a nice slider.

It's nice here.

I haven't tried it.

So you have the advanced view there but if you click advanced view yeah.

Yeah.

Oh.

Just a slider.

Yeah.

Very pretty, very colorful.

Yeah that idea was to like reduce it to like one dimension even though there's multiple dimensions right.

Try to project it onto a single dimension for the user.

Yeah.

Like you know something from that represents like you know speed and efficiency on one side.

Yeah.

And then like sort of like quality and thoroughness on the other side.

I am just puzzled that it uses Sol so much.

Like the lower I think the slider if I'm not mistaken is Terra.

Oh it is.

See?

So they preset Terra to only be the light one.

I see.

But like I think a lot of people actually more people should use Terra.

One because Sol keeps running out of capacity.

I'm the reason you know.

Here's 10 minutes of our retirement calculator.

Oh that's the Excel thing working for you.

Oh my god look at that.

This is work and then Codex is still cooking so we'll get back into it.

I think it'll be interesting to actually see the thought process, the reasoning and also you know I guess this is eight minutes on work Codex is still cooking.

Yeah and by the way do you know Gabriel he showed me this and I was like pretty shocked that this looks like Excel.

It edits Excel files.

You never paid an Excel license.

Right?

But somehow this is like kind of workable and it's agentic Excel.

Yeah I mean one of the big like pushes that we made for this launch was like artifacts.

Right?

Like both on the model side like I think if you compare this with 5.5 and 5.4 before that you'll see that there's been pretty dramatic improvements in the quality of these artifacts and then also on the product side.

The UX side is also crazy like hosted sites and whatnot no longer needing to host your own little webpage.

Oh I have a story about that.

I can do a separate thing.

I'll need to take the visuals here but we'll cut to that later.

Was there co-training I guess because you were making this big move and you launched 5.6 on the same day as ChatGP's work was there influence between the model training teams and the harness teams or did the launch days just happen to line up the same day?

I think that we collaborate heavily with the research team and I think that's one of the most magical parts of the job the most fun parts of the job.

But yeah just using artifacts as an example a lot of what you're seeing underneath the hood there's a lot of work that went into making sure that we had the right infra to be able to train the models to get better at this and then on the product side had the right experience for users to be able to collaborate with the model on an artifact like this.

In fact this whole viewer the intuition here is that it's not necessarily that you wouldn't need an Excel license this is stage one right?

This is probably not what you meant when you're making a retirement calculator you want to iterate and when you're seeing it and if this thing is high fidelity to what you'd actually see or what your co-workers would see if you were to send this to Sean that I think makes it so easier and makes you trust the product in terms of iteration.

When you say co-workers would see do you see a multiplayer multi-team collaboration with artifacts?

Any things you guys think about?

It's something that we're actively thinking about one thing that we've noticed internally without talking too much about the roadmap is that there's many times when someone will ping me about something and I will ask the question and then I'll ping them back the answer and then I'll be thinking The simplest would be the three of us are just all on one hosted.

Exactly.

And I'll think about was I required in this loop?

And then maybe it was I'd rephrase what they were asking or pulled from certain context or whatever but when I gave them back the answer that process was also lossy.

I gave them just my interpretation of what ChachiBG work cooked up but underneath the hood there's so much context in the rollout and stuff that could be interesting.

So the answer was preemptively respond to every inbound request?

No it's just literally this is what I do sometimes as my job.

I know you copy paste and then you just a message forwarding service from AI to AI.

I think it's interesting right?

It helps people understand the capability of what you can ask and delegate that oftentimes people don't realize until they try or someone shows you and then you're like oh okay okay I see.

I think there's also a light security issue where basically you're the permissions layer like yes I could query everything that you query and I could get an automated response but maybe I'm not supposed to see it.

Yeah.

And there's no way I would know because I'm not supposed to know what I don't know.

Especially as like you know which IWD work for asking you to connect your plugins and you know it's pulling from your local files and stuff like that.

The amount of context that the agent has access to is like deeply personal and like that's something that we need to preserve so that'll be definitely a challenge.

there's Excel there's PowerPoint there's docs you know the grand trio of work what other formats of work do you think about you know like obviously you worked on Airtable is there a future where there's like open AI Airtable like you know like what what does that look like if you ever ended up doing it?

That's a really good question I think I mean one that you didn't bring up was sites and I think that was a core part of this launch there's one side of sites that I think people commonly talk about especially on Twitter and stuff or X of like you know this sort of like prototyping tool and actually like we saw that happen with this launch even the model slider that you guys were referencing earlier like that was developed almost fully in a site like you know the collaboration between design and engineering and product on that was like on a site where we play with you know the affordance and figure out how it feels and all of that but the other aspect that I think is a little bit less talked about is like sites as like an artifact for knowledge work I was actually talking to someone the other day who was on like our corporate finance team and like they were mentioning how like now when they have these reports that they're working on as a team month to month historically those things were in slide decks and in spreadsheets and now they're just in sites and like sites is the mechanism that they collaborate across the team the reason is because it's like it's like somewhat higher bandwidth like you know these tools like PowerPoint and Excel are like infinitely flexible but at some point you reach the boundary of like either as a human you may not know how to use some feature or something or the product itself doesn't support it but with the site you can kind of do anything you ask for anything and you can get that once people see that magic I think it's been really valuable yeah let me show you my case study this involves all the hot topics including chat GPT work but also 5.6 token billionaires and token maxing and sites and auto research I'm a fan of this game called Strata it's basically it's like a little board game that you play with physical blocks that come on top of it like that so over the weekend I took like 30 photos and just threw it into chat GPT 1.7 billion tokens later out comes this site with a fully playable thing with 3D block placement and everything because it requires physical blocks and I needed friends to train on it so they can get better so I can play against them but also I could also do things like train an AI on it and that's your auto research that gets into auto research so you want to train your own AIs and then make sure they self play against each other I need to set both AIs so this is AI versus AI and they're going to self play obviously the AI started out bad and then you want to define a loss function and get good I wasn't going to supervise all this I was at Daudi San Mateo attending a conference what I ended up doing was auto researching on this and creating benchmarks and there was just way too many parameters for me to read so I started asking it for a site and it's created this this this lab panel where is there is there a shortcut for a site that is created you should be able to go in the sidebar to sites top of the sidebar the last sidebar this one oh left yeah I just scroll all the way to the top oh oh it's the sites oh there you go yeah so it creates the sites I don't I don't think this is a it is exactly what I what I wanted but let me let me show you what it popped up right like I think as a as a research artifact it is very important to communicate exactly what is being done outputs this this thing which I eventually started publishing so I moved it off of sites because I wanted more database and infrastructure than sites afforded me but this is this is like research output that you can start to mess with and like try to think about like what hyperparameters are you tuning for training your AIs and like I was trying to make like scaling laws and everything and doing all sorts of like game optimization stuff and the fact that you can just kind of fill this up as a research artifact like I no longer need to read chat gpt output I read site output but then there's also a huge sprawl like look at how long this thing is there's so many numbers it is pretty overwhelming so then I have to start putting it from there but it's an interesting transition from markdown yeah actually that you're putting out to you're putting a whole functional site I think markdown just isn't that optimal for people to read right might as well just write HTML website and I don't know I think you can do a lot with customizing this right you have your skills that explain what you want like I noticed they're quite verbose I don't need a lot of this information it's very verbose so and then the nice thing of having a site side by side is you know you just iterate on what you want and what you don't right yeah I don't know if that triggers any stories for you of how it's run internally am I doing this right yeah I mean I think that this is like a workflow that we're seeing like all different types of teams use where like the canonical artifact that was previously a deck or something is now becoming a site and like with a site you because it's just HTML you can like it's infinitely flexible and so you know if you want to give more prominence to a certain thing that like in a slide deck would you know feel like it was buried like you can do that you can have it be like the hero image right and so I think that like people are starting to see that there's obviously more work to be done to make these things like much more easier easy to collaborate on you mentioned that they're very they're long and verbose could be broken up I'm sure they're super long yeah yeah but I think we're starting to see that like there is this aspect of this is a really interesting format for people to use that's like much more flexible than what they had before I think your job also becomes kind of meta you're not designing the products you're designing a product to make products and I'm curious how you manage that I think one thing that we've been like when we look at the UX like that we've been thinking a lot about is how can we balance like simplicity with capability like if we if we're designing a product like you said that like is made to build other things right you can build so many different things but we can't put that all in front of you because you'll get overwhelmed and so we had similar problem or similar challenges even with ChatGPT but especially now like when there's so much that can be done I think the balance that we're constantly trying to strike is like how can we give the user enough of a UI surface where you know they can be expressive they can tell the agent what they need they can verify that it's using the right tools it's pulling from the right sources etc but then it gets out of the way and then how can we build the right system such that we can show them instead of telling them what can be done because so much of this is going to be like how do they discover the next use case and the next one after that if they really want to be super powered yeah it's interesting I feel like everyone else just has a different way to do it right I made a similar version of this same game I didn't take any pictures of board or rule game I threw an at goal 18 minutes 53 seconds later a lot of tokens later I've got a similar version obviously not with all the auto research and whatnot but you know you had to do all the latest trends and yeah I did it with codex not work but it's interesting right yeah and this is obviously GPT image generating the avatars very good for game design like a lot of game designers were like really into GPT image I will say like the broader takeaway probably is the reason that we do this is more so just to test the tools right like this was also a test for 5.6 came out I had done the game on 5.5 right the ability for me to no longer need it to I had to feed it the rules it's a pretty niche game it couldn't find how to do this on its own oh yeah 5.6 it is auto distribution that's why I'm also very keen on testing the 5.6 capability but you know this is just as work comes out as new things come out these are just our sideways to test things right yeah it's some kind of private evil I guess that is not all this private but also valuable because now you can send this to your friends and I mean I learned about this game through seeing this it's a hard game he's very good it's good when no one is competing with you but yes it's a classic RL problem of like self-playing bootstrapping your game AI yeah you see how easily work becomes personal and personal becomes work because the thing I do for personal it actually directly informs people I work with because I showed it to them they were like oh you can do that with GPT which I imagine is the growth strategy yeah the show not tell is a big piece that you know I think we're not still not fully cracked of like you know showing people all the things that they can do with the product versus like trying to teach that to them like you know articles or onboarding or whatever yeah meeting them in the moment it's a career risk for me because I used to be in developer relations right where your job is to show and then you're like what do you mean you don't need actually your job is to tell and then but the product people are like well we don't need you if our product isn't intuitive enough so yeah I mean that's the magic of the models so you can tailor the telling or the showing to like specifically what the user needs like what they care about what they've done in the past exactly where they are on the adoption journey so I think that's like going to be a super big opportunity seems easier and easier now to tailor custom showing right people have different use cases as much as you said you don't want to segment different people into different buckets right it's also not that hard to for people that are in different categories but the question I guess is you said your team is more broadly on what was the term you used productivity which is now work basically is it work is there another distribution that we're not hitting is there a group of people that will have something different than chat GPT codecs or work is there is there more that the mass isn't targeting I see it as like a sequencing like you know the vision is like bring useful agents to everyone we started with like developers like developers historically are like early adopters that are willing to put up with more friction set things up etc like that's where you know codecs started I think the next opportunity is like sort of what we call general knowledge work you know all the other functions around developers I think when you go from developers to this segment like there's inherent challenges obviously with like you know this show not tell thing that we're talking about making the product more understandable bringing in new capabilities that matter more for this cohort than matter for developers things like artifacts things like computer use etc and then I think like the same learnings like similarly how we took the learnings from developers and brought it to you know general knowledge work the next stage will be like taking the learnings from general knowledge work and bringing it to everyone no matter what they're doing in their lives and we're already seeing that a little bit like this game example that you have is you know something that's like on the border of like fun and personal life to you know your professional life I use chat GPT work full time at home for everything like for whatever I'm doing I used it the other day to come up with a meal plan and like you know save that on on the like computer environment that it has and something that I can continue going back to like is everyone doing that yet probably not because the things work on it but eventually you know we want to get people there chat GPT life yeah exactly chat GPT cooking but I think there's a lot of there's a lot of opportunity there but I see it as like you know we built the foundation in software engineering and we're going to take the same learning so we take software engineering to knowledge work knowledge work to everyone do you have any power user advice I feel like there's a group of people that will live it use it for everything stay on it 24-7 yeah and then there's a bit of a gap between that crew and people that you know okay I use it for work I use it occasionally sometimes I pipe questions any advice any learnings anything you recommend or just you know takeaways that you found that help bridge that gap I think a couple things that I've seen is like one that it really helps to broaden your imagination of what's possible and this has been a learning even for me like you know the technology has progressed so fast that you know something that like even three months ago I was like no way no way that the malls can do this like now it's like wow it's like you actually can like you have an example we're going through right now that are like review cycle internally and people always talked about this as like kind of a thing that the the models are good at and like you know there's a cliche of like okay like no one wants to be writing reviews and like we just use AI to do it but I mean in all seriousness you have to evaluate it as well yeah exactly in all seriousness before it was like just like slot basically and like I think it was helpful but you know not super productive now I've found that like the model can do a much much better job than me especially in this environment of like pulling context on like what people are up to how they've like the things that they've done to make a difference highlighting like you know wins that they've had that like I may not even have seen you know has access to like everything right like the code like you know things that they've caught reviews slack everything and so it's like incredibly powerful in that domain and like just like six months ago the last time we did this like I didn't even I tried using it but it was not at all helpful and this time it's been like incredibly helpful and like so I think continuing to push the frontier of imagination what's possible even if you tried something before I think is maybe my biggest piece of advice the other I guess thing is like the more the more you put in especially in this environment where like you know the model has access to everything on your computer or in Chagibity work like you can create you know artifacts over time and save them in your library and like the model will continue having access to those like the more information you give it about whatever domain you're in whether it's your life or your work the more valuable it becomes and it'll become both valuable in like ways that might surprise you like it might pull from context in a way that you know may be proactive and that you might not even have thought about but it needs to have access to those those tools or that context first one thing I just want to talk about the review stuff because I still that's a very sensitive thing and you're a founder you've managed people you've hired people as manager myself I'm very reticent to put out any LLM generated things especially when it comes to people because it feels like you don't care presumably at OpenAI people are obviously more open to being basically rated by GPT but are there any unofficial rules around this like what's the etiquette oh I mean I think the etiquette is that like I would never write something via like solely via AI and like present it as like a review for someone what I was talking about is more like gathering context that's the place where it's incredibly helpful It's just search it's agentic search like agentic search but you know that you can tailor and steer much more capably than you could before and because like the thing is it's all there's sort of a flywheel happening right because of codex people are able to do and because of Chagibur people are able to do so much more now than ever before and if you're able to do so much more it's easy to miss things as well and so like I think we need to use these same tools to keep up with all the impact that people are having and understand you know where it can be helpful I think that the thing like obviously I run a small company so easy to search but at the scale of OpenAI with the amount of messages that you guys put in Slack do you think that it misses things?

Probably but I think that I also miss things Like it doesn't matter I think sometimes It needs to be human level It's all relative right Yeah Sometimes it's nice when it finds things you wouldn't right Like right now my codex system prompts they're set up in such a way that every project I have has a separate notes MD and it just writes learnings to there and then the global one can pull from all these so sometimes it'll be like oh there's this project you did like four months ago here's a note that we had and it randomly pulls it back in the context that I would never do I haven't thought about and I'm like okay this is quite superhuman right like stuff that would and you know it'll save like hours on chunking of stuff or find something that's already been done and I'm like as much as it might miss stuff I would too but it's very useful when it finds stuff and I have like a very you know non super engineered solution to this it's just markdown files that get pulled whenever they want yeah I actually have a funny anecdote about this like recently gearing up to this launch you know the team has been you know really cooking on it for a couple months and over that time like there's so much conversation and chatter going on in Slack and Docs and elsewhere and one of the members of the team set up this scheduled tasks like automation to like look at everything that's going on and like come up with the best memes and then post it in one of our shared channels and like there's two cool things about this like the first is like I think the models are you know over time like actually starting to become like funny or it was like you know a year ago like that was not at all the case the second is it was what you were saying like they find things that in surprising ways that you may not have thought of and like create connections that you may not have thought of and that really helps with like the meme generation because then you can see something that you know genuinely surprises you and is funny in that way so yeah I mean obviously that's like not like the most productive use of this technology but it does it doesn't cover this like this capability that's emerging which is just like defined information that you otherwise would not know talking about the launch I think I have pretty much said this is the most successful launch in a long time I think even more successful personally than 5.0 and you're announcing 10 million users does it feel different you've been through a lot of launches I think it feels like a culmination well I think two things one it feels like a culmination like I was mentioning earlier like this like vision mission that we've been on for a long time like I said we saw the magic of Codex internally and then we're like extremely excited to bring this to many more people and to see it working to like see us reach you know the distribution goal I mean numbers that you mentioned like I think that's like huge and super exciting the flip side of that is like there's so much more to do too like that's also really exciting like you know ChatGPT as a whole like this product that you know everyone almost equates to AI and like loves you know has hundreds of millions of users and so like 10 million is really cool but like we need to get this to everyone like we need everyone to feel this magic and so that's the next step from here but yeah I think extremely pumped about how it's going so far and the opportunities awesome I did want to also because I've been tracking the number closely it transitioned at some point from just Codex users to Codex plus ChatGPT work obviously because the same harness the whole point is that you don't you can't content separately you have roughly a billion ChatGPT users why did it just jump to one billion right away like isn't that the default on ChatGPT or no we don't default you into ChatGPT work if you're on ChatGPT if you're free yeah it's also only available to paid users right now and I think there's like a process of educating users of what is the value of this product having them try learning from their feedback and making it better over time but I mean the goal is to get as many people who love ChatGPT today to feel the power of ChatGPT work but I think it'll be a journey yeah Codex will still be alive as a brand for the foreseeable future yeah and we'll just toggle between them as needed for UI stuff yeah I think it's an even stronger point than that I think we fully intend to treat developers like developers have been you know a core market for us for so long and like there's there's so much more that we can do to make Codex great specifically for software development and we'll continue to do that this doesn't take away from that at all if anything it should increase the utility of something like Codex because now you can move seamlessly between writing a diff to creating an artifact or you know doing a search over your character I do wonder how much this terminology leaks to the non-technical user like do they have to learn to say artifacts if I want artifacts or you know it's funny like we call artifacts internally because that's what the team's called but like externally like no one says that no one calls it an artifact but I think that people like often like describe things whatever they're used to right so if you know ChachiBD work is good at creating slides they'll say ChachiBD work is good at creating slides and that's actually what we want one big another I mean it's July of 2026 one big thing that also happens for OpenAI was OpenClaw and that's I think a lot of people's first time really maxing an agent for personal stuff but also crossing over to work in some same way as far as as far as I understand OpenClaw is still independent but did you go through your own OpenClaw moments were there any lessons you took from OpenClaw to Codex or back whatever I think there's a lot of inspiration I did go through my own OpenClaw moment yeah tell the story me and my my wife like set up an OpenClaw to like try to manage everything in our house not that there's like a ton but it was like actually quite useful we gave it a calendar and started you know creating events for us and stuff at some point the laptop that we were running on it died and I never got a chance to pick it back up but there's a lot of inspiration there like you know in ChachiBD work in web and mobile like you get access to like persistent computer environment where you know you can store files and those files stay around between sessions and the idea is to be able to enable use cases like this one of the members of our team actually uses ChachiBD work for what they used OpenClaw for before and I feel like it has like completely transitioned which is like workout planning and like meal tracking which again it's like a worky thing right it's like not work necessarily but it's like in personal productivity space but it has all the same primitives so it has scheduled tasks it has the ability to store files on a file system it has the ability to like reference those things over time and so you start to see the same types of use cases emerge which has been really cool Is there a point that ChachiBD work completely replaces OpenClaw obviously they're independent so yeah I mean I'm not close to it so I can't speak to the OpenClaw roadmap but I don't think so I think that there's going to be you know there's always a need for like this like incredible like open source technology that team has built and I think that we can draw inspiration in the product and you know ChachiBD I think many more people have like heard about and used ChachiBD than have used OpenClaw and if we can take the magic from OpenClaw and bring it to them I think that'll be a success I think that like one thing on the ChachiBD work side that we feel strongly about is that like the core experience is that you come to this product and you have a conversation start a session whatever you want to call it with this agent and the magic of the product is that you can do anything in that moment and we would like to create a product where you don't have to click a button or to go to a different place whatever and you can get whatever functionality exists in you know your finances app or any other product like in this one place and so that's the goal it's like we want an extensible system with plugins where you can connect to the tools that you need in order to be able to accomplish like a financial task where you can you know if you're doing science work like we have an ability to like extend the system and such that you can like write the tech and it performs well there will always be like products that we support that are best in class at those things but we want as much of the magic as possible in that core experience yeah do you think that you can do everything you used to do with well friends in chat chpd finance I actually tried it I mean like chpd doesn't yet custody cash and assets for me so that part no not yet but I mean there was like a whole component of like retirement planning and sort of like financial planning and budgeting and stuff that we were looking into when I was there and like with the finances plugin like that's all possible so I feel like at least that component is replaced for me I haven't really plugged it in yet I'm somewhat scared to look at the answer like that's honestly like the same reason for health and finances like I'm like no no it's really good I mean it's really cool how I mean we were talking about like the agentic search aspect a little bit earlier but like it's really cool how like you know in conventional the more power you want to give to a user the more like knobs and bells and whistles you need to add like you know for like these finance and budgeting apps like there's always like a bunch of the different filters and like search bars and stuff like that but like now like with the right connectivity to the right data you can have whatever you want you can ask any question you want and into that box and get the answer and I think that's super powerful I think it's also nice to just have it centralized in one space right you have different health apps I have one for a smart scale a watch all these different things it's just nice to centrally co-locate it which is you know part of the whole thing of open claw right like that you would have a personal OS which presumably chat GPT wants to become I do think that just relying on just in time pulling of data for let's say via MCP CLI API whatever you do still not enough I come from a bit of a data engineering background you still want a data warehouse or some kind of caching or semantic layer do you feel that or do you already have that I can't speak to all the details on how everything works but I think it depends on the access pattern right like if you want an answer immediately then yes it's very difficult to do that you need to pull from all of these sources but a lot of the use cases that we want to enable in chat to be work aren't necessarily something that you need immediately it's more like a task that you want the agent to go and do and that's going to take a certain amount of time and with things like programmatic tool calling and stuff now and sub agents is also parallelizable and so it's possible I think it's very possible that the ceiling on what can be done with MCPs and calling out to these services has been raised substantially so we're really excited about that you mentioned some agents I got a double click on that ultra is a new mode you have special affordances in chat GPT itself to show off the agents can't really do much with them to be honest just watch what have been your experiences any design issues that you would call out to other builders building with sub agents I think it sort of goes back to the balance that I was raising earlier about showing builders the power of the tool but also creating enough of an abstraction to not overwhelm them I think with sub agents the thing we wanted to show is that you can take a task that has many parallel tracks or is complicated in a way that some agents can handle and this product is for you the model can accomplish those goals and so that's the point of showing them the product and that's where we've gone with the design there's another iteration of this where you can see exactly what they're doing and things like that which I think could verge on overwhelming with information and so this is the trade-off we made for now you do display quite a lot of transcript right right I think it's hidden by default right some people could want more so I'm one of those people that will basically throw a lot of stuff at goal and pretty much every goal I'll tell it to use sub-agents seems redundant right but every time I'm like okay use sub-agents where possible and I have a lot of people a lot of friends that recommend and do the same whereas I'll sometimes talk to people that are like okay this is where I want you to use sub-agents for this sub-task and I'm sure they would appreciate seeing into how they're being used for me it's primarily like two things right one is net time efficiency so span out across sub-agents two is probably cost right don't use big expensive model offload to a lot of smaller cheaper models and some people want that level of control so if you have repetition in what you're doing right say I want something built where I wanted to consistently do this every day I might want to go in and fine tune sub-agents here sub-agents there so you can see both but I think if I'm not mistaken it's hidden by default there's a drop-down that goes a lot where I'm like okay I'm just going to keep using you can change the model that they use I know I tell them to be steered I'll say my I know Anthropic offers this in Cloud Code you can tell Fable to use Sonnet or Opus to use Sonnet as sub-agent so pretty trivial thing you know you tell it to span out sub-agents with Sonnet you know it's cheaper faster I would assume if it's not there it could be built there but I think there's a side of too many toggles it's not a toggle actually it's just you tell it in chat the way I do it is prompt it right and I think this is something that gets abstracted unless it's something you built for repetition right so if I'm building something say that's podcast prep right research into people do a very very deep extensive research that I might want to configure to cheaper faster model just for web search right I can see a world in which you want both I think the default is actually pretty good right now where it's hidden but you can drop down and get some more info into what's done I know people talked a lot about it on 5.6's launch this thing loves to use a lot of sub-agents and causes the chat GPT app to just crash because it's so processor heavy but for what that's not my experience yeah I mean you know I haven't had a crash from sub-agents I haven't either we both have big laptops I know people brought it up there was a topic of discussion that we didn't see the same but it is another vibe eval right people are like okay the amount of sub-agents slowly spawning is crazy and I'm like I think this is okay I think it's good but just stuff people bring up I think when we launched the product too we weren't as opinionated about who is ultra for and when should they be using it and since then we made some changes to require you to turn it on and find it in the advanced setting because that's who it is for it's for power users who understand what's going to happen because it because it also depending on your use case can use more of your limits as well so that's where I think a lot of the feedback was coming from that's okay reset the limits always reset the limits today we're resetting because of this I want to change topics to one last piece of the harness memory a lot of people are commenting on memory recently Chad is not very good and then this guy also basically the same thing and Samir who you presumably work with talking about memory what can you say there I think Samir and the team have made a ton of updates and improvements over time I think when I talk to friends family members about what they love about Chad GPT the fact that it knows them they feel like their Chad GPT is their Chad GPT I think comes up probably number one and Chad GPT work in the cloud by default all conversations are inherent from your Chad GPT memory so you'll know context about you and they'll also be able to write back to this memory with like a small text write you tell me when you're writing right no it's part of the same memory system that we launched so I think that's been really powerful because going from Chad GPT to Chad GPT work feels like an extension of what I've already been doing with the product for sometimes many years so it's been awesome and it's awesome to see people are recognizing the improvements here so it's basically a retrieval problem right like are you retrieving the right things are you over focusing on the wrong things is there like more false positive or false negative you know if that makes sense like what's the bigger problem so I don't work on memory directly so it's hard to say what the bigger problem is with like certainty but I think you're right I think that like you know there's two sides of it it's like you know making sure it knows things about you but then so I think it's a very challenging problem but something that I think we feel very is a huge opportunity to get right which is like why we've made big investments in it how do you see the side of okay when you're building chat2pt for work different than the regular chat app different than codex managing memory across different projects collaboration and whatnot how do you see the side of what's separate from the harness right so if I have four threads on one project any learnings on how to build memory systems there you know for background as well I guess to steer it a bit is when you do chat style applications I'd say you have a lot of one offs right when you switch to work it might be something you're doing for a month something you do a lot right now as I add more sessions there's a lot more than just single threaded right and there there might be memory there I mean I think first I challenge that like the depth of the memory or the value of it is like fundamentally different across chat and work like it is true that like you know there are a lot of like shorter sessions on chat but I think you know the chat the product has had this technology has been around and people use it for worky like productivity related things already today and so I think we found that there's a lot of value I mean I found this with my personal usage like all these one-offs add up over time into something like quite durable and like quite a good representation of who I am I know like from time to time something will go viral on X about like you know chat telling you everything it knows about you and people are always surprised like how how deep that is the fun roast me you know exactly so like I think that's all to say that like I think there's a lot of depth there in the existing chat product and so that's why I think it's valuable to bring into the work product but the other reason I brought that up is because I think like hopefully we can use some of the same fundamental primitives and systems to extend memory here as well and I know this is something that the team that focuses on this is like working through right now I wanted to bring up one element of memory which I honestly don't really use much and I'm curious if you do Chronicle which is up on screen right now it's kind of a super memory or like what is it I think the idea is that it can learn from how you're using your computer and it's another input source into memory and I think it's experimental right now and something that isn't default off but I'd recommend that you try I think that it's quite interesting how it goes back to a conversation we were having earlier on you were asking can chat gbd miss things on Slack when it's searching does it miss things because there's such a volume of stuff you can ask the same question about everything you're doing on your computer is going to know everything you're doing is going to capture the intent and stuff like that probably not but it probably will find things that you might not know about and if it can surface those to you in relevant times in proactive ways when you're doing tasks you're trying you're trying so mostly for insights and longer term yeah exactly like insights and it builds context that can make you more productive on certain tasks but it's hard to describe without feeling it I will say you can feel it pretty well like the idea of what they're saying here right just check through my memories or check through my logs and add skills yeah pretty underrated that's automations you can repeat that using a chron job checking through your memories and creating skills yeah I think the creation of the memories from chronicle itself is like what's different it's like you have much deeper memories because you have chronicle on it's there I don't use it much but maybe I just I need more examples I imagine you guys use a lot of it internally so I'm always fishing for use cases yeah I would just try turning it on and then like it just auto works yeah and seeing where it might start helping you I think you'd be surprised yeah amazing I think that was about it in terms of the overall coverage of chat I think there's been a lot of good progress and discussion on building and all these things there's a lot of ex-founders in the community and in open AI as well do you think that things have changed a lot your overall reflection of building pre-AI and post-AI I mean I think things have changed a ton I think it's super exciting to see how quickly you can go from idea to something real today whereas even before I think 5-10 years ago it was fast if you were scrappy and rolling to build the minimal viable thing but now the extent of what you can build is much broader and I think that also what we've seen internally building is that gives you an opportunity to validate much more quickly to talk to users to talk to internal doctors etc and make sure you're on the right track and that loop I think has become more closed than ever before and that's a win for product development and I think it's a win for consumers and users too because ideally that means they're getting much better products out the gate Does it mean your team is smaller?

I think there's much more to do now so I think people can accomplish more individually or in a small team than they were that would require more people than before but at the same time there's also more to do so I think the teams are much more ambitious Have you seen any changes in scopes of roles and building teams and how we used to have teams say a few years ago versus what ideal teams look like now?

I think we've seen a blurring in the lines between the typical product development functions between EM, BM, engineer, designer, et cetera I want to bring up this quote there will be only four jobs left in tech there's AI slot cannon people who just burn a bunch of tokens and then there is SRE people who are more responsible there's grownups who sell things and then there's hot people This is an interesting take I think my suspicion is that there's everything everyone will be T-shaped in a way and that AI will enable everyone to become a generalist things that I never would be able to come up with a design before and even now I don't have the visual taste required but I can iterate on something with the help of AI but then people will have a specialty and that's the straight line in the T or the upward line in the T and so you can have a specialty that you're interested in with the help of AI you can go deeper and become better at over time but then you'll also be a generalist and so with that foundation what you can accomplish is almost limitless What are you bottlenecked by in terms of specialties?

Do you need more designers?

Do you need more slop cannons?

Do you need more hot people?

I think the bottleneck becomes sort of like ideas and taste I guess I think because anyone can build now I think it really is the era of bottoms up ambition and because there's so much to be built you're always going to be bottlenecked by the amount of ideas and the amount of things that you're doing at a given time I have the example of like I have the front end design skill that's like they give me four drastically different examples of what this looks like sure it burns a lot of tokens but you know and then I'll mostly just condense down okay I like this part I like this part let's draw these together and it's like yeah I had a vision but like I don't know I would say that the one automation that I would love to work and it doesn't work is bring me new ideas right somehow LLMs which is not it one interesting part about ideas is like they're not like in a vacuum it's like not they usually come from somewhere and like you know in product development like they're coming from talking to users or reacting to you know friction that you're seeing or feedback building on some foundation that you already had planned out before whatever and so I think that's where like there will always be value in these generalists that we talked about like you know closing that loop and having coming up with those ideas that are grounded in that feedback or talking to users or whatever it is cool you lead the productivity team how do you define productivity I think our mission is to make it possible for people to do things that they weren't able to do before and right now we're thinking about it from the perspective of knowledge work and so when I look at knowledge work I think about people are no longer siloed by their roles they're no longer siloed by maybe the background or training that they have like no matter what function you otherwise might not be able to interpret etc and then I think that extends to your personal life where we want to give you leverage at the end of the day and we want the models and the product to be able to give you leverage so that you can create time for yourself to do the things that you love does that also translate to a way to measure productivity like what is the end of the leverage I think we haven't figured this out yet part of the reason is it's so diverse everyone has different goals and really the true measurement is like their ability to achieve that goal did we help you or did we not yeah and it's very difficult without knowing what that goal is up front and also tailoring your career individual and the thumbs up and thumbs down from chat doesn't give you anything right you don't know if they're thumbs downing the content of the answer the vibe of it whether or not it helped them with their goal I think that's difficult but it's something that I think we will need to figure out and the industry at large will need to figure out because that's how we measure success if this is what we're for do you think it's changed productivity and how you measure it basically you said there's a lot more work that can be done a lot more scope has it changed I think it was always true that what you really wanted to measure is like you know was your team was the individual was you personally able to hit the goal or are you closer to hitting whatever your goal is right but I think previously we used proxies for this so like you know code commits lines of code or whatever story points yeah exactly story points and like they're coming back by the way maybe but that is sort of part of the change and like I think with AI now those proxies are starting to fall apart like you know the number of tokens you use or the number of pull requests you make or like no longer like maybe as hyper correlated that is your team able to hit the goal or are they on track to hit their goals so I think we'll need to come up with new measurements for the managers listening give them one thing to try I think for me what's important is like at bats are we as a team building the muscle to have not just quantity of at bats but quality like are we able to go all the way from like generating an idea building it out getting the feedback reacting to that feedback actually validating or invalidating the hypothesis going on to the next idea are we able to do that really efficiently like that goes to like you know the actual like code that's being written or the designs that are being made or the specs that are being written whatever but also the culture of the team like do we have the humility and and and are able to like go through that process many many times and stay motivated and excited throughout that so that's the thing that like I think is important now especially when we're on the frontier of this technology and like there's so much to build there's so much to do that's probably the most most important thing that we look at any traps people fall into around measuring productivity what your team work on I feel like there's a lot of okay we added a lot of LMs we have dashboards for this and that but not much has changed right that is the trap yes and you know the broader source of the question is for for the managers and teams building you know how how should they approach this I think maybe the trap is like conflating motion and progress I think motion is much easier now than ever before because of the tooling that we have but progress requires you to be like very prescriptive and deliberate about like what you're actually trying to achieve and it goes back to our question of measurement right like you wrote we were talking about like can we open AI like figure out how to measure productivity for our users that's that's a very hard problem because of the diversity but like as a team like you should have a really prescriptive and deliberate view on like what progress looks like for you and for your team and if you don't have that then it's very easy to conflate these two things I think at-bats is a really great thing I'm really glad I like the discussion between motion and progress I think that's a quote that we're going to feature on the write-up you've been very generous with your time thank you so much and congrats on 10 million yeah thank you for having me next one I had 100 in two months thank you