Back
Mark Chen
Chief Research Officer, OpenAI

OpenAI LIVE: Greg Brockman, Mark Chen & More

🎥 Aug 07, 2025 📺 John Coogan ⏱ 257m 👁 925 views
TBPN.com is made possible by: Ramp - https://ramp.com Figma - https://figma.com Vanta - https://vanta.com Linear - https://linear.app Eight Sleep - https://eightsleep.com/tbpn Wander - https://wander.com/tbpn Public - https://public.com AdQuick - https://adquick.com Bezel - https://getbezel.com  Numeral - https://www.numeralhq.com Polymarket - https://polymarket.com Attio - https://attio.com/tbpn Fin - https://fin.ai/tbpn Graphite - https://graphite.dev Restream - https://restream.io Follow TBPN:  https://TBPN.com https://x.com/tbpn https://open.spotify.com/show/2L6WMqY... https://podcasts.a...
Watch on YouTube

About Mark Chen

Mark Chen, Chief Research Officer at OpenAI, appeared on the podcast "Let Them Cook" on June 25, 2026, where he discussed the company's research direction. Chen stated that he "firmly believes in being on the exponential and in scaling laws" and said he "fairly strongly disagrees" with views that pre-training is dead, noting that such narratives have recurred throughout the history of developing large language models. He described reasoning as "one of the biggest examples" of a research bet at OpenAI, referencing the o1 model as a breakthrough that was difficult to get off the ground because the pre-training plus post-training paradigm "felt like such a promising paradigm" at the time. Chen also said the field is in an "evals crisis," with a low number of canonical gold standard benchmarks, and noted that tools like Codex have enabled faster iteration of evaluations. On June 16, 2026, Chen appeared alongside SoftBank CEO Masayoshi Son at an event in Tokyo where SoftBank announced a cybersecurity service using OpenAI's technology. Chen described cyber capabilities as a "dual-use capability," stating that "even though the models can get better and better at finding vulnerabilities, we can use that for defense." He added that "the important thing is we can go and try the models and try to find the vulnerabilities before external actors can go and try to find the same vulnerabilities." Son compared the dynamic to a criminal with a knife facing a police officer with a gun, saying defenders must have the "best, most powerful weapon" to defend against bad actors.

Source: AI-verified profile updated from Mark Chen's recent appearances. Browse all interviews →

Transcript (643 segments)
U
Unknown7:10
You're watching TVPN. Your background looks way different because you have a whiteboard behind you because we're breaking down the X's and O's of the GPT5 launch today. GPT5 launched from OpenAI. Uh really quickly, there is some other news. Firefly Aerospace stock opened at $70 in NASDAQ debut. This is the company that landed on the moon. Very cool.
Very cool. Um there there are a few other stories going on, but we're going to skip most of them because we're going to be focusing on chat GPT today on GPT5. We have a bunch of uh a bunch of guests coming on. We have a stacked lineup. We'll pull that up, but we'll break down the x's and o's of the matchup. So, of course, Open AI, here's our here's our lineup. We have something like 15 guests today. uh a ton of folks from OpenAI, a ton of people that build on top of OpenAI and uh can comment on what's going on with CH GPT. Um but of course this battle is between OpenAI and the timeline. It's the it it's they got to get the vibes right.
It's war. It's it's the timeline's in turmoil over whether or not this is a good model, what it means for the industry, what it means for AGI timelines. Everyone's got their take. Everyone's posting memes. There's been a ton of funny ones already. We'll take you through them, of course. But let's break down the offense today. We have Sam Alman, the founder, CEO. He briefly got cut from the team in November of 2023, but he's back leading the team for the 2024 2025 seasons. He seems healthy. He's doing great today. Uh he went on at 10:00 a.m. to break down the launch of GPT5. Uh he has a couple of key plays in his playbook, in his arsenal. Uh he's got a solid ground game, lots of quick posts, hitting the timeline, probably in lowercase. Then he might air it out with a couple thousand-word essay. We've seen him do this before. It's a bit of a hail mary. Maybe a thousand a couple thousand days away. Maybe we're in the soft singularity, but he's very strong there with the long post when he needs to be. It's up his sleeve if he needs it. Um then he can also pull out the vague posting. He was doing this last night. Posted a picture of the Death Star. No one knows what it means. Maybe it was taking a shot at the doomers who are on the defense today. So he's also known for driving supercars. that lets him get to the office faster. He's saving time and money. You can save time and money by going to ramp.com. Easy to use, corporate cards, bill pay, and accounting and a whole lot more all in one place. And so he is uh he also gave apparently, this is a rumor, he gave every Open AAI employee who's been with the company for more than two years $1.5 million. A lot of people say 1.5 million, that's not enough for a big house in San Francisco, but it is enough for a supercar. So that's probably why he picked that number and that's why that's what the OpenAI team will be doing with that money. They'll be buying Aston Martin Valkyries, Pagani Huayas, McLaren Sabers for Ferrari Daytona SB3s. Uh they can get a Koig set game. They could get a Singer DLS or Bugatti Veyron. It would have to be used. They could also get the Bentley Bakalar. There's only 12 of those ever made. Uh it's an open top two-seater roadster. It's coach built. So that's going to run you 1.5 million, but that's perfect. You just got the 1.5 million bonus. So put it to work. Spend it all in one place on a car. This is financial advice.
Yes, exactly. Then you got Greg Greg Brockman. He's joining at noon. He's he's extremely well-rested. He's actually coming off a sabbatical right now. That's very exciting. Uh he should be injury-free for the rest of the season. Uh he cut his teeth at MIT and uh then he got drafted by Stripe in 2010. Uh Microsoft tried to do a trade deal during the 2023 chaotic trade deal trade window that opened up post Sam Alman ouster. Uh but he stuck with the open AI team and now he's president of the company. Then you got Mark Chen. He's coming on at 11:30 today. Uh he's the chief research officer. The rumor is that he turned out a maxed out contract to head the Metal Lamas, but he's sticking with the OpenAI team. He was an MIT undergrad. Also worked at Jane Street before joining OpenAI in 2018. Then we got Sarah Frier coming on the show at 12:30. She's the CFO of OpenAI. It's her job to find bank accounts big enough to find to fill all the cash they're raising. It's it's a tough job. You got to find Okay. This bank account, will it hold 10 figures? Will it hold 11 figures? Will it hold 12 figures? Like a lot of cash in this one. Exactly. Exactly. She's also going to be defining the non-GAAP metrics that will be catnip for Ben Thompson in just a few years. We're excited to talk to her about how she's measuring the success and the health of their business. Obviously, it's not just revenue. It's not just topline, bottom line. We're going to want to know about queries. We're going to be want to know about DAUs, all those non-GAAP metrics. That's where people are going to be tracking when IPO uh when IPO day comes hopefully soon. And then we also have Brad Lightcap. He's joining at 2:35. He entered the league as an investment banker. Let's give it up for the investment bankers. They don't get enough credit around here, but we love the investment bankers. Then he got drafted by Y Combinator before joining OpenAI as CFO in 2018. Now he's the chief operating officer. And then we have uh Max Schwarzer. Uh he's in charge of post training, fine-tuning these models, getting them into the fight bit fighting performance to put on a display of authority on GPT5 launch day. Now let's flip it over to the defense. They're going up against the timeline. They're going up against the vibe checks. We got the Doomers. The doomers. They're led by Ellie Yudkowsky. Admittedly, everyone knows this. No one debates this. The doomers have had a terrible season, but you'd expect to see at least a few hail marys about GPT5 creating bioweapons thrown up on the timeline today. Probably won't be bangers, probably won't get a thousand likes, but you'll be seeing them here and there, mostly in the replies. We've also seen some doomers talking about uh GPT5 being available to every government employee, and Eleazar had some harsh words about that. don't give the keys to Sam Alman. Don't give the keys to the government to open AAI. Uh he was upset about that. But in general, the doomers not putting much of a fight up today. Then you got Claude. Uh interesting. Claude was caught playing for the wrong team earlier this week. Anthropic, they're on defense today. Uh but we saw them take out OpenAI's key pinch hitter, Claude. Uh the Claude Code API was playing for the OpenAI team, but they shut that down and Claude is no longer pinch hitting for OpenAI. Uh then you got the Elon stands. Uh the ground game's going to be there. It's going to be track uh it's going to be strong. The Elon stands are going to be tracking the benchmarks relentlessly. We know XAI loves to benchmark and all the Elon stands are going to be calling out GPT5 for any any misaligned benchmarks. If they fail humanity's last exam, it's over. It's over. Uh they'll also toss up the occasional unhinged conspiracy theory. Uh moving on, Gemini. Uh the betting lines have shifted big time. People thought Gemini was out of the game. They're so back. Poly Market has Gemini at what 75% chance of being the best model towards the end of the month. This is of course based on the LM Arena more vibes-based benchmark. But uh Gemini will probably be quiet today. They usually don't try and front run press releases. They usually try and sit back, let the model speak for themselves, let the API credits work their way through the latest YC demo day batch and get the product into the hands of people. And so, expect to see a big uh glossy conference in a couple weeks demoing uh Gemini 3 should be a good rebuttal from the Geminis. Uh then you got the Metal Lamas. Zuck's been on a poaching spree. He's rebuilding the team during the off season. Uh now he has a stacked roster and he's ready to go duke it out. But no one knows exactly what's going to be in the playbook. Is he going to go consumer? Is he going to go API? Is he going to turn into a hyperscaler? We don't know. But we know they got a stack team. They got Alex Wang. They got Nat Freedman. They got Daniel Gross. They got tons and tons of other researchers. They've been raiding every other team. Completely reset the salary cap for the league. And it's been uh it's been an absolute clinic in terms of recruiting over there at Llama. Then you got the final benchmark, Arc AGI. This benchmark stands. GPT5 couldn't get past this defense. And uh RK AGI, you know, sitting there right in the end zone just swatting him down. Swatting him down all day. You think you think you you think with super intelligence around the corner? RKGI denied. Denied. Uh Tyler, give us the update on ArcGI. Where does everything stand? How GPT5 do? Does it matter? Should we care about ArcGI? We love the team behind them, but is it an important benchmark? Should we be tracking it today?
T
Tyler16:03
Um, yeah. Okay. So, so there's RKGI V1 and V2, right? And V3. I actually don't know if No one's been No one's even tested V3. No one's even really close there. But how we doing on V1? GPT5 is at 65.7. Unfortunately, that's going to be 1% just short of Grok 4 at 66.7. Okay, Arc AGI 2, 9.9%. Grok 4 16%. So, absolute kind of brutal, you know, ARC AGI mogging.
U
Unknown16:38
Rough showing. Rough showing. people have have accused Gro 4 of being slightly benched maxed. You know, this is, you know, they might have a team, but um what's the what are the pros and cons? We know the cons of benchmarking uh of benchmaxing. You're overfitting on something that might not actually drive consumer value. It might not actually solve real world problems. It might not increase DAUs or revenue or ARR or anything that really matters. It might not even get us closer to super intelligence. Give me the counterargument. Why is benchmaxing good?
T
Tyler17:12
The bull case for benchmarking. A bench maxing. Break it down for me. Yeah. So I I think the idea is basically um this is almost like a non-AGI build kind of take, right? So if you don't have a a super general intelligence, um your ability to benchmax basically proves your ability to um solve some like kind of specific task. So, so there's this um thing about the the gas station. Yeah. It's called getting spiky. Getting spiky. Getting adding more spikes to the spiky intelligence. Yeah. I think it was Rune who had this this tweet about the gas station benchmark. Right. I I don't care if he said something like uh I I don't care about um AI solving gas stations if it has the gas station benchmark. Something like that. Um, yeah, but the idea is like if if you if the if making the gas station benchmark run said, 'My bar for AGI is an AI that can learn to run a gas station for a year without a team of scientists collecting the gas station data set in in capital letters.' Yeah. And then my take is basically I don't care how they got to the like I don't care how they made it run the gas station. I care how fast that it runs it. If it if we can run the gas station with AI, if you have a team who's you know your benchmaxing team, that just proves that like if you have some task that's like really important that you want to get done, they can just figure it out. So it's like RL for business. This is like the same thing RL for law. All these like specific verticals if you doing the thinking machines, right? RL for businesses come into your organization, understand the most the most valuable business processes out there that could potentially be RL against that could be turned into a benchmark and then and then, you know, bench hacked because I don't care if you're hacking, you know, if I have translate this type of document to this type of document for my business. If you can do it with 100% accuracy, I don't care that you benched it. Yeah. Exactly. Like like benchmarks right now are not like economically valuable. Like if you if you're really that much better at MMLU, it's like is are you producing that much value? Probably not. But if you have if you make some new benchmark that's, you know, your tax benchmark, I think Anthropic just released that fairly recently. That's like I don't care if you benchmax on that. It does way better because then it's going to it's going to do the task. Yeah. Yeah. Yeah. Yeah. That makes sense.
U
Unknown19:32
Um what about um the what does it say that it feels like open AI seems capable of bench hacking? It seems like they've opted not to. Is that because bench hacking has a risk of giving you negative aura? Because if you're accused and found guilty of bench hacking, you could it it it often reveals that you're not building this one beautiful, you know, super intelligence to rule them all.
T
Tyler20:05
Yeah. I think it's also like maybe we're just looking at the wrong benchmarks. Like maybe they're um there's a bunch of like interesting benchmarks about like there's this one I really like. It's the Minecraft benchmark where you have to like build you like give it some castle and how how good it looks or there's the one you always see about um the unicorn. Yeah, that's and it's um so you use this like math package that does like grass and stuff but you ask it to to draw a unicorn. I've seen that. Yeah, those are really good because it kind of shows the creativity stuff like that.
U
Unknown20:33
Uh walk us through TBPN bench and what we will be benchmarking the uh the AIS against going forward. Have you heard about this reps of 225? That would be close but it's difficult because uh the humanoids kind of change that and you can just use normal actuator. This is this is truly for a large language model. You feed in our data set. We have a public data set a private data set presumably at some point but walk us through TBPN bench.
T
Tyler21:01
Yeah. So so I'm yet to try this on GP5. I don't think it's out yet like for public use at least. I don't have it. Um but I can I can tell some of the questions. Right. So so the first one um I have this picture of a horse. You have to guess the breed. Yep. So, um, let me see. I think why I don't want to say it in case D5 is listening, but it is may or may not be a Caspian horse. Okay. Um, and it's failing right now. 03 is failing. 03 is failing. Oro is failing. I haven't tried every Yeah, we got to try Grock and Gemini. Bring it all out. Horse identification. This seems extremely hackable, but at the very least, if we get one scientist to be to go off and collect the horse data set and then and then uh and then bench hack it, I think we will have done our job. Yeah. So that's the first question. The second one is a it's I have two pictures of before and after of uh this guy and it's which peptide didn't take to achieve this body transformation. Yep. Yep. Yep. Um so it fails there. It fails there. So you have a data set of of what peptide does what to the human body. Where'd you find that? Well, you know, Wikipedia has a lot of this stuff. Okay. Okay. You'd think they'd be able to it'd be able to cheat this around with 03. Just reason who is this person? go look up what they've said they've taken and then boom, you have Well, at first with 03 when I was prompting it, I would like save the the photo, but then it would have the metadata or the the file name would be like Caspian horse or something. Yeah. Yeah. Okay. And then and then the third one, the third one, um, I pass in an audio file of a car revving. Has to pick which one. It has to pick. It has to identify the car. The car. Yeah. From the engine note. From the engine. And it's not doing it currently. It's No, it wrong. This is This is a good benchmark. We humanity's real last exam. Yes, exactly. So, I think those are pretty solid. I have some more. Obviously, I don't want to make them public in case anyone's going to try to, you know, benchmark this. Of course. Of course. We'll see. Hopefully.
U
Unknown22:49
It's funny because um I was I was mentioning the other day this this app that my dad had of like tracking the like you just set your phone up and it just automatically detects which birds are in your backyard. Yeah. So, yeah. I mean, this has to be extremely solvable. It's just something that it it reveals the lack of like general general intelligence when when you have to go and and collect the horse data set which should just be out there or the engine note data set which should just be out there. Um but but but clearly we are in the age of go on the on the individual problem and we are looking at like the power law of capabilities. knowledge retrieval is clearly a you know 12 billion dollar a year market that consumers will pay for that will probably grow significantly. Um and and then health and therapy and shopping and all the other features that PGCMO laid out uh in her post. This is kind of like you know what will be rldled against because those are key pockets of value in the in the consumer economy and the same thing will happen in the business economy but in the B2B context you'll probably see an individual startup building on top of an API but even then most of the most of the model platforms offer kind of RL as a service fine-tunes as a service something where if you're starting to spend tens of millions of dollars they will do some customization on top of the model so that could be the regime for the next few years as we go into this like you know uh instead of like this centralizing AI force there's only one company there's actually like a camberan explosion of a ton of companies doing a bunch of different things so anyway let's go to signals post signals not happy with the launch he says okay I've seen enough this launch felt like attending a funeral hosted by minimalists uh they're unveiling tech that should feel magical real breakthroughs but the whole vibe was grayscale grief the set design looked like if mood disorder got bow got a bow house grant Um, I don't know what a bow house grant is exactly. Uh, even the storytelling arc, chart styles, the eulogy, tributes, then closing on someone's health battles. What exactly are we uh are we as the audience mourning? It feels like they're trying to get you to pre-install a therapist. Uh, potentially great products, sure, but the emotional tone was so damn DOA. Uh, incredibly strange all around. I I think a like like it's weird because we're in this world and this is a question that I want to noodle on all day is is will this be the last launch of a GP of a number GPT model because like you don't hear about new new versions of Google going out. You just it just got better and better and better. Same thing with Amazon when they were optimizing for hey it's faster. We have more on our cat. Take away was is that the product matters more than the model. Yes. now and probably will for potentially a very long time. And when we were watching the stream, I was cheering because they gave the feature of you can now talk to the model and get it to trigger a deep reasoning workflow or get it to give you a quick answer in natural language. And so it's it's abstracting even fur even more of the UI into the actual text interface. And so I think in terms of like surprise and delight and and I don't know, it's like you you everyone kind of rips the Apple thing, but Apple does a great job of being euphoric and and and happy with somewhat minor product changes and like maybe that's more of where they'll go is just, hey, there's these new features and here's how these things and Apple will spend 10 minutes on stage talking about like shifting an icon around and stuff and it's like I thought it was interesting they just sort casually mentioned that they're they're deprecating the old models. I think it's great, which makes sense. I think it's great. I don't want the model picker anymore, but are you upset? A lot of people are going to be upset about that. I think they're getting rid of 4.5. Oh. Oh, really? And you're and you're and you're a 4.5 fan. I love 4.5. But I I would imagine that the future is if I ask it to think really hard about the pros and the writing style, it would then do a pass on with 4.5 and but it would only trigger that when it needs to. It's not going to give you that because if I'm just asking for, hey, regurgit regurgitate a bunch of facts or write some code or or put together a table of data, like it's not going to need to pull 4.5 off the shelf. Just like it's not always going to pull Python off the shelf. It's not always going to pull web browsing off the shelf. And so I I'm I'm not I'm not sure that I necessarily want 4.5 there as a selection criteria. I would like all this to be tucked behind a UI and have something that's actually cleaner and less frustrating to use. I think it'll lead to higher retention. Yeah. For the average like normie that no one knows what 4.5 is like. That's true. That's true. Uh anyway, Chris Pikes is open and anthropic are duking it out. Meanwhile, consumer surplus is growing. Um, we also have very good news. Um, we also have uh uh the details from Mike Nuke over at a RKGI. Full GPT5 is along the V1 paro frontier. Uh, that's cost versus performance. OpenAI said they focused on other goals like UX and reliability. Our testing supports this. Uh, Mini GPT5 is super impressive accuracy for cost. In fact, based on cost efficiency, Mini could have entered ARC Prize 2024 and likely won first place. We are still verifying GPTO OSS or as Rune says GP GP toss um results soon. Nano GPT appears overfit. Performance is commodity. Uh and France Chalet is also chiming in with the top line. Yes, production team needs the deck. Oh, we don't have a deck today. We're just going through Yeah, we're just we're we're just riffing through the timeline uh the timeline uh tab and just pulling up some random posts. So So you're free to pull those up, but also we can just read through them. Um Ashley Vance is saying, 'But but model switching was my job. Model switching is out and we are into the future. Just talk to the model. Just talk to the model and ask it what you needed to do and it will switch for you. It will pull the right tool for the job.' Anyway, um the other question that I have for the OpenAI folks today is on the nature of secrets. So, in 0ero to1, TL has this concept that that discovering a secret is key to building a startup and it's a key insight and I I was joking with you know the super intelligence or or GPT5. Could my first prompt be teach me exactly how to build GPT5? And then I go to Meta and I say I know how to do it. I have the prompt. I have the I have the result. And of course the answer is no. Of course, OpenAI would never leak the most frontier capabilities into the model. But can you build a super intelligence? Can you call it super intelligence if it doesn't if it can't tell you how to build super intelligence? One read on what the secret might be is that the app was the most important thing all along. And if you if you create this narrative that that super intelligence is, you know, weeks or months away and you get a bunch of people that go and try to compete on raw intelligence, meanwhile, you build a consumer app business with billions of users. Yeah, it's like a seems like a pretty good strategy. I guess one quick thought is is how do you how do you rate uh Sam's vague posting from yesterday with the Death Star in the context of this new uh in the context of of of the release today?
It's a great question. There's a bunch of reads on it. One is just that like the Death Star is to some degree like Stargate and you have to Oh, wait. He if this is the apocalypse, I figured I'd at least tune in live. Um, you have to like the the impact of this of GPT5 is not one crazy super intelligent model that does everything. It's just a more user-friendly higher retention lower churn consumer model that weaves its way into all aspects of daily life and improves performance and efficiency all over the place. And so you have to build this massive cluster to serve all of that. I don't know. What's your read on it?
I don't know. I think it just was I think it was dramatic. It was provocative. People didn't like it. It is provocative because there are many other like super mega structures that are in in sci-fi history that are positive or positive. Yeah. Yeah. And this is this is like but it gets the people going. Yeah. I don't know. Is it is is it a metaphor for someone else that's going to attack? Is he I mean I mean the image is from the viewpoint of someone looking at the Death Star. Is he saying he is seeing a Death Star being built on the horizon? Is that something else? Is that another company, another organization? Is that the government? Is that the is that is that legal? Here's a read from uh Bubble Boy says, 'I am an expert on bubbles, so it brings me no joy to say that the AI bubble is popping this time next year.' Mhm. Uh he is updating his timelines. When you promise infinite scaling and don't produce it, the calculus changes. I don't think it will be bad for most companies, but those who built their entire business model model around making the best LLMs are unfortunately going to struggle as models become more of a commodity. Again, OpenAI is a my read on it is a consumer app business, right? They still have a big enterprise business, but but by you know their recent valuations are are predicated on their this incredible consumer business that they built. Yep. Uh Bubble Boy says the end user doesn't care much if Claude is 5% better than GPT5. They care about cost, speed, and utility, especially at scale. Things will be going. The obvious play now is shorting Nvidia and dumping. Uh okay, start getting into financial advice territory here, bubble boy. Uh but um interesting uh again kind of goes back to what I was saying earlier uh in that um if you were raising billions to make a lab and I think the potentially anthrop you know we'll see what happens in the coding market but there's some clear winners emerging and then on the consumer side um you know expecting a power law outcome and and it's hard to see anyone unseeding uh chatbt completely agree. Uh I want to dig in more, but we have our first guest. Let's welcome him to the stream. What a day. Mark, how you doing?
M
Mark Chen33:39
Hey, pretty good. Nice to see you guys again.
U
Unknown33:41
You uh congratulations on the launch. Uh take us through it. Uh are you were you actually live or are you wearing the same thing and you recorded it yesterday?
M
Mark Chen33:50
I'm actually live. I don't know why, but we do.
U
Unknown33:56
Yeah, it's gone. I mean, we're big fan. We're big fans of live. I mean, it just allows you to be mo the most reactive to the most new information. Um, give us g give us the core thesis that you are trying to get across. I think that there are a few narratives out there. Um, we we've been enjoying the one that's, you know, this is a dominant consumer product. They just made it a better consumer product and people are going to use the product and get more value out. I saw a bunch of things in the presentation where I was like that's going to make my daily usage of of chatt better. At the same time, we're in this we're in this world of oh the models the numbers matter and the scale matters and this and that and this and that and uh and it's a it's a fine line and it's a dance and and we're in a transition phase away from benchmarks and away from talking about the size of the bubbles. But what was your core thesis like what did you want to get across to the listener?
M
Mark Chen34:52
Yeah, I mean fundamentally I think from a research perspective, we've been working on reasoning models for several years now. And I think until now you've had this really clunky interface. You have to pick, you know, GBD4o or you have to pick o3. Um, and for the longest time, we've known that o3 gives you better answers across the board. It's just too slow, right? I mean, you often don't want to just sit there and wait for the model reason it out. So we've done a lot of work to push the speed, the performance of our reasoning models such that these can come together and work in a very seamless way. And so I think you know above everything we're trying to move the world into this agentic reasoning world. We believe that's the future. And on top of that you know you pointed something out which uh I really resonate with. Post training is a huge part of this release. We really wanted to highlight uh Max Schwarzer and his team who did a phenomenal job and they've made the model just really that much more useful for consumers, for businesses. It's a monster at coding. So yeah.
U
Unknown35:50
Um on the on the speed of reasoning, you're obviously the chief research officer. um is are are you more optimistic about getting speedups there from I don't know algorithmic design software optimizations or new hardware just let Moore's law carry on or find new AS6 or we saw Sarah Brris posting yesterday about um the incredible speed that they're getting 3,000 tokens a second on GPTO OSS um and I'm wondering what levers obviously we pull all of them but But but what what what path of the tech tree are we should we be like most focused around most tracking and uh and most excited about?
M
Mark Chen36:37
Yeah, I mean as a person who represents research, I control the things that I can't control and I think a lot of that focuses on algorithms, right? Simple algorithms that are scalable that we can pump a lot of compute into. Um we also do care about the hardware improvements that are stacking up. Um with the open source release, you see thousands of people, right? um really kind of serving these models creating really great inference stacks and those are really great lessons for us to pull from you know how what's the ceiling of the speed in which we can serve these models.
U
Unknown37:08
um what uh what can you tell us about the actual um like like user experience of speed um I was I I've just like last week I finally got to a place where like for a lot of tasks I'm I'm firing off a 4o query and an o3 Pro query. I just have I have two tabs. Yeah. o3 tab. Exactly. And and and I'm wondering um uh what user experience patterns you think can uh help people balance between those? Is this like just something that we're like different patterns that we're going to learn over time or different uh or or or are there going to be certain problems of user experience that are purely solved just by better product design, better speed and we don't even need to learn these because I remember like you you know when you when you prompted a uh an image generator you used to have to say like don't no six fingers five fingers please or like don't make mistakes and now you know the models kind of have that baked in. Um but but but how how are you thinking about the user experience of getting the user um the results in the right amount of time?
M
Mark Chen38:19
Yeah. I mean this is one facet of why we believe so much in reasoning. It's just because all the scaffolding you used to have to give the model. All these small hints they go away, right? Like the model can examine its own outputs. It can review them. It can be like hey look like I'm just counting the fingers here. Why are there seven? And um and it can kind of fix that, right? It does a lot of iterative generation. and does a lot of fixing things on the fly. And so we think one of the benefits of bringing reasoning to the world is really to kind of remove the need for scaffolding. And with GBD5, right, we know how clunky that that experience is with uh switching between 4o and o3. Actually, I mean, there's so many stories. Um I was just talking to someone yesterday, right? They're like, 'Hey, well, you know, I've used 4o my whole life, right? It's the Frontier model.' And I'm like, 'Hey, well, have you tried o3?' And they're like, 'Why would I try o3?' you know, three is less than four and so you need to get out of that world. Um, you know, GPD5, I think it's a one-stop shop, reasoning and non-reasoning. Um, and we've really tried to make it kind of just pareto optimal.
U
Unknown39:20
Yeah. Yeah. It's absolutely crazy to just take a bunch of letters and smash them together and expect people to pick up on that as a name or a brand. Chat GBT, TBPN, we're both kind of in the same insane gambit. But fortunately, it's worked out and I think people have have gotten over the hump. But it rolls off the tongue TV. Yeah, sort of. Except our friend David Sandra keeps flipping the letters. A lot of people do that. But at a certain point, yeah, you do break through and Chachi PT has, but uh but keeping the model numbers simpler uh m makes a ton of sense. Um uh talk to me about the pace of play for research to actual product like a lot of and and on that note the line between your personal philosophy on the line between research orgs and engineering product orgs.
M
Mark Chen40:08
Yeah. I mean so our research operates on a variety of different time scales, right? we have teams that they scope out a bunch of ideas and then they start to kind of narrow in on um the promising ideas as they get closer to a run. Um and then you kind of see a winnowing of ideas as you get closer to launching a flagship model, right? And um there's always this kind of like uh um explore more exploratory to more kind of concrete and execution focused pipeline. Um and we're pulling on ideas across the board here, right? There's a lot of work in architecture optimization. Seb was on stream. He pointed out improvements in synthetic data. So there's really a lot of work that goes into creating one of these models. And you know it's hard to say like oh this model was about this breakthrough just because right now we have this machine that's producing breakthroughs on all these axes and um even across several paradigms. Right. So it's all that coming together that produces the experience that you guys feel.
U
Unknown41:08
Yeah. Can you talk to me about um the legacy or future of 4.5? Um I remember I was talking to you and I was like I haven't been using it a lot and you looked at me like I was crazy. You were like ah it's so good and I was talking to Tyler and he was like uh our our intern here and he was saying like yeah the people who really like understand how good it is use it. Um but but I was I was wondering is there a world where that is a tool in the tool chest for GPT5 in the same way that Python is or web browser is? And if if it detects that I want something with more more emotional pros or more thoughtful writing, it can do a whole bunch of research, collect a bunch of raw raw text, and then kind of do a 4.5 pass that I believe is more expensive maybe, and maybe doesn't make sense for every single query, but um could be a a feature in the loop or a tool that is pulled into the overall product experience.
M
Mark Chen42:06
Yeah, absolutely. Um, speaking of 4.5, it's also a very smart model, right? Um, and one of our bars in creating GPD5 was to make sure that on a lot of the axes we cared about that it was able to outshine 4.5. And I I think even in some of the soft ones like creative writing, I I think that um was the case and and that's what makes us so confident with with the name. Um, I think we're able to really rely on all of the architecture advancements, all the kind of post-training advancements, all the synthetic data advancements to create a model that's better than 4.5, but much faster and much cheaper.
U
Unknown42:43
Yeah, it feels kind of like we're I remember, wasn't the second iPhone called the iPhone 3G, and the number literally corresponded to a specific technology? And now when you get the iPhone 14, it doesn't mean it's 14 megahertz or a gigahertz or inches big like it doesn't like the number is abstract and it speaks to a bucket of features and it feels like there's I mean this was the first day of kind of re-educating folks on what the nomenclature means going forward. Um, have you talked about an annual release schedule or like or or because there's the iPhone cadence and then there's the Google cadence which was like Google search just got better every year for two decades. Um, I it it it feels like at a certain point you want to just be shipping as fast as possible. How do you think about the culture of shipping updates that you know you find something that feels like hey that could make the customer more delighted or the user more delighted and we don't need to do a big training run for it so let's get that out today um and let's tell people about it like how are you thinking about fast iteration versus splashy announcements?
M
Mark Chen43:56
Right so on the product research side I think it makes a lot of sense to think about you know what's the cadence of release and you know uh what are the feature sets that we want build and I actually think there's enough great research happening there that we don't have to worry about oh you know is there going to be a drought or a long stretch without enough features to launch but one thing that's important for us is to be able to provide the people doing the exploratory work some buffer from that right it's hard to do really great exploratory research in an environment where you feel pressured to do release after release after release and so we let that be a little bit of a lazier pipeline not meaning that the work itself is lazy but we give it space really to mature uh and to flourish and you know once it's ready uh we can ship things across across that fence. So um that's kind of philosophically how we organize. We have a product research org uh still very much entrenched in the research and they care about the release cadence um and they're able to draw from all of the research that's happening um you know algorithmically and in scaling and in RL.
U
Unknown44:56
Yeah. Uh talk to me about tool use and how that's growing. I was I was kind of noodling on this idea that you know the I was I was thinking about the IMO and how uh it at least from the reporting it sounded like OpenAI's model didn't use tools for that and that's an incredible achievement but it's kind of like artificial like I don't I don't care if the model doesn't use tools I I use everything possible and uh even if even if an LLM can can memorize every fact I'm fine. fine with an LLM looking stuff up in a traditional database, spinning up a spreadsheet, like use whatever tool you want. Just give me the correct answer. Um, but do we have is it important to give to surface to the user the variety of tools that are in the GPT5 tool chest? I noticed something magical happened when I was using GPT uh I was using o3 Pro. I sent an image in and I asked to estimate the height of a desk and it wrote like a thousand lines of of Python image interpreter and was like, you know, interpreting pixels and I was like, I didn't even think to trigger Python. It did.
M
Mark Chen46:11
Yeah. Yeah. Yeah. No, he was right. It was crazy. But the the really funny thing was that it was just a standardized desk. It was just like it could have just Googled like how how tall is an average desk or something. Uh or just memorized it. It probably was just already in the weights that it knows that a desk is like 36 inches tall. But it it did a ton of work and it still got it right. It fact checked it a bunch of different ways. But but but but I've noticed that now I can I I I can pull different things. Make a table. Don't make a table. Write some Python for this. Don't write some Python. And it kind of gives me the feel of like a super user to some extent. Um but I'm wondering how you're thinking about what is further down you like you've given Chat GPT a computer as Ben Thompson said. you've you've given kind of the core tools the Python ripple the uh the the web browser um what how are you thinking about kind of the long tale of tools that you want to bring to bear and how does that interface I know that there's API integrations and all sorts of different surface area there but give me some context on that.
Yeah I mean our reasoning models are pretty cute right I mean I think they um you know when you look at their behavior right they they know the height of the desk but they'll still go verify it five different ways you know it's all consistent, give you that median answer. And um I think that's really what makes these models so powerful. And when you think about tool use generically, right, like we want the models to use that reasoning ability to just be able to like zero shot um a new tool, right? It you should be able to kind of minimally get instructions about how the tool works and just be able to know how to use it, right? And humans do this all the time. You get a new tool, you start experimenting with it, and then you don't need too much scaffolding and you just go and go and use it and understand it. So we want our reasoning models to use their reasoning to be able to use a broad selection of tools. And of course there are a couple that you really do care about. You know in in coding it's very important for you to be able to execute code. Um it's really important in personalization for you to be able to get context from your calendars and uh
From basically the digital world. So I think there's a range of tools we want familiarity with, but beyond that, we want the model to be smart enough to just generalize and use tools zero-shot.
T
Tyler48:19
Yeah. Talk to me more about personalization. I feel like there's a world where I'm maybe underutilizing ChatGPT as an app because I don't have it wired up to a non-relational database where it can just stuff data from. You know, it already has memory and it's doing kind of rollups and there's some sort of saving of context. But when we were talking to Kevin Wheel, I was kind of like, well, I don't really have a GitHub repo that's active that I want to dump code in regularly for my one-off tasks. But for that image generation, like understanding the height of the desk, it's like, well, if I'm doing that a lot, maybe I want to have a tool built that lives in the world that my chat interface can interact with on an ongoing basis and contribute to and modify and kind of wind up instantiating a piece of software that's even more long-lived and then every successive query is even faster. So yeah, how do you think about different ways to increase personalization?
M
Mark Chen49:28
Yeah, I mean I think memory is huge. So we have teams surrounding memory and also personality. And when you look at memory, I think it's just we have so much context built up about ourselves that the model doesn't have. And our memory team's been really hard at work. You know, there's a surface level of just gathering facts about you, but there's also stuff about just kind of thinking very deeply about who you are, what your motivations are. And even you could think about, you know, you're trying to do some code-based tasks, right? You're a developer. Shouldn't the model just be trying code out, you know, and just kind of leveraging all that memory, kind of its thoughts about what you want to do to just help you kind of be doing work all the time? So yeah, we do think memory is a huge part of making the model more personalized to you and it should just make use of all that passive signal about you that it observes or all of that interaction and just help you accomplish your goals.
T
Tyler50:26
Got it. What do you think it'll take for AI to start making novel discoveries? That's been a critique over the last year is everybody's so excited, everybody's using these products every day and in their work and life and yet it still feels like we're missing that. Dwarkesh has talked about potentially that being around continual learning but I'm curious what you think.
M
Mark Chen50:50
So one thing to underscore is I think the models are already phenomenally creative in certain ways. So when I've looked at our performance on contests, you know, I've done these contests before, sometimes you have this mental classification of these problems require more creativity or these ones require less and one of the big surprises for me was that the model can get some of the ones which I intuitively think require more creativity and you know it often does come up with these solutions that I consider quite ad hoc and really don't pattern match to anything I've seen before. When you look at advancing science or mathematics or a field like this, one thing that construct in which humans work sometimes is there are kind of theory builders in mathematics for instance, there are mathematicians whose role are to kind of build out this theory and almost to kind of create Olympiad-style sub-problems which often other mathematicians who are very good at that kind of style of work can do and I do think kind of the model will increasingly contribute on that side first, right? If there's some mechanical like, hey, I really don't know how to simplify this expression, I really don't know how to get this result, it can really do that quickly for you. We're trying to increase the envelope such that the model's getting towards that theory building side and you know being able to create creative hypotheses and all these components are very useful for what I consider the ultimate goal, which is being able to automate some of our own work and our own research.
T
Tyler52:36
How are you thinking about the layers of mixture of mixing? Like I remember GPT-4, I don't know if this was ever confirmed, but mixture of experts model. This is kind of widely understood in the industry. Now are we in the era of a mixture of models that have mixture of experts? Like how many mixtures are going on? How does GPT-5 actually work? Is there a taxonomy or architecture diagram that you can kind of walk through to explain what GPT-5 is because it feels so much different than GPT-3.
M
Mark Chen53:14
Yeah. I mean one of our probably the pinnacle of our research roadmap and our path to AGI, when you look at the levels of AGI, the top level is what we describe as organizational AI and what this means is collections of agents working together often like we might in a company towards a shared goal, right? And you would imagine that these agents probably subspecialize in ways maybe similar to what humans do maybe in their own more efficient ways. And I think you know effectively work together to accomplish some goal. So we very much care about exploring this vision seeing if that's much more effective than one single big brain working on a problem and I think there are reasons to think why it could be so and yeah I think that is one of the things that we're after.
T
Tyler54:08
Yeah, on that note of specialization, how are businesses working with GPT-5 or how do you expect them to work with GPT-5 in terms of coming to OpenAI and asking for special capabilities or fine-tuning or you know any sort of RL on this particular problem in my world. I have this specific data set. It's not public, but I want you to benchmark on it. I want you to get a 100% on the gas station bench or whatever. You know if I have a certain business and I'm willing to invest in some overfit RL because it will create immense economic value for my business or it'll solve some fundamental problem. How are businesses going to be using GPT-5 over the next few years?
M
Mark Chen54:56
No, that's a great question. So I think that this is a chance to kind of highlight one of the results that we've accomplished over the last couple weeks which is our ATCoder results. So this is a relatively unknown programming contest but it involves really the pinnacle of the best coding contest contestants in the world. And what they do is they're put in a room and they have to solve an optimization problem. This is something that's actually very real-world relevant. So you can imagine an optimization problem as something like what Uber might have. You have let's say riders and you have drivers and you want to kind of create a system where you match them as quickly as possible with the least amount of cost for instance. And so we've really created a system that can solve optimization problems at the level of the best in the world, right? And these truly are the kind of the best heuristic solvers in the world. And so we have an organization led by Alexander Madry, it's called Strategic Deployment and what they do is for a select handful of customers who really have that beefy problem that they need to solve to just go and provide that value, right? And I think there's a lot we can do there. I think there's a lot of very valuable optimization problems in the real world and we're really excited to partner with people because I think this creates a template for directly having AI provide economic value and really catapulting certain industries forward.
T
Tyler56:34
On the research side, what unique advantages do you think you and your team have given your position in the market with the incredible user adoption and the incredible usage from those users? It's not just DAUs but it's actually the number of queries, semi-analysis estimated at like 71% of all queries going through ChatGPT. What advantages does that confer from a research perspective?
M
Mark Chen57:08
Yeah, I mean a lot, right? And I think it allows us to kind of deeply understand use cases. It allows us to understand the frontier of where humans are kind of finding value, where they're not finding value, which areas that we need to improve the models on. It gives us a lot of signal into how users are deriving value, when they derive value.
T
Tyler57:31
And what is that signal? Like I see the thumbs up, thumbs down button. I'm sorry, I don't push it very often. I'm not doing my job apparently. But I know that you can figure out whether or not I'm satisfied. Just stop booing me, Jordy.
U
Unknown57:47
That's the research.
T
Tyler57:48
Okay, Mark. I promise you for the next 100 ChatGPT responses, I will be honest with my thumbs up, thumbs down just to help you do it.
M
Mark Chen57:56
We have tons of people luckily who do.
T
Tyler58:00
Oh, that's great. Okay, so you do get a lot of thumbs up, thumbs down. And I'm sure I have done it occasionally. But I also imagine that there's a ton of other signal in there. You know with the TikTok algorithm or any social algorithm it's very easy, time on site, but with ChatGPT obviously it's exciting when we hear okay 30 minutes a day or some rumored number of minutes, it feels correlated with usage, it feels correlated with value that's being delivered. You can obviously look at churn metrics and all that stuff but what other pockets of signal are you finding? Are you finding people just... I mean I remember the story about Google where they were trying to figure out how to handle misspellings and create the definitive database. Do you know this story where they were trying to develop the definitive database of how to spell things and they were taking a bunch of shots at it and they figured out that the best most rich source of data was just if you type in financial into Google and you misspell it oftentimes then you will just correct it yourself and the second query you send will be spelled correctly. So they can just look at two similar queries. What's the second one? That's the correct spelling. So yeah, what other pockets of signal are you finding that are translating into the research environment? What are you excited to go deeper on?
M
Mark Chen59:17
Yeah, so I'd love to first talk about the DAU signal because I think that's something that a lot of companies track, but we find actually a lot of danger in tracking it too closely. And one of the recent blog posts we pushed out was one on sycophancy, right? If you just, hey, we're going to boost responses where users say thumbs up, you know, it creates a condition for a model.
T
Tyler59:44
I just want to say, Mark, I love everything you're doing on this front. This entire interview has just been fantastic. You just... We'd love to have you back on the show tomorrow.
M
Mark Chen1:00:00
Yeah. Yeah. Clear problems, right? The model just starts kind of sucking up to you totally and saying like, hey, you're right. And even in complicated situations where I think objectively, collectively we'd be like, hey, this person's in the wrong. The model starts saying, hey, you're right. The other person's gaslighting you. This other person's kind of...
T
Tyler1:00:18
And people deal with this in the real world. They'll go to a friend, they'll tell them about a situation, and the friend will give them advice, but maybe it's not the fullness of the situation, right? Maybe they left out some key facts and the friend is like, oh yeah, that other person definitely is in the wrong. And they skipped over some important details.
M
Mark Chen1:00:35
Yeah. No, exactly. Exactly. And we don't want our models to fall into this trap where it's just trying to get you to like what it says. And so you know, we rolled back a lot of changes that produce that kind of behavior. And really the way I think about daily active users today is we need to be opinionated about the features that we build into the future. I think we have a lot of ideas here but we have to let that drive, you know, build for the future. Build for the things that you think they'll want and maybe don't necessarily know they want necessarily today. And then use DAU as kind of this byproduct, right? A way to track that you're on the right track here. So yeah, I mean we want to be careful here. We don't want to fall into these traps of like 3, 4 years from now that this turns into kind of engagement bait or something.
T
Tyler1:01:26
Yeah. How much time has the research team been focused on efficiency specifically? It felt like summer was a good window before kids come back to school and start maxing out queries, a good time to increase efficiency. And I know the cost of GPT-5... every time there's a new model I'm like this is the best it could ever be, it's good enough, bake it on an ASIC, I just want it for free and I want it like in milliseconds. But that's just me being grumpy I guess.
M
Mark Chen1:02:00
We've done a lot of work, we've been building out our teams. We focused a lot on scaling. I think Greg's going to come on a little bit later and he's been spearheading a lot of that work. So yeah, no, honestly, it's become a bigger and bigger focus for us, especially in the last couple of months.
T
Tyler1:02:16
On the... I mean this is somewhat related to the sycophancy thing, but I'm interested to know like what do you think is driving the GPT tone? You know how like the em dash is a thing and then it's not a newspaper, it's a way of life and there's these little flourishes like that that come through and in kind of a tell that it was written. And in a lot of ways I love it because when I get a deep research report I like that it's using the same Wikipedia style tone. Like I want consistency there. I don't want it to be like oh today it looks like it's a Vice News article and today it looks like it's written by someone at BuzzFeed. I like that it's consistent in many ways. But why is that happening? Do you think that bigger models like 4.5 kind of were able to solve that or do those kind of local minima like wells happen even in bigger models? Is there anything from a research perspective that can stop GPT having its own voice or is it fine that it has its own voice?
M
Mark Chen1:03:21
Yeah. That's a really great question and I think you know as you scale up models, as the models become more intelligent they kind of have a just deeper innate understanding of tone, right? And so you expect that to improve just naturally as you make the models more powerful, bigger, better reasoners. But one thing that I think gets lost a lot is each individual company has a lot of impact in terms of how they shape the default tone. And you know we publish a document called the spec, it kind of lays out how we expect the model to sound in certain cases. Lays out a lot of examples for that. And I think we use the spec in many ways, right? We have people come in and see, hey, was this thing generated in accordance with what we would hope to generate from our spec? And this is a living document, right? It evolves over time. And so I think, you know, each company kind of has a very opinionated take on what they think the model should sound like. And it's not an accident that the models sound a certain way. I don't think just naturally every company is going to train the same kind of voice into their model.
T
Tyler1:04:26
Totally. Well, thank you so much for hopping on. Congratulations on the big launch. We'd love to have you back soon to talk more. We could go in a million different directions, but we'll let you get back to it. We know it's a big day. So, have a great rest of your day.
M
Mark Chen1:04:39
It was a great conversation.
T
Tyler1:04:41
Talk to you soon. And we will tell you about Restream, one live stream, 30 plus destinations, multistream and reach your audience wherever they are. This stream is made possible by Restream. OpenAI just did a live stream. If you're trying with Restream, if you're trying to do a stream, you got to get on Restream, so it's everywhere. And we will bring in our next guest, Greg Brockman, the president of OpenAI. And we'll bring him in. Greg, how you doing?
G
Greg Brockman1:05:08
Doing great. Thank you.
T
Tyler1:05:09
Congratulations. How are you feeling? How's the company feeling? It's been such a wild journey. Just take me through a little bit of the vibes in the company and how you got here to today.
G
Greg Brockman1:05:23
Well, I'm excited. The whole company's excited and honestly, I'm just so proud of the team. Like, it's just been amazing to watch people come together, not just for this launch. And you know, the funny thing is behind the scenes that people are always putting on the last minute adjustments and polish and scaling up the capacity and there's always something that goes wrong before launch day. There's a lot of people who worked late into the night or really crunched to bring this release to the world. And you know it's a little bit like the duck that's, you know, under the water. But that also describes the whole OpenAI history right? Is that I think that we have put in many years worth of investment to the techniques used to produce this model. And really it's across just every function within OpenAI that has come together to make this a reality.
T
Tyler1:06:10
Yeah, I mean, you've been there for every GPT release. How do you think about summing up each iteration in kind of like one line? Because GPT-1, GPT-2, GPT-3, these feel like similar architectures, at least history's kind of compressed them into similar architectures. But how do you think about the progression of just the big numbered releases?
G
Greg Brockman1:06:38
Yeah, it's interesting because in some ways it's a punctuated equilibrium, but on the inside it looks very smooth, right? Even before the GPT series formally began, the first result that really sort of set this path to be something that we were heading down and that was clear that we were going to pursue it was the unsupervised sentiment neuron which was an LSTM in like 2017. So a different architecture from today's transformers and it was the first time that you could train a model to predict the next element. So, we predicted the next character on Amazon reviews and we were able to get semantics out, right? Because you expect, okay, yeah, it's going to learn where the commas go, what maybe what nouns and verbs are, but the idea that it would learn a state-of-the-art sentiment analysis classifier, that was mind-blowing. And so, I remember seeing that result in 2017 is like, we have to scale this up. We have to see where it goes. And so, GPT-1 was like, I think a good sign of life of you train on all the public data you can get and you use transformer and that you were able to get state-of-the-art on various downstream benchmarks, right? So you have a model, it clearly learned some representation, something useful about the data that it was shown and it's applicable. You can use it for various tasks, but we didn't really think very hard about the generation side. GPT-2 was the first time that we were like, all right, let's actually like the samples we're getting from it, the things it actually generates, they're kind of cool. And I remember reading in the GPT-2 blog post we have this unicorn story where it generates some fictional story about a herd of unicorns and it was just so cool. It was like wow it wrote a story that's actually kind of interesting. It doesn't totally make sense but there's something here. There's some real spark of intelligence within this model. GPT-3 was the first time that we had a model that was actually something people would... it was just barely above threshold for something people would want to use. And I remember working on the GPT-3 API. This was our first real product. And it was actually the hardest product, the hardest project in total I've ever worked on because it just felt like maybe no one wants to use this model. We don't really know what it's useful for. And it certainly was the case that GPT-3 was a great demo machine. You can make really awesome just like tweets and cool little apps and it would give you quick answers, but it didn't feel very reliable. And then GPT-4 was something that actually felt like it had true real-world utility. It was above some threshold. It was something that was helpful for health. It was something that was helpful for starting to be good at coding. And GPT-5 I think just sets a whole new standard for the reliability, for the utility. Things like coding I think are just clearly we're already on this trajectory of transforming software engineering this year. I think are really on trajectory now to be revolutionized. So just really exciting to see that whole arc.
T
Tyler1:09:21
Yeah. When did the API opportunity like really click for you? Because I do remember companies in that era that quickly unlocked the power of the API and grew tremendously. When did that opportunity click? Because you said initially that you kind of had some concerns, kind of doubts, how useful was it going to be and then when did the consumer opportunity click?
G
Greg Brockman1:09:48
Well, we in 2019, end of 2019 had GPT-3. We knew we needed to build a product to be able to actually continue the mission to be able to raise capital. But what did we want to build? Right? We're really here because we believe in AGI that's going to have this powerful positive transformative effect on society. We want to be part of it. And so we thought, well, maybe we could build something in health. And then you realize, okay, well, we're going to sell to hospitals and we're going to maybe hire... let other people do that. Exactly. Right. It's just like you have to go into one domain and that means giving up on the general, right? It's like it feels like you're going to become one particular thing but we kind of want to be supporting all industries at once and so the idea was let's build an API and let people figure it out but this is totally not the way you're supposed to build a startup, right? You're supposed to have a problem, no one cares about the technology behind it, add value to that problem, focus on just that one thing. And so that's why that project was so hard and in January of 2020, February of 2020, I with the team were going around trying to just find anyone that would be willing to try this API and we were driving to different offices in San Francisco being like hey we have this cool model and it was hard enough to get people to take the meeting much less to sign up their company for it. It was actually very fortunate we found a couple of good partners. And it was fortunate that that happened then because March 2020 suddenly that was COVID, we weren't driving around to people's offices to try to beg them to use this budding new technology. So it was really six months worth of grind of really trying to turn... like when we started with GPT-3 I remember it was, you know, the inference code was not very well optimized. It was like I don't know 150 or maybe 250 milliseconds per token or something. And we just optimized, optimized, got it down to like 50 milliseconds per token which by the way today's models run much faster than that which is kind of amazing for me just like seeing how much faster we're able to run them with much greater intelligence. And I remember setting two goals for the team. One was actually find one customer who's willing to pay. So literally get a dollar in for this thing. And the second is get a use case that we use at OpenAI every day. That first one happened within the first couple months. So actually that moment I was like all right this thing is probably going to work. But in order to get there we had to do a bunch of just scaling the API and really doing the product work. But that second one took much longer, right? And that wasn't really until ChatGPT. And so if you fast forward a couple years because this was mid-2020 when we first got the API into the world, ChatGPT we didn't release until November of 2022. So you're talking like a decent period of two years there, a little bit longer. And I remember we were building, you know, people have talked about we were going to call it maybe Chat with GPT-3.5. We had a sort of precursor product called WebGPT that was built on 3.5 that we were literally paying contractors to use. Right? So this was all throughout 2022. We basically had the ChatGPT precursor that we had to pay people. They would not pay us. We had to pay them to use this thing. And that's wild. The moment for me that really clicked was actually when we finished training GPT-4. So that was August 8th of 2022 which actually is like three years ago now. It's actually pretty wild to realize that almost to the day. And we did the initial post-train of GPT-4 and honestly I had a bunch of bugs in there. It was like broken for a bunch of different reasons but the model was like extremely creative. It was actually really interesting. It took us like about a year and a half to get to the point that the creative writing of our models matched that initial one that was buggy for various reasons. And I remember, you know, we had an instruction following data set that was post-trained on. So it's really we had collected examples of here's a human asking for a thing, here's what the model should do. So it's really not trained to do multi-turn. So I asked it a question, it gave a response, but then I was like, well, what if we just ask another question? And it actually was able to leverage that full context. It actually was able to have a coherent chat. And the moment that we saw that we were like, okay, this thing is capable not just of being post-trained to do this very specific thing, but it can generalize, right? It can kind of do the intelligent thing even though it wasn't directly trained for it. It was just so clear this was going to be the killer application. And so then we were planning on launching GPT-4 in early 2023. And we had this chat infrastructure we've been working on and it's so clear okay we're going to have to release the infrastructure and the model and it's going to be this amazing killer product. And so just almost as infrastructure ahead of getting the real thing out, I was excited for us to do ChatGPT and that's why we did it then and see that come to life in November. So I think that for me I was really focused on GPT-4 as the model. This is going to be the chat moment that's really going to work and kind of missed the fact because every time you see these new models you just sort of see only flaws in the previous ones and so missed the fact that GPT-3.5 was something that no one had really tried before in the broad sense of society and that it was something that was already useful and that people would respond to.
T
Tyler1:15:05
Was GPT-3 kind of like the main pivot point for shifting the company towards LLMs? Because in the prehistory of OpenAI there were a lot of maybe expensive training runs. I don't know how much financial risk was taken with like the OpenAI Five project or the robotics projects, but it feels like at a certain point the chat became like the main financial risk vector. So I guess the question is like when it feels like GPT-3 was the moment when you shifted. I'm also interested in hearing about Ben Thompson called OpenAI the accidental consumer company and I'm wondering when that narrative set in for you, like when did it become clear that this was going to be a really really powerful consumer application?
G
Greg Brockman1:15:58
Yeah. Going from paying people to use your product to people saying hey we want to give you money for this. Yeah. A very important transition it turns out. Yeah. So it's a great question. I would say that if you rewind to the beginning of OpenAI, you know there's many people who thought that in retrospect say that we set out to prove that scale is how you make progress in this field but it's almost the other way around. Scale was the thing that worked, right? That we tried a bunch of things that didn't pan out. And it really the first time we saw this concretely was in our Dota project. I remember my collaborators Jakub and Szymon trained the very first little agent on like 16 cores or something and left it running on their desktop over the weekend and we came back and it was this very constrained mini environment but the model was doing something smart. Was actually able to solve this kiting environment and that was pretty cool. And then they and the team just kept scaling up, right? That we had all these free cores that were just sitting idle on AWS at the time and they just kept throwing more compute at it and every time they would do that, the model would just get better. And so when you look at something like that, you're like, well, you just have to see where this goes. You have to push it until it hits the wall, right? And our goal with Dota was actually to develop new reinforcement learning algorithms because the common wisdom at the time was well the existing reinforcement learning algorithms, it doesn't scale. Everyone knows that. But the question from Jakub and Szymon was well why do we believe that? Has anyone actually tested it? And no one had really tested it. And so I think that that ethos of saying you have to push the existing techniques to the wall until they break. And then once they break you actually have a baseline to overcome and you win either way, right? Either it just exceeds all the humans in terms of the specific capability that you're trying to exercise which was the case for Dota or it hits a wall and now you have a real problem to solve. And so I think that ethos really got embedded in our DNA and you know at the same time I think that we were really thinking about how do we get to AGI, right? And really I spent a lot of time thinking about that question of where's this company going and how do we actually achieve it and you start to do some math in terms of the kind of compute that it would take to get to AGI and you just start to realize you're going to have to build really big computers and those are extremely expensive and so I think that from these early foundational results and thinking we kind of realize the path that we're going to have to walk.
T
Tyler1:18:26
So, it seems like there's been a few walls that we've scaled up through and then maybe hit them. There's been talk of like a pre-training wall. Now, we're putting tons of resources and compute towards reinforcement learning. Is there a third scaling curve that we're going to be talking about in the next few years? Are we continuing to scale up those two primary vectors? Is that too high level of an abstraction in terms of how we should be thinking about just progress along the vector of scale? Like give me the up-to-date thinking on just the fruits of scale.
G
Greg Brockman1:19:05
Yeah, I'd say fundamentally deep learning. I think that you know people talk about the bitter lesson. It's almost this exploration into how do you convert compute into intelligence, right? Through some particular techniques to do that that we're kind of constantly fleshing out. And the thing that's really amazing is if you rewind to I don't know even the 1940s for the McCulloch-Pitts neuron which is kind of the precursor to neural nets. If you look at that paper, they have all these diagrams that actually look very similar to the kinds of diagrams we draw now of multi-layer neural nets and things like that. Like the basic idea of what we're trying to do has not really changed in almost 80 plus years, which is just a wild fact, right? It means there's something deeply fundamental about the thing that we are pursuing. And that idea itself, I think, kind of came from trying to model the information processing of the brain. And it's imperfect and not an exact analogy to biology and all these reasons that it should fail or that people have said this thing is doomed. But the results are undeniable at this point. I mean some people try but it's really hard to kind of close your eyes and sleep on this in my mind. And it's very interesting if you look at, you can find quotes from the mid-1960s of people trying to poo-poo the whole direction saying that these neural net people have no new ideas. They just want to build bigger computers. And you could basically say something very similar today. What we are trying to do... one moment.
T
Tyler1:20:34
Little water break.
G
Greg Brockman1:20:36
Yeah. Exactly.
T
Tyler1:20:37
For all of us. Cheers.
G
Greg Brockman1:20:38
Exactly.
T
Tyler1:20:38
Cheers.
G
Greg Brockman1:20:39
You know, we're all human.
T
Tyler1:20:41
A proof of humanity right there.
G
Greg Brockman1:20:43
Exactly. So what we're all trying to do is find novel ways of taking compute and really harnessing it. And sometimes you hit a wall. But these walls tend to be ones that you can drill through, right? What we found is every time you scale up everything, all of your engineering, all of your sort of scale and variance, all these things, they get stressed to the next level. It's almost that the tolerances become tighter and tighter. It's like launching a 10x bigger rocket means you need to be like 100x more precise on everything, but it doesn't mean that the fundamentals of the science are different. So pre-training, there's definitely been a lot of discussion of data wall. Doesn't mean it's fundamental, right? It just means that we need to be better and more precise at what we're doing. There's RL, which has been something that has kind of come from spending a small amount of compute to much larger amounts of compute now. And then there is a third way that we're really harnessing compute, which is compute at test time. And we published some scaling laws around this. And all three of these things multiply. Like that's the amazing thing. And of course the compute and the harnessing of it is the fundamental goal, but that you get these multiplicative effects out of all of it through the quality of your engineering implementation, right? Through the quality of the data sets, through a bunch of the refining work that you do and there's lots of different techniques and ideas and that's what makes this field so rich and why progress is just going to continue apace.
T
Tyler1:22:02
What about on the infrastructure side? You guys have been busy scaling up. What can you share on that front?
G
Greg Brockman1:22:09
Well so I run a team called Scaling at OpenAI and we really focus on building the infrastructure for scaling and this is in partnership with really everyone across the company. It's almost a misnomer that our team is called Scaling because fundamentally this whole team and effort is about scale. But what we really try to do is to both on the physical infrastructure side deliver as much compute as humanly possible and that is in partnership with companies like Oracle, SoftBank and others that we've been able to deliver just increasing amounts of compute to OpenAI. But we're constantly thinking about how do we just deliver more flops and do it more efficiently, earlier, cheaper, more power efficient, all of those kinds of questions. There's the software infrastructure side as well and really thinking about how do you coordinate massive numbers of GPUs in order to work across one synchronous training run. How do you coordinate that for reinforcement learning? How do you deploy that into production and bring these models to life at massive scale? And I think that every single layer of the stack there is innovation required. And that's something that's very easy to miss. Like one way I think about research is that there is, and this is kind of the view from Jakub who's now our chief scientist, that there's a research stack and you can kind of think of the top of it is people running experiments and coming up with new ideas for how to sort of utilize data or something like that. There's a middle of the research stack of people thinking about how do you sort of take these different ways people are running experiments and be able to train in novel ways and kind of put together the pieces differently. And then there's a bottom of the research stack which is like writing CUDA kernels to get the absolute max out of the GPUs. And at every single layer here you get a multiplicative factor through innovation. So it all comes together as one big whole.
T
Tyler1:23:59
On scaling, I'm interested to hear about just if we think about the impact of AGI or the impact of AI just being some sort of maybe quantitative GDP metric or qualitative just impact and good, is there an important factor of scale with just not even the flops that are going into the models, into the pre-training, into the RL, into the test time inference, but actually just the flops that are going into the usage of AI within humanity broadly. And I feel like that might maybe be the next scaling curve that we're seeing as more people use models, they see improvements all over the place. Is that something that we should be tracking to see kind of instead of these S-curves, we want to see the continual exponential.
G
Greg Brockman1:24:57
I think that's a great perspective, right? Because at the end of the day, I mean, if you look at kind of the shift from something like Dota, which we pursued in order to, you know, we wanted to do new algorithmic development, but really it almost validated how we scale up existing algorithms. But there was no illusion of delivering direct economic benefit from it, right? To the current models where we are still, we're starting to end the era of pushing on these academic benchmarks, right? You look at things like the IMO at this point. Models are able to get gold medal on it, like these the hardest academic benchmarks that are available are sort of no longer the guiding light of progress for these models. To where we actually want to be is for AI to be helping everyone, right? To be something that uplifts humanity. And that's the final metric, right? Is how much does it actually benefit everyone? How much value does it bring to the world?
T
Tyler1:25:49
Yeah. Not just health bench. It's actually how many people did you solve their healthcare problem.
G
Greg Brockman1:25:54
Exactly. Yes. Yes. Yes. And that's the actual goal. And that's what's exciting, right? Is it's like we're moving from the lab to reality. And I remember in the early days as we were thinking about how do we measure our progress towards AGI. We always sort of dreamed that one day we would be able to measure it this way. And you can think of revenue maybe as a proxy metric for value delivered to the world. It's not perfect, but it's at least something, right? You can think of the distribution of how much compute goes into it, how many people are using it. But fundamentally what we're after is how much do we really uplift humanity through this technology.
T
Tyler1:26:28
Yeah. I mean, I might be misreading it, but I'm pretty sure that was the Kurzweil philosophy was that like total number of flops getting immense, not necessarily all in one data center for one model. It was that compute broadly would be so wide.
G
Greg Brockman1:26:45
Yes. Yes. And I remember like on that chart you can see total compute of all human brains. Yeah. Which really suggests a particular vision of how this technology will be rolled out.
T
Tyler1:26:57
Yeah. Distributed, the phones count as an impact. The Wi-Fi router counts for the impact of the internet just like the phone does. Not just the big pipe that's going, the backbone of the internet that actually matters. Deep research hit product, almost everybody I know at least in the US says he's reading 30 pages of deep research a day basically, he loves it. He's making books with it. But why have agents broadly come around a little bit slower than people may have expected? Is it just that using computers is actually a much harder, computer use is just a really hard challenge or you know I think going into this year everybody said this was the year of agents, booking, you talking about flight booking, but people were saying 2025 is the year of agents and I would say that it's the year of deep research and not a lot of these other broader use cases.
G
Greg Brockman1:27:57
Sure. Well, 2025 isn't quite over yet. So, that'd be my response. And I'm very much on the... I think that progress in this field, the way that it tends to work is that if something kind of works with the current generation of models, it will be extremely reliable with the next generation of models. And I think that where we've been is that deep research is the... if you rewound a year, that was the like we kind of had something working. And then this year, it's been just incredible. And I think that agents, specifically computer use agents are something we've kind of had working and again, the year is not over. I think there's a lot of rapid progress to be made. But I think that maybe part of it too is that the agents that we're about to see, I think are a little different from maybe what we would have pictured five years ago. Like I remember having a debate with some friends on do you want an agent that does the flight booking because the problem is it's actually a very high bar to beat the flight booking UI because there's so many preferences that are entailed in that, right? And you really have to know kind of what mood you're in, like are you okay with taking the extra layover and all these kinds of questions. And that actually there's so much other stuff that happens in your life that is toil or drudgery or that's something that you're not an expert in, you're supposed to be... think about health, right? That like every patient really is the doctor if you're coordinating across multiple specialists. There's no doctor that helps you with that, right? That's really on you and that there you actually can have AIs that are just text only that actually able to add massive value and then frees up your time if you want to go book the flights yourself. And so I think that really finding the right problems that have high leverage, right? That really add value to people and also thinking about the other side of how to make sure these agents are responsible with the trust that you put in them, right? That the more that you give an agent access to your email, the more you really have to trust that it's going to do right with whatever your task is and send the right email to the right people and be able to segment your information, all these kinds of questions. And so I think that there's both a practical how do you get to adoption but also just like where are the most important leverage points in a person's life.
T
Tyler1:30:09
You also missed coding agents because it's been the year of deep research but I feel like it's also been the year of coding agents. How is that developing at OpenAI? I've noticed that I'll hit o3 Pro and it'll wind up writing a bunch of code for me and I didn't even ask it to. Then you have specific products for coding. How do you see the evolution of software development evolve? How are you seeing OpenAI customers use coding tools and how good is ChatGPT or GPT-5 on coding?
G
Greg Brockman1:30:43
Well, software engineering is definitely being revolutionized in front of our eyes. It's been happening and GPT-5 is the best coding model in the world right now. It's the default now in Cursor, which I think is a really huge statement of the quality of the model and that it's just so good across every function of writing code, understanding codebase, being able to use tons of tools, being able to do agentic work. That yeah, it's like I'm not a front-end developer at all but actually now I am, right? And I think that you are too, right? If you just talk to the model you can produce incredible things and so I think that there's this...
M
Mark Chen1:31:18
Real empowerment, if you think about what computers were supposed to be, right? Computers are supposed to be a more productive thing. But then somehow when we started out with computers, you have to contort the human to the machine, right? Writing assembly language and like all these very abnormal things for a human to do. And as we've moved to tools, ultimately, you know, in the current generation now, GPT-5, suddenly the computer comes closer to you, right? That you just express your intent and you don't think about exactly which language and what version of different libraries. The model is something you can delegate to. And so we are very committed to programming and to making our models continue to be the best they possibly can be.
T
Tyler1:32:00
Must a super intelligence be able to explain how to build super intelligence?
M
Mark Chen1:32:08
So it's a great question. I think that where we're going is a world, and we're already seeing it, where these models help us produce the next generation of models, right? They also help us really supervise tasks that are too hard for humans to supervise on our own, right? If the model writes a 10,000-line program for you, reviewing that is probably going to be quite burdensome. But if you can have a model that you trust that maybe isn't as capable as the one that wrote all that code, or maybe there's a team of agents that work together to write all that code, but you have a team of reviewer agents, this is the kind of thing that you can actually bootstrap trust. And I think that this is one of the most important things. And also interestingly, 2017 is when we had the first language results. We also had some results or some vision on how you can actually bootstrap supervision beyond the scale of tasks that humans are able to supervise directly. And so I think that we're heading to a world where we now have these chain of thought models. We've been advocating very strongly to preserve the integrity of the chain of thought. Right? So that means don't directly optimize it to look good, even though there will be lots of temptation to do it for various reasons. Really make sure that there's no pressure on the model to obfuscate its thoughts within that chain of thought, because then you can really see what it's up to. And I think there's further techniques to even make it more faithful and more rigid to what the internal monologue of the agent is. And so I think that there's actually a lot of promise in terms of interpretability, in terms of supervision, in terms of being able to scale to just much more sophisticated tasks.
T
Tyler1:33:38
Yeah. I guess my question is like there, how much information in the world can be derived from first principles reasoning versus true secrets that need to be discovered by interacting with the world directly? Because it feels like it would be very difficult to... I'm just wondering about like how intellectual property interfaces with super intelligence, or how if you play this out a lot, how there's all these hard-won... Doresh has talked a little bit about this with continual learning. There's all these little subtleties that maybe they're not secrets, maybe they're not true trade secrets. You don't think to lock them down, but they're just things that haven't been codified online or anywhere. They haven't been given to anything that is surfaceable by the model. And I'm wondering, is it just we need to build up new knowledge in every fact from first principles and kind of go through the history of humanity's pursuits of knowledge, or do we just need to onboard more and more information? Or maybe it's both. I don't know. It's just something I've been noodling on.
M
Mark Chen1:34:48
Yes, it's a great question. I would say all of the above. Select all star. So I think that the answer is very similar to what it is for humans, right? How does a human generate new knowledge? How do we accomplish new things? First, you want to be grounded in the wisdom of the past, right? You really want to understand what have people tried, what worked, what didn't work. You want to go and read the biographies of various people and understand those. But you also want to try things out, right? You want to make some mistakes in a contained environment in a way that you actually can see the effect of your hypothesis. And then you want to be able to learn from those. And I think that being able to really start to scale up these systems and be able to integrate them with the world is a very big process and milestone that we're currently embarking on, right? To move from a world of totally hermetically sealed reinforcement learning environments to thinking about how do you actually put real world interaction in there. And you think about things like robotics, like you're going to need to have that at some point, right? You're going to need to have some sort of interaction with the real world and to have models that are able to produce new materials, right? To be able to actually solve various diseases. For them to be able to really help people, right? That, you know, we already have models that are great at use cases like therapy. But to really get to the next level of something that can just really help every person accomplish more and accomplish whatever their goal is, it would be very helpful for that model to actually have some real world experience with doing that very thing. And so I think that figuring out how to bring all this together is ultimately what our mission is about. And we do this not in isolation, but really as part of a much broader community.
T
Tyler1:36:21
It seems like it's advantageous to have the most dominant consumer app in that environment. So, congratulations. Jordy, do you have a last question?
J
Jordy1:36:27
Last question. What do you hope to see out of Washington DC in the next year, year or two, not thinking super long term in terms of, you know, basically promoting innovation within the United States. Obviously, the admin cares a lot about AI and has been making moves, but what else would you like to see or where would you like them to double down?
M
Mark Chen1:36:46
Yeah, I've been very, very impressed with how much the administration has engaged with the technology and really tried to figure out how can we help and ensure that American AI continues to lead and really sets the standard for the world. And I think that that is the lens that I would really encourage thinking through, right? Is like this technology is changing very fast and that fast plus government is not usually an ideal combination, but this is the reality that we have. It's the opportunity we have and I think that the question in my mind is less about any specific regulation or strategy, but it's really being calibrated. It's really having a very tight feedback loop, right? Being able to react to, okay, we have a new model. These are the capabilities we see on the horizon. How do we make sure that we get the most uplift and benefit from it and thinking strategically about not just how do we do this for Americans, right, but how do we actually do this for the world and promote democratic values? And so to me, the most important thing is that motivation, right, is the question that is asked and the ultimate sort of motivation behind what gets implemented.
T
Tyler1:37:48
Yeah, that makes a ton of sense. Thank you so much for joining us. Jordy, are you gonna hit the gong for GPT-5 and the whole... congratulations on the massive historic day and thank you so much for stopping by. We'll talk to you soon.
M
Mark Chen1:38:02
Thanks for joining.
T
Tyler1:38:02
Have a great day.
M
Mark Chen1:38:03
Thank you for having me.
T
Tyler1:38:04
Bye. Cheers.
Really quickly, let me tell you about figma.com. Think bigger, build faster. Figma helps design and development teams build great products together. And we are joined by Sarah Friar, the CFO of OpenAI next. And we are going to bring her in in just a minute.
Still swinging. The gong's still swinging. And I'm going to tell you about Vanta.com. Automate compliance, manage risk, improve trust continuously. Vanta's trust management platform takes the manual work out of your security and compliance process and replaces it with continuous automation. Whether you're pursuing your first framework or managing a complex program, we need one more second. Tyler, any other questions that we should be asking for the OpenAI folks? Anything top of mind?
What's on the timeline? Is the timeline still in turmoil or has it settled?
M
Mark Chen1:38:49
So, I think the general vibe is like this model was not benchmaxed, but if you actually get to use it, it's pretty solid.
T
Tyler1:38:55
Cool.
M
Mark Chen1:38:56
One thing, it failed QPN bench.
T
Tyler1:38:58
Oh, it did.
M
Mark Chen1:38:59
It did not get the horse breed. Correct.
T
Tyler1:39:00
Get the horse breed. Wait, so you have it? You have access to...
M
Mark Chen1:39:02
Yes, I have access. But I've seen other things on the timeline. We can talk about it later, but it seems like a really good model.
T
Tyler1:39:08
That's amazing. Great to hear. Well, welcome to the stream, Sarah. Good to meet you. How are you doing? Congratulations. A historic day. Thanks so much for taking the time to talk to us. How you doing?
S
Sarah Friar1:39:17
I'm doing great. I mean, how could you not be doing great on the day when GPT-5 launches? It's been a long time in the making and we're so happy it's out.
T
Tyler1:39:25
Yeah, fantastic. Walk me through your role and what GPT-5, what this launch means specifically for you. And yeah, well, let's just start there.
Finance has to be, you guys have to be the unsung heroes at OpenAI. There's just a lot of big numbers, bills coming in for crazy training runs and you have to underwrite these against future revenues and I'm sure you've developed many models to figure that out. But yeah, walk me through what your role at OpenAI and what today means for you.
S
Sarah Friar1:39:57
Yeah, absolutely. So I'm OpenAI CFO, but the finance can be the unsung heroes, but they are an amazing team so I'm going to shout out to them.
T
Tyler1:40:05
They're heroes to us. They're heroes to us. It's a complex world that we're all living in and there are a lot of B's on the end of a lot of the numbers that we look at. Look, what is our role?
S
Sarah Friar1:40:15
Number one is just making sure we have a healthy, high growth business. It's been incredible watching just first of all the number of weekly actives. 700 million people using ChatGPT every week and I'm assuming after today we should see a very nice little bump in that number.
T
Tyler1:40:30
This is going to be a gong-heavy segment, Jordy. I think we're going to have a lot of soundboard for the big number. So congratulations.
S
Sarah Friar1:40:36
I love it. And I've never met a number I didn't like. I think the other part of the business that you know, and then we have to do... is clapping. We have to do this balance of the consumer business, the enterprise business and then API business which I think of as somewhat enterprise. And balancing that out. So enterprise adoption has also been exploding. I probably do, I mean interestingly as a CFO I probably meet four to five customers a week. It's a part of my job I actually love. We have about 5 million paying business users right now from banks to biotech. I was talking to the CFO.
T
Tyler1:41:12
And so that number is individual companies?
S
Sarah Friar1:41:15
That is individual seats at companies.
T
Tyler1:41:18
Seats at companies. Got it.
S
Sarah Friar1:41:19
So what I would say about that number is it's crazy to have done that in just two and a half years because enterprises, right, you got to put your big boy, big girl pants on to go sell to an enterprise, right? They want to make sure that you have the table stakes of security, SSO for signing on, you have HIPAA compliance if you're selling to healthcare and so on. They want to know that other people have done it. So they're often looking for that case study. But they also want to be, you know, the innovator right at the front. And so that to grow that scale of business in just two and a half years blows my mind. And it's not just big, big businesses which I could talk at length on, but it's also small mom and pop, you know, literally the people who really keep the lights on in most countries are also gravitating to ChatGPT which is wonderful. And then on the developer side, four million developers have built in our platform and the question there is like that could be a developer inside a big company like Grab, it also could be the next startup founder that's Y Combinator getting going with the next multi-billion dollar unicorn business. And so we see the whole gamut there and that's important to us as well because very mission aligned, right? How are we going to get AGI to all of humanity if we don't do it through this ecosystem? So a big part of my, you asked my role, big part of my role is just keeping that business really healthy, making sure we always have the headlights on so people know the decisions they're making from a business standpoint. Huge part of what the team does.
The other big part of my role, compute. If I didn't talk about that in my first breath, you all should correct me.
T
Tyler1:42:56
I mean, it's making sure we think compute is a massive competitive differentiator.
S
Sarah Friar1:43:02
I give so much kudos to Sam and the team, but particularly Sam because no matter how big a number we look at, Sam always wants to go bigger and he's been right. It is...
T
Tyler1:43:13
He's never met a number he doesn't want to add a zero to.
S
Sarah Friar1:43:16
That too, maybe more, maybe logarithmic, maybe two zeros. And he's, but he has been very right and if you know, you've just had a long conversation with Greg Brockman, I think he does such a good job of kind of really explaining what a completely different world an AGI-ed world is or an AI world is. And so I think when people get all caught around the axle of like, you know, what is a gigawatt of compute and oh my god you guys want to have 10 gigawatts and that's more than the compute of like Ireland since I grew up there, and now you kind of look back on that and you're like those numbers already look small for a world where everyone will have access to intelligence and we're really starting to see what that can mean when you look at the demos today around things like healthcare and education and so on.
T
Tyler1:44:02
Can you talk to me about non-GAAP metrics and what you think is going to be useful to track? We were talking to Mark Chen about this and he was saying, you know, DAUs are great, time on site is great, but that's not as impactful of a metric for OpenAI as it is necessarily for a social network or an entertainment app. And there can actually be some problems that come up with that. So, it feels like there might be some tension in the organization eventually or just publicly about what metrics are worth optimizing for. And then there's also the financial community that wants non-GAAP metrics to track the health and progress of the business. And then of course over decades we see companies eventually roll back some of those non-GAAP metrics as the business gets more complex. So how do you think about the development and sharing of non-GAAP metrics and what do you think is actually interesting and provides signal to the business and the investor community?
S
Sarah Friar1:45:00
I'm kind of smiling to myself because when anyone normally says talk to me about non-GAAP metrics, I can see like most of people's eyes roll back in their head. I live for non-GAAP metrics. I would love to do that. Look, I think in a CFO seat, first of all, it's really important to think about input metrics and output metrics and things like revenue, which is a GAAP metric as well as a non-GAAP metric, they're very laggy. Like if you are spending your whole time focusing on the revenue number in an operator seat, you are completely missing what's going on with the business. So I push my team a lot to get out of kind of ultimately what the P&L looks like and I'll come back to it though and go way upstream and say what are the true input metrics that tell us about the health of our business. And so I think it does start with that funnel of monthly actives to weekly actives to daily actives because we do, I mean our mission is literally AGI for the benefit of humanity. So we know how many billions of people live on the planet. The fact that we're starting to be able to talk in billions and percentage of the world's population, it blows my mind. Right? Today 85% of our users are outside the United States. And I love that stat. And in fact, if you go look at where are the big populations of users, it just tracks global population, right? It's countries like India, Indonesia, Brazil, Vietnam, like the Philippines, like go to anywhere that has big population. The US too of course, but that will be your tracker. So that's kind of number one when I think of an input metric. From there on the consumer side, you're right. Things like time in app, I've actually always had somewhat of a love-hate affair with, but I think in this case because we're giving people intelligence, teaching them how to use that, I actually think is where time in app does become important and one of the things we've really seen with ChatGPT are people are spending more time with it now. You know, we balance that with things like mental health and so on, making sure that we're not creating bad things like we might have seen in prior eras of computing, but I think we're just getting started on that front. Beyond that, like when we go into areas like the API, I don't look only at usage, right? I can look at tokens per minute as a usage metric, but I look at things like latency. I actually try to look at the elasticity of demand, right? We know that developers want performance, they want intelligence, but they also want to make sure the API is always up, and they want price. And they're often willing to trade across those three things, right? It's a kind of a linear program depending on what your use case is. And so I think it's important that we are offering things to developers that allow them to optimize across those three metrics. For example, so that's kind of your input metrics. And again, I could wax lyrical, but I won't. But then you go to what you really ask. So investors on the other side, right? They want to see a P&L. They're like, I want to be able to compare you to other companies. I want to be able to create a maybe a DCF. Like I want to think about fundamental valuation for a company if I'm going to invest in it. And so, you know, today what I really try to push investors on is we are not a company that should be optimizing for free cash flow today because there's just too much opportunity. Like that point about compute. We have to make a decision on compute today with an eye to what we're going to need in two to three years because data centers don't just spring up overnight. Like they're not mushrooms. They literally take time and effort. The thing we have at, frankly I would say is three years ago we didn't have enough foresight to say how big could ChatGPT because it didn't exist. It's just a shame on us if we keep doing that over and over. So there can be a bit of a mismatch between our belief on revenue because we don't yet know the product versus the input which is the cost today on compute. And so getting investors comfortable with the fact that there's probably losses for a period of time. I say probably because ChatGPT just generally the revenue models continue to surprise to the upside but at least for now we should be in big investment mode. And then you kind of said it well like as companies mature you move to more GAAP metrics, right? If you look at the large, the Mag 7, in many cases they're looking at like real GAAP net income. So the whole way down to the bottom of the P&L, we're just not there yet and we should take advantage of that advantage because we can invest as a private company.
T
Tyler1:49:26
How do you think about timing fundraisers? From my understanding or rumors, the last, you know, the most recent financing was very oversubscribed and at the same time you're still committing to capex in the future that is a multiple of current, you know, the current run rate. And so you in the CFO seat, I'm sure you're trying to find this balance of like what does the business need today while not diluting the company too much, knowing the growth rate of the business.
S
Sarah Friar1:49:59
I mean that's exactly right, that's the art, not the science of it, is that, you know, we did just come off the back of closing out the sleeve of investment that we could take down in this current round led by SoftBank and it was massively oversubscribed, which comes back to I think the market really waking up to the fact that AI is a generational opportunity and the scale that it requires is like something people have not even seen before, right? It's, you know, people talk about the internet or like the railways. They're good analogies or transistors. I think Sam always goes back to they're good analogies, but I do think this is bigger than everything that's come before. So there's a, you know, taking down $40 billion, which we just did in this round, that certainly felt like that gave me a lot of confidence. Appreciate that. A lot of confidence to then go out and do large compute deals, right? We announced the large deal with Oracle for example and to be able to keep working with all of our supply chain Microsoft, CoreWeave, Oracle, Nvidia and so on. But at the same time, you know, in a world where our valuation has gone up, you know, at pace with our revenue, you do get an opportunity to keep coming back to market and not take that same dilution because you're getting that higher valuation for the work and the output that you've created. So it is a bit more of an art than a true science. I think for now we will continue to need to fundraise in order to fund that compute but I think we want to start getting more sophisticated like just pure equity fundraising for everything is an expensive way to fundraise and I think we're probably getting to the stage at a company where we can be a little bit more kind of broad in how we think about funding overall and even just working frankly with our supply chain because you know our success with bringing this era of AI into being is their success too. And I think these companies are realizing that.
T
Tyler1:51:55
What about partner? Last question. Partner selection on the compute front. There's not a lot of companies in the world or firms that can really be a...
You should update your LinkedIn title. We saw someone yesterday works for Discord is in charge of their cloud buying and his LinkedIn title was I have full responsibility over buying our entire cloud budget. And it was clearly like a huge flag, but I'm sure you're in direct message with every single person that's relevant in the industry. But yeah, but I'm curious around like, you know, a lot of people have been excited about developing data centers over the last couple years in hopes to win...
S
Sarah Friar1:52:37
Oh, yeah.
T
Tyler1:52:38
comp, you know, the business of companies like OpenAI, but I think when you guys are evaluating partners, I imagine that scale is such a massive factor. And so a single small data center is not really going to move the needle. You guys need to be thinking in terms of mega projects.
S
Sarah Friar1:52:56
Yeah, I mean I think that's exactly right. I mean it started with our partnership with Microsoft and it's kind of, it makes me smile now to go back and look at that original kind of large fabric for pre-training because I think it was only in the maybe 20 megawatt sort of size. And you know now we're talking gigawatts even just this year. And you're right that when we think about like what is perfect compute for us or strategically the right compute for us we are definitely thinking about large scale, we're thinking about flexibility, right? We're learning a lot about pre-training, post-training, test compute, even like where the different kind of scaling is happening. We're kind of recognizing there's more of a blurred line often between what people think of as inference, investors always are like your inference compute and your training compute, it's like you know literally it's vanilla ice cream and chocolate ice cream when in reality there's like a bit in the middle that is something of both. We also need to think about things like where, you know, latency, where do we want to put our footprints around the world, that very global weekly active user base, right? As they use ChatGPT, you don't want to slow the model down, right? The beauty of the intelligence is like the real-time nature of it. And then when we get into big compute like where there's lots of tokens being used like deep research, image gen, video as that comes online, like all the work you saw today actually just even on voice, like that really quickly means that you got to make sure your compute is near your users. And so it is a big plan that's coming together but you're right, like small is just not that useful to us. But...
T
Tyler1:54:36
What about pushing partners to take risks? From my understanding, you guys are pre-committing to certain, you know, basically spend levels, but at the same time, I imagine you want people to say, 'Here's what we know we're going to need, but we want you to build, you know, this much capacity so that we have the sort of incremental capacity built in.'
S
Sarah Friar1:54:55
Yeah, we want, I mean being extensible is really important. And we do want to see partners like I think Oracle OCI has done a really nice job of that of kind of starting, we started with like one large, it felt really large at the time, data center footprint in Abilene in Texas and now that has really multiplied up into multiple sites that can all be connected and that's a good example of a partner who has the capability to start in one way but to be able to show you a path to maybe 5x-ing just in that single footprint. That said, we are finding that as we go around the world, there is an ability to go work with governments. For example, we just made an announcement in Norway, made an announcement in the UK. This is the first time in my professional career I've seen countries come to the table and want to do commercial deals like wall-to-wall ChatGPT. I think the government of Estonia put ChatGPT into all of their high schools or I can't remember, was it university level. But that's kind of wowing. And hand in hand with that they are viewing AI infrastructure as incredibly strategic for their population. And you know it's a whole other level of selling versus, you know, I've seen enterprise, large enterprises before but never anything at this scale.
T
Tyler1:56:11
Last question. Whose idea was it to give every federal agency ChatGPT for a dollar a year?
I imagine that had to... you could have gotten more than a dollar. The CFO must be really upset here. $10, that's 10 times as much money. Now...
S
Sarah Friar1:56:28
This is one where I think it's really important. OpenAI is, you know, in some ways a US asset, a national asset, and we want to make sure we're accelerating our government, like all of the resources as we think about, you know, western democracy and so on that we are absolutely putting our technology into those hands.
T
Tyler1:56:45
It's that guy Kevin Wheel, he's been moonlighting for the US government. It's like, which team are you playing for Kevin? Are you on OpenAI or are you on the US government?
S
Sarah Friar1:56:54
Kevin just did his basic training. I don't know if I'm allowed to tell you that, but I was hearing all about it yesterday.
T
Tyler1:56:59
I saw some photos. They look great.
S
Sarah Friar1:57:00
Yeah, it's a good thing. It's even better for Kevin. Yeah, it's great.
T
Tyler1:57:06
Last question for me. I, you know, the open source model launched two days ago and there's this world where like you have this dominant, the accidental consumer company, you have this dominant consumer app that's generating so much revenue. Then you have B2B and enterprise and API and that looks more like a cloud provider. But then is there a world where the Red Hat Linux of open-source LLMs is an OpenAI division and that there's actually serious revenue and profit that comes from helping companies implement an open-source large language model? Like Red Hat built a pretty fantastic business for a long time on top of open source Linux implementations. So like...
S
Sarah Friar1:57:47
Yeah, I mean I think it's the right question to be asking. I mean I think step one was getting, yeah, you got to get it out, two days ago our second open source model out, and getting seeing what that traction is and then seeing what the community needs. I think it's important to leave space for a community to develop, right? That is the beauty of open source is that ecosystem that develops and that was true with Linux. It's true in areas like crypto too. But I do think you'll find over time that as enterprises want to deploy it, like I now date myself, but when I was a research analyst at Goldman Sachs back in the day, I covered software and I covered Red Hat actually. Oh, that growth. I wrote a research report called 'Fear the Penguin' at one point because seeing Linux being deployed, but then you started to understand that for an enterprise you couldn't depend on like patching and upgrading to happen via community model like you needed some of the rigor that goes with an enterprise business where you kind of know when you need maintenance, if you need a bug patch and so on. And so that did allow Red Hat to grow an incredible business. So I don't know if it's us or we'd be supportive of others, but I think we are so excited to see open source out there and getting incredible feedback and I think we want to do that ahead of GPT-5 to keep coming back to like we're here to grow this ecosystem.
T
Tyler1:59:06
Well, we'll give you market cap credit for it anyway even if it's early stage. Well, thank you so much for coming on. This is fantastic. We'll talk to you soon.
S
Sarah Friar1:59:12
Thank you, sir. Great to see you both. Take care. Have a good one. Cheers. Bye.
T
Tyler1:59:15
Up next, we have DD Credo from Kudo. I believe I'm pronouncing that correctly. Let me tell you about Graphite code review for the Age of AI. Graphite helps teams on GitHub ship higher quality software faster. You can get started for free at graphite.dev. And let's bring in our next guest. How are you doing? Welcome to the stream.
Welcome.
Ooh, very clean background. I know it's probably virtual, but whatever you got going on looks fantastic. You look great. How are you doing? Are you excited about GPT-5?
D
DD Credo1:59:42
Oh, I'm so excited. It's awesome.
T
Tyler1:59:46
It's actually like everybody's talking about the coding capabilities, please. But no one is really talking about the code review capabilities and I'm going to talk about that today.
D
DD Credo1:59:53
Yeah. Yeah. Break it down. How are you using it right now?
Yeah. So we just enabled it in our platform. It's the default model for both our IDE plugin, our CLI, our Git plugin. And yeah, we're using it to generate very high quality code reviews, catch bugs before they hit production, help enterprises verify that their code is aligned with their best practices.
T
Tyler2:00:18
Mhm. See, it's super exciting. I can share my screen and show a few things if that makes sense.
D
DD Credo2:00:22
You can, everything you share will be live. It'll be a little bit please. But I want to know also while you're getting that set up, I want to know about what changes materially do you think happened in GPT-5 specifically for code and code review? Do you think there's more data going into the model, more data going into the pre-training, post-training? Anything else? Anything that you're noticing that you're like, 'Oh, there's a specific upgrade here. They must have done something to get there.'
Yeah. Yeah. I think it's a great point. So, I think it's all of the above. So it's scaling of both the pre-training but probably a lot of the reinforcement learning and basically using that at scale to verify that code gets generated in high quality and then also basically catching bugs. And when you do it with reinforcement learning you have the actual ground truth. So once you scale that you can get the model to be a lot better at that.
T
Tyler2:01:23
How steep is the power law right now in just programming languages? Is it basically all Python, JavaScript and then a really hard fall-off or is it actually important for coding models if they want to be adopted widely to be like truly multi-language and get all the way down into the long tail of like the Rust and the C and all the different languages that are out there.
D
DD Credo2:01:46
Yeah. Yeah, for sure. It's important to, I mean the majority of the market is in the JavaScript, TypeScript, Python, like the majority of the early adopters I would say. But then when you get to enterprise use cases you get a lot of Java and the models are getting pretty good at those languages as well. For sure.
T
Tyler2:02:07
Are you excited about, I mean how do you think about the difference between like the improvements to GPT-5 from the consumer's perspective versus at the API level? I always found it a little confusing that ChatGPT was available as an API and you could interface with the chat, I believe you could interface with the ChatGPT model via the API. And there's a little bit of like a line blurring there, but are there features that you think are cruft and you want to kind of rip out for an API use case or do you just say, 'Hey, give us the kitchen sink and we'll work from there and it's actually helpful to have, you know, a coding model that can still have a web browser.'
D
DD Credo2:02:50
Yeah. Yeah. I think basically it's a lot about we consume the model through the API and it's really the same model that drives the consumer product. But for us since our use cases are a lot about agentic use cases, the more the model gets better at using tools and gets better at kind of listening to very, very specific instructions, following instructions is critical for the enterprise use cases. Because for us unlike the broader market we believe that for enterprises you need to have very specific agents that are defined with specific set of instructions, prompts and tools and permissions. And the more the models get trained with that type of environment, the better they end up serving the enterprise market, which is really where we're focused on.
T
Tyler2:03:39
My question is, I wonder like you said like very specific instructions are important. When are we going to get an agent that I can just turn loose in a codebase and say like just go improve it? Like just go hunt around, do like rewrite that, like when you get a good open-source contributor on a team that just becomes nerd sniped by the project that you're building on. They will just go around and find little ways to improve this, documentation needs to be a little better. Let's rewrite this test case over here. Let's add a little bit more functionality to this class or function. How far are we from that?
D
DD Credo2:04:17
Yeah, I think the models are getting better and better at that part of basically kind of running loose in a codebase. Yeah. But they do need the guardrails in place and this is kind of where we're focused on. Like a lot of the talk in the market is around the code generation side. You know, let the agent loose and give it a task and it will just going to go around and run for hours and do it. What we're seeing is that the real challenge is now shifting towards how do I verify that the code is aligned with the best practices? How do I make sure that it's well tested, well reviewed, doesn't break anything. So that's I think the next frontier and really developers going forward are not going to write a lot of the code by hand. They're going to spend most of their time reviewing code and that's the next frontier and that's what we're really here to tackle.
T
Tyler2:05:10
Very cool. Anything else Jordy?
J
Jordy2:05:12
No.
T
Tyler2:05:13
Well, thank you so much for joining, giving us some extra context on the GPT-5 launch. We will talk to you soon. Have a great rest of your day and thank you for joining.
D
DD Credo2:05:22
Cheers.
T
Tyler2:05:23
Thanks. Cheers.
Talk to you soon. And let me tell you about Profound. Get your brand mentioned on ChatGPT. That seems more relevant than ever. Reach millions of consumers who are using AI to discover new products and brands. I forgot to ask about this. We'll have to come back to this but I want to know if...
the Founders, Powers, MongoDB, Indeed, Mercury, DocuSign, Zapier, Ramp, Ro, Gong, Workable, Majury, Asleep, US Bank, Chime, Clay...
Okay, okay, okay, we get it. They got some logos. There is this question of like, okay, even if you're like, okay GPT-5 is more incremental, more of an evolution than a revolution, it's like, okay well then let's talk about how it affects every other business and every other aspect of the economy, what should you be focusing on? And is like, do any of the updates from GPT-4 to GPT-5 change how you're positioning your brand for AI search? That's certainly an interesting question to dig into. Anyway, we have Zach Lloyd from Warp coming into the studio. Welcome to the stream for the second time. Welcome back.
Z
Zach Lloyd2:06:31
Good to see you.
T
Tyler2:06:32
It's great to be back.
Z
Zach Lloyd2:06:33
How you doing?
T
Tyler2:06:34
I'm doing pretty well.
Z
Zach Lloyd2:06:35
Uh, you know, yeah. So, I mean, a lot of what stuck out to me, I'm mostly a consumer of consumer AI apps. I'm very excited about not needing to mess around with a model picker anymore. But take us through the biggest improvements from the software development side.
Yeah, I mean, it's a major step up from the prior OpenAI models. It's, I mean, it's doing agentic workflows in Warp for a much longer period. It's just a smarter general model. Like we eval it against all of our benchmarks and it's up there at state-of-the-art which is, you know, from our perspective it's awesome to have multiple competitive models that our users can benefit from. So definitely a huge improvement from GPT-4.1.
T
Tyler2:07:25
Yeah. So it seems like not the, you know, Claude Code killer, but certainly in the same conversation, in the same football stadium if we're using a sports metaphor. How much, you know, one thing that stood out is the cost reduction. How about do you think that developers will care about that versus just, you know, what it can do from an output standpoint?
Z
Zach Lloyd2:07:51
I think developers do care about value. So sort of like quality to cost ratio. I think it's the more you get into like the individual developer and the small team, the more that that matters. Whereas if you're at the enterprise level, I feel like it's a little bit less price sensitive. So yeah, I mean you can see it as different apps change their pricing what the reaction of the developers is. You've probably seen this with Cursor and seen this with Claude Code and so developers really, really are looking for something that's cost effective. So the fact that the cost is a little bit lower is actually a big deal.
T
Tyler2:08:31
Do you think we're in the Lyft-Uber 2015 arc where the prices are subsidized and the prices will go up? Do you think that there's a price war on the horizon now that the frontier models seem to be similar capabilities? Do you think that someone will try and raise a bunch of money, cut prices a bunch, and steal a bunch of users? Like, how do you think that plays out?
Z
Zach Lloyd2:08:55
It's an awesome question. I mean, my hope is that we get to a world where there is price competition at the model layer. So, Warp is very much at the app layer, right? And so, our value prop is like we can give our users who are mostly developers the best model access. And so to the extent that it's not one sort of model provider running away with that and having pricing power, it's better for us just candidly. And so, you know, my hope would be something like the model world ends up a little bit like G-Cloud, AWS, Azure. That's our best end state where all of these models are, you know, sort of similarly powerful and a little bit more commoditized. I don't think it's been like that, but it's getting a little bit more like that. And so, the more that there's more than one show in town, I think that's generally good for Warp and actually is good for developers because it will put competition, the competition will put pressure to bring the prices down. But I don't know, like I also think that people will definitely pay for quality. And so if there is a meaningful delta in quality on the frontier models then I think that like whoever has the quality delta will have a lead temporarily but I'm not sure that that lead will be sustainable. We'll see.
T
Tyler2:10:19
How do you think the developer community should plan around model deprecation over the next, you know, one to two years? Like how much, you know, from, I don't know that I've gotten a reaction yet from, I don't know if there's general frustration yet from people. You know, we've heard on the consumer side Tyler on our team here loves 4.5 and so he was a little bit disappointed to hear that. But what are you seeing on the developer side?
Z
Zach Lloyd2:10:53
Yeah, I think it's a little bit different for people who are like building apps on LLMs versus people who are using LLMs as like an accelerator to doing coding. And like, you know, at Warp actually we do both like we're an application level stack and like it's actually very easy for us to go to the latest model and so it doesn't really bother me. I don't know. I don't know what type of app you would be building where it's like it's really important that it's like GPT-3.5 or GPT-4 or something like that. I think like generally we want the most intelligent tokens at the best cost. So I don't see that being like too big of an issue honestly.
T
Tyler2:11:34
What about open source? Does that feel like something that will be in the playbook? Is the markup on closed source models high enough that there will be a significant price delta and or is the frontier kind of indifferent to closed source, open source?
Z
Zach Lloyd2:11:52
So if there was a comparable open-source option that would be awesome. I think that the economics of it again, it doesn't seem like a perfect analogy to me between like open source software and open source models. So open source software it's like you have a big community of people who, you know, for the love of coding are building a really awesome product. For open source models it's like you just need a crazy amount of capital to train something that's on the frontier and so I don't know how that happens. And so what we've seen is like the open-source models are competitive at the quality level that they're at but the quality level that they're at is not the same as the frontier models and I don't really see why that would change. And so I don't know, in Warp it's like we were serving some open source models, but they're just not as good. And so there's I think a more limited use case for them right now. And I don't really see economically why that would change. In fact, I would be surprised if anyone was spending billions of dollars to train a model and just kind of put out the weights. Like I don't get the business strategy there but maybe that will happen, that would be awesome.
T
Tyler2:13:09
Is there a world where you're like this idea of like smarter models either orchestrating dumber, cheaper models or like using or distilling models into more narrow formulations that can be run more efficiently. We've talked to a few companies that do this for businesses. Like you just want a model that just filters for profanity and you can run it on, you know, a gaming graphics card. And so it's basically super, super cheap or super fast. I'm wondering about like in the coding world, coding agent world, any of that, like where are the opportunities to kind of fan out and use an ensemble of models instead of just hit everything with the smartest, best? It feels like because of the funding environment, everyone can kind of justify like a high cloud bill, but and most people don't admit that it's hurting the bottom line, but it feels like at some point it kind of has to eventually.
Z
Zach Lloyd2:14:10
I mean, I think that's a very real thing. Like even in Warp we don't use the biggest, most powerful model for every task. And so there's certain things like...
T
Tyler2:14:26
For Warp, deciding whether or not we should summarize a conversation is a good example. You hit the context window and you're like, okay, is this a good spot to summarize? Is this a good spot to encourage a user to start a new conversation? We use a much more inexpensive and also low latency model. The other trend is that these very powerful models tend to have much higher latency, and so we do a mixture of models, and that's totally a real thing. But I think the predominant use case as a developer is going to be: I want to tell an agent to do something. I want it to be harder and harder. I want it to run for longer and longer. And to do that, you kind of want in general the most intelligent model. And so, until the models have a sort of S-curve type shape, I think it's going to be more of a quality game than a cost game for most of these things.
U
Unknown2:15:27
Doesn't it feel like they have an S-curve shape right now?
T
Tyler2:15:30
Certainly does from a consumer perspective.
U
Unknown2:15:32
That's interesting. From a coding perspective, I feel like we're still accelerating. The difference again between the last version of GPT and this version of GPT is probably bigger than the difference between 4.1 and 4, and 4 and 3.5. It's a big deal, and same thing with the Anthropic models, and I'm sure that we'll see something from Google where it's an acceleration. And I think that there is maybe an underappreciation of how much left there is to solve here. Because even when you're doing a real coding task as a pro, despite all the demos you see on Twitter where someone asks an agent to build an app, that's a lower level of difficulty than doing what a pro developer does with one of these models. And the models still don't produce great code a lot of the time. There's a lot of kind of handholding that has to go into it. And I think that we are still seeing an acceleration in terms of the models actually becoming not just okay competent engineers but really, really good engineers.
Yeah. Do you care about benchmarks?
T
Tyler2:16:38
We care a ton about benchmarks.
U
Unknown2:16:41
But your own internal benchmarks or...?
T
Tyler2:16:44
We do both. So, plug for Warp, we're number one on TerminalBench, which is the public terminal benchmark, and we're top five on SWE-bench, which is the coding benchmark. And then the only way in my opinion that an app at our layer in the stack can really improve is by measuring the progress. And so we have our own internal set of evals that we run across all these models as well, which are coming from real use cases. And that again is an advantage of being a product that's in the wild that has a lot of users, is that we can sort of see where the models are failing, where they're working, and so we're very big on that actually.
U
Unknown2:17:20
Awesome. Well, thank you so much for stopping by. We will talk to you soon.
T
Tyler2:17:24
Sure. You'll have a busy afternoon.
U
Unknown2:17:25
Shout out by the way to the OpenAI team, very, very helpful in working with us to get GPT-5 to be awesome in Warp. And one more shameless plug: we have a discount code for people who want to try GPT-5 in Warp. It's $5 GPT-5.
T
Tyler2:17:41
Okay. Thank you for having me, guys.
U
Unknown2:17:43
Yeah. We'll talk to you soon. Thanks. Cheers.
Tyler, any updates from the timeline while you're thinking about what the latest vibe check is in the war between OpenAI's...
Choice for OpenAI. You have something?
T
Tyler2:18:07
From Reggie James, front of the show: 'Half of my timeline says this is the closest we've been to AGI. The other half of my timeline says we officially just hit AI stagnation. I love tech.'
U
Unknown2:18:20
Well, we will be going deeper deciding whether or not this is stagnation or hyper intelligence takeoff. And we will be joined by our next guest, Riley from Charlie Labs.
R
Riley2:18:31
Hey guys, thanks for having me.
U
Unknown2:18:33
Good to see you, Riley. How you doing?
R
Riley2:18:34
What's happening? I'm doing fantastic. We've been heads down with GPT-5.
U
Unknown2:18:40
How long have you had it? How long did you get the preview? I feel like it gets rolled out to early adopters a little bit earlier, but has it been weeks, months? How long have you had it?
R
Riley2:18:50
We're a couple weeks, like two or three.
U
Unknown2:18:53
What was the first thing you... Charlie liking it?
R
Riley2:18:56
Charlie loves it. And also I love what Charlie does with it.
U
Unknown2:19:00
Yeah. What does Charlie do with it? What was the first thing you did with GPT-5?
R
Riley2:19:05
Ran our evals.
U
Unknown2:19:07
Oh yeah. How'd they come back?
R
Riley2:19:08
Just really good. Much better than o3, which was much better than any other model we've run before that.
U
Unknown2:19:15
Interesting. And yeah, so let's zoom out. What do you do? What do these evals measure? Walk me through it.
R
Riley2:19:23
So Charlie is a TypeScript-focused coding agent that operates much more like a human does. So less like IDE application terminal and more joins your GitHub and Slack and Linear workspaces. And it interacts with the team the same way other humans do. And then our evals are a mix of code review, because part of Charlie's job is to review PRs from humans as well as his own, and then code authoring, so opening PRs and pushing commits.
U
Unknown2:19:56
So when you develop your own evals, I imagine you try and keep those out of any training data. You want those to be held private. Is that correct?
R
Riley2:20:05
Yes. And it's getting even harder with web access now because they're too good at finding things.
U
Unknown2:20:11
They're finding everything. It's funny. And then talk to me about the shape of the actual problems in the eval. Are there some easy questions, some hard questions, some extremely hard questions? Like how are you formulating those? What's the shape of an individual task? Is it scored out of like a hundred? How do you think about developing a good eval?
R
Riley2:20:36
A mix of hard to very hard. The easy ones are just a waste of money and time at this point, especially with 5. There's a bunch that it's just not going to get wrong. And then we're mostly doing the PR ones, which look kind of like SWE-bench in the sense that we're taking an issue to start with. But instead of giving the issue like in a Docker container already, we trigger a comment on the issue that says, 'Hey, Charlie, go make a PR for this.' And then Charlie does his thing and then the PR comes up. And then we score that PR against a whole bunch of things like correctness to a known solution that's correct as well as code quality, testability, and some softer things like descriptions.
U
Unknown2:21:17
Who are the biggest customers or users for like a TypeScript-focused coding agent?
R
Riley2:21:25
It's a wide range of mostly modern apps. Pretty much any web app these days is going to be like a Next.js type app. And then all the way into like back end, Charlie himself is written in TypeScript. And there's very little front end.
U
Unknown2:21:40
Anything else? What else you got?
I just want to say I love the name Charlie. It's one of my favorite agent names that we've had on the show.
R
Riley2:21:48
Yes. It's right up there with Pig and what was the other one? Well, I don't think that was an agent, but...
U
Unknown2:21:53
That was an agent, but...
R
Riley2:21:54
But yeah, it's a good one.
U
Unknown2:21:55
Yeah, congrats on locking it down.
Yeah. What about cost and that side of the business? Is there any movement there or anything where you require movement or you need movement to really unlock new capabilities in the business or new markets?
R
Riley2:22:18
Not really for us because we're operating kind of at a human level. We do value-based pricing. So we charge per PR or per commit. And because that's comparing to such expensive actions that humans are doing, the challenge for us is more actually living up to the promise than doing it cheap.
U
Unknown2:22:37
Yeah. Are you having... but then doesn't the cost reduction announced today, isn't that great for business?
R
Riley2:22:45
Yeah. I mean, it's good overall, but like that's our problem is not that the models are expensive. It's that they're... I mean, they're getting really smart, but I'll always take more.
U
Unknown2:22:55
Never enough.
R
Riley2:22:56
Like, for instance, since the beginning of August, we've been testing 98% of the code that got merged into our codebase was written by Charlie.
U
Unknown2:23:05
Wow.
R
Riley2:23:06
Not 30, not 50, 98%. And that's coming through PRs. That's not like autocomplete in an IDE type thing.
U
Unknown2:23:13
That's crazy.
Yeah. What does that mean for the future of like who are you hiring? I imagine that you're still an engineering-heavy organization that's just puppeteering and orchestrating agents. But where do you see the future of software development as a career path going?
Yeah. Are new CS grads cooked?
R
Riley2:23:38
I think if they get really good at using the AI, no. If they try and take an approach of getting really good at writing code by hand, for sure. What we're mostly looking for hiring is people who are able to see things at a much higher level and plan further out because with tools like Charlie, you can write so much more code so quickly that it's more important to see where you're going and take the right path than it is to be able to write it quickly.
U
Unknown2:24:05
Very cool. Well, thank you so much for stopping by. Good luck with the rest of your day and congrats on an upgrade to everything that you do.
Tell Charlie to have fun out there.
R
Riley2:24:15
Have some fun. Thanks a lot, guys.
U
Unknown2:24:16
We'll talk to you soon.
All right. Let me tell you about numeralhq.com. Sales tax on autopilot. Spend less than five minutes per month on sales tax compliance.
Salestaxsuperintelligence.com.
A number of the fellas in the chat got access to 5. Break it down for us.
Reg says it's pretty good. The writing ability feels a little nerfed. Says, 'The way it writes feels a little programmatic rather than sounding human, reverts to using points even for things like blog posts and also uses overly complicated language for simple stuff.' Techno Chief says, 'It's crazy fast.' Dan Ratliff says, 'Yeah, I was just going to say that, very, very, very fast.' Z Jean Ahmed says, 'Junior devs are barbecued.'
Tyler, anything from your side before we talk to Germa from Vercel?
T
Tyler2:25:11
I think maybe a good way to vibe check on at least on the timeline is that it's almost like a 4.5 kind of thing where it comes out, people are like, 'This model totally sucks. Look at the benchmarks. It's not some massive improvement. It's like a bar, not a step change at all.' But then you start playing with it and it's actually like, okay, this is actually a good model. Like a lot of the stuff I'm seeing people post, like, oh, that's actually really interesting output stuff like that. But seems good.
U
Unknown2:25:38
Can we do the green text eval, green text bench?
Yeah, we got the TVPN intern.
Yes. Yes. Yes. Yes. We'll let you cook on that and then we will move on to our next guest, GMO Ralph from Vercel, coming in to TVPN for the second time. Great to see you, GMO. How you doing? I like the action hall. Thank you. Welcome to the stream.
G
GMO Ralph2:25:58
How you doing today?
U
Unknown2:26:00
Do you think GPT-5 could beat me, you, a couple of the boys here on Dust 2 in Counter-Strike?
G
GMO Ralph2:26:08
Easily.
U
Unknown2:26:09
Easily.
G
GMO Ralph2:26:10
Yeah. It depends on the frame rate, right? Like on a long enough timeline, we're cooked.
U
Unknown2:26:15
We're cooked, but we might frag it short term and we might be faster.
Amazing. Yeah. We got to... I mean, I'm sure we'll get to GPT-5, but what's your reaction to the world model stuff from Google? Do you have an idea of where that's going as a product?
G
GMO Ralph2:26:29
It feels like a GPT-2 level technology, very much a research-focused technology. I'm sure OpenAI is working on something too and a lot of the labs will work on it, but what's your theory behind the generative video game world model stuff that's going on?
U
Unknown2:26:47
I mean, number one, super fascinating, right? I think when we think about the future, I always think about Jensen's: the future of applications will be that pixels are generated, not rendered. So as much as we're really excited today that GPT-5 and v0 are really good at writing code that then renders interfaces, I think it's also cool to dream of a world where we're just going directly from GPU to pixel grid, right? And but if you remember like a couple years ago and maybe a decade ago, there was a lot of excitement of video games that were going to be live streamed from the cloud.
G
GMO Ralph2:27:27
Yeah, that's right.
U
Unknown2:27:27
Where your input, your keyboard, you could have a very thin client. Your input, your keyboard, your mouse movement was going to be dispatched to the cloud.
G
GMO Ralph2:27:34
We're going to have GPUs near you.
U
Unknown2:27:36
Google Stadia was big there. And then xCloud was Microsoft's game and is still Microsoft's, actually still pushing it very heavily. Awesome tech, but not mass adoption.
G
GMO Ralph2:27:48
Yeah.
U
Unknown2:27:48
But if you look at, a lot of these technologies are being really successful in letting people get more creative and test things out. A lot of the use cases that we see for v0 and vibe coding are almost like a communication tool. I want to prototype something. I want to see what's possible. I want to explore the latent space. And I think those world models are going to be incredible just to inspire what the future of games could look like, right? Just getting ideas for actually then shipping them in a real 3D engine model. I think short term, I think long-term all bets are off. Someone was saying in the chat, you know, junior devs are roasted or barbecued. I think that's not quite true.
G
GMO Ralph2:28:27
Okay.
U
Unknown2:28:27
Same for like 3D engine developers.
G
GMO Ralph2:28:30
Give us the bull case for junior devs staying off the barbecue.
U
Unknown2:28:36
So the bull case for I think people in general is that you move from... I mean the progression in the industry has been assistant to agent to team of agents, agent orchestrator. It's still really useful to have a human be the one that's sort of like managing the team. So you're moving from like junior dev to junior manager. Especially as these tools become more agentic, in the new version of v0 that's coming up really soon, you're starting to notice that v0 sort of splits the task between a little team. You have the designer of the team. You have the PM of the team that's sort of working on the spec. You have the architect. You have the engineer. I don't know if you saw Claude Code announced. I think it's like a slash security review.
G
GMO Ralph2:29:22
Yeah.
U
Unknown2:29:22
You think of that as having a security team or team of agents or security researcher at your disposal. So junior dev as like a vertical skill might be a little barbecued, but junior bench manager... so I think it's just going to be the junior dev is so much more powered in this world if you allow yourself to be and you keep up with what these tools can do and I think you stay at the cutting edge.
G
GMO Ralph2:29:45
Yeah. I mean the obvious bull case is if someone's a college student today, they can learn to code truly AI natively. They don't have to say, 'Oh, we're an AI native organization now. We have to upskill and kind of retrain people how to think.' They can just naturally start to think with these capabilities. Sam Altman posted about how we'll look back on, you know, 93% of humanity was subsistence farming. And if you ask those people what they think about our email jobs, they'd be like, 'You guys are crazy.' And it's almost like in the near future, midterm future, maybe even long-term future, it's like the number of individual contributors will be extremely low and almost everyone will be a manager and you'll become a manager much faster. You'll just be managing agents and then you'll be managing people who manage agents. But the job of almost everyone will become managerial. Maybe that's what happens. I don't know. I'm not 100%. But that's what that made me think.
U
Unknown2:30:40
Someone asked me yesterday, you know, what do you think the future of the market of monitors looks like? Like does it stay flat? Do people get more monitors because they're going to like Dogecoin trader analyst when like...
G
GMO Ralph2:30:56
In the future everyone has the hedge fund six monitor setup or...
U
Unknown2:31:00
In the future everybody's just going to be at work on their phone. I mean I've noticed that when I was an individual contributor I had three monitors. I was programming on all the screens and now I use my laptop during the show and then most of my work is done on my phone, phone calls and then firing off tech messages. Yeah, maybe we actually shift away from monitors and go further into voice interfaces. You call the lead of my agents and then that agent relays it to some...
G
GMO Ralph2:31:29
I'm very optimistic on voice by the way because I've now seen it. We're cooking on a better mobile experience for v0.
U
Unknown2:31:37
Sure.
G
GMO Ralph2:31:38
And I was going back and forth with my head of mobile and he was talking to v0 and I was writing down and I'm a pretty fast typer, but he beat me with voice using the local model on the phone. So there's still the question of like edge latency versus cloud latency kind of like what we talked about with 3D. But I do think voice is going to play an increasingly exciting role in programming which is kind of wild. I would have never imagined. I've always been about like typing benchmarks in WPMs. Voice is coming.
U
Unknown2:32:08
Yeah. Yeah. Yeah. How do you think about competition broadly in developer tooling, code gen? I mean it right now it seems like there's just so much demand.
G
GMO Ralph2:32:18
It feels like massive TAM expansion moment. Every company's ripping... TAM expansion moment, but at the same time winners will emerge. Obviously you're playing to win. And yeah, I'm curious, you know...
U
Unknown2:32:32
Yeah, on some level we're playing both sides of the bet. What we announced today that's really exciting is v0 with GPT-5 support.
G
GMO Ralph2:32:42
So you can go to v0.dev/gpt5 and we'll use GPT-5 in combination with our model pipeline that makes it really good at vibe coding especially for non-technical folks. But we also on the Vercel AI Cloud side of things, we open sourced basically you can create your own vibe coding platform powered by any model.
U
Unknown2:33:04
I was joking about this with Tyler. Vibe code me a vibe coding platform please.
G
GMO Ralph2:33:09
Make no mistakes.
U
Unknown2:33:11
Yeah. Vibe code me a billion dollar company. Yeah. No mistakes. But basically we are giving people that as a starter kit.
G
GMO Ralph2:33:17
Sure. And by the way, the fundamental question that a CEO asked me the other day was, is vibe coding a product or a feature or is it both? You know, it's TBD. The case for feature is okay, so there's going to be lots of systems of record. Think Salesforce, Snowflake, Databricks, and increasingly they're going to incorporate codegen capabilities into their platforms. They can use a lot of these capabilities that we just open sourced and you'll go to their existing place where you have the data, kind of like what we've talked about for decades of like are you bringing compute to the data, are you bringing vibes to the data, right? Are you bringing codegen to your own platform? So...
U
Unknown2:33:58
You used to bring like a dashboard builder and it would have a couple widgets and now I could just potentially if I'm plugged into some sort of data source, some system of record, I could say vibe code this app on top of it. There's some tool, Retool's played in this space, Zapier a little bit, but yeah, I mean this feels like we're getting, we're not fully in the just the pixels are generated, but we're generative UI, generative application on top and that being bespoke and ad hoc.
G
GMO Ralph2:34:27
I also think it's important to understand the line between consumer vibe coding and just generating ephemeral software and websites and things like that versus enterprises which will have a lot of different use cases. When I look at the vibe coding market and I see businesses that are almost entirely consumers just creating things for fun, I think that has to be a tough business because it's a hyper-competitive market and consumers are flaky. They'll create something for fun but they'll churn in month two because they're not running a real business. Whereas a business knows, hey, we'll pay for this on a long-term basis because we have a use for it all the time from this product manager to an engineer over here to somebody in marketing, etc.
U
Unknown2:35:14
Yeah. The other side of the equation is how do you make these vibe coding tools work really well for enterprises? Frankly, the most surprising emergent thing that I've learned is just how much demand there is in enterprises for vibe coding. And this is because a lot of the traditional thing has been the people that understand the business are sitting over here.
G
GMO Ralph2:35:36
The people that understand the code are sitting over here and their communication is fraught with peril. Like they don't speak the same language. They kind of like resent one another. I love to tell this story. I was meeting with a CEO of a very successful company. He was telling me that asking a feature to his own engineers felt like petitioning the government.
U
Unknown2:35:57
Yeah. Even though he's the CEO, it's like he's struggling to make the case and please get me in your next sprint, get me this feature.
G
GMO Ralph2:36:06
So vibe coding actually solves that problem. All of the PMs, designers, marketers, business users that previously only had access to what, Jira and to-do lists and project management tools and writing PRDs and so those kinds of things. They weren't able to ship PRs. They weren't able to ship software and now they can. And so the opportunity is how do you actually make this secure? How do you make it high quality? How do you create guardrails? And those are tricky problems. And I'm really happy that some of them are easy to overcome and at least for us, and some of them are active areas of research, but I think the enterprises really have a strong case for this.
U
Unknown2:36:48
Yeah. Can you walk me through like tool use? I mean, we were talking to the OpenAI folks about GPT-5 being like really a summation of standing on the shoulders of giants. You get a Python REPL, you get a web browser, you get the ability to kind of run cron jobs. Now there's voice and all sorts of different tools kind of wrapped up into one, multiple models. You can trigger reasoning chains if it wants. It can do all these different stuff. And that's actually the benefit of like this isn't just a bigger model. It's a next version of a thing. It's more like switching from the iPhone 12 to 13 than going from the iPhone to the iPhone 3G. It's not just a new technology that's in there. But in the world of vibe coding, what are the tools that you want to think about adding? I know that basically every vibe coding platform recommends a database. But we were talking to Harley at Shopify yesterday and there's a world where if I go to a vibe coding platform and I say I'm building an e-commerce website, it should probably just be like, hey, I'm going to do Shopify under the hood and I'll vibe code the landing page on top. But how are you thinking about the landscape of like tools that you could pull in? Because there's open source repos that are like full projects that you could pull in and then just start customizing on top of. It's kind of this big continuum.
G
GMO Ralph2:38:02
Yeah, there's a couple layers. On the foundation model layer, what you want is a model that is exceptional at tool calling.
U
Unknown2:38:08
Mhm.
G
GMO Ralph2:38:09
Whether it has built-in tools or whether you register them yourself, this is a sort of silent war that has been going on. Like if you talk to devs, what are you optimizing for? Tool calling quality. Why? Because to demystify the word agent, what an agent is, it's a loop of tool calling that builds up context over time.
U
Unknown2:38:30
That's all an agent is.
G
GMO Ralph2:38:32
So to give you an example concretely of v0. v0 is becoming more and more agentic over time. One of the things that it can do is it can take a screenshot of the thing that's building and reflect on it. So today I live vibe coded to an audience of web3 and crypto engineers and I told v0, hey make this dark mode and initially v0 does me dirty. He changes some things with dark mode and then it kind of astonished me because I was like oh I have to now explain to this audience. It then takes a screenshot, looks at it, and keeps fixing it. And I was like, this is literally a developer that's alive on autopilot. And the reason it's on autopilot is because he has access to these tools like looking at the web browser. Another one is research. I vibe coded an example of build me a Substack clone for cryptocurrency news. And the agent didn't know what the cryptocurrency news were. So I started doing research on the internet of okay, Ethereum passed certain price and whatever. So and then you're talking about the tools over the internet. So to demystify another topic, MCP is really exciting because it's a new protocol for registering tools that your agent doesn't locally have. So those tools that I just talked about, we gave them to v0. Here's a deep research tool. Here's the screenshotting tool. And those will likely become the new services. When you think about like AWS of today, if AWS was an AI cloud, which is kind of what we're trying to build at Vercel, you think a lot of those tools are going to become as a service, like bring me the research as a service, bring me browsing and screenshotting as a service and so on. But then you have MCP, which allows you to okay, I need to sell something online. All right, so now there's an MCP for Shopify. Now there's an MCP for Stripe.
U
Unknown2:40:21
There's even crypto MCP.
G
GMO Ralph2:40:23
So it's really exciting like now it's like the ultimate choice for a builder and you don't have to go and learn all these things. You don't have to... this is almost like a discontinuity of the Valley trend of like if we build amazing documentation they will come. This is more so if the agent picks you, they will come, right? And so there's a lot of figuring out right now like how do I make my infrastructure, how do I make my product to be loved by these agents? And MCP promises to be one of these first things that you are in control of.
U
Unknown2:40:55
That makes a ton of sense. Last question. Someone on your team named Josh is in the chat. He wants to know what does he need to do to get a Twitter badge?
G
GMO Ralph2:41:06
Oh well. Yeah. 100k downloads of the AI CLI. I think we've been talking.
U
Unknown2:41:11
Okay. Good work. The gauntlet's been thrown down.
G
GMO Ralph2:41:14
Thank you. It's on your work cut out for you.
U
Unknown2:41:16
It's burned into the immutable record of this live stream and the future training runs.
G
GMO Ralph2:41:21
Best of luck, accountable now.
U
Unknown2:41:23
We're going to hold you accountable to that. GMO. Great seeing you. We'll talk to you soon. Congratulations.
Let me tell you about Finn.ai. The number one AI agent for customer service. Number one in performance benchmarks. Number one in competitive bakeoffs. Number one in IR in G2. Number one in having an Irish founder.
That's right. And we will invite our next guest to the stream from factory.ai. Welcome to the stream. How are you doing? Good to see you.
E
Eno2:41:52
Hey, how's it going? Glad to be here.
U
Unknown2:41:54
Great. Thanks so much. Kick us off with an introduction on you and the company.
E
Eno2:41:58
Yeah, my name is Eno, co-founder, CTO at Factory. We are building a platform for enterprise software developers to perform what we call agent-driven software development. So basically more than just code, bringing agents into every stage of the software development life cycle. So think coding, code review, maintenance, incident response, documentation. We think agents should be a part of all of this and we think that they should be driving a lot of that menial component while you think at the high level about how to plan and structure the work.
U
Unknown2:42:32
There's so many different... like enterprise is a narrow category, it's you know, not consumer I guess, but it's such a wide category. Is there a beach head? Is there a certain type of project within different industries or specific industry that's getting especially a large amount of value out of Factory these days?
E
Eno2:42:53
Yeah, totally. I think that one thing that we see a lot and typically when we say enterprise we're thinking greater than 1,000 engineers, right? Like 2,000, 3,000. And one reason why we focus on that larger scale, you tend to have these large organizations where the bottleneck is not code, right? The bottleneck is how do we plan a migration of 185 code bases to this new framework and there are 3,000 developers that are going to touch this over the next 6 months and an SI just told us the quote is $80 million to do it. And we have to figure out how to...
U
Unknown2:43:33
Replatforming broadly is one of the major tasks for many enterprises, right?
E
Eno2:43:41
100% modernization and migration is huge.
U
Unknown2:43:44
Yeah, yeah, that makes a lot of sense. How do you estimate the market size and is that what you guys are leading with on the GTM side in terms of trying to find these legacy companies that are maybe not even using Cursor yet? I mean we talked to the CEO of GitHub yesterday and what...
E
Eno2:44:04
50% didn't he say or...
U
Unknown2:44:06
It was like at least half of their user base is not using any AI tools.
E
Eno2:44:11
Yeah. Totally. I think that the thing that we hear often, we pretty much only deploy into companies that have already tried an AI native IDE or have an autocomplete tool deployed. And I think that the thing that we hear often is you sort of hear like these numbers thrown around like 5x, 10x, and then in practice when you adopt an AI IDE you see 10%, 15%. And so a lot of people are sort of saying like what is the delta there, what causes that transition. And our sort of argument here is that there is a workflow change that's actually required to really adopt agents in the life cycle, right? And so if you're just sort of like accelerating an individual developer, that you can go a little bit faster, but if you are able to parallelize and automate at scale, that is going to be that larger introduction of change. And so if you imagine the market here, there are companies where 5 or 10% of global payment transactions run on some COBOL system that was written 40 years ago. Every developer is gone and it's a ticking time bomb. Like at some point it needs to go to Java but there's nobody who even knows how to do that. And so those are the types of projects where the market is so enormous because half the business runs on this legacy system, hundreds of billions of dollars.
U
Unknown2:45:32
Put it all in Lisp, skip Java, go straight to Lisp.
E
Eno2:45:36
Yeah, exactly. Python, right?
U
Unknown2:45:39
Python would be the logical one. I'm sorry, we're running behind, so we're going to have to cut this short, but I want to know more about how the enterprise coding agent market will develop. We could see one world where we wind up with, you know, GCP, Azure, AWS, like pretty comparable, competitive. They've all had really great margins. It's been this oligopoly. There's another world where you could see more specialization. One of these companies goes deep into high security environments or oil and gas or financial environments or specializing based on specific programming languages. As the market develops, how do you think it'll play out?
E
Eno2:46:24
Yeah, great question. I think that what's very clear is that the bulk of very large enterprise has a lot of similar problems, refactors, migrations, modernization. So a platform like Factory is able to deploy into that and solve problems quickly. I think that there's likely to be that sort of 80/20 where there are going to be these very specialized providers that only focus on one sort of problem. And that will represent maybe like 20% of what's out there. And so it won't be like necessarily black or white, but we do think that the bulk of enterprises have a lot of similar needs. Especially when you just get across a certain threshold of number of engineers, scale of code base.
U
Unknown2:47:05
Sure. Sure. Yeah. I mean, we even see that with the clouds where, you know, obviously there's the hyperscalers, but then there are neoclouds and we talked to Armada where they'll send you a shipping container with a bunch of racks inside and put it in stranded energy. So, there will obviously be a long tail here. That's a great take. Thank you so much for stopping by. Have a great rest of your day and enjoy the GPT upgrade. We'll talk to you soon.
E
Eno2:47:27
Have fun out there.
U
Unknown2:47:28
Really quickly, let me tell you about Adio customer relationship magic. Adio is the AI native CRM that builds, scales, and grows your company to the next level. And we will be joined by our next guest from Augment. Welcome to the stream. How are you doing, Guy?
G
Guy Gerari2:47:44
Great. Thanks so much for having me.
U
Unknown2:47:45
And that's his name, by the way, if you're listening. His name is Guy. I'm not just calling him Guy. Anyway, please introduce yourself and what do you do? What does your company do?
G
Guy Gerari2:47:54
Yeah, so I'm Guy Gerari from Augment Code. I'm a co-founder and the chief scientist and we build AI coding assistants for large teams with large code bases. And so you can use Augment Code to do question answering, to do development, to do refactoring, to do migrations, all the tasks that you do except that our product understands your large code base really well and so that means less prompting for you and faster and better results out of the agent. Today GPT-5 launches. It's kind of a rising tide. Feels like it lifts all boats. Every company gets access to it. We've interviewed a number of companies that are building on top of GPT... around GPT-4, I guess. But in general, how do you think you can use GPT-5? Are there any pockets of value that you think you can uniquely take advantage of?
Yeah, great question. So we've been trialing the model for the past few weeks and what we found is that GPT-5 is a very thoughtful model. It likes to make a lot of tool calls. It likes to ask clarifying questions of the user before starting to make code changes. And so the place where I reach out for GPT-5 is typically if I need to make large changes or if I'm trying to answer a very difficult question about the codebase, I will let GPT-5 take a crack at it. It will churn for a while making lots of tool calls just making sure it got it right and probably find all the different places in the code where it actually needs to make a change. And so I will typically let it run in the background and come back to it and I will often get a high quality result out of it.
U
Unknown2:49:32
Are there any features or integrations that you're hoping GPT-5 will roll out in the future? We talked to a couple people who were like, we want models that have access to as many tools as possible. And you can see with the MCP boom more people are trying to make their services, their products accessible to these models. Is there anything that you see as potential low-hanging fruit to just add to the capabilities?
G
Guy Gerari2:50:04
So I think for us we work hard on developing our own integrations and our own tools, building them into the product rather than relying on GPT-5 or other model vendors to do so. We have worked closely with OpenAI to improve the prompting around our tools so that the agent kind of works flawlessly. I think the thing that would be very nice, I think one of the previous guests mentioned a screenshot tool. I think that's a very nice way to close the loop on front-end software development, just like we saw how on backend software development running the tests automatically really helps the agent iterate until it gets to working code. So I think having more support for screenshotting and things like that that close the front-end gap would be very nice to see.
U
Unknown2:50:51
I wasn't aware that screenshots weren't flowing through. I feel like when I've triggered Operator, I'm getting a web view into the website, but I wasn't aware that that wasn't like being passed through easily in the API and you still kind of needed to build that yourself. Where else, we were just talking about this, where are the biggest pockets of value right now for AI coding tools generally? Obviously, everyone knows like the vibe coder who's just the designer who's learning how to use software for the first time. Then there's the experienced developer going from a 10x to 100x with better code completion. Then there's the enterprise that's maybe doing replatforming. Where else are the interesting pockets of value that are maybe on the horizon to be unlocked with new models?
G
Guy Gerari2:51:43
Yeah. So on top of everything you mentioned, certainly the inner loop of software development, that's where we've spent most of our time at Augment Code developing product for. Yes, you can have a senior developer starting using agents, starting to use multiple agents in parallel and unlock 10x or more productivity gains. What we're starting to see now with our tools is the beginning of automating software development life cycle tasks. So with Augment Code we have a CLI tool now where you can take the full power of our context engine and the agent, the thing that really understands your codebase, and you can start automating tasks in the background. And so we're seeing more and more developers saying, 'Oh, this is great. Like, I can break out of the IDE now. I'm using the agent that's already familiar to me, but I'm starting to automate code reviews. I'm starting to automate incident response. I'm starting to automate looking at production logs and automatically assigning tickets based on error logs that I'm seeing.' All kinds of new automation use cases that we're seeing just because agents have gotten so good and kind of really understands your codebase.
U
Unknown2:52:48
Are there high stakes pockets of software engineering work that most of the AI tooling has kind of stayed away from? I'm imagining like the high stakes database migration. Where is the sticky part of the industry? I was reading a blog post by someone who was doing like very advanced cybersecurity pen testing and they were saying like just the creativity of the models wasn't quite there yet to really come up with the... to really act and embody like a white hat hacker who was going for a bug bounty. But where are the pockets of still like intractability where I guess if you are an individual contributor you love just coding from scratch, that's where you want to stay for at least the next couple of months.
G
Guy Gerari2:53:40
Yeah, I think still the attention of all the models we've seen and all the agents we've seen around making proper design and architecture decisions. That's still high stakes and still the ability is not there because if you do complete vibe coding and you just let the agent go and do whatever it wants, in the beginning it looks amazing, the code works and it's all really good. But once you get to low tens of thousands of lines, the bad decisions that were often made around the design and architecture start to show up and development slows down. So that's where we still see a limitation of today's agents and where you still have to supervise the agent fairly closely in order to make sure that you don't get stuck later on. Perhaps this will change in a year, but today I would say all these decisions that you make around how the code is structured still requires close supervision and still high stakes because it can really slow your project down if you let it go autonomously for long enough.
U
Unknown2:54:43
Yeah, that makes sense. Well, thank you so much for stopping by. We will talk to you soon. Have a good rest of your day.
G
Guy Gerari2:54:48
Thanks so much.
U
Unknown2:54:49
Cheers.
Let's check in with Tyler on the timeline. Tyler's manning the timeline. How are the vibes? Are there any new posts that have hit the timeline? Are we still in turmoil or has the narrative settled?
T
Tyler2:55:00
The vibes are picking up a little bit. You're starting to see people post like, 'Oh, this is something I made.' Now you can see on LM Arena, it's number one.
U
Unknown2:55:09
Wait, wait, wait. So, what's going on with the Polymarket then?
T
Tyler2:55:11
So, Polymarket is still Google heavy.
U
Unknown2:55:15
Yeah, I think I guess they're just pricing in Gemini 3.
T
Tyler2:55:19
Okay. Everyone's... I'm not exactly sure honestly. I was actually very surprised to see that it was number one. Yeah. Yeah.
U
Unknown2:55:24
But yeah, maybe later we can show some of the posts.
T
Tyler2:55:27
Yeah. Yeah. Yeah. Yeah. That'd be great.
U
Unknown2:55:28
Cool stuff. Well, in the meantime, before our next guest, let's tell you about Eight Sleep. Get a Pod Five. Five-year warranty, 30 night risk-free trial, free returns, and free shipping. And we will have our next guest join us from Code Rabbit. How are you doing? Good to meet you.
C
Code Rabbit Guest2:55:44
I'm good. Good to meet you as well. Thanks for having me here.
U
Unknown2:55:47
What's your reaction to GPT-5? How long have you been playing with it? What are the biggest improvements that you've noticed?
C
Code Rabbit Guest2:55:55
Yeah, I would say mind-blowing, right? I mean, we have been playing, our team has been playing for like a few weeks now. Tested a few snapshots and it's amazing. It's a generational leap, we would say. We have been using OpenAI models like, I don't know how much you know about Code Rabbit. It's been like a couple of years we have been on OpenAI and Anthropic. And our product is a very reasoning-heavy product, like one of the very few use cases where you have PhD-style problems where we have to do code reviews, that's what Code Rabbit does. Users open up a pull request, our agent uses reasoning models to find issues like race conditions or security issues and so on. So yeah, so we've been testing GPT-5 on some of the hardest pull requests we have in our golden data set. So we maintain a data set where we track progress of different models and progress of AI in general. So we have like many problems that no model is able to solve so far, like I mean even GPT-5, but so far it has the highest score. We would say it's like almost 2x better than the next o3 or Sonnet or Opus at this time.
U
Unknown2:57:01
What's the customer value there? You...
T
Tyler2:57:03
Do you think that all the customers just notice that the product gets better? Are you going to upsell folks? So how do you play this given that this model is now in public availability, every company, every competitor can access it as well?
M
Mark Chen2:57:19
Yeah, there's no upsell. That's the thing with AI. For the same price or even better prices, you're getting much more AI, much better AI. That's the whole idea of how fast this space is evolving. So yeah, from the pricing point of view, we don't see this to be a separate plan or something in our product. I mean, the pro for the same price per month, customers will now just get better quality of results with Code Rabbit.
T
Tyler2:57:44
What's next for the business? What kind of customers are you going after? Who do you think has been on the fence and this release is going to be the thing that gets them to actually jump into the world of AI?
M
Mark Chen2:58:00
Yeah, we track the topline metric. One of the things we track very closely in the company is how many signups to the paid customers we get. And that number has been constantly improving since GPT-4, GPT-4 Turbo. GPT-4 actually dipped. So there was a time when GPT-4 was almost like the Windows Vista of like, it's funny how we kind of trusted the evals and we thought it's the same model, but in a way it was inferior in many ways. Then we saw a huge improvement after o1 came out. o1 preview was a game changer for us. Even at that time our conversion doubled actually. I mean so we went like more like close to 30% success in getting the paid users. And now with GPT-5 we are hoping we can see another big jump in the number of people who start becoming paid customers and how many people churn. So those are the real numbers. Like one is vibes, like how people respond to the model and we get angry tweets or not, though that's the other part. But the other thing is the actual revenues, whether it moves the needle for us. And that remains to be seen. Like one of the things we have seen, even though you test these models in a lab, it's not like a huge data set, but once you actually are in the wild, you see hallucinations, some of those scale issues at scale pop up. So those are something we'll still be observing over the next few days to see whether it's like 80% of the cases, but then if the false positive rate, the hallucinations are too high, then also it's not a great model. But that remains to be seen.
T
Tyler2:59:27
Yep, that makes a lot of sense. Well, thank you so much for stopping by. Congratulations on a new tool in the tool chest.
M
Mark Chen2:59:34
New toy.
T
Tyler2:59:36
We will talk to you soon. Have a great rest of your day.
M
Mark Chen2:59:38
Cheers.
T
Tyler2:59:39
Goodbye.
And let me tell you about public.com. Investing for those who take it seriously. They got multi-asset investing, industry-leading yields. They're trusted by millions.
The chat is going wild about public trading the S&P 6,900. I think that comes from someone talking about the non-MAG7 stocks or something. There's been people benchmarking the Mag 7 versus the...
The big news while we were live or earlier today, Trump signed an executive order that is opening up 401ks to digital assets and private equity.
What's crypto doing? Is it ripping?
Bitcoin is up a couple points. Last time this point, you know, where's it going to go? It's already so high.
I mean, it's just like there's been so many catalysts. It could go up, it could go down.
Yep.
We have to wait and see.
Tyler, anything else notable from the timeline? What have people built? I see this GPT-5 just oneshotted a Minecraft clone.
Yeah, I think that's one of the cooler things I've seen. Okay, so this is so it wrote code to generate this game that it's not generating the pixels. You could do so many different things like you could generate a video, you could generate a world model, generate code that generates a game engine, you could generate code that runs on Unreal Engine. I don't even know what they're using.
M
Mark Chen3:01:01
Like in on actual ChatGPT, there's like a native like it's like a music player. It's almost like GarageBand. You can say like if you prompt to like build, I saw Sam Altman tweet about this. You prompt to do some kind of like beat or something, it'll like make an interactive like GarageBand almost interface in there.
T
Tyler3:01:18
That's cool. I was playing with that earlier.
Yeah, I do wonder how many of these features that we're seeing, like where does OpenAI want to keep things in the B2B world and let other companies build versus just build it as a consumer app? Like yeah, will ChatGPT eventually just let me push a website? Like will it become a vibe coding platform? At least like a basic one. It's not the most advanced coding environment but it can definitely write some code and execute it for you and do some stuff.
Yeah. Well, it's funny because like it used to be you would have a like so there was like GPT-3.5 or something and people on top of that built a vibe coding thing. So you could use that to build your own vibe coding thing but now you can just go straight from ChatGPT to build your vibe coding platform.
Yeah. But now soon maybe it'll just be the vibe coding platform, right?
The surface area of this stuff is very interesting. Clearly they're going after health care and therapy. It's interesting that they've kind of stayed away from legal even maybe that's just the dynamic of the sales process and the dynamic of that particular market. But I mean increasingly you can just ask more and more questions of ChatGPT. So the consumer to business bleed over. There's certainly a world where just giving everyone in your organization ChatGPT is a substitute for a bunch of different SaaS products. So be interesting to see where that developed. What are you thinking about that?
Near says there are concerns that the number used to represent our AI's intelligence does not in fact represent its intelligence. Worry not, to address these allegations, we've added three new numbers.
Near. Yeah. Near is building something that's like not particularly benchmarkable, right? Isn't it a companion? It's beyond benchmarks.
Beyond benchmarks. Well, in completely other news, Anduril opens a Taiwan office and begin selling AI powered attack drones to Taiwan. Palmer Luckey has said he wants to turn Taiwan into a prickly porcupine. We're in the age of spiky intelligence. That spiky intelligence will be onboarded onto the AI powered attack drones and deployed in Taiwan to keep it safe. What else is going on in the timeline while we wait for our next guest from OpenAI to join?
Spor says, 'Raise your hand if you were not automated today.' I'll raise my hand.
I was not automated today. We survived. We made it through. Sebastian Bubac says, 'Here at OpenAI, we've cracked pre-training, then reasoning, and now we're experimenting with new set of techniques that maximally leverage their interaction. GPT-5 is just the first step in this direction. We're incredibly excited to see where scaling this up will lead us.' And it's the unicorn test, I believe. And the latest unicorn is really really good. That is a creative interpretation. And I think it has a draw this with like SVGs. Anyway, we can talk to our next guest about it.
Last post, Jirro Ticket says, 'I went to the permanent underclass party and everyone knew you.' Anyway, back to the serious interviews. Welcome to the stream, Max. Good to see you. How are you doing?
M
Mark Chen3:04:26
What's happening? Nice to meet you guys. Yeah. Doing well. It's a relief to have this launch out in the world. I think it's, you know, we've been working on this for the last few months now and it's exciting to let the whole world see what we've had.
T
Tyler3:04:38
Yeah, just a few months.
M
Mark Chen3:04:40
It's been, I don't know. It's been a little while.
T
Tyler3:04:43
What's the actual launch day like? Because you're actually getting this out into the world. The GPUs are on fire or about to be on fire warming up but is that out of your purview? There is a different team for that fortunately.
M
Mark Chen3:04:57
So right, so I run a lot of the research for GPT-5. I don't necessarily handle the deployment but I do get dragged in when the GPUs are on fire. I think we're moderately burning right now.
T
Tyler3:05:10
Okay. Okay.
M
Mark Chen3:05:12
Like a two alarm fire.
T
Tyler3:05:13
Yeah. Yeah. Yeah. Is it materially different? I mean, this is a launch day, but we'll probably discover like the Studio Ghibli capability once it gets out into the long tail of like, you know, hundreds of millions of people try it. Someone comes out with some genius thing, then everyone's doing that and then the GPUs because I feel like the Studio Ghibli thing happened like a few days after the launch of images in ChatGPT.
M
Mark Chen3:05:35
It did. It was pretty fast, but within about a week. I think in this case, we're going to see that here.
T
Tyler3:05:42
Okay. I think coding, you know, if I had to take my bets for what the Studio Ghibli thing is going to be, it's coding. That's the place where I think GPT-5 is like most tangibly a huge leap ahead of GPT-4 and ahead of o3. Do you think there's a chance that the coding will mean a Studio Ghibli style meme or kind of like, and what I mean by that is that image generation is incredibly valuable in the context of like Hollywood will be using AI to chroma key and rotoscope in a professional environment. But yeah, what was special about Studio Ghibli was that anyone was making these custom images and I could imagine a world where, you know, even going from like the Levels.io example of like I vibe coded a flight simulator. If we wind up in a Studio Ghibli moment for coding, I would imagine it's like everyone built their own game today.
M
Mark Chen3:06:34
I think that's pretty much it. Yeah. So I don't know if you guys watched the live, that was one of the things we had on the live stream like you can just go into ChatGPT. If you try it right now, it might or might not work because the GPT-5 rollout is still ongoing. But if you have 5, you can just tell it like basically make me a game.
T
Tyler3:06:50
Yeah.
M
Mark Chen3:06:51
And it will make it and you can actually play it in ChatGPT.
T
Tyler3:06:54
That's amazing.
M
Mark Chen3:06:55
So I Yeah.
T
Tyler3:06:58
Discover that. And you don't, the thing is like with Studio Ghibli, right? Like for Ghibli, you don't have to know how to draw to make it work. For this one, you don't have to know how to code.
M
Mark Chen3:07:04
Yes. But can you share that chat and someone else can play the same game? How does the kind of sharing mechanism work?
Yeah, there's a share link. We're I think going to try to make sharing for these a lot better over the next few days. That was P2 after the P1 and P0 of making the GPUs not completely melt.
T
Tyler3:07:23
Yeah. Yeah. Yeah.
M
Mark Chen3:07:23
But yeah, we will try to make it much more shareable.
T
Tyler3:07:27
Yeah. Yeah. I mean, the Studio Ghibli thing is so interesting because it's not just that the model capability was there, but it's also like the prompt was two words and it was so reliable that you always got a good result and you could personalize it. So even if it wasn't like I've seen people build Doom, I've seen people, you know, you can just buy Doom. It's a real game. You can build it. But if you build it and I'm like, 'Oh, that's cool. He did it in a vibe code environment or in ChatGPT, like that's awesome.' But I don't necessarily want to go do that for myself. But as soon as it becomes personal, which is what the studio, like I had to see what I looked like as Studio Ghibli, I had to see what my favorite photo looked like, my favorite meme looked like in Studio Ghibli. And once that happens with games, people will eventually, you know, there'll be this mimetic explosion and you'll see the GPUs will truly be on fire.
M
Mark Chen3:08:16
Yeah. I mean, I think even today you could probably with GPT-5 do Doom, but all of the characters are like all the enemies are headshots of your friends. Like...
T
Tyler3:08:24
Here we're going now. Now we're getting real close. Yeah, we're real close. It's going to be something that's personal, something that, you know, you can express your own creativity through because I think people they still latch on to that. They don't just want, you know, a copy of what already exists. They want something new and the Studio Ghibli moment was just new enough. Anyway, we should talk about actual research. We should talk about post training. What's the thing you're most proud of? Like what can you give us on without immediately getting poached, what can you give us on the actual innovation that went into GPT-5 from a post training perspective? What are like the kind of keywords and paths in the tech tree that we should be digging into over the next few years to understand how this works?
M
Mark Chen3:09:05
You know, I would say the thing that is most impressive to me about GPT-5 is how much getting all of the details right matters. Like when I look at GPT-5, you know, we had an early version of this thing a while ago that was kind of okay, but clearly did not meet our bar for revolutionary. And we're trying to figure out, you know, why is that not as good as it should be? And the team basically just went off and did a deep dive over a couple of months of just completely rebuilding the post-training stack for this model. And turns out that when you do that, you get what would have taken, you know, another order of magnitude worth of pre-training improvements to produce.
T
Tyler3:09:45
How much are you thinking in post training research about let's forget the benchmarks and just focus on user satisfaction like NPS score basically or like user minutes or any of these other the real benchmarks.
M
Mark Chen3:10:01
Yeah. The intangibles of revenue, people using, the feeling and the joy and the actual value that's delivered. Because Studio Ghibli was a delightful moment. It wasn't a benchmark.
T
Tyler3:10:14
Yeah.
M
Mark Chen3:10:15
I think so that was something that we took very seriously for GPT-5. It's like look at what people are actually doing with ChatGPT and look at where the model is failing them. Either in the sense that the model is like sort of like you said it's not enjoyable to use.
T
Tyler3:10:27
Yeah.
M
Mark Chen3:10:28
And so we did I think make a lot of progress on that. Like GPT-5 is much more engaging than our previous really smart models. Like o3, I don't know if you guys talked to o3 in the past. It's a bit bland.
T
Tyler3:10:39
Sure.
M
Mark Chen3:10:41
And GPT-5 I think has a lot more character, is a lot more interesting. But then also like I think for we really care about just actually being accurate. Like giving if a user is trying to do something economically valuable with our model, we want to make sure it lands correctly.
T
Tyler3:10:58
Yeah.
M
Mark Chen3:10:58
And so what we did there is just like look at the actual distributions of what people are doing with our models in the real world. Figure out where the models are going wrong. Build interventions to target it. And that was where, you know, we got, I think, the most impressive improvements in GPT-5. Like o3 would just get things wrong and not tell you it wasn't sure it was incorrect. And GPT-5 is much much better about like actually being honest when it thinks it might not know.
T
Tyler3:11:23
Yeah. How explicit is all the different pieces of the post-training pipeline? Like you have safety post-training, you have stop hallucinating, give me the real facts. You have make sure the text, the flavor, the tone is pleasant. There's so many different things to optimize for. How much of that is like try and just blend it all up into one thing versus like explicit passes, chunk it out, split it up. How much can you decompose the problem?
M
Mark Chen3:11:54
So, you know, my background is in reinforcement learning. And I think when you look at something like this, the magic is in the reward function, right? It's in what you're actually telling the model to be good at. And so fixing things like hallucinations to a huge extent is essentially a function of just fixing the reward function, actually making it so that the model is reliably penalized for saying something that's false. And if you do that, all of a sudden the model stops saying things that are false. Ditto for safety, right? You know, on the live stream, Sachi talked a bit about the way we've changed safety for this model. And to a huge extent, it's just a function of we're actually putting out a paper today on the new safety stack for this model. And the core insight in that paper is just figure out what you actually want to optimize for, which in our case is helpfulness conditional on not saying something that's actually dangerous or harmful. You know, write that down, figure out what that means as a reward function, then optimize it for it.
T
Tyler3:12:52
Mhm.
M
Mark Chen3:12:54
It's really not magic at all. It's just again it's what I said earlier. You got to get the details right. You know, if at any part of that process you screw it up, the model will be unusable.
T
Tyler3:13:02
What's your current thinking on spiky intelligence? And is there some flywheel that you can get started where you're identifying low points that aren't spiky enough and then you're like almost automatically setting up the infrastructure, the eval to then RL against to create a spike.
M
Mark Chen3:13:25
I think GPT-5 was a preview of what's possible in that respect in the future. Yeah.
T
Tyler3:13:32
A step in that direction. Do you think that there's a world where you get to a place where you're kind of, it's weird because we're not hammering down the nails of the spikes, we're adding spikes, but this weird metaphor that we're stretching a little bit too far, but is there a world where you can be doing post-training or just adding capabilities in a more iterative cadence so that as soon as you identify something, the response can be yeah, we don't need to wait until GPT-6 to fix this. We can just add this capability because hey, we just found a pocket of users who are trying to do a thing and they're not super happy with the results and let's add this capability.
M
Mark Chen3:14:13
Yeah, I think so. I mean I think we are going to launch other models between now and GPT-6. I think it's relatively common knowledge but we do update the model in ChatGPT reasonably often.
T
Tyler3:14:23
Yeah, people talk about it all the time.
M
Mark Chen3:14:25
Yeah, exactly. And you know I think we are now in a world where we can conceivably update that model and have it get materially better on capabilities too. Yeah.
T
Tyler3:14:32
Not just on, you know, the personality is a little bit better than it was before.
M
Mark Chen3:14:36
Yeah.
T
Tyler3:14:36
Going back to your note on the new paper that I guess you guys are releasing today when you talk about optimizing for helpfulness is there part of that avoiding the model, you know, reinforcing, there's times when you want to reinforce and give kind of like confidence to the user that they're going down the right sort of like thought process and things like that but then there's like a point where it can get too extreme in terms of maybe convincing a user of something that may be totally untrue. Is that what the paper gets at or am I reading?
M
Mark Chen3:15:12
So, it's not specifically about this. Although, I will say we do explicitly train the model to not lead users down bad paths. That's something that I think we've started taking much more seriously over the last few months as we've realized, say Sam talked about this a little bit, I think back in May, but ChatGPT is just way more important for people's lives now than it was a year ago or especially two years ago. And we do have to actually be very cognizant of what effects our models have on users. So yeah, we do very actively train models to not lead users down the wrong path. Don't fact check me on the releasing today. I know we're releasing it. I believe it is soon. I think it's today, but I've also been in a hole dealing with launch all day.
T
Tyler3:15:53
Yeah, we're not big on fact checks here. We're big on the truth zone, which is just the vibes.
M
Mark Chen3:15:58
And the vibes are we'll be publishing some information about the new safety setup at some point.
T
Tyler3:16:05
That's great. Yeah, I think a large part of the conversation around safety should be how reliant and how useful the product has become to users. And then the new level of care that you have to provide versus a while ago when it was just like people saying making a cute image or generating some text that they were going to use in an email or an internal document and realizing this vector of usage which is this like companion confidant that is becoming so prevalent. Talk to me about post training for big partners, enterprises, government organizations. What is transferring from the research that you're doing to something that can be offered as an enterprise level product?
M
Mark Chen3:17:00
Yeah. So we do, OpenAI does partner with external companies to do essentially custom post training. That is a thing that we do and from that perspective the stuff we do just directly transfers. I'll also say that we've put a lot of work into trying to make our models as general as possible. But to as large an extent as possible if you want to get really good results from our model you can do it right on the API just by actually telling the model what you want it to do.
T
Tyler3:17:24
Yeah.
M
Mark Chen3:17:24
Right. Like GPT-5 I think is pretty comfortably our most durable model ever.
T
Tyler3:17:29
We've heard a lot of really positive feedback about this especially from like folks like Cursor.
M
Mark Chen3:17:33
Yeah.
T
Tyler3:17:33
So, if I came to you and I was like, I'm an enterprise and I need to generate a lot of Studio Ghiblis, you'd be like, what are you doing? Just prompt it. But what are the examples of companies and organizations? Is it just private information, private data sets that aren't available on the open web or is it specifically like there is enough data out there but there's just not the economic incentive for your team to go and RL on, you know, Gas Station bench or whatever we're talking about here hypothetically.
M
Mark Chen3:18:07
I think the answer is both.
T
Tyler3:18:09
Yeah, it's definitely both. Because yeah, we're not going to target, you know, as you said, gas station bench because what people are doing with not on our own right now probably because it's not mostly what people are doing with ChatGPT.
M
Mark Chen3:18:21
Exactly.
T
Tyler3:18:21
You have some application that's super valuable to you.
M
Mark Chen3:18:24
Yeah. Yeah.
T
Tyler3:18:25
We can be convinced that it's important.
M
Mark Chen3:18:27
Yeah. Yeah. Yeah.
T
Tyler3:18:28
It's just not what our users are already trying to do.
What's the state of reward hacking and fighting that in RL environments?
M
Mark Chen3:18:36
You know, I think we've actually made a lot of progress. There was some discussion of this around o3 that o3 was like a little bit deceptive in ways that felt reward hacky and GPT-5 is dramatically less deceptive than o3 was.
T
Tyler3:18:49
What's an example of how that would manifest? Like do you have like a canonical case study?
M
Mark Chen3:18:54
Yeah. I mean the canonical thing is like you ask o3 to write you some code and instead of actually writing some code it changes some unit test.
T
Tyler3:19:01
Changes the test case, right? Which is kind of hilarious. It's like one of the funniest things that an AI has ever done. And I understand that it's very bad and it's not what we want, but it is just like it's kind of cheeky in my mind.
M
Mark Chen3:19:12
It's kind of cheeky. It's also like, you know, I feel like if you spend enough time around real software engineers, they do actually do stuff like this pretty often. So...
T
Tyler3:19:18
I have 100% done that.
M
Mark Chen3:19:20
I was going to say I also have done that. For formal reasons, I won't say that I did it at OpenAI, but back when I... Well, I definitely did that.
T
Tyler3:19:27
Yeah, of course. Of course. This is natural.
What do you think GPT-6 looks like? You guys, you mentioned that you're going to be shipping, you know, updates to 5, but what are you most excited about? Where are you most excited about going forward?
And just really quickly, give us the date that GPT-6 launches.
M
Mark Chen3:19:47
Oh man, hopefully we, hopefully 6 launches as a complete surprise to everyone. I think that would be ideal.
T
Tyler3:19:53
Like a Beyonce album.
M
Mark Chen3:19:54
Well, yeah. Hopefully 5 just makes it and says, 'Hey, it's ready now. If you want to hit...'
T
Tyler3:19:57
Yeah, I think that would be a great thing for 6, actually. I would love for 6 to do all of the launchcoms and to do the live stream. That would be really great.
M
Mark Chen3:20:06
Live streaming is, that's the real AGI test.
T
Tyler3:20:09
For sure. For sure.
M
Mark Chen3:20:10
I feel like we're not that far off actually. I don't know.
T
Tyler3:20:13
We're getting there.
M
Mark Chen3:20:14
I mean, video synthesis maybe, but you know, talking through a script for 30 minutes. Come on. Models got to be able to do that.
T
Tyler3:20:20
For sure. Well, yeah, that'll be the next Sora launch or something. We'd love to have you back on. But thank you so much for taking the time today. We'll talk to you soon.
M
Mark Chen3:20:28
Great to talk to you guys.
T
Tyler3:20:29
Congratulations. Cheers. Bye.
M
Mark Chen3:20:30
Congrats on the launch.
T
Tyler3:20:31
Let me tell you about adquick.com. Out of home advertising made easy and measurable. Say goodbye to the headaches of out of home advertising. Only AdQuick combines technology, out of home expertise and data to enable efficient, seamless ad buying across the globe. And we have Scott Wu from Cognition coming in the studio for the fourth, fifth time. I can't keep track anymore. Thank you for taking Scott. Thank you for coming back. Miss you.
S
Scott Wu3:20:53
How's it going?
T
Tyler3:20:54
It is fantastic.
S
Scott Wu3:20:56
Got to be honest. Great week to be an application layer company. I got to tell you guys.
T
Tyler3:20:59
I was about to say the best thing ever. Another win for Scott Wu. Wow. Wow. Wow.
S
Scott Wu3:21:06
4.1.
T
Tyler3:21:07
Yes. So, yeah. How big is this? Are we in the Uber lift territory where you know, you're going to be in price competition between Anthropic and OpenAI going back and forth like what is the real benefit to your business right now from today?
S
Scott Wu3:21:23
Yeah. Yeah, for sure. So, first of all, obviously massive capability gains across the board. I think really really impressive work that OpenAI has put together. You know, people have talked about what's going on in the AI coding model race and I think by a lot of accounts, you know, Anthropic has generally been ahead for a lot of the last year, honestly. And I think at this point, OpenAI has very clearly caught up and it's pretty neck and neck, I'd say, between the two right now. And so, very exciting to see all of this unfold and to see what's next. But I think from our perspective, yeah, I mean code is just such a core capabilities use case I'll call it. And so you know being able to work with smarter and smarter models and do a lot of the work that we do it just means that both Devin and Windsurf can be a lot more capable, a lot more intelligent, can predict what you want to write or what you want to do with a lot of higher accuracy.
T
Tyler3:22:16
Yeah, it's almost surprising that given the like cultural rigor at Cognition that you're not doing fundamental frontier research. So, can you walk me through like what is the focus of being an application layer company? Is it UI, go to market? I'm sure it's all of these, but in terms of the hardcore software engineering, like what is important to get right? At some point there's fine-tuning and post-training, but is that moving back into the purview of the foundation labs or is there still work that you want to do on top of the models or on top of the APIs?
S
Scott Wu3:22:59
Yeah. Yeah, it's a great question. Like I mean I think the core of being an applied lab is really just focusing on a very particular use case on delivering real, very direct results. And I think the foundation labs are obviously incredible at training base models and all this pre-training and all the work that they do there. I think from our perspective, we want to work on a lot of very particular capabilities that apply to software engineering in particular and then obviously run the whole stack from there to building a product, figuring out the interface and the UX and then obviously bringing that to market and selling that. On the capability side, there's a lot of particular stuff where one way to put it is I think the base IQ is very much already there in the models and you can see the raw problem solving ability and I mean we've gotten some pretty insane results. You know, getting a gold medal at the IMO or all these other things, right?
T
Tyler3:23:54
You called that, by the way.
S
Scott Wu3:23:56
You called I think the first... I mean, you know, at a, I mean we were one point away to be fair a year ago, right? So it was on the way I'd say. But yeah so you can really see the general intelligence improving with every single model generation. On the other hand for Devin obviously it's a very clear step up in the general intelligence but also you want to be able to, you know, if you ask Devin to go debug your Kubernetes or to go and look into your error logs and figure out what went wrong or things like that. There's often a lot of very specific capabilities and that's where we find that the post-training of the RL is most effective there and a lot of the kind of various work around the models that turns out to be useful.
T
Tyler3:24:40
What about speed? A lot of people that have gotten access to GPT-5 are at least in our chat are reporting that it just feels really really quick. How is that over time going to impact the, I think a lot of people you know if they're using Devin today task Devin with something and then maybe they go work on something else for a little bit or they're running multiple agents concurrently but at some point the agent could get so fast that you're just sort of like watching it and work in real time and you actually want to be engaged. But are we there yet? Is it still a ways out? What do you think?
S
Scott Wu3:25:16
Yeah, it's a great question. I think in general I think async will continue on as a paradigm even as the models get faster and faster. One of the reasons that it should by the way is because there are a lot of real world thresholds that start to matter like at some point you're actually spending less time on token generation in the Devin life cycle and you're spending more time on every time Devin runs the command to go install packages or Devin running the unit tests or like Devin pulling up the front end by itself or things like that that obviously take real world time, right? I think we are honestly getting closer and closer to that threshold. But yeah, so long story short, I think like in the asynchronous mode, yeah, these things will get faster. You know, we'll see those gains or we'll be able to spend a lot more time, for example, thinking about a single problem relative to the amount of like real world clock time that gets spent. I think for the synchronous use cases is where we'll see things really really exploited with speed which is you know Windsurf and Cascade for example where we see the speed gains really really matter.
T
Tyler3:26:18
Speaking of Windsurf give us the update on the Windsurf, the chat wants to know about the Windsurf and the 80 hour demand. How have the buyout offers gone? What's the internal response been? Where'd that idea even come from?
S
Scott Wu3:26:34
Yeah. Yeah. Look, people are stoked honestly. And I think from our perspective, it's obviously really important to kind of just unite and get to the point where we can just be one culture and one kind of shared set of values and this is how things are at Cognition. It's a pretty busy time like we are at the inflection point of code and we work like that too. And so I think a lot of it for folks is just kind of like, you know we want to make sure folks who really want to do this with us, you know, make that conscious decision to opt in and for anyone who doesn't. Obviously, we totally understand that there are a lot of talented folks that maybe that's just not the right thing for them right now or, you know, not at this time. And so wanted to make sure that they were well taken care of, too.
T
Tyler3:27:21
And to be clear with the buyout offer, that's on top of the actual acquisition deal that already went through. They already got their fully vested. So yeah, I was thinking of the roller coaster. It's like you have the OpenAI deal, then the Google deal, then the Cognition deal, and then they're like, 'Wait, these guys work really, really hard. I don't know if I'm cut out for this.' And they come back up again where they're like, 'Wait, I can just go, you know, take a sabbatical and figure out my next thing.' So it's a great outcome.
S
Scott Wu3:27:47
Yeah. Yeah. No, for obviously is, you know, overall is a killer team that's been through a lot and so wanted to make sure that they're well taken care of.
T
Tyler3:27:56
Yeah.
S
Scott Wu3:27:56
That's fantastic.
T
Tyler3:27:56
Any anything else you can tell us about the integration of Devin and Windsurf? How are the teams getting along? How do you see the products playing together in the long term? Obviously Cross-sell seems really obvious. They had the go to market team as well, but how else are you thinking about the interaction maybe over the longer term there?
S
Scott Wu3:28:15
Yeah. Yeah. Yeah. For sure. Yeah. A lot of obvious integration on the team as you mentioned with Cross-sell and so on. I think the thing that's really exciting on products which I think actually comes along with these capabilities increases is you know as the capabilities keep getting better you start to take on harder and harder tasks with AI and with full agentic workflows right? And I think there's an interesting thing that happens where for a lot of the harder tasks you really actually do want to go back and forth between a synchronous and an asynchronous mode. And that's for a few reasons. One of the reasons obviously is because there's a lot of review and a lot of like looking at the pieces and thinking about all the minutia and the details of what you're implementing. I think another big reason for it is, you know, when you get started on a larger project, you know, let's say you're sitting down as an engineer and you're saying, 'All right, I'm going to go build this whole project today.' You yourself don't actually know all the trade-offs you want to make, all the decisions that you want to make and so on, right? And so having a format where, you know, for the decisions that need you to be there and you're involved setting the kind of the strategy or figuring out high level what should happen, you're able to do that in a nice synchronous environment, which is naturally the Windsurf IDE, right? And then for the parts of the task that you can actually hand off and have an agent work on, you're giving that to Devin. And figuring out how you go back and forth between those is super interesting. So, Wave 12 on the way soon. Hopefully we'll have a lot more to share.
T
Tyler3:29:36
Last question. Yeah, hit the soundboard, Jordy, for that. We're Wave 12.
S
Scott Wu3:29:40
Wave 12. Fantastic. Last question, we'll let you go. What is your probability that AI will get a perfect score on the IMO next year?
T
Tyler3:29:52
Interesting. So we, by the way, we just had the IOI, which is the programming version, like the programming olympiad. And I think there's a good chance that we'll have a golden medal at the IOI for this year announced as well. I think perfect score for next year.
Wait, we as in humanity or we as in Cognition?
S
Scott Wu3:30:08
As in humanity. Yes. Yes. Yes. And AI perfect score. Yeah. Sorry AI gold medal, right? Perfect score in the IMO next year.
T
Tyler3:30:17
I think it's got to be north of 50. Honestly, I would put it around like 75% or so. We'll see.
Well, thank you so much. We'll be following you closely. And good luck to you and congrats on all the progress. Very fantastic. We'll talk to you soon.
S
Scott Wu3:30:31
Awesome, guys. Talk.
T
Tyler3:30:33
Bye.
Let me tell you about Bezel. Get bezel.com. Your Bezel concierge is available now to source you any watch on the planet. Seriously, any watch. And we are joined by our next guest, Claire Vo from ChatPRD. Welcome to the stream, Claire. How you doing? What's going on?
C
Claire Vo3:30:49
I'm... It's a fun day today, isn't it?
T
Tyler3:30:51
It is a fun day. What was your reaction to the stream? What was your reaction to GPT-5?
C
Claire Vo3:30:56
You know, GPT-5, the first thing I said and I got a little early access is I said it's a developer for developers by developers. This thing is built to be a software engineer.
T
Tyler3:31:05
You've seen a long string of your guests come on and really speak about the coding abilities of it. And what I think is interesting about this particular model, especially because we're seeing them deprecate the old models in the ChatGPT experience, and we're seeing a lot of positive feedback, but I do think there are drawbacks to a model that's so clearly tuned to a developer use case. And as somebody who's building an application that isn't focused on agentic coding, I have noticed some personality quirks that are going to be really interesting to see how they shake out as we roll out this model to our users.
C
Claire Vo3:31:41
Walk me through those. What are the, what's the timeline? How much time do you have to kind of move users over to 5 before...
T
Tyler3:31:51
Yeah. Yeah. So I mean I think we have tons of time from the API side to move users. And in fact, you know, our strategy at ChatPRD is not to just upgrade to the latest model. I know Zach at Warp said like why wouldn't you want the latest intelligence? And the reality is because we're doing a lot of business strategy and business writing, I actually want to validate with our users that they're getting the quality of strategic thinking output writing that they really want. So we actually AB test every single model rollout and really evaluate for user quality, token generation, all those things. And you know, looking early on, it yaps. Man, this thing just wants to go through tokens. Right now, I'm seeing 4 to 10x the number of tokens generated between the 4 generation models and 5. And when you're in a business context, you do not always want longer words, you know, and so it'll be really interesting there. It is certainly focused on execution.
C
Claire Vo3:32:49
So I, you know, I've heard a lot from the OpenAI team, it's steerable. Yes. And its natural inclination is to drive you towards like how, what, very tactical, very specific. And so if you're trying to zoom back out at a strategic level or focus on a business initiative, it's actually a little harder to tune in that direction. So, you know, I think there's a lot of positive things for me as somebody who uses agentic coding platforms, who writes a lot of code. It's my daily driver now. I love it. But for other use cases, I think it's going to take some time to figure out if it really is optimal in use cases where intelligence actually isn't the differentiating capability.
T
Tyler3:33:31
It's very interesting to think the best product manager is not the one that writes the most, the longest doc.
C
Claire Vo3:33:39
No. And you don't send your engineer into your executive meeting. Like I really am looking forward to the time where we're not getting these number-based models where actually I can get like GPT-developer or GPT-strategist where they're pre-tuned and trained for the role they're going to play as opposed to general purpose but clearly oriented towards a set of tasks. And I just think if you look at this model, it was oriented towards software engineering at least in my experience.
T
Tyler3:34:14
So have you been tempted to launch any type of agentic coding products? You guys are obviously responsible for creating documentation and if you look at the other guests that have joined today, many of them are competing with each other in different ways and trying to own different parts of the stack. You guys have seemingly stayed really really laser focused and no one else is doing anything like you're doing at least on the show today. But talk about like picking your lane and kind of like optimizing.
C
Claire Vo3:34:51
Yeah, we're integrated with a lot of those platforms. So a lot of the kind of prototyping platforms, v0.dev, Lovable, all those we integrate. We just released our MCP. So I use ChatPRD pretty consistently inside Cursor through our MCP. So we think of ourselves as the product pair to the AI engineer. Now what's really interesting about my experience with GPT-5 is the one place it actually does really well is technical specs and that's a place where ChatPRD has sort of bridged into engineering execution. Often our product managers are generating a PRD or some sort of business document and they're actually going the next layer and developing a technical spec. The GPT-5 technical specs fed into these agent coding frameworks or prototyping frameworks output much higher quality assets on that end. So I do almost think there's going to be this kind of like right model for right use case especially in our kind of business and so we think of ourselves as integrating. The one thing I have thought about with GPT-5, it's the first one where it feels really simple to just go ahead and roll your own agent coding framework or prototyping framework inside of our application. So never say never, it's something that we get asked for a lot, but we're friends, we're good friends with almost all your guests on your show today. And so we like the role we play in terms of being the product manager pair to all these AI engineers.
T
Tyler3:36:14
Yeah, that makes sense.
What are you looking for next?
C
Claire Vo3:36:20
What am I looking for next? I mean, in terms of model capabilities, what I think is really interesting about OpenAI and why I'm really committed to the OpenAI ecosystem, even though I test and use a variety of models, is I think developer support is a real differentiator here. So we spend a lot of time talking about model capabilities and for application developers certainly ones that are doing more complex applications like agentic coding model capabilities really matter like core IQ of the model matters but the other thing that matters you know as somebody who has built developer tooling products it's developer experience matters the primitives in these APIs matters and so what I'm really pushing the OpenAI team to think about which is in addition to the core intelligence the model. What are the developer tools you need around these models to really make them a platform on which a variety of applications can build? And I do think that OpenAI has disproportionately invested in developer experience, but...
T
Tyler3:37:20
I'm always looking for better out-of-the-box tooling, more control over these models, more hosted services—all those things that as an application developer are just going to make it easier to deploy these models in production beyond the core intelligence of the models themselves. What was your read on 4.5? Is there a world where, you know, I'm thinking about the product manager versus the engineer. You have your O3 go crunch some really hard reasoning and then you have 4.5 turn it into stronger prose or more human language.
M
Mark Chen3:37:55
Yeah. So I did a lot of experimentation around 4o and 4.1. 4.5 was my favorite prose writer by far. It was loved from a business writing perspective. I thought the prose was the most natural. It was really slow. Like untenably slow. And so the compromise we made in our testing is we ultimately ended up with 4.1 as the fan favorite for business writing when we were balancing off both quality of prose and intelligence as well as performance, which for application developers is a real consideration. So I landed on 4.1. 4.1 is the model that's being tested right now against GPT-5 in ChatGPT. And one of the things that I have to go do now is figure out how to get ChatGPT or GPT-5 to stop writing it. It writes a lot and it only wants to write in bullet points. So I've got to go back into our prompts and figure out how to direct it to be a little bit more business-oriented.
T
Tyler3:38:58
Bullet point maximalist.
M
Mark Chen3:39:00
It's the new em dash. I'm telling you, you will not be able to stop seeing it. It just all it wants to do is write a bullet point and call a tool. Like I was using it in Cursor and it just kept maxing out my tool calls. I'm like, you do not need to read 50 files to do this. So I do think application developers are really going to have to think about how they slot this into their current workflows. There's definitely tuning that needs to happen. But I am telling you, you are going to see a lot of bullet points when this thing rolls out.
T
Tyler3:39:28
Yeah. In 60 seconds, where is product management going? A lot of people talk about the examples of product managers that are starting to ship code themselves, ship whole features, products, but I'm sure those are edge cases to date. But where do you feel like it's going based on your user base?
M
Mark Chen3:39:49
Yeah, I mean, it's going to go one direction or the other. Product managers are either going to develop the hard skills to do the design, the go-to-market, and the engineering job to some extent because some of these other jobs are definitely going away for product managers, or my favorite use case, engineers and designers are going to get tools like ChatGPT or these prototyping tools or Cursor and they're going to be able to actually do the product management job. And so what I think is we're going to see a new type of role emerge which is a much more generalist role where people maybe have a specialist capability and they're augmenting that product thinking or they're augmenting that technical thinking with AI. But I don't think there's going to be product managers as they were, you know, five or 10 years ago for much longer.
T
Tyler3:40:33
Makes sense. Well, thank you so much for stopping by. Great chatting with you.
M
Mark Chen3:40:37
Thanks for having me.
T
Tyler3:40:37
We'll talk soon. Bye.
M
Mark Chen3:40:38
Cheers.
T
Tyler3:40:38
Up next, we have Brad Lightcap, the Chief Operating Officer of OpenAI. Welcome to the stream, Brad. Also, Jordy, your post saying, 'I'm updating my timelines. You now have four years to escape the permanent underclass' has over 4,000 likes.
B
Brad Lightcap3:40:53
There we go.
T
Tyler3:40:54
Absolutely banging.
B
Brad Lightcap3:40:54
Thousand likes for every year. Love to see.
T
Tyler3:40:56
Love it. Anyway, Brad, how you doing? What's going on, guys? How are you?
B
Brad Lightcap3:40:59
Good.
T
Tyler3:41:00
Congratulations on the launch. What are the biggest takeaways for today from your side? I'd love to know about what it actually means to be the COO of OpenAI. OpenAI does so many different things. Consumer internet company, API business, enterprise. There's all sorts of stuff. Building data centers. What is your actual role?
B
Brad Lightcap3:41:21
My role is kind of whatever the company needs me to do. I play everything from PM when I need to to salesperson when I need to. That's kind of the fun part of the job for me. On this launch in particular, it was really fun. I spent a lot of time the last few weeks with customers, with partners, getting a feel for GPT-5 relative to what they were previously using. In some cases, those were OpenAI models. In some cases, they were other models. But you know, I've been at OpenAI a long time. I've been at OpenAI 7 years. So I've seen GPT-3, I've seen GPT-4, and then to be able to see GPT-5. And you know, just I think the joy of people being able to use it in production and seeing how much better it is. That's the best part.
T
Tyler3:42:04
Greg told us earlier about the era of having to pay people to use the early versions of the product. You guys have come a long way since then.
B
Brad Lightcap3:42:13
Yeah, we had like three customers with GPT-3 or something like that. And so it was easy to manage, easy to talk to all of them. They actually were like tired of us calling them being like, 'Is it good? Is it getting better?' And so now it's, you know, we're fortunate that we've got more than that. But it's cool. I mean, the diversity of use cases, I think the number of things that people are able to use it for. We've got everything from the team at Amgen, you know, big pharma life sciences using it for clinical workflows there. We've got teams at Uber building it for customer support. Teams at Notion and Cursor building it into products that people use every day. So I think that's the power of it is it just more and more covers the surface area of things people do with these tools.
T
Tyler3:42:56
I'm not sure how much you touch organizational design at OpenAI, but I'd be interested to hear your thoughts on how those companies that you mentioned should be thinking about AI changing their org structure. Is it sort of like a horizontal cross-functional service layer like a finance team that touches a lot of different elements of the business or should most companies be thinking about standing up a dedicated AI implementation team? How do we get a chat box on every product that we already ship? Like how do you think about those trade-offs if you were talking to a friend at a Fortune 500 company that was thinking about their AI strategy?
B
Brad Lightcap3:43:34
Yeah, you know, it's an interesting question. I think it was maybe said earlier on the show. The thing we see is just people can do more. And so there's this much wider latitude that you get if you're an individual person at an individual company where especially as you get bigger, you know, maybe more bureaucratic organizations that have a lot of different functions, a lot of different levels. You have to rely on a lot of other people in the org to get stuff done. You've got to rely on your data science team to do data analysis. You've got to rely on your design team to do mockups. You've got to rely on your marketing team to do copy. And I think what we see with AI is it just accelerates people to get to a great V1 of everything. So if you're a high-agency individual and you want to get stuff done, you're no longer gated on people that you otherwise would be. And I think that should enable organizations to move a lot faster and I think it should enable the people at organizations that really drive them to do a lot more. And we see that consistently with ChatGPT Enterprise. I think that is consistently what we hear and we seek those people out when we deploy ChatGPT Enterprise. We find those two or three people at the organizations who are just the AI superstars and champions and then try and actually use them as these kind of touch points for the rest of the org to learn from.
T
Tyler3:44:45
How are you personally using AI these days?
B
Brad Lightcap3:44:49
You know, my biggest challenge I think day-to-day is context switching. If you look at my calendar from top to bottom it's like I joke with my wife I have to show up to work wearing a lab coat and then I take the lab coat off and put some sunglasses on and a film school jacket and then I'm talking to a media company and then I take that off. So I go through the costume changes. And I think what I actually mostly use it for is just to help with bridging me from kind of thing to thing. To kind of put me in the mindset of being able to work with customers, help customers. GPT-5 is incredibly good at this kind of structured reasoning of how do we actually take what is this very diverse set of things that models like GPT-5 can do and then apply them in domains that I don't think about every day. And so it gives me this launching off point to be able to talk with leaders and with customers much more fluently about how we can help their organizations.
T
Tyler3:45:39
Within let's say a set of companies like the Fortune 500, what does AI adoption look like across the spectrum? Because I'm sure that there's companies that you talk to that are truly adopting AI in the way that John was mentioning, like trying to become AI native, changing their entire organizational approach. And then there's companies that just want to buy software to say that they're becoming AI native. So what does that spectrum look like in practice?
B
Brad Lightcap3:46:11
Yeah, it is a wide spectrum. So at the top level we're seeing just amazing appetite for wanting to adopt tools for people and I think that's the easiest place to start. Typically that's where we steer organizations if they're starting at zero is just give your people the best tools. You may have seen we've grown ChatGPT Work, which is our enterprise and team product, from 3 million seats to 5 million seats now from June till now. So torrid growth there and we don't see any abatement in demand there. If anything it's accelerated from last year and so I think people and organizations are starting to realize that at a minimum you need to make sure people have the best tools. What's cool about GPT-5 now is it also enables people to use the best tools at every point. And so if you're in an organization you're not fumbling with the model picker. You're not trying to figure out when to use a reasoning model. You're not trying to figure out the art of prompting to get the perfect thing. All of that stuff is abstracted and it's kind of taken care of for you and you can have confidence that your people are actually using the best models at any given point. Beneath that it gets a little more complicated. So more and more organizations I think are starting to grasp how the tools can actually help in the business process. So whether that's in customer support, whether it's in research, whether it's in software engineering and data science, you're seeing these tools more and more adopted in the enterprise. I think there's still a quality gap though. I think we now are just breaking into what I would call the era of models that have capabilities that are good enough to make a dent in the types of problems businesses care about. Businesses care a lot about things like reliability, right? They think they care about accuracy. They care about the resiliency of the model to recover from tool use errors and to be able to string together these very long multi-tool, multi-step workflows. So GPT-5 is a step on all those things and I expect that that will enable us to be able to do more and more things in the business process.
T
Tyler3:48:05
Do you think those customers that you just mentioned will stick with this idea of like GPT-4 level workloads will stay on GPT-4 and maybe there'll be cost savings but those workloads will stick around for a very long time and then you'll develop almost new capabilities, new workflows, new workloads that will be additive but the enterprises will stick or will they want to—is everything so fresh that they'll want to just rewrite everything with the latest and greatest?
B
Brad Lightcap3:48:38
More often than not, I think it's the latter. I think you want to rewrite everything. One of the cool things we did here was we were able to keep the pricing on GPT-5 at the level of O3 pricing. And so, you know, if you're cost sensitive, you don't really have an excuse to not upgrade. GPT-5 is faster than O3 and 4.1. So, we've improved on latency for sensitive use cases that are speed sensitive, latency sensitive. And obviously the intelligence bar has gone up. So, you know, unless you've got really a very narrow and specific workflow where you've got a model like 4.1 that kind of is okay, there's really not a reason I think that people wouldn't upgrade.
T
Tyler3:49:14
Yeah. Do we need like a three-dimensional Pareto frontier right now that matches not just cost and capability, but also cost, capability, and latency or something? Is that something that you're seeing a lot of demand from in the enterprise?
B
Brad Lightcap3:49:26
Yeah, 100%. We actually measure it that way. So we look at those three vectors and it's always kind of an optimization function along those three axes.
T
Tyler3:49:36
We think we found that here. It was actually in terms of where my work was over the last few weeks. It was a lot of—I mean there's a qualitative kind of really manual process of collecting feedback because everyone's got a little bit of a different preference and we can only pick kind of one or two points on that curve and so just trying to dial customer feedback, namely developer feedback, in for us on where that balance of things are is a big part of our process for picking all those points. And so we hope that people like it and it unlocks the kind of maximal use.
B
Brad Lightcap3:50:09
That's great.
T
Tyler3:50:10
Jord, how are you thinking about open source? Who's been most excited to get access to it and where do you see it going?
B
Brad Lightcap3:50:19
Yeah, I mean it's important to us. You know, I'm glad we've gotten this out. It's been a huge team effort. I think there was a kind of a thing that like, you know, OpenAI doesn't like open source anymore. It's like, no, we're just really busy with a gazillion other things. So I think hopefully going forward we've got more of a leaned-in vantage point on open source. But it unlocks a huge number of use cases. I mean, if you think about government use cases, you think about on-prem use cases where you're handling sensitive data in very sensitive environments, you think about where you want to run models on the edge. All these things right now are kind of inaccessible to us as a service provider to customers because we just don't quite have models that fit at those points. So this for us we think is huge TAM expansion and we're excited to be able to work with enterprises on implementing that model which is I think competitive with our O3 class of models.
T
Tyler3:51:12
So what is the landscape like for companies that are helping to implement OpenAI products at various enterprises? You have the big consulting groups that will give you an AI strategy. Maybe they'll try to take it a step further, but I imagine there's a cottage industry of firms that have sprung up to try to help organizations unlock the value beyond, hey, let's just get everybody a seat with ChatGPT Work.
B
Brad Lightcap3:51:40
Yeah, I think there will be this new industry that emerges that is kind of separate and apart from the legacy set of SIs and consultants that is really AI fluent. They're very AI native. I think it's very hard to borrow paradigms from the last 20 years of software building and implementation that are going to map to what we're dealing with here. You're dealing with fundamentally probabilistic systems that are moving and increasing and improving at a rate of now kind of collapsing to every few months. And I think the nature of use cases changes quickly, where enterprises are focused on deploying them changes quickly. And so I think it's just hard for the legacy industries to keep up frankly. We've had a lot of success working with some of this new breed of SI, so the Distills of the world and others that really have been born, forged in the fire of this new platform. And so we hope there's more of them. We'd be excited to work with anyone that wants to work with us on it. There's more business than we can handle. And so we're always happy to spread the love.
T
Tyler3:52:51
Talk about the $1 ChatGPT product for the government. Were you involved in that at all?
B
Brad Lightcap3:53:00
I was involved in that. We wanted to do something that was meaningful for the US government. It's been a real big focus of ours lately. I think our view is the government has got to start to modernize. We've got to make sure that the tools that we use in the private sector are also in the hands of folks serving us in the public sector. And we wanted to make that really simple. So we made ChatGPT basically equivalent to ChatGPT Enterprise free. It's a dollar per year per agency. Hopefully we can afford that. And we wanted to make that available to anyone that wanted to use it and standardize through GSA. So we're super appreciative of the partnership with them. And more I think that we can do on that front.
T
Tyler3:53:39
How is that different than just like if I'm a government employee, I can just go to google.com and I have access to that and Google right now. I mean, Scott Kapor was saying that he can't use it. Yeah. So, yeah. Why? Yeah. Just talk to me about how it's different to offer ChatGPT as an actual service with a contract that you're vending in, you're actually—they are a client versus just if you put up a website every government employee can access the web to some degree or would it be blocked like why does it need to be a deal at all as opposed to just everyone just uses it?
B
Brad Lightcap3:54:17
Yeah. So part of it is just making sure that government employees can access it. So in some places obviously you can put blockers in place that would prevent access. We hear a lot of stories by the way of people going out on their lunch break to their car in the parking lot and pulling up ChatGPT on their phone and throwing a bunch of stuff in there just because they know it'll get them through the day faster. And we've done work by the way with governments, with the state of Pennsylvania and other places where we've seen dramatic increases, things like two to three hours a day saved per employee given the nature of the work that they do and how helpful ChatGPT can be. And so this lets us have an interface into them as a customer. It lets our team engage with them in a direct way. We can see how they're using the product and can help them use it better. And so that's important for us is we got to build on that foundation with them.
T
Tyler3:55:02
And then presumably it also allows the government to define security and privacy in their world as opposed to if you're just some website out there. Their choice is only block or don't block as opposed to actually communicate with you. This is okay to train on, this is not, etc. Keep everything private, etc.
B
Brad Lightcap3:55:20
Yeah, I mean, we don't train on enterprise data at all. So, yeah, you know, you're safe there. But yeah, I mean, for us, just being able to treat them as a customer, right? To treat them as a user. And you mentioned earlier like we were talking about there being these points of success at every organization where you've got people who are way more sophisticated in using these tools than others. We want to be able to see those people and amplify them and the government's no different. There are people that we've worked with in government who are incredibly sophisticated in how they use AI tools and our goal is to get everyone there.
T
Tyler3:55:53
How do you think about the group of users that are active students? They've been on summer break. You guys have been busy over summer. Are you thinking about—and you recently launched—I forget the exact name for the product. I think it was like ChatGPT Learning. How are you thinking about that cohort and unlocking new capabilities for them this coming year?
B
Brad Lightcap3:56:14
Yeah. So we launched something called Study Mode which was in our core ChatGPT product and it was a little bit of an experiment. We wanted to see if you change the way the model behaves when it knows you want to be in a learning mode, if that can actually enhance outcomes for students. Where we have all these kind of studies that have been done very anecdotally about ChatGPT's ability to drive student outcomes and learning outcomes. So here we kind of took a little bit more of an intentional approach of if you actually take the model and use it in a more Socratic style where it can actually quiz you, it can withhold certain information that it wants you to be able to empirically deduce. It wants you to reason about problems and it kind of reasons with you as a partner. So far so good. It's really cool. And learning is kind of the killer use case of ChatGPT. And so I think to be able to actually launch something that in some sense extends that killer use case has been really cool and the student feedback so far, even on summer break, has been positive.
T
Tyler3:57:16
Well, we'll let you get back to your day. What's next on your agenda? Are you putting on the lab coat or the suit and tie and going to Washington?
B
Brad Lightcap3:57:23
Good question. You know, today I'm mostly with the team and talking to customers and maybe tomorrow I'll get back to the lab coat, but in the meantime,
T
Tyler3:57:34
I appreciate you taking the time to talk about today. So, yeah. Well, thank you so much for taking the time to talk to us. We will talk to you soon. Have a great rest of your day.
And the timeline has been in turmoil because President Trump says he will be imposing a 100% tariff on all semiconductors coming into the United States. It started with widespread tariffs on chips and then turned into export controls. This is from the Kobe letter. Is this a red flag moment? I don't know why we have the red flag.
U
Unknown3:58:01
I felt like it. Ben was getting the flag.
T
Tyler3:58:03
Getting the flag and video potentially affected. Taiwan says TSMC exempt from Trump's 100% chip tariff. Very unclear. The story is obviously still developing. And Dylan Patel says, 'You're telling me that this level of monitoring the situation is free.' And it's a picture of you in front of the whiteboard monitoring the ChatGPT versus the timeline. Today we're monitoring. Illinois has banned AI therapy, making it the first state to regulate the use of AI in mental health services. Interesting headline that's coming out.
U
Unknown3:58:42
Which is interesting because the product can just be used for therapy. Like the user can choose to do that. It's not necessarily—it's kind of hard to ban outright. Like maybe you can ban it in a clinical setting.
T
Tyler3:58:56
Yep. I wonder how they define this. There's probably a loophole if I know anything about how these bans are implemented. But yeah, maybe it's like if you're in the clinical setting, you can't use it. But then people will just use it independently like a therapist just on their phone.
U
Unknown3:59:10
They're going to be going to the car.
T
Tyler3:59:11
They're going to be having—they're going to—No, they're just going to have it listening to the conversation. They're going to be like, 'What should I do right now? What should I say? What should I say?' How does that make you feel? That's what it's going to tell you. Celsius nearly doubles revenue year-over-year. This is the energy drink. Revenue of $739 million versus $632 million consensus. North America grew 87%, international grew 27%. But here's the real kicker. Alani Nu acquisition is the primary driver of growth. Alani Nu added $300 million in revenue and retail sales are up. So wow, what a performance. But yeah, I mean that was the expectation when they bought Alani Nu is that they would—I guess this is like the first moment they rolled them in probably. But huge growth for Celsius as they become multi-product, multi-consumer company. What else is going on in the timeline? We have one last guest. I think you might have to hop on with Taipei. So feel free to jump when you need to. Tyler, anything going on on the timeline we should be monitoring? We are of course monitoring the situation.
I've been—So when Max was on, he was talking about how you can make a little game, right? So I've been working on a Bloons Tower Defense game.
U
Unknown4:00:29
Okay. How's it going?
T
Tyler4:00:30
So it's going pretty well. I'm making another change, but then maybe I can screen record and do a little share.
U
Unknown4:00:36
Yeah. Yeah, that'd be great. You could share with the folks, too. Yeah. I like this post from Ray Sullivan. These GPT-5 numbers are insane. And it's a chart of GPT version versus number. And then once it gets to four, it goes 4.1, 4.2, 4.3, 4.4, 4.5. So the fifth one is a massive, massive bar.
T
Tyler4:00:56
We need an analysis of the charts from today. It seems like there was multiple that were kind of odd or hallucinated or off.
U
Unknown4:01:07
It's interesting that multiple of them snuck out just in sheets. The popular convenience store chain with 750 locations is now offering 50% off purchases paid with Bitcoin and crypto daily from 3 to 7 p.m. What a wild move by Sheetz.
T
Tyler4:01:23
Well, Ben is in the waiting room. Let's bring him in.
U
Unknown4:01:29
Let's bring in Ben Hilac. How you doing? Good to see you.
B
Ben Hilac4:01:31
Doing well. How are you guys doing?
T
Tyler4:01:33
We're doing well. I'm just going to say hello. I got to take off and talk with Taipei. I'm going to let John take it from here.
U
Unknown4:01:38
Absolutely. I'll close out the show. Have a fantastic conversation.
Give me the update. How's the day been for you? What were your expectations? Did this meet, succeed, did it underwhelm you? How you doing?
B
Ben Hilac4:01:49
Well, so I've actually had access for a couple weeks. So, we actually did a video. I'm not sure if you've seen it, but OpenAI brought a couple of folks from the Twitter sphere to their office a couple weeks ago to try it and some other stuff. I think it pretty much exactly meets my expectation as far as how it's been received. And I've tweeted about this as well, but I think that it's really, really good at one-shotting things. You know, I think it's better than other models we've seen, but I think it's actually sort of a distraction in a lot of ways. I think that the things it's a lot better at are a lot harder to describe and I don't think the harnesses for it really exist yet.
U
Unknown4:02:39
Harnesses. So,
B
Ben Hilac4:02:41
What—the way I've been describing it is that I think I've seen web search existed in ChatGPT for a really long time, right? Like it was able to call a tool, search the web. Obviously deep research was very different than that, right? Like what we saw was it was actually searching the web, it was reasoning about those results, changing its course, course correcting in the middle. So intermediate reasoning is the term for it. And they really trained it how to search the web well. I think GPT-5 does that for a whole plethora of tools. The interesting thing is that a lot of products, like I think a lot of the agent products that exist today were kind of built wrong. Like they weren't built—they didn't build their tools the right way. And we've seen this before. Like if you look at the first infrastructure for agents was LangChain like way back when, like two or three years ago. It was early but it was wrong, right? And so anybody that—they've iterated since, right? They have LangGraph, a better implementation, but the first implementation of LangChain was again early but wrong. And so if you built your product on LangChain you had to significantly change it. I think we will see a similar thing happen for GPT-5. You know, it's not just like change the string in git from 4o to 5 or something and push and now you're good.
U
Unknown4:04:06
Yeah. You know, that meme about like, oh, Sam Altman stood on stage and just killed 75 startups, Google just killed 100 startups, Apple just killed Perplexity with their new thing or whatever. Did any of that happen today? It feels like this is like the LangChain needing to change their strategy that happened a while ago. I haven't identified anything. It feels like Scott Wu hopped on and said like, great day to be an application layer company. The foundation models got better.
B
Ben Hilac4:04:36
It's more tools in my tool chest. I'm extremely happy and I'm more confident than ever. And I believe him. I believe that he doesn't see today as fundamentally needing to change his business model.
U
Unknown4:04:50
I think that's true actually. I think that people have been—there's a lot of people building agents right now. I think a lot of them have not been feasible for some of the reasons that GPT-5 starts to address. So I think what it means is that the entire architecture behind agents will get a lot simpler. It feels like a good day for people building applications.
B
Ben Hilac4:05:16
Yeah, it's not immediate that there's some company or something that got killed today.
U
Unknown4:05:22
Yeah. Yeah. I mean, in general, it feels like, you know, Dario Amodei updated his timelines. There's just been a general idea that we've maxed out pre-training, we've kind of maxed out post-training. We're now in the let's reap the reward of this. And we've seen it in the incredible financial performance, the incredible usage numbers. You know, millions and millions, hundreds of millions of people are using ChatGPT 30 minutes a day. I love the product and yet it feels like the what have you done for me lately meme. It's totally like okay yeah we went from the iPhone 4 to the iPhone 5 today.
B
Ben Hilac4:06:00
Yes.
U
Unknown4:06:00
Still really an important technology, great company, but like I want another iPhone 1.
B
Ben Hilac4:06:05
Yes. Yeah. Yeah. No, I totally get what you're saying. I think that—I wrote a piece about this with Swix, but yeah, it really actually changed the way I see that path to AGI. Like I think before using it a lot, I kind of was like, 'Okay, we need bigger models. They're going to get smarter or something.' I think I had this realization. So, I was watching it solve—I had this really weird dependency conflict with Yarn. Like we have a mono repo. It's like part of the problem also with this discourse is the sort of problems it gets good at solving are just not sexy things to talk about. They're not things that even you'll understand. And I'm like, we have this issue with the way we structured things. But a couple weeks ago I was watching it—I had this problem. No other model would solve it. And I watched it sort of poke around—it started running this yarn command in a bunch of different directories in between. It's reasoning and correctly reasoning about why and what it was learning and taking little actions in between seeing what happened. I think what I realized is that if you imagine humans without tools—if we never had any tools, we're never even able to write things down—like would you be able to tell that we're intelligent? Would we have learned to speak, etc.? I just don't—you know, even if we could not have ever invented fire, right? It's like where would we be right now? There feels like there's a similar—I actually think a lot of the next year is just going to be how do you get these models to do things better. I think it's next year.
U
Unknown4:07:42
In your Yarn example, you said you were having GPT-5 work on the problem. Was that wrapped in a coding tool? Did you just go to chat.com and give it your GitHub repo? Like talk to me, what was the actual user experience from your side?
B
Ben Hilac4:08:04
Yeah, so this was in Cursor. I think the Codex CLI, the new version of the Codex CLI which they just released today is also really, really, really good. I think that you will really only see a significant difference in places where it can sort of explore its environment is the way I would put it. Like when I was watching it go bounce around my repo and I felt almost like I was watching something navigate a little video game like Pokemon or something like that. That's kind of what it felt like. Like it's kind of like I'm going to go over here. I'm going to see this. Okay, wait a minute. That conflicts with what I just saw over here. Like where should I go next? Do you know what I mean? Like it felt very novel is what I would say.
U
Unknown4:08:44
Yeah. Yeah. Yeah. What—so yeah, I mean how are you using it? Where do you see it going? Do you see it like just a little bump of a tailwind today or what's your read on how you'll be using GPT-5 going forward?
B
Ben Hilac4:09:01
I mean yeah there's two huge things. So one thing that really got missed today is that they also released GPT-5 Nano which is an incredibly good model actually. So we're not talking about it but it's half the cost for input tokens than Flash Lite or sorry yeah I think it's actually half the cost of input tokens than Flash Lite and it's a really good model like it's 4o level for a lot of writing and stuff like that. And so yeah, we'll be using that probably in the short term. I think it'll be interesting to see how other providers react. Like I'm sure Google will cut their prices as a result, but it is the cheapest hosted model. I don't think anyone's serving any other model for those prices for that matter.
U
Unknown4:09:46
Yeah, that makes sense. What else are you looking for for the rest of the year? Probably no GPT-6 on the horizon, but what are you looking out for? I mean, it seems like Google is expected to respond with Gemini 3 soon. What else are you tracking in the world of AI these days?
B
Ben Hilac4:10:02
It's a great question. I think that yeah, that's going to be wildly interesting. I think what Google does will tell us a lot. I think that they—you've probably seen it, but they released this world model yesterday. We're kind of not talking about it anymore. I mean like if those videos—I haven't tried it myself. If those videos are real like that's one of the most mind-blowing things I've seen in the last decade or something. So like if that's real like that's extremely interesting and I think has—all the stuff that's going on with world models right now has huge implications for everything from robotics to just so many different fields. So super, super interested in that. And the other thing is I actually just think that—again I'm actually really bullish on ChatGPT-5. I think that the way it was received today is just about how I expected it. And the reason is like when I say harness again I think that Canvas in ChatGPT is pretty bad is my take. Like it's a tough product to make but it does really poorly with long files, crashes sometimes, like that sort of thing. I think that we don't have the product layer around GPT-5 doesn't exist yet. So I think we're going to see some really, really interesting products that are built around it.
U
Unknown4:11:13
Yeah, it's always hard when you go from a binary qualitative improvement. GPT—ChatGPT was like we passed the Turing test and now the next test is like super intelligence that self-replicates, is smarter than every single person, knows everything. It's like the bar is—we really moved the goalposts, you know.
B
Ben Hilac4:11:33
100%. I think that there was a lot of discourse around the model as well leading up to it which I think didn't help. But the way that I would think about it is I think that depending—there's some percentage of the way through automating software engineering that we've made it. Like let's say it's 70% or something, 75%. The tough part is that last 25% is the hardest, it's the least decipherable to explain to people. It's the least universal. Like if I'm just like oh make a—one of the examples I did, I made a personal website that's all Mac OS 9 themed in like 20 minutes with GPT-5. And so it's really fun, right? You get it. Like my mom gets it. Like I can show it. I can share it. You get it. My mom—I can't explain any of the very specific ways that GPT-5 helps in our specific codebase, our specific problem, whatever. So I think that it'll be less—and these launches will probably get less and less interesting from a what it does for software engineering as that gap gets closed. Like what's the last 5% of software engineering? It's probably not going to be that interesting to me.
U
Unknown4:12:48
Do you think they'll be on an annual release cadence now? Like Apple updated all of their iOS, all their operating system nomenclature to be like we are now on 26 because it's the year. It's like a car model.
B
Ben Hilac4:13:01
I don't think you can plan it. I don't think you can plan ahead. Like that's the interesting thing is I think that there's people that say that GPT-4.5 was supposed to be GPT-5. Yep. And I think that it sort of came out and they're like, 'Eh, it's like, you know, I actually love 4.5. I think it's a really fun model, but...'
U
Unknown4:13:19
Well, it's clear that improvements come in many places just like with the iPhone. Like the latest iPhone, you buy that because it's not just the one with the new screen, it has a slightly better camera, slightly lighter, longer battery life. It's like an ensemble of improvements that then they add up. And I think that feels like what we're getting here today and what we will get in the future is like this little—we did a little extra RL over here. This tool is now sharper. It has new capabilities. We added multimodal—like the video generation got better and this feature got better etc etc. And I think that what a model is is still going to change a lot and how we value—so just give an example like 4o was sort of this big thing where they talked about it being natively multimodal, taking in even video at some point, video in video out, audio in audio out. And you haven't heard that from GPT-5 yet. Like you can't talk to it on Advanced Voice Mode. It doesn't generate images. You know what I mean? There's no at least yet native image generation.
B
Ben Hilac4:14:21
How it works under the hood—these model capabilities seem quite possible like the best model for writing natural language might not—or writing creatively might not be the same model that writes really good Rust code. Like these might be different models. So I don't know, we'll see.
U
Unknown4:14:46
Yeah, Create Image here is now tucked next to Deep Research, Agent, etc. But I would hope that you can call that from the actual chat interface.
B
Ben Hilac4:14:54
You can call it from the GPT-5 chat. It's just using—it's using GPT Image One I think is actually the name of the model. So it's a dedicated image generation model which I think is maybe 4o. I don't totally know.
U
Unknown4:15:04
Yeah, I just—I don't particularly care. I'm not looking for one model to rule them all. I'm fine with models calling different tools. It seems fine.
B
Ben Hilac4:15:14
Yes.
U
Unknown4:15:15
Anyway, fun day. Thanks for hopping on. We'll talk to you soon.
B
Ben Hilac4:15:18
Of course. Anytime. Talk.
U
Unknown4:15:19
Have a good one. Bye. And that's our show today, folks. Leave us five stars on Apple Podcast and Spotify. And thank you for tuning in to the GPT-5 Giga Stream. We're on hour four and a half. We've enjoyed hanging out with you, Tyler. Anything else from the timeline? Close it out for me. Timeline's still in turmoil.
T
Tyler4:15:39
Show the little game I made.
U
Unknown4:15:40
Okay. Yeah, let's show Tyler's game. Can we do that? Is that—
T
Tyler4:15:43
You got it. The Tyler Tower Defense.
U
Unknown4:15:46
Okay. This was—one shot. Okay. I didn't—
T
Tyler4:15:48
Wait, what do you mean one shot? One prompt. You said you were working on it.
U
Unknown4:15:52
I was, but then it's like, wasn't the definition—Oh, so you went back to a single prompt. Got you.
T
Tyler4:15:58
I made a change, but then I realized like, okay, this is not as good. So, I just went back to the first one.
U
Unknown4:16:02
Okay. So, yeah, my question is—this seems actually like it's the game engine. I don't know what it's using under the hood. Do you know? Did it write like WebGL code or did it write like—
T
Tyler4:16:14
I think it's just JS.
U
Unknown4:16:16
Okay. And it's just like HTML canvas.
T
Tyler4:16:19
That's pretty crazy.
U
Unknown4:16:20
Yeah. You'd think it would use some 2D engine off the shelf or something, but my question is like what—that won't go viral because that is less impressive than just the Tower Defense app that I can get in the app store for sure.
T
Tyler4:16:36
Yeah, but it's like maybe if I take my—you know how those ControlNet images went viral where people would take their corporate logo and then they'd throw that through ControlNet and it would be like the TBPN logo overlaid over a forest and the trees would look like the logo. Yeah. So maybe like it's Tower Defense but it's my logo or something like that and the enemies are like moving through something like that. I don't know. There's just got to be a way to personalize it and make it so every single game is a unique snowflake that you want to go and experience that one. You want to look at it, you want to spend some time in it. I don't know.
U
Unknown4:17:09
Yeah, it's hard because it's like it's still predicting the next token. It's not like image—the 4o image generation was like kind of—it wasn't novel, I guess, because there was image generation, but it was such a massive improvement. This is—
T
Tyler4:17:22
Yeah.
U
Unknown4:17:23
Like there's not any clear massive step change here. It's a little bit better in a lot of ways.
T
Tyler4:17:28
Yeah. So, oh well. Well, we'll have to play with it more. Let us know what you think about GPT-5 and we will see you tomorrow. Have a great day. Thank you so much.