CEOInterviews.AI
Start App
Mike Krieger
Co-founder of Instagram, Instagram

How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic

📅 Aug 27, 2026 AI Engineer 26 MIN 54 SEGMENTS · 2 SPEAKERS
Over a single weekend, Mike Krieger had Claude port a few hundred thousand lines of Python to TypeScript, verify it, and churn ...

What Mike Krieger said

Written from the verified transcript and checked against it. Every figure links to the moment it was said.

Mike Krieger discussed his shift from Chief Product Officer to an individual contributor at Anthropic, driven by FOMO from watching others build with models. He described moving from task delegation to expressing end states, and noted Claude is 'way, way smarter' than him. Krieger advocated for being 'unreasonable' in AI usage, citing a weekend port of a Python codebase to TypeScript. He emphasized the importance of pre-measuring everything and using feature flags, lessons from Instagram's early scaling. On Tasks, he said it's how most of Anthropic's code is written, enabling multiplayer, async, proactive work. He acknowledged bottlenecks in code review and human conceptualization, leading to Claude Code Artifacts. Krieger detailed labs' 'persevere or pivot' two-week reviews and flexible pod structures. He discussed unshipping Styles, blurring lines between Claude Design and apps, and advised startups to focus on verticals and user understanding. He concluded with mental health advice: carve time off, verbalize emotions, and maintain perspective.

Key takeaways

  1. Krieger ported a couple hundred thousand lines of Python to TypeScript over a weekend using a dynamic workflow with Claude.
  2. Tasks is how 60-something percent of Anthropic's code is written, enabling multiplayer, async, proactive delegation.
  3. Anthropic labs run two-week 'persevere or pivot' reviews, shutting down projects every cycle.
  4. Krieger unshipped Styles, a prescriptive feature, in favor of Skills.

Numbers and commitments

FigureWhat it refers toTypeAt
60-something percent of Anthropic's code written via Tasks metric 7:59
4 to 5% usage of individual features at Instagram metric 15:38

Chapters

  1. 0:00Role shift and model usage
  2. 3:01Being unreasonable with AI
  3. 4:17Porting Python to TypeScript
  4. 6:45Scaling lessons from Instagram
  5. 8:23Tasks and multiplayer AI
  6. 10:10Code review bottlenecks
  7. 11:54Labs structure and reviews
  8. 14:08Claude Design future
  9. 15:38Deleting product complexity
  10. 17:44Advice for startups

Questions asked in this interview

12
  1. 0:47How has your model usage changed as you've seen models internally grow?
  2. 2:40He said, "Be unreasonable." In what ways have you been more ambitious with your prompting?
  3. 5:16I think a lot of people are also like, "Well, it's a compiler, it's a runtime, it's got lots of tests, easy to do." Can you port Instagram, which you would know very well, to PHP like that, like a product?
  4. 8:10I don't know, it's for this segment of the population"?
  5. 10:01Is there a world in which you just merge it in?
  6. 11:30How are you structuring the labs?
  7. 14:00It's one of your biggest launches this year. Where does this go?
  8. 15:16What would you delete in AI, or more spicy, what would you delete in Claude?
  9. 17:17Why bother starting any other company?
  10. 19:41Any potential for Claude, obviously a lot of Excel spreadsheets?
  11. 21:41How do you advise people who are working 996 to avoid burnout?
  12. 24:43Has anyone, any coach or mentor, said something to you that you repeat to yourself that gets you through the tough times?
Host 0:12 ↗
Joining us on stage is the co-founder of Instagram and a member of technical staff at Anthropic, Mike Krieger.
Mike Krieger 0:36 ↗
How's everyone doing? I mean, good morning.
Host 0:39 ↗
Nice. Mike, thank you for releasing Claude just in time for us.
Mike Krieger 0:44 ↗
Exactly. For the conference, we timed it.
Host 0:47 ↗
We're so glad to have you. You are one of the preeminent builders and you're leading labs at Anthropic. How has your model usage changed as you've seen models internally grow?
Mike Krieger 1:02 ↗
Yeah, I mean, for me, it's been like both the model shift and then my role shift. For like the first two years I was at Anthropic, I was Chief Product Officer, and then I kept seeing people build with the models, and the FOMO just kept increasing. Because you use the models as much as possible, but for example, on product strategy, I would write a strategy doc and then have Claude critique it, and maybe you can use a workflow, but it's not quite the same as building in that pure way. I was spending all my weekends trying to build with it, and I realized, okay, I actually just need to shift. It's like way too interesting a time.
It's actually an interesting trend I've seen now, like several people that were CTOs at other places are now joining as ICs at Anthropic and other places. But I made a role shift, and it was actually right around the time where we started getting sort of internal snapshots of what became Mythos and Claude. What was really interesting watching that sort of shift was that kind of change between: I have an idea, I'm going to sort of break it down in my head much more how I would do engineering normally and then iterate through these different steps, to moving to much more of the paradigm of: I'm going to describe the goal, go off and work on it, and then we can talk about what trade-offs you surface, some questions along the way, but then figure out where you landed and where we can go from there.
I find it's hard—I don't know if people have this experience where, and I know Claude's only been re-enabled for a couple of days, Claude is definitely way, way smarter than me. So sometimes it'll finish work and be like, "Here's the trade-offs I made." I'm like, "Can you explain it to me like I'm a little dumber than you are, because I need you to sort of break this down for me." But that's been one sort of big change, moving from that task delegation to expressing the end state and then having it go and cook on it.
Host 2:40 ↗
Yeah, we're all learning how to delegate better. Tariq did us a huge favor yesterday. I want to read it in the newspaper. We have write-ups of talks now in the next day's newspaper. He said, "Be unreasonable." In what ways have you been more ambitious with your prompting?
Mike Krieger 3:01 ↗
I love that framing. We actually just hit this. One of the labs initiatives I have is this internal product, and somebody was like, "Hey, it doesn't work the way I want it to, and can you make some changes?" And I realized I'm just going to go ask Claude to do this. Like, why don't you ask Claude? And this was a non-technical person. So I actually think as an industry, or even as a product team, we have to teach people to be more unreasonable in their usage. And it's sort of hard to imagine.
If I can digress for a second on product design, I think right now the kind of first generation of AI products, we put them too much in a box and constrained their sort of access to tools or degrees of freedom, which means it was much harder to be unreasonable, right? When you say, "Do this thing for me," and then it would be like, "Whoa, I can barely—I can write code, but I can't really run it," or, "I can kind of introspect my environment, but not really." As you see our own product progression, even with things like Cowork, does every single knowledge worker need a virtual machine that can write bash? On the face of it, no. But then when you realize, oh, actually, that way it can remediate an issue where, "Oh, I tried to parse a PDF using our built-in PDF..." I hit this yesterday, and it was like, "Ah, I can't parse it this way." "Well, okay, I can probably write a script that can do this as well." So I think that's it.
My most unreasonable thing, though, was one of our labs projects I wrote in Python, near and dear to my heart. All of Instagram was in Python. I think they're finally converting it to PHP now that they have models, but tokens. For deployment, I realized that Claude Code had figured out a better deployment story with Bun, and I was like, "Okay, I need to port this whole thing from Python to TypeScript." If I put on my 2010s engineering hat, or even my early 2020s, that's a dumb idea. Who would ever port, at that point, a couple hundred thousands of lines of code? But I was like, "I think this is doable now." I basically created this dynamic workflow setup, and over the weekend had it port the whole thing, verify it, double-check it, then read both code, basically churn and churn and churn, and then came back Monday to a completed workflow that was a ported version of that thing. So that probably ranks on the more unreasonable things: "Yeah, just port this entire Python codebase to TypeScript. Get it working, get it deployable in a weekend."
Host 5:16 ↗
Yeah. I mean, a lot of people are talking about the Bun Zig to Rust version. I think a lot of people are also like, "Well, it's a compiler, it's a runtime, it's got lots of tests, easy to do." Can you port Instagram, which you would know very well, to PHP like that, like a product?
Mike Krieger 5:33 ↗
Yeah. I think on the product side, it's even—I don't know if easier or harder. One of the things we did at Instagram, this is when Python 3 came out and we were able to add type hints for the first time, people had a lot of internal conversations: "Are we going to run out of steam on Python?" My perspective was always, "I think we can take this way further than we think we can, but I think types are going to help us not sort of be in our own way." We built this thing called MonkeyType, where we basically captured runtime types that were actually getting used in production and then mapped those back to the types in the codebase. Because of that sort of pattern, I think there's really interesting ways in which, if you're doing sort of conversion or cross-compiling using LLMs, you can also lean on production data a lot more or run sort of segmented tests. There's a lot of things you can do there. The sky is the limit there as well. I think the hardest part is always finding the boundary around where you can start doing it incrementally without trying to boil the whole ocean and swap it overnight.
Host 6:29 ↗
Yeah, your users are your test ultimately. I also read another article in the newspaper about how you can just use rollouts, and sometimes you don't really know what you're going to need it for, but when that infrastructure exists for you to experiment and to roll things out, it enables so much.
Mike Krieger 6:45 ↗
Yeah. I mean, I always found this was advice we got. When we launched Instagram, it happened to be the first week everything melted because we didn't really know what we were doing on the backend side of things. Coincidentally, that week there was a lunch that one of our investors just scheduled, not even for us, it was just an infrastructure lunch. We totally monopolized that conversation because everybody had their own opinion about how we could fix our scaling. The two pieces of advice I got there in 2010 that I will forever retain is: basically, pre-measure everything that you think you might even remotely need, because the worst thing is an outage where you're like, "Is this number normal or is it high?" and like, "Oh, I don't know because I don't have data until I just added this metric." And the other one is being really thoughtful about knobs and feature flags.
Even early Instagram, we had a very simple but really effective way in which you could do ramp-ups and rollouts and dynamic config, too, where a lot of our runtime configurations had to be changed in a matter of seconds so that we could handle load. Being able to do that in a first-class way was really important. I'm seeing that definitely in AI as well, where we're making all sorts of different trade-offs, and having that kind of runtime configuration is super key.
Host 7:52 ↗
Yeah. My favorite scaling story from Instagram, by the way, is your launch day when you DDoSed yourself with email.
Mike Krieger 7:58 ↗
Yes.
Host 7:59 ↗
Which people should look up that story if you haven't seen it. I wanted to go into tasks, a very major ship. It's how 60-something percent of your code is written today.
Mike Krieger 8:09 ↗
Yeah.
Host 8:10 ↗
How did you square that with everything you just said, where it's like very dynamic, like you don't actually ship one app, you ship one app with 3,000 flags, and like, "Well, what are you working on today? I don't know, it's for this segment of the population"?
Mike Krieger 8:23 ↗
Yeah, I think there's a bunch of things. With Tasks, I was really excited. I was talking to Swyx earlier. I'm really excited that we have Tasks out there because it is how we've been working for a while. I would get up on stages and people would be like, "How do you work at Anthropic?" and I'd be like, "Oh yeah, we use these things that are not quite Claude Code, but you know..." It's hard to describe it, but if you got to poke into Anthropic, you would see, of course, Claude Code usage for things that are more interactive or if you're kind of iterating on a particular specific thing where you want a lot of high-bandwidth back-and-forth. But most usage is actually much more delegating via tagging and via Tasks.
You can say, "Here's the..." And the reason it's really interesting is how multiplayer it is. It reminds me sort of like Midjourney, the fact that everybody was on Discord seeing how other people were using it. I think actually, to your earlier question, it really helps with that unreasonableness or ambition. Where the first time you see somebody tag Claude and be like, "Hey, don't just fix this bug, but now you are responsible for this part of the codebase, and I want you to monitor this feedback channel and proactively take on tasks and then fix them, and then also take, if this API changes, do that." I saw somebody do that, and I was like, "Oh wait, I've been totally underutilizing this thing. I've just been using it as a glorified Claude Code in Slack." That's definitely a whole new version of it, right? The more advanced version is really trying to start thinking of it as a teammate that actually holds context, has memory, and can be proactive. That's just really changed how we operate internally. It's much more like this multiplayer, async, proactive way than it is most people off in their own CLIs.
Host 10:01 ↗
Are you bottlenecked by code review and Git? Obviously there is Claude review, but someone usually still looks at it. Is there a world in which you just merge it in?

19 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Krieger, M. (2026, August 27). How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic [Interview transcript]. AI Engineer. CEOInterviews.AI. https://ceointerviews.ai/interview/1269405/

MLA

Mike Krieger. "How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic." AI Engineer, 27 Aug. 2026. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/1269405/.

BibTeX
@misc{krieger2026_1269405,
  author       = {Mike Krieger},
  title        = {How Anthropic Builds: Lessons from Labs — Mike Krieger, Anthropic},
  howpublished = {Interview transcript, AI Engineer. CEOInterviews.AI},
  year         = {2026},
  month        = {aug},
  url          = {https://ceointerviews.ai/interview/1269405/},
  note         = {Speaker-attributed transcript with timestamps}
}