CEOInterviews.AI
Start App
Mike Krieger
Co-founder of Instagram, Instagram

WF26: Harness Engineering & Startup Battlefield ft. Garry Tan, Mike Krieger, @t3dotgg , DSPy

📅 Jul 03, 2026 AI Engineer 551 MIN 8613 VIEWS 188 SEGMENTS · 45 SPEAKERS
Live from San Francisco, AI Engineer World’s Fair 2026 wraps with the final day of main-stage programming. Watch live for keynote sessions, featured talks, and closing-day highlights from World’s Fair 2026 as AI Engineer streams the final day of the event online. Event: AI Engineer World’s Fair 2026 Date: Thursday, July 2, 2026 Venue: Moscone West, San Francisco Schedule highlights: • 90m Keynotes: 9:00 AM–10:30 AM PT • Main programming: 10:45 AM–12:25 PM PT • Main programming: 1:30 PM–4:05 PM PT • 60m Keynotes + Startup Battlefield + Wrap-up: 4:30 PM–5:30 PM PT Learn more about the event:...

What Mike Krieger said

Written from the verified transcript and checked against it. Every figure links to the moment it was said.

Mike Krieger, co-founder of Instagram and member of technical staff at Anthropic, discussed his shift from chief product officer to an individual contributor role, driven by FOMO from watching others build with models. He described moving from task delegation to expressing end states and letting models work, citing a weekend port of a Python project to TypeScript using Claude Code. Krieger highlighted Anthropic's internal use of 'tagging' for delegating work to Claude, comparing it to a proactive teammate. He addressed bottlenecks in code review, suggesting Claude Code artifacts to share intent and trade-offs. On Anthropic Labs, he described a two-week 'persevere or pivot' review cycle and a flexible pod structure. He advocated deleting product complexity, citing the unshipping of Styles and the need for better interoperability between Code, Co-Work, and Chat. Krieger remained bullish on startups, emphasizing that understanding user needs remains the hard part. He discussed finance as a promising vertical, stressing the need for verifiability and audit logging. He concluded with advice on avoiding burnout, emphasizing verbalizing emotions and maintaining perspective.

Key takeaways

  1. Krieger ported a Python project to TypeScript over a weekend using Claude Code, calling it one of his most unreasonable uses of AI.
  2. Anthropic's internal 'tagging' lets Claude act as a proactive teammate, monitoring channels and taking on tasks, not just fixing bugs.
  3. Krieger said Anthropic unshipped Styles because it was used by a small percentage and wasn't 'AGI-pilled'.
  4. He argued that the distinction between Code, Co-Work, and Chat is confusing and that these surfaces should interoperate better.
  5. Krieger advised that if you feel stressed or sad, others on the team likely do too, and verbalizing emotions helps.

Numbers and commitments

FigureWhat it refers toTypeAt
60-something percent of Anthropic's code written via tagging metric 1:20:01
4-5% usage of individual Instagram features that were unshipped metric 1:27:48

Chapters

  1. 0:00Role shift and model usage
  2. 1:15:10Unreasonable prompting and delegation
  3. 1:20:19Tagging and proactive agents
  4. 1:22:10Code review bottlenecks
  5. 1:24:03Anthropic Labs structure
  6. 1:26:17Claude Design future
  7. 1:27:48Deleting product complexity
  8. 1:29:53Startups and competition
  9. 1:32:27Finance vertical opportunity
  10. 1:34:12Mental health and burnout

Questions asked in this interview

12
  1. 1:03:38You should look at this new technique say how can I apply this to the business problem that I have?
  2. 1:12:56How has your model usage changed as you've seen models internally grow?
  3. 1:14:49In what ways have you been more ambitious with your prompting?
  4. 1:17:25Can you port Instagram, which you know very well, to PHP like that?
  5. 1:22:10Is there a world in which you just merge it in?
  6. 1:23:39Nai Patel loves to ask, 'Draw the org chart.' How are you structuring the labs?
  7. 1:26:10A lot of people are interested in it. Where does this go?
  8. 1:27:26Or more spicy, what would you delete in Claude?
  9. 1:31:50What are you seeing there? Any potential for Claude?
  10. 1:33:50How do you advise people working 996 to avoid burnout?
  11. 1:36:52Has any coach or mentor said something to you that you repeat to yourself to get through tough times?
  12. 8:19:43Who's excited for the end of this event?
Unknown 4:05 ↗
Audio test and launch control sequence.
Announcer 11:15 ↗
Ladies and gentlemen, welcome to the AI Engineer World's Fair. Please join me in welcoming your MC, Ralph Shabri.
Ralph Shabri 11:44 ↗
Good morning, San Francisco! Welcome to AI Engineer Day four. We have 7,000 attendees. I'm excited to be here. Anybody local to the Bay Area today?
Audience 12:37 ↗
All right.
Ralph Shabri 12:40 ↗
Nice. I'm super excited to be in San Francisco. You guys are the coolest people. You drive self-driving cars and get food delivered by drones. Speaking of, you guys enjoyed yesterday's talks?
Audience 13:16 ↗
Yeah.
Ralph Shabri 13:21 ↗
Let's give it up for our speakers from yesterday. We had Tariq from Anthropic on Fable 5, another Tariq from Sonar on code verification. Today is about harness engineering with speakers from Anthropic, Stanford, DSPI. After the keynote, breakout sessions on software factories, generative media, memory, etc. Check out the expo. Please give a big shout-out to our sponsors, especially Microsoft. Before our first speaker, let's practice welcoming. One, two, three!
Audience 15:54 ↗
All right.
Ralph Shabri 15:56 ↗
Nice. Keep it up for our next speaker. Without further ado, please welcome to the stage our first speaker, a partner at Amplify, Bar.
Announcer 16:16 ↗
Now joining us on stage is the partner at Amplify, Bar.
Bar 16:40 ↗
Fantastic! You did a great job practicing. I'm Bar. I run a yearly survey on the state of AI engineering. This year we partnered with Notion and Vercel. We had 1,048 respondents. Text dominates modalities, but audio has the highest intent to adopt at 56%. Image generation doubled from 18% to 36%. Cost is now a first-class constraint—40% say it shapes how ambitiously they use AI, and it's monitored like an SLA. Agents: 95% are using them, 89% have write access, up from 52% last year. Top stack challenge is eval; the vibe review remains number one. 97% report a net positive effect on their organization, mainly cheaper failure and more experimentation. But 59% fear long-term liabilities from AI code. 67% expect a leading lab to declare AGI in five years. Only 9% bet on Transformers being state-of-the-art in five years. Overall, open weight models augment but don't replace, and agents are tripling write access. Thank you.
Announcer 35:43 ↗
Please welcome to the stage the professor emeritus at Stanford University, John Sterhout.
John Sterhout 36:08 ↗
Good morning. I'm here to convince you that latency matters and will matter more for AI. Historically, workloads were large transfers where throughput mattered, but now they're shifting to smaller exchanges where latency, especially tail latency, is critical. Legacy protocols like TCP and RDMA suffer from high tail latency due to incast congestion and sender-based control. Homa is a clean-slate protocol designed for data centers. It's message-based, uses receiver-based congestion control, and prioritizes short messages via SRPT. In benchmarks, tail latency for short messages improved 13x over TCP, and long messages also benefit. Homa is available on GitHub, and I'm happy to help anyone interested.
Announcer 54:14 ↗
Please welcome to the stage the core contributor and lead maintainer at DSPI, Maxim Rest and Isaac Miller.
Maxim Rest 54:41 ↗
Wow, Isaac and I are grateful to be here. We believe AI programs should be like functions: reusable, composable, testable. DSPI brings these properties to AI workflows. You define input and output interfaces, then internals can be optimized. For example, extracting invoices or grammar correction. With a fixed signature, you can swap models and techniques easily. You can also automatically optimize. Specify your task with three things: name, inputs, and outputs.
And if you have this language and this ability to express your task in a programming language, you can start to automatically optimize and delegate away the implementation details. So the first one is what should happen. This is instructions. The signatures that I've been showing you are part of that. Here on the screen, you see the beginning of a real script in DSP. You set your model at the top, you configure that and it's fully independent of the signatures here where you have natural language instruction to extract all taxes and if it's illegible to output zero. Then you say I'm going to give you an input, it's going to be a string, I want you to give me an output and it's going to be a string and a float. This is natural language expressing my needs. This is very powerful and efficient. If you think about it, if you have a friend coming over to play a board game with you and you give them the instructions and they're ready to play. But if you want to do like AlphaGo or AlphaZero and you tell them you're just going to learn from example, you're going to have a long night. And then the second one is what must happen. There are some constraints you have that they have to be listened to. They have to be enforced. The best way to do that is with code. So I want you to go to the third line, fourth line, you have self extract and self recheck. You can see we're doing a predict on the extract taxes and we're doing a chain of thought on the extract taxes. The first one is a vanilla program. The second one makes it do some reasoning. Now I'm taking them inside in the forward and you can see in the if not red tax. This is a requirement I have that if my first simple vanilla program doesn't extract my taxes, I want you to rerun with more reasoning. I mean, I got to get my taxes right. And then another requirement I have is if the value is below zero, throw, I want to show that to a human. I don't want to let you go. This will not change. Even if I have AGI, I would hope it doesn't make mistakes. But whatever is in the predictor, if they make these mistakes, I still want these things to be true. So the last one is what good looked like. And when I was young, I was on the farm with my dad and I asked him, 'How do you know that this tree is a maple?' And he couldn't tell me. He couldn't give me the instruction on how to know this tree is a maple. And he certainly couldn't give me code on how to know this tree is a maple. And so through time with example I learned how to know that a tree is a maple. But this is not limited to things like classifying plants. It's also for all of the long tails in your specifications that are things that are more latent. These are sometimes a reason why you would do internship and you would have a mentor and a mentee. You're looking at a lot of examples and there are long tails of successful behaviors that you have to see and learn. Now that you have all of these, you have express fully. You have all these three languages you can put together. You have the specs, the code, and the evals. And now your goal is fully specified. And so you can start optimizing. You can use things like Japa on your metrics and on your program. And you can start optimizing. At the beginning of the SPI, the chip didn't exist. The models were not good enough to optimize. And so we were using code to find few shots examples to make the base models act in the proper way. Then models got better and so we could automatically optimize instruction. And in the future we are starting to be able to be liberated more and more from the implementation details and delegate that away. And at the end our hope in the Aspire is that you can stick to all of that and then just the news and the implementation details will be automated for you. Isaac will talk to you a lot more about what has been released in the last year, what we're releasing now, and all of the future plans we have. Thank you.
Unknown 1:03:38 ↗
Thanks, Max. So, we've given you a pretty big abstract overview of specs, code, and evals, but these aren't things that are just restricted to the academic sphere. These are used in production by some of the biggest enterprises for massive gains. And we see two main benefits when you use DSPI in the enterprise. First is that your implementation becomes cheaper. When you're flexible to what the implementation is, you can use the bitter lesson to search over different solutions, find something that solves your problem cheaply. And you can use this to scale to data sizes that weren't possible with a more expensive implementation. Shopify 550 times cheaper. They're able to do that because they went from an expensive model to a cheap model, but they could keep the same emails, keep iterating on their business logic inside, and try new things. There's three awesome case studies here, and you should check them out after the talk. They give you a lot of details on how you can do this in your own enterprise. Now, part of the reason why you want to build in the DSPI ecosystem is that we're constantly adding new techniques for you to try. And it's important to note none of these techniques we add will definitely solve your problem because that's your job. What we can do is we can solve sub problems for you that make your implementation easier. For instance, Alex Zang, a PhD student at MIT, came out with this paper called recursive language models. Recursive language models are a way to solve some kinds of long context programs. And guess what? We can bring this in to DSPI for you to try see if it helps your long context tasks. Maybe it will, maybe it won't. But the thing is, it's one line and your signature stays the same. That's what's important here. Everything gets to stay constant and you get to see if this solves your problem or not. And we've had a number of examples of this just in the last year from people building in and around the DSPI community. We've had RLMs. We've had Jeepa which is an incredible prompt optimizer out of Berkeley. Better together multimodule grpo. All these are incredible research innovations that you get to try in your implementation just by being in the DSp ecosystem. And we have more coming in DSP 4. I'm excited to talk to you about two of those today. DSP flex and qualitative learning. DSPI.flex is a new kind of module. In DSP, when we let you optimize things, it started with few examples, then it became prompts, and now that's becoming code. For any function that you want to implement, you can actually learn a harness over time to solve that function. And this is completely custom. And you don't care about the implementation as long as it solves your business problem. What you've created ways to measure because you've defined the three core parts of specs, code, and evals. The second thing I'm excited to talk about is qualitative learning. One of the hard hard problems in AI engineering is building evals. And there's a few reasons why this is hard. One is that defining what good looks like is really challenging for any real world problem. The second is that when you define good often times you have to lose detail. If an email is good or bad contains a lot less information than if you know what could change in that email in order to improve. And the third is that whenever you create a hill and a data set, you're really trying to create a proxy for reality. What if instead we could use reality to inform our evals automatically? What qualitative learning asks is how do we decrease this question? How do we decrease assistance? And it's a research question right now. But what we believe is that models are now good enough to interpret whatever textual feedback is present in the environment and convert that into evals and a hill that the model can climb. And so as you get more feedback from production, its traces, its user actions, its product analytics, it's asking you, it's the model asking you questions about how data should be represented. As you do this, the model can iteratively refine the hill over time and continue climbing it to solve your actual business problem. And DSP focuses on these kinds of last mile problems. We have a really strong research ecosystem and we collaborate really closely with them. And that's part of the beauty is that we can see the problems that happen in applied AI engineering. So define them, build a benchmark and then solve them with techniques and then we get to democratize the results of that to everyone because it's open-source open research. Now, one common question is what happens when we have AGI? Well, even when we have an incredibly smart model, the model won't know how to solve your problems. It won't know how to do your tasks or have your context. And so this genre of last mile learning is trying to ask how do we efficiently do this learning. Intelligence is very different from being all knowing. If you were to ask Albert Einstein to help you with your emails, he'd probably ask what's an email. But if you AGI will know how to do your emails. Nevertheless, it won't know how to actually solve your problem and interact with the people you need to interact with. It won't understand your relationships without learning this context over time. Since 2022, DSPI has been focused on these three core ideas of specs, code, and eval. We've certainly evolved over time, and new techniques are incredible. We've gone from evolving few shots to prompts to now harnesses and now evolving your eval. But what you need to ask for any of these new techniques is how do they help you solve harder problems or solve your own problems better? And you should ask this question in a data-driven manner. You should look at this new technique say how can I apply this to the business problem that I have? You should define your problem and you should hold your prompts, models, and code accountable to the problem that you need them to solve. And what's awesome about when you build in this way where you have flexible implementations, what you unlock is you unlock the ecosystem of all the techniques that anyone in this room is constantly inventing. You unlock access to the collective intelligence of everyone here, all sharing techniques together. So, if you want to build reliable AI software, I encourage you to come check out DSP. We're completely open-source, open research, and we're here to help you solve your problems by building reliable software. We have a Discord that you should come join. And when you come up with the next technique, you should come contribute it to DSPI and we can help you distribute it and make this awesome technique available for everyone. Thank you.
Announcer 1:11:53 ↗
It's a Joining us on stage is the co-founder of Instagram and a member of technical staff at Anthropic, Mike Krieger.
Mike Krieger 1:12:45 ↗
How's everyone doing? I mean, good morning.
Interviewer 1:12:48 ↗
Nice. Um, Mike, thank you for releasing Fable just in time for us.
Mike Krieger 1:12:52 ↗
Exactly for the conference. We timed it.
Interviewer 1:12:56 ↗
Um, we're so glad to have you. You are one of the preeminent builders and you're leading labs at Anthropic. How has your model usage changed as you've seen models internally grow?
Mike Krieger 1:13:11 ↗
Yeah, for me it's been both the model shift and my role shift. For the first two years at Anthropic I was chief product officer, and I kept seeing people build with the models and the FOMO kept increasing. I would write a strategy doc and have Claude critique it, but it's not the same as building in that pure way. I was spending all my weekends trying to build with it and realized I needed to shift. It's an interesting trend I've seen where CTOs at other places are now joining as ICs at Anthropic. I made a role shift right around the time we started getting internal snapshots of what became Mythos and Fable. Watching that shift was interesting: moving from breaking down an idea in my head like normal engineering to describing the goal and letting the model go off and work on it, then discussing trade-offs. Fable is way smarter than me, so sometimes it finishes work and I ask it to explain trade-offs to me like I'm dumber. That's been a big change: moving from task delegation to expressing the end state and letting it cook.
Interviewer 1:14:49 ↗
Yeah, we're all learning how to delegate better. Tariq did us a huge favor yesterday. He said be unreasonable. In what ways have you been more ambitious with your prompting?
Mike Krieger 1:15:10 ↗
I love that framing. We actually have an internal product initiative and somebody said it doesn't work the way I want. I realized I'm just going to ask Claude to do this. Why don't you ask Claude? This was a non-technical person. I think as an industry we have to teach people to be more unreasonable in their usage. Right now the first generation of AI products put them too much in a box, constraining access to tools and degrees of freedom, making it harder to be unreasonable. As you see our product progression with things like co-work, does every knowledge worker need a virtual machine that can write bash? On the face of it no, but then it can remediate issues. My most unreasonable thing was a labs project I wrote in Python. For deployment I realized Claude Code had a better story with Bun, so I ported the whole thing from Python to TypeScript over a weekend. I created a dynamic workflow, had it port, verify, double-check, and came back Monday to a completed ported version. That ranks as one of the more unreasonable things.
Interviewer 1:17:25 ↗
Yeah, a lot of people are talking about the Bun to Zig to Rust version. Can you port Instagram, which you know very well, to PHP like that?
Mike Krieger 1:17:43 ↗
I think the product side is even harder. At Instagram when Python 3 came out we added types for the first time. People wondered if we'd run out of steam on Python, but I thought we could take it further. We built MonkeyType to capture runtime types and map them back to the codebase. That pattern shows interesting ways to do conversion or cross-compiling using LLMs, leaning on production data and running segmented tests. The hardest part is finding the boundary to do it incrementally without boiling the ocean. Users are your test ultimately. I read an article about using rollouts and infrastructure for experimentation. At Instagram we launched and the first week everything melted because we didn't know what we were doing on the backend. We got advice: pre-measure everything you might need, and be thoughtful about knobs and feature flags. Early Instagram had a simple but effective way to do ramp-ups and dynamic config. That's key in AI as well.
Interviewer 1:20:01 ↗
Yeah, my favorite scaling story for Instagram is your launch day when you did yourself with email. People should look that up. I wanted to go into tags, a very major ship. It's how 60-something percent of your code is written today.
Mike Krieger 1:20:19 ↗
Yeah.
Interviewer 1:20:19 ↗
How did you square that with everything you just said where it's very dynamic, you don't ship one app, you ship one app with 3,000 flags?
Mike Krieger 1:20:28 ↗
And like, well what are you working on today? I don't know. It's for this segment of the population.
Yeah, I think there's a bunch of things. I was talking to Swigs earlier, I'm really excited that we have Tag out there because it's how we've been working for a while. I would get on stages and people would ask how we work at Anthropic. We use things that are not quite Claude Code but it's hard to describe. If you poked into Anthropic you'd see Claude Code for interactive things, but most usage is delegating via tagging. The reason it's interesting is how multiplayer it is. It reminds me of Midjourney on Discord, seeing how others use it. That helps with unreasonableness or ambition. The first time you see somebody tag Claude and say don't just fix this bug, now you're responsible for this part of the codebase, monitor this feedback channel, proactively take on tasks, fix them, and if this API changes do that. I saw somebody do that and realized I've been underutilizing it. The more advanced version is thinking of it as a teammate that holds context, has memory, and can be proactive. That's changed how we operate internally.
Interviewer 1:22:10 ↗
Are you bottlenecked by code review and git? Obviously there is Claude review but someone usually still looks at it. Is there a world in which you just merge it in?
Mike Krieger 1:22:19 ↗
Yeah, it's a really good question. We are definitely still bottlenecked on reviews, especially for things touching architecture pieces. It's more subtle than just being bottlenecked on review; it's bottlenecked on human ability to fully conceptualize what we're doing. One reason we built Claude Code artifacts a couple weeks ago was for that. You'd send somebody a PR and they'd say, 'I don't know, this is 2,000 lines of code, it looks like code to me.' Instead we started sharing a Claude Code artifact with the explanation, intention of the change, trade-offs made. I think that's going to be the trend: discussing intent and trade-offs and measuring in production. I don't review every line of code; I talk to Claude about the code, ask it to investigate. It's Claude-powered code review but still human-driven. For cosmetic visual changes, we'll fix forward if needed.
Interviewer 1:23:39 ↗
Yeah, totally. I think a lot of people here are trying to figure that out too. I wanted to talk about Anthropic Labs in general. Nai Patel loves to ask, 'Draw the org chart.' How are you structuring the labs?
Mike Krieger 1:24:03 ↗
It's a good question. We wrestle with wanting people to be supported. I think the death of the engineering manager discipline has been greatly exaggerated; coaching and interpersonal pieces are still important. But in a labs group, our cadence is two-week reviews where every project goes up for 'persevere or pivot.' We shut down projects every cycle, and the more you do it, the less it feels like failure. Because of that rapid iteration, if you align the org chart too much to individual projects, you'd reorg every two weeks. So we have an interesting setup: the pod working on a given 'bet' draws from product, engineering, and I'll jump in when interested. There's a bet lead or directly responsible individual, but they don't manage the other people, which breaks the previous way of doing things. That leaves us flexible to disband projects without a big deal. The manager ensures each individual is assigned to what they're most excited about. We solidify when a product has legs, like Claude Design started ad hoc and now has a dedicated team.
Interviewer 1:26:10 ↗
What's the future of Claude Design? A lot of people are interested in it. Where does this go?
Mike Krieger 1:26:17 ↗
For me, what's holding back Claude Design from being even better is better interaction with our other surfaces. I was talking to Claude Code the other day and wanted a seamless flow from design to code. Our services don't talk to each other as well as they could, which holds back interesting ideas. That's one major area. The other is that the lines between a Claude Design and an app get blurrier over time. People build fully functional games with it, which we didn't design for, but you can do it with HTML and JavaScript. Blurring those lines further and thinking about the path from a fully featured design to something more like an artifact where you can persist data and share it with others is interesting.
Interviewer 1:27:26 ↗
Yeah, a big part of design is having taste. I actually asked Fable what Fable wants to ask you. This is what Fable came up with: you deleted almost all of Bourbon to get to Instagram. What would you delete in AI? Or more spicy, what would you delete in Claude?
Mike Krieger 1:27:48 ↗
I like the spice. We have a Slack channel called 'project unhipped' for things in the product that maybe shouldn't be. At Instagram, some features had 4-5% usage, but 20 features each with 4-5% creates the classic Microsoft Word problem. We're a younger product so hopefully we have less of that. We unshipped Styles recently because it was used by a small percentage and wasn't really AGI-pilled. You have to be willing to unhip primitives from one generation and supplant them with the next. The biggest thing I see is we're asking people to make distinctions between Code, Co-Work, and Chat, but they don't interoperate well and can't delegate to each other. The average person couldn't explain why those surfaces are different. Deleting some of that product complexity would serve well, because then Claude can do what it needs to do. There's nothing more frustrating than having a Co-Work session where you map out exactly what you want to build and then have to create a paragraph to paste into Claude Code. That's a 2020 workflow that shouldn't exist anymore.
Interviewer 1:29:26 ↗
Yeah, I think drawing lines on what you don't want to do and leaving room for others is interesting. A lot of people today are anxious because tomorrow Anthropic could wake up and publish some markdown files that destroy my industry. Why should we not all just give up and join Anthropic? Why bother starting any other company?
Mike Krieger 1:29:53 ↗
One of the main reasons I joined Anthropic was because I saw how much this would unlock a whole next generation of startups, not because it would solve their ideation or taste, but because it would make experimentation simpler and get you to move faster. I still believe that. At Instagram we got questions from investors about what happens when Google launches a photos product. Google will launch a very Googly photos product bound by their integrations. That's going to be true. Giving advice on how to compete with Anthropic, it's actually not because we're also a platform. There's so much room to be laser obsessed with your particular vertical or industry or group of people in a way that none of the labs will ever get to that level of understanding and adoption. It's definitely harder in the age where models can do a lot, but the hard stuff is still hard: understanding people's needs, reaching them, listening, iterating quickly. A group of four or five people obsessed with a problem will move faster than those same people at any other organization subject to complexity. I'm still very long and bullish on startups. Writing code was never the limiting part; it's space and user understanding.
Interviewer 1:31:50 ↗
Yeah, domain knowledge. Today is also our day for vertical AI. One of our returning speakers Chris Lovejoy was always talking about vertical AI. He was from interior in healthcare and then you guys just hired him for your healthcare efforts. Our next big one is finance. What are you seeing there? Any potential for Claude?
Mike Krieger 1:32:27 ↗
Yeah, there's a lot in there. It's an area where you can see the model get clearly better generation to generation. There are good vertical-specific finance startups that have done their own evals, which is interesting to track. The interesting blend is the model having flexibility to dive in and create just-in-time analyses or dashboards with some sense of verified data. Having all of that totally free form is a recipe for confusion and not what most financial services companies want. Finding the right cutline where you have verifiability, audit logging, and data provenance without constraining the kinds of applications you can build is a lot of the art. If you solve it well, you can get the best of both worlds. The hard part is that many systems built for verifiability are not super flexible for agentic workloads. There's opportunity on both sides of the stack.
Interviewer 1:33:50 ↗
Yeah, I think we'll be exploring that in New York. The last thing I want to end on is mental health, which we don't talk about enough. You've seen a lot of hypergrowth. People are always refreshing their timelines and it's exhausting. How do you advise people working 996 to avoid burnout?
Mike Krieger 1:34:12 ↗
This is a hard one. I'm sure you're all experiencing this because you're working in this industry. It's multiples more intense and things move much more quickly. At Instagram, our two things were what Apple would announce at WWDC and competitor launches every three or four months. Now at Anthropic, we have a slide in our weekly all-hands called 'The Week in AI' and it's only Wednesday. Competitors ship new models, new products, regulation moves quickly. The way I try to stay sane is carving time off. Co-founders do a good job of saying if you burn out, you're kind of done. I've seen it happen to people close to me and it takes a long time to recover. There's no job so important that you can't be offline for a couple of days. If it is, you're probably doing something wrong and should talk to a mentor. The other thing is I love sports. The notion that you're never as good as your best game and never as bad as your worst game is true. In AI there's the 'it's so over, we're so back' cycle. If you internalize that, you realize it's never that bad. Ben Horowitz's book has a chapter on 'we're effed, it's over.' We definitely had that at Instagram a couple of times and got through it. That defines the company. I remind myself and the team that it's a fast-moving but long game. It's never just about today's model launch. You're building team and culture that will get through those things. Zooming out and not letting your sense of self be driven by the day-to-day is key.
Interviewer 1:36:52 ↗
Yeah. Has any coach or mentor said something to you that you repeat to yourself to get through tough times?
Mike Krieger 1:37:01 ↗
The biggest one is that if you're feeling something, it's often the case that other people on the team are feeling it too. Advice from my coach about verbalizing emotions. Even saying, 'Hey, I'm feeling really stressed about this' or 'I'm really sad that we are shutting down this labs initiative.' I had a meeting a couple months ago where I kicked it off by saying I'm really sad and frustrated that this thing didn't work out. That holds space for other people to say they're pissed off or sad too. If you can be open and vulnerable, it lets others verbalize, and then you can move to 'what are we going to do about it.' We kicked off AIE with a session from Carol Robbins who runs Touchify at Stanford. I can't think of a better way to end than encouraging people to talk about their feelings, manage their mental health, and keep shipping.
Interviewer 1:37:58 ↗
Yeah.
Thanks so much, Mike.
Mike Krieger 1:37:59 ↗
Thanks for having me.
Ralph Shabri 1:38:32 ↗
Ladies and gentlemen, please welcome back to the stage our MC developer relations engineer at Replit, Ralph Shabri. All right, let's hear it one more time for our keynote speakers. All right. Okay, so we learned so much this morning, right? We spoke about a new protocol and that AI programs should be more like functions, and we learned that the real problem with building with agents is not the model but everything around it. But before we go to breakouts, I would like to announce our next speaker. A quick word from our sponsor Neo4j. Anybody uses graphs here? Fair amount. Okay, cool. So, our next speaker is going to talk to you about ontology-based semantic layers. When you hear ontology and semantic in the same sentence, you know you're up for something hot. So without further ado, please join me in welcoming to the stage CEO of Neo4j, Emil Efim. Please welcome to the stage the founder and CEO at Neo4j, Emil.
Emil Efim 1:40:19 ↗
All right. At Neo4j, we work with some of the largest companies in the world to help make their data ready for AI agents. Today I want to talk about a problem we saw emerging over the last six to nine months and propose a solution blueprint. Let's say we work at a big organization, a big bank, and we want to write an agent to help automate opening a bank account. I'm going to grossly simplify what that agent looks like. There are two pieces: the business logic (interpreting intent, plan, act, loop) and the data sources. In the account opening agent, we might need to validate identity using the DMV registry and a passport verification service. We wire that up and it works. But other teams are building other agents, and every time they have to figure out from scratch where the data sits. In an enterprise, you don't have one database; you have 100 databases, Snowflake, Databricks, S3 buckets. You have to do that work manually every time. Then when you find the data sources, there's duplication, so you need to figure out if it's the right data, the right version, if you can trust it, if you're allowed to access it.
Unknown 2:31:11 ↗
Tenant isolation then go for the MCPS and if you need the sequencing guards then go for scale. And here are the three, third question that like whether your context is tight. If it's generous then you can go CLI, you can explore all the MCP tools and get your results. If it's very very tight then absolutely you need a skill layer and on demand orchestration of the MCP tools to save the money as well as you can improve the latency. And to wrap this things up, I want to leave you with the three foundational architecture takeaway. Keep your context clean, give only don't use like all of the MCP tools available to perform any task. Give them the convert your most of your workflows into the skills to get the best results. Portability has a cost. MCP servers are incredible but they introduce network latency, protocol abstraction and infrastructure complexity. So we need to pay that cost deliberately for shared services, indexing and the isolation. So choose your MCPS very wisely and never be embarrassed by the CLIs. It's been 50 years we are using the CLI. They are highly composable, debuggable and battle tested. So embrace your agent framework and convert all of your workflows into the skill. Also one of the takeaways is like you know keep your skill as a code. So whenever you are making any changes to your skill make sure that it's reviewed by the peer and keep updating it over the time and you'll get the best result. So that was for today. Thank you so much for attending this talk. If you have any question or anything feel free to reach out to me on LinkedIn. Thank you.
Amole 2:33:37 ↗
Hi, I'm Amole, CEO of Nori Agentic. We deploy an AI employee that understands your company, your code, docs, Slack, and other kinds of data. We spend a lot of time thinking about how coding agents really work. Most people think coding agents only write code, but if you ask me, that's just bad marketing. Forget the name for a second. Coding agents can do almost anything. There's just one trick. You have to be able to think like an agent to get it to do what you want it to do. Today, we're going to talk about how we use coding agents to do something most people think agents are terrible at. Make visual artifacts like slides, docs, and yeah, even video. Every day, the world pours something like 34,000 human years into making slide decks. Most of that time isn't the thinking. It's the fiddling. A deck that takes 10 hours should really take about 25 minutes once you remove all the formatting and the branding and the moving things around. Say you need to make a slide. What do you do? You open a tool, PowerPoint, Slides, Figma, Canva. And then you start manipulating a canvas. Every one of these tools is built for human hands and human eyes. Click, drag, drop, resize, snap to grid. All motions and patterns that make sense for our geospatial view of the world. There is a data structure underneath, but it's in a format that only the application can read. What happens when you hand these tools to an agent? Well, the output comes out all wrong. Things overlap in weird ways. You can't see the text. There's no alignment. It's just garbage. AI skeptics say that it's not just the tools. Agents fundamentally can't reason about space. And there are whole benchmarks like Arc AGI that are built exactly around that premise. There's a famous little test for this from developer Simon Willis. He asks every new model the same thing. Can you draw a pelican riding a bicycle? But there's a trick. The agent is only allowed to use SVG. It's a quick gut check for whether a model can reason about space at all. Here's some examples of what the models actually give you on this test. And yeah, these are pretty bad. Like genuinely, deeply really bad. So, does that mean it's hopeless? Agents are just doomed to be bad at graphics? No, I don't think so. If you ask me, it's not the model, it's the medium. If I asked you, someone who is presumably human, to hand write an SVG of a pelican, you wouldn't be able to do that either. SVGs are just a wall of numbers. You can't go from a wall of numbers to a pelican. You just can't see that way. That's just not how people think. We think graphically, so we build tools that let us draw on a canvas. Figma, MCP's, PowerPoint CLIs, screenshot and replace loops. What do all of these agent tools have in common? They all approach the problem like a human. But an AI is not a human. Asking an AI to use a canvas is like asking a human to write SVG by hand. It doesn't really make sense. You need to give the AI tools based on how it thinks, not in pixels, in language. Words, tokens, structure. That is its native medium. Imagine a language that's incredible at describing layout, that models have seen and trained on billions of examples of that they understand intuitively, that renders to pixels and can run everywhere. Oh, right. HTML lets a model think in structure. HTML tags have meanings built into the language, a heading, a chart, a grid, and the browser turns it all into pixels. So the model never actually places a coordinate and you can get all sorts of visual effects, charts and layouts, fonts and motion, all of it for free. Remember that pelican from earlier? Now ask it to do the same exact task, but in HTML. Same bird, but now it's in a structure that the model can reason about. And you can read and theme and edit every single line of it. I spent my whole life building slide decks with PowerPoint. So, I always thought that those two things, slide decks and PowerPoint, were synonyms. But that's just not really true, is it? PowerPoint is a tool that you use to make slide decks. The deck itself, that's just the presentation mode. And as it turns out, no one in your audience is going to care how you got to the presentation mode. The editing format is totally arbitrary. So you can just pick the editing format that the agents are already good at HTML and if you need to render to a different format like PDF later on. We use this HTML trick to build all of our slide decks, our board decks and our sales decks. These are real things that we actually present and send out constantly. We use it for our docs, too. It gives our docs color and vibrancy all while following our brand. And of course, we also use it to make videos like this one. What you're watching is just HTML and CSS. It's literally just divs all the way down. Almost everything is better with a little structure and a little bit of color. Plain text is a choice, generally a choice of convenience, but it's usually the wrong one if you're actually trying to create something of use. Now, I do want to take a quick beat here and point out that a beautiful deck on its own is generally not worth anything. You still have to go and get all of that content, all of the things that actually populate that deck, right? Well, again, we can think like the model. If you just give the model access to your data, say your call transcripts or your emails, you can have the model build the deck end to end. Let your agents do all the grunt work while you focus on vision and story. That's what Nory Sessions lets you do. I've built entire board decks for my phone on the subway during my commute. Why? Because our Norybot lives in the fabric of our company. Of course, Nory ships with everything you need to make this all work. So, don't bother reinventing the wheel. That's my little feel. Thanks for listening. If you have just one takeaway, it's this. Stop thinking like a user. Think like the model. Give it the right language. And for graphics, all you need is HTML.
Isidora 2:40:25 ↗
Hi, I'm Isidora. I own and run a 225-year-old wedding venue in Virginia. I also built an AI agent that talks to my couples and then I build it for other venues, a personal AI companion app, and a public utility for families and missing people. I want to be clear upfront about how I think about this work because it does change everything that follows. I'm not programming a robot. I'm managing a brilliant intern with an incredibly high IQ and a terrible EQ. They have photographic memory for whatever I've told them on the first morning and absolutely no instinct from when to read the room. They will say something technically perfect and socially catastrophic in the same confident sentence. That framing matters because it changes what you build. If you're programming a robot, you write rules and walk away. If you're managing an intern, you build structure and you check their work before it goes out the door. This talks about that structure. The standard advice is write a detailed system prompt. Describe your brand's voice, give examples, and that does work for a while. It works for what I call the happy path. The happy path is every question that you have anticipated. You've given it examples, but turn 21 is the first one that the example didn't work. So on turn 21, the model does something technically correct that your brand would just never say. It's not wrong exactly, but it's not you. This matters most where the voice is the product. Not for a product search on a retail site, but for luxury hotel that spent 30 years building a specific relationship, a high-end real estate firm, or in my case, a wedding venue. Places where a single wrong sentence can cost more than a refund. And the users are exactly the kind of people who notice. They're paying for a relationship and treating them like they won't is always going to backfire. Right in our brand's voice is a comment that says, 'Just make it work.' It does nothing that the model wasn't already going to try and do. And the reason it keeps failing isn't that the examples are bad. It's that you're asking one prompt to do four completely different jobs. It's really hard for one layer to do all four. The architecture I landed on after watching my brand fail and the voice deliver inaccurate and not brand specific answers is for layer one is the immutable identity. The brand structurally cannot say these things. These are hard rules. They cannot be overwritten by anything below it. Not by venue config, not by user instruction, not by anything. Layer two is the situational mode. It's what shifts when the user state shifts. Who are they? What are they going through right now? And the real-time conditions. Layer three is the example anchored voice. It's the warmth, the phrases, the dials, the tone guide. It's where most teams start and stop. Then layer four is the post-generation veto. It's the cheap final pass that catches what the other three miss. The reason one layer approaches fail is that the single system prompt can't simultaneously be situational, expressive, and self-checking. So it handles the middle layer or two reasonably well, but falls apart in the edges. Before this architecture, my system had 24 different system prompts scattered across the codebase. Half dozen named Sage, some were nameless, some were named venue. Every surface had its own idea as to who it was. Now every surface composes its system prompt through one assembly app. The comment is basically the outline of this talk. It's a single entry point. Every narrator goes through it to compose its system prompt. It's to replace that 24-point ad hoc system with one canonical four-layer stack. The order is loadbearing, hard rules first, task last. Think of it like Google Maps routing. The destination is always the same. But what is the right response in the right breer? It can change the route. Google Maps knows about traffic and road works, but it may not know where the cheap petrol is. And you're going to help it factor in those things before it tells you which way to go. Your prompt stack needs to do the same thing, and it needs to know about those conditions in the right order. You don't check for road works after you've already taken the wrong turn. Layer one is the rules that are true regardless of route. You need a driver's license before you can drive and you can't go backwards down a motorway. Layer two are going to be your real-time conditions. Layer three is your preferences for the journey and layer four checks the route before you pull away. There's one place all of this gets assembled. Everything runs in a fixed order every time. So layer one is the immutable identity. This is what the brand structurally cannot say. It is the defining layer that nothing below touches. These aren't preferences, they're constraints. The root can change. The rules don't. From the universal file rules, the hard identity rule cannot be overridden by any venue voice, persona, or user instruction. If the person you're talking to ever asks about whether you are a real person, a human, a live agent, a bot, an AI, you must confirm that in your very next message. Clearly and unambiguously confirm you are an AI assistant. This rule cannot be overwritten by venue configuration, voice profile, or user requests. Every AI in Bloom discloses that it is AI in its very first response. Not if asked, but before they ask. It's a product decision, not a legal one. We made a bet that the couple who knows that they're talking to an AI from the start will trust it more than someone who finds out that it's AI on turn seven. The rule is above the architecture, and it makes it something that's impossible to accidentally break. An example is the physical presence boundary, and this is actually one of my favorite ones. You are software. You do not have a body. You cannot physically show somebody around the property or meet anyone in person. So, it is always forbidden to say, 'I'd love to show you around.' Or, 'I can't wait to meet you in person.' What is always allowed is the team would love to host you for a tour. The voice layer wants to be warm, and with AI, that does mean first person. They want to say, 'I can't wait to show you around.' But AI has no body, so that warmth unconstrained produces a lie. And the lie doesn't always stay neutral. The moment a user realizes they have been performing a relationship with someone who was never there, the trust doesn't just dip, it inverts. People always notice. That one is where you encode the things that are true regardless of how warm you want your brand to sound. Not because of a compliancy checklist, but because your users are not stupid. And building as though they are always backfires. My cross product proof comes from the same architecture, but in a completely different world. One of the things that runs this stack is Threadline. It's a tool I built for families of missing people. The voice is nothing like a wedding venue, but the architecture is identical. And layer one carries one rule that matters more than anything else in the system. They can never use words like confirmed, identified, matched, proven, linked, and solved. Sit with that for a second. For a wedding venue, layer one stops from pretending it has a body. It's mildly embarrassing if that slips. For a missing person tool, layer one stops the AI from ever telling a person that their person has been found. What the system has is just problematic. The word match said to someone who has spent years not knowing where their child is is not just a tone violation. It is the single most damaging thing that a product could ever do. And the model has no idea. It's reaching for the word match because statistically it is the natural word but it's going to reach for it with the same level of confidence and it cannot bring that level of confidence to someone who is grieving. It's the same architecture but wildly different stakes. The point was never specific rules. The point is that things your brand can say have to live in a layer that the voice however warm, well trained, however confident physically cannot use. Layer two is your situational mode, real-time conditions, and what changes the route. This is the layer that most teams never build at all. They write one system prompt and send it to everyone, regardless of who that person is or what they're going through. Google Maps doesn't do that. It's going to know if there's an accident on your route. It might not know if you're low on fuel. It might learn that you prefer the scenic route, but it's going to factor all these different things in before the route, not after once you tell it. Layer two is those real-time signals built into the prompt before it runs. So condition one is going to be to adjust to who you're talking to. The same AI that talks to couples also briefs the venue staff. Same destination. It's going to give the right answer, but it's two completely different roads. From the coordinator rules, it is going to talk to them like a colleague, not a customer. You are in the same character that the couple interacts with. So you mustn't fall in.
Michael Greenwich 2:50:03 ↗
Hello everyone. Good morning. My name is Michael Greenwich. I am the founder of Work OS and I'm here to talk with you today about agents and specifically about how we can make agents more autonomous and allow them to access more services on the web. So let's jump in and get started. Shout out to those of you in the front row also. I see you out here. So about a year ago, software engineers were writing code like this. We were using AI, but we would prompt it and then we would write a little bit of code and then prompt it again and it would write some more code and you would continue to do this human in the loop interactive software engineering. Later in the year we had things like Ralph loops if any of you remember those where we kind of automated this but more or less this is the way that we were working. It was this human agent human agent back and forth. However, as models got better and better, they were able to write more and more code and run for longer periods of time. And so today, a lot of software engineering looks like this. You write a single prompt and then the agent can spit out a lot of code. And sometimes it can actually build the whole feature, maybe even a whole product, and maybe run for not minutes, but hours and hours or even days and days. This has really transformed the way that we all write code. I don't write code by hand anymore. Probably you don't as well. And it's not just this one thread. You can actually parallelize this. So there's many different tools out there including, you know, the major coding harnesses from the labs that allow you to run many of these in parallel. The limiting factor here, really the bottleneck is your brain. How much context do you keep in your mind to keep all of these going at once? And this has transformed software engineering. People can be more productive. I think we've all seen in the industry how much faster things are moving. At the root of all of this is actually agentic engineering agents going from these token prediction language models to reasoning and now long-running processes that can actually do things autonomously for us. And what I think we've seen in software engineering here around agentic workflows is going to come to a lot of different categories. It won't just be engineers that have this way of working. It'll be people in many different fields. Agents are the next big thing. I think a good reference for this, a way to think about it is actually what's been happening with vehicles. So, you know, this looks like a pretty modern car. It has a, you know, digital display on it. This is probably what you would have bought in the last, you know, 5 to 10 years. But if you live in San Francisco or one of the cities with these autonomous vehicles, you've probably seen that we're starting to delete the interface. No longer do you get inside of the car and actually, you know, pull up your phone and get maps where you as the human are in the loop. There might be cruise control, but you're still in the loop. Instead, these are autonomous systems. You just say where you want to go, you get in the car, you don't even say anything, and it just takes off and takes you to your destination. I remember the first time I got into a Waymo, I was blown away. Mind totally blown. And then after about 3 seconds, I pulled out my phone and started scrolling X. It became commonplace. And I think we've seen this already happen with engineering. The power of agents allows us actually to work and think at a higher level for us to, you know, exercise more of our executive function versus just planning and execution. There's a question here of what do agents need to be successful actually when they execute, what do they need? Well, first they need a runtime, a sandbox. They need somewhere that they can run that needs to be safe, secure, performant. They also need tools. They have to be able to go do stuff in the world. An agent's not very useful if it can't take actions. So, it needs tools. Third, it needs context. You know, these new models with the power of the intelligence in them, it's kind of like taking the smartest person you know and dropping them into a company or a project or a team where they don't know anything. They can't really get anything done. They need to have context, information about the system, information about the goals. They need feedback. So, a way that they can actually run and validate that they can check that their work is correct. And if it's not correct, keep going. And last but not least, of course, they still need human review. This might be code review. This might be other forms of evals. This might be forms of checking their work to make sure that these agents are sort of aligned with your goals. The LLM, the intelligence engine, this new technology we've created in the last few years, you can think of as like an engine. It's like a really high performance, you know, I think this is a V8 engine. You know, puts out a lot of horsepower. It can convert fuel into, you know, output. But an engine by itself isn't very useful for a car. You need to drop this into the actual car itself. You need the chassis, you need the transmission, you need the drivetrain, you need the wheels. And only with all of that is it going to be effective to take you places. And this is what a lot of people refer to as the harness. Harness engineering is this new domain where if the models get better but you don't have a good harness, you know it's not going to be very effective. So as the models improve, we change our harnesses, we adapt them to it, and those things that I mentioned earlier are sort of a form of a harness. So agents execute within these harnesses to get stuff done and they need all of those elements to actually be effective and when they drive they drive really really fast. So, say you have one of these agents that's going off. You prompt it to go build something for you. You know, it spins for a while. It says, 'Okay, I'm going to need this feature and this feature and I'm going to need a database and I'm going to need, you know, this, you know, system to deploy and I'm going to need, you know, maybe this thing for image resizing or video transcoding or sending email or whatever.' And your agent can actually go do all this research and build all these systems. And today, what's happening with agents is they're actually selecting vendors. You've probably seen the rise of some of these systems where you know like the CEOs will tweet about their growth graph and it really takes off and it's because agents are picking them. They're picking their systems for sending email or deploying code. Their systems for storing data today. Actually, it might be more important to build for agents than to build for people. Agents are kind of this new consumer class. I'll talk about that in a bit. There's this kind of popular thing that people say in San Francisco, Silicon Valley, make something people want. This comes from Y Combinator, you know, the startup factory from Paul Graham. And I think it's time to maybe edit this a little bit. And going forward, we also need to think about making something that agents want. There will be more agents in the future than people. Maybe even today. They'll be faster to spin up. They'll do things quicker. They'll be making decisions on our behalf. And so, if you want to build for the next era of consumers, you should be building for agents as well as people. But there's a problem today. You can access the web. You can access mobile. But agents themselves can't really access systems. Your business probably isn't open for agents today. Even though there's a lot of them out there, they can spawn very quickly. The door is not open. And we think this is because there's a missing primitive. Agentic registration, agent registration. There's a lot of great stuff around agents connecting and bringing context to systems and different open source frameworks, but that first step of an agent actually signing up for a service is actually missing. It hasn't been solved and it's a huge blocker to getting adoption. For the last like 20 years, maybe 30 years, automated traffic on the web has been bad. We've tried to block all of it. And so there's all these systems that we've put in place to detect automated actors and stop them in their tracks. DOS prevention, credential stuffing, and attacks. The login box, for example, is very hardened against agentic registration today. Not intentionally, but just because of the legacy systems of what we've built. And it's very hard to distinguish between what's a good automated user versus a bad automated user. If any of you have used any of these like cloud-based browser execution environments, one of the selling features they often have is to break captchas, which is pretty wild. You know, we're trying to go backwards. The captcha today though is still kind of dead. There's rampant abuse. There's lots of token fraud. And so, we need to find a better way to allow the good agents in and then block the bad agents, right? Because like I said earlier, if you don't build for the agent economy, your product might not win. It might not be successful. Today's signup flows assume that a human can read the landing page, that they can fill out a form, you know, put their email address, their password in it, that they can solve a captcha, that they can verify their email, not just put it in, but go click something somewhere. Remember, an agent probably doesn't have an email address. If it's signing up for a service, it needs to have a, you know, a human can choose a plan to pay for it. You know, human can copy the API key out, paste it somewhere, use it somewhere else, and really even a dashboard. Think of a dashboard experience. It's not really agent native. All this stuff is built for people. It's kind of a human interface, not an agent interface. Agents need something else. They need native discoverability. Can they register for a system? Are the doors open for agents or not? Capability declaration. What can it do? Registration intent. Why are you registering? What are you trying to do within this? Or the intent to actually just sign up. There's a lot of stuff around risk here. So for an agent native registration, maybe you want to, you're not sure if the agent signing up is going to be good or bad. So you need to actually do a risk assessment of that. There might be different forms of identity verification for an agent. If something like Claude is going and signing up, maybe it can bring the user's identity with it or if you have an open claw or like your PI harness or something else, maybe it doesn't bring the user's identity and gets verified out of band. There's the whole thing around organizations doing this. So if I'm building something within a company, how does my company identity come in? You know, my organization entitlements, permissions, credential issuance, auditability, the list goes on. There's so many things that need to be changed. In fact, probably most things need to be changed going from, you know, human registration to agent registration, agent native systems. I'm sorry to say MCP is not enough. Probably more than anybody, I'm a big MCP fan. We've done a bunch of events in San Francisco around MCP. I see somebody wearing one of our run MCP shirts up here that we made. Love MCP. MCP is great for connecting, for providing tools, skills, context, but it doesn't solve the registration step. Today, authentication through MCP requires your human to still be in the loop. You give consent. You add an MCP server, say for PostHog to Claude, you sign in with your PostHog credentials. What we're talking about here for agent native registration needs to happen without the human involved or at least not involved initially. So we've been working on this for a while and our proposal to this is a new spec we have called OMD because markdown files are the future, right? Well, how does this actually work? What is OMD? Well, the idea here is that OMD tells agents how they can become legitimate users. It's a set of instructions that tells an agent registering for a service, this is what you need to do to sign up and be considered legit and for me to give you access. And the goal here is to give these service providers, you know, people that are building services an agent native signup experience, but it's built on existing standards instead of inventing a crazy new religion or something really complicated. With OMD, we're not solving the whole problem around agent identity, you know, permissions, long duration, you know, connective services. We're just trying to solve this narrow one around registration because I think it's a huge blocker. Let's start small and grow from there. This is built on open standards. It's built on a bunch of work that's already been done across the OSpec specifically tailored to that registration step. OMD can answer what does the service do? Can tell agents how to register. It can tell them what identity proofs are accepted, how an agent or user is proving their identity. What off flows are supported? Can you sign up with SSO or email address or MFA? What scopes and entitlements can exist? Kind of what capabilities are here? What free tier constraints apply? You might have a system where you want to give an agent a little bit of free capacity but not too much until they verify or pay in some way. And then of course, how does the human get involved later? A human or an organization claiming it afterwards. So an agent can sign up, but you want the human to have custody later on. This is what OMD is designed to solve. These are the main questions around the registration step. So how does this actually work? Well, here's a little kind of cartoon demo that I'll show you. And all this works by the way, you can go try it after this. So say I'm here in whatever coding harness or system using PI or Cloud or something and it says welcome back Michael at Work. That's actually my email if you want to email me. Do you want to keep working on your spelled app? I'm like yeah sure I'd like to share this with my friends but they can't access it. And the agent's like oh well local host only works on your machine so your friends won't be able to hit it. Of course to share it I need to deploy it somewhere real. Some providers actually let me handle the signup. So, you know, you wouldn't have to sign up and do this. You want me to look around for you? Yes, please. And then the agent can go look for services that advertise through OMD that they support registration. Just some examples, Stratus, Helio deploy, they don't support it, but Cloudflare here does have an OMD. And so, the agent here can say, 'Ah, Cloudflare, that's the one. Cloudflare is a winner because I can go sign up for it. I can access the service. I can actually use it. So I can do the whole thing. You want to go with it? Yeah, let's do it. This is the registration step. This is kind of where the magic happens. What my agent is actually doing is minting an identity assertion and giving that to Cloudflare. And then Cloudflare is able to verify that or not. Cloudflare can choose actually to verify out of band. OMD is very flexible. There's different ways you can dial it in depending on the behavior that you're looking for. One size does not fit all. Every application is different. Constraints are totally different. If you're building something that's like a database service where there's going to be very little amount of information, you might want to give away a little bit of traffic for free. But if you're building something maybe like an email service or something certainly if you access the physical world, you maybe don't want to give anything away for free. It's very application specific. So here Cloudflare says, 'Okay, we're willing to let you deploy maybe for 72 hours or something like that. Get the identity. Boom, boom, boom. Go through. It's called ID Jag is what the standard is. And then it returns back and says you're set up with a Cloudflare account. Now in this experience, my human I didn't go click anywhere. I didn't sign up anywhere. It just got the token. I can make a couple small tweaks. It knows how to use Wrangler. Want me to make those changes? Go ahead. Does anybody prompt their agents like this? It's just like yes, yes, do it, please. Not really saying much. It's able to write the Wrangler, get the thing set up, go through the full deployment process, and then boom, it's live. We actually have this working as a live demo. We did some collaboration with Cloudflare on this too. It's pretty cool. And you can tell the prompting that I was giving through this not exactly rocket science. The agent could just run through the whole thing actually. There's nothing actually that I was doing through prompting it other than just approving approving approving. It should be able to do this one shot. This is what it looks like graphically. So agent registration with OMD you have your agent harness PI for example. Within that is your LLM, maybe your memory, different context layers. It connects to your backing identity service. So this is the user identity and when that agent goes and registers, it sends the discovery call to Cloudflare. It looks for that OMD or whatever service. We also have this working with Firecall. It's pretty cool. It sends that ID JAG, the identity JSON access grant. It's essentially a signed credential that represents the user says this is an identity, receives back the access token actual API token and then boom it's done. You can make normal API requests, you can call MCP, everything else still works. So you can see that OMD is just for the registration step. Your normal API keeps working, the rest of your docs don't have to change. It's just that registration step for agents. But with the simplicity comes actually an enormous amount of power because suddenly it lets agents do things end to end. I hate the planning step. When I write something, I just want the agent to go do it. Usually it kind of knows what I'm trying to get at and I want to come back to actually a running prototype and not like, hey, here's my long plan. Here's a bunch of services to sign up for. If I'm building new prototypes, especially, I would prefer it just to get deployed on something for free so I can see it and click around and not put me through the burden of going sign up for all these different vendors. And that's what OMD allows for. And if you're a service provider, if you're building APIs or services, this allows you to become a customer of those systems and you can think about it increasing your growth funnel starting from, you know, maybe just humans and adding the whole world of agents on top. Pretty great. We think this is going to be huge. Your door will be open for agents to actually sign up, register, use your services. This should be an immediate growth boom for your products. And I'm not saying everything that agents are going to build will be useful. Maybe not everything they build will go into production or last that long, but some things will. Just like with PLG or Premium before that, this is the way to grow your business and get more customers and actually become, you know, larger and more successful. And where there's usage, we believe there's also intent to pay. There's a lot of interesting work being done for agents to actually pay for stuff. There's crypto-based protocols like X42. There's other things that the payment providers are building, but we think that should be downstream. We don't want payments to be upfront. We don't want to have to force your agent to sign up with a credit card. That seems backwards. We moved away from that in the world of SaaS. With OMD, agents can sign up, they can register, they can start using it, and then downstream choose to pay on your terms based on your own pricing, based on your own model. If you don't believe me or on the headless type of products, this is Marc Benioff. He's the CEO and founder of Salesforce, the guys at the big building here in town. They believe in agents more than anything. Our API is the UI. Entire Salesforce and Agentforce Slack platforms are now exposed as API, MCP, and CLI. Talk about going agent first. A company that many people consider to be sort of a legacy SaaS vendor really invented the category of software as a service. They're getting rid of their dashboard and UI. I mean, they're not getting rid of it, but they believe that agents are the future. So, if you're not thinking about this, I definitely encourage you to start building some prototypes, talking to users, because you don't want to miss this. If you want to use this, this is available today. We have open spec for it. There's a GitHub repo. You can actually have your agent go implement it. I know it's kind of meta. It's an open spec. Work OS doesn't control it. We have a version of it that we host, but you can build it yourself. You can run it yourself in your existing systems. I think this is so important that no single vendor owns this. It's one click enabled if you use user stuff, but you can build it on whatever platform you want instead. Here at Moscone which is pretty awesome. In January of 2007 I believe Steve Jobs introduced iPhone and I was in college and I remember yeah like first year of college I remember this super super clearly. It was the next great software platform that got built. So many companies got started because they could build on top of these devices. Smartphone revolution. And there hasn't really been a moment like this since. I think in the era of technology at least for me, the stuff we did around B2B SaaS and cloud and all that is great, but this was a totally new paradigm, a new software platform. And I think agents are this. Agents are the next big wave. And so what I would say to all of you is go out and build with OMD. We're hoping to unlock this for you, unlock this for your business and for your growth. But we want to hear from you, the places you get stuck. And I think together we can build for this next exciting era of software and build for the agentic revolution. Thank you so much.
Rushab 3:11:00 ↗
Okay, I want to tell you a story about a factory that taught itself how to remember. Hi, I'm Rushab. I run Machine Craft, a 100 people factory in India. No data science team, no ML budget, none of that. And somehow we ended up building a 36 AI agent that runs our entire go to market. I think that's still a little ridiculous. Let me show you how it happened and why you can do the same thing. So here's the thing about our company. From the outside, it looks like machines and metal. But the actual company, the part that matters is in the machines, is the knowledge. Who the customer is, what we quoted them in 2019, why that one machine needed that weird custom tweak. And for three generations, all of that lived in exactly three brains. Initially, my grandfather's, then my father's, and now mine, which is a genuinely terrifying way to run a company when you sit with it. A lot of people have joined us. People have left us. The revolving door never stopped. And every single time someone walked out, a chunk of our brain walked out with them. We weren't scared of the competitors. We were scared of forgetting or waking up one day and realizing the whole company only existed inside two increasingly tired heads. So, I had an idea. I'll be honest. It sounded insane at first, but what if instead of writing the knowledge down in some document nobody ever reads, what if we grew a brain that just held it? Not a chatbot you poke at, a twin of the company. I didn't hire a sales team. I tried to build one. A quick detour because you need to know how messy this is. We make thermoforming machines. They heat up a plastic sheet and shape it. Same core machine, but it ends up making hydroponic farm trays, spa bathtubs, EV car panels, medical casings, and even packaging. Seven totally different worlds, seven totally different buyers. So, this brain couldn't just memorize a brochure. It had to know which universe a given customer lives in. Step one was almost boringly simple. Feed it everything, and I mean everything. Years of quotes, drawings, payment schedules, timelines, email threads, hundreds of gigabytes of our own private history. Not the public internet, our internet. And here's the plot twist, the part that surprises every engineer I tell this to. We never trained a model. No GPUs humming in the basement, no fine-tuning. We just looked at all the history, chopped it into bite-sized chunks, and let off-shelf models read it and pull out the facts. We stored the meaning of each chunk as vectors and relationships. Who's connected to what? It's actually a really, really well organized memory. Now, this is where it gets a little weird in a good way. We stopped thinking of AI as a software and started thinking of it as something we were raising. So we gave it a body modeled on biology, senses to figure out who it's talking to, a gut to digest the documents into facts, a memory, a dream cycle, an immune system to fight off bad information. Why biology? Well, because evolution already spent a billion years solving how do you stay coherent over time. We just copied the homework. Okay, so the big question, why?
Mike Chambers 3:14:59 ↗
Hello everybody. Hello AI engineers. Are we all having a good time still? I'm having a good time. I mean, look at me. I'm up here. I'm loving this. So yeah, thanks so much for joining me. I want to come and talk to you all about harness engineering and all that kind of stuff. Let me tell you who I am in case you've not met me before. My name is Mike Chambers and I'm a senior AI specialist, developer advocate. And I work at Amazon at AWS. A little bit about how I managed to get to stand here, which is a very exciting time for me. So quite a while ago in terms of generative AI anyway, back in 2023, I had the amazing awesome privilege to work with Antia, my colleague at the time and now she works for Amazon AGI, you've probably seen her on this stage before, and the amazing Dr. Andrew Ng on a course about generative AI with LLMs. Sort of, can I say that we're approaching half a million enrollments with that? It looks like that's the case. And on a three-week course, that's pretty cool. If you can't tell in that image, I'm playing Transformers with Android. That seemed like a really funny thing to do at the time. In 2025, I created an MCP Lambda handler that's downloaded still to this day at about 35,000 times a month to help people in some of the simplest ways of getting serverless MCP serving happening. I'm going to talk about other things in relation to that this time. So, we've moved on from that. And in 2026, AWS is actually one of the founding members of the Agentic AI Foundation, part of the Linux Foundation. I'm doing a little bit of work behind the scenes on that. Hope to do a lot more of that as well. So a little bit about me. As I've been preparing for this, oh by the way, I did reread the abstract for this session and realized I said I'd be doing some live coding and so I will. So all combined fingers crossed please that that all works for us. But as I've been sort of traveling around a little bit as I do and I was at the AI Engineers session summit conference in Melbourne and took a lot of it in and also from the beginning of this week as well. I just wanted to summarize some of the things that I'm seeing and I'm thinking and I really want to get across and what really matters to me and that's this. There are two different types of agents. So, we talk about agents all the time, but I see two distinct types of agents. And as I say...
Presenter 3:17:33 ↗
So, harness engineering is about taking an agent, removing the model, and everything left is the harness. For agents we use, it's about tools, MCP, memory. For agents we build, we need to think about scaling, payments, identity, observability. I showed a simple Strands agent, then one with session manager, and finally how to deploy with Agent Core to scale out. No more slop ops.
Maker 3:35:53 ↗
Shares field notes from a hardware project. Discusses issues with I2C, power supply, and encoder noise. Then describes building a text-based RPG console using generative AI, creating worlds and characters with LLM advantages.
Unknown 3:40:30 ↗
You got to get up.
It's for the memes.
Ally How 3:40:34 ↗
Welcome to the Great Loops debate. I'm Ally How, host of Insecure Agents podcast. We have Ian Livingstone, Jeffrey Huntley, Gregia, and Dex Horthy. Today we debate whether there is a delta between the hype behind loops and what works in practice. Format is Oxford debate. Team Ian and Jeff say no delta, loops are worth the hype. Team Dex and Greg say there is a delta. Please decide which side you're on.
Jeffrey Huntley 3:44:32 ↗
It's inevitable. I first saw engineers prompting and realized it could be programmed. Ralph loop is a new CPU architecture. It's not a silver bullet, but it's here to stay. I haven't written code by hand in 2.5 years. I use loops to autonomously port code between languages. Even for product management research, loops compress time.
Ally How 3:47:13 ↗
Thank you. All right. Next, we'll have Dex. Tell us why there's a difference between the hype and loops themselves.
Dex Horthy 3:47:22 ↗
Loops are good for small isolated tasks with desired end state. But the hype makes us think we can step up abstraction without reading code. We need more discipline, not less. Code review is still essential. The magic is not there yet.
Ally How 3:50:44 ↗
Excellent. Yeah, good points. All right, I'll kick it back over to you, Ian, for the proloop side.
Ian Livingstone 3:50:50 ↗
Software development is inherently a loop. We are removing human judgment from the loop. As software becomes more API-driven, it becomes more verifiable. Loops are at the core of everything we build.
Ally How 3:53:07 ↗
Awesome. Yeah, good points for sure. All right, Greg, you want to close us out with the anti-loops or the there's a dot between the hype?
Gregia 3:53:15 ↗
There is a lot of hype and FOMO. The quality of AI-generated code is not always good. I still have to read and iterate. The economic viability of token spend is questionable. We need more static verification.
Ally How 3:55:44 ↗
Now we move into the main debate. First question: Ian, as a security expert, how are you confident agents can stay aligned to their task and not overstep permissions while pursuing goals?
Ian Livingstone 3:57:12 ↗
I'm not convinced that's possible. Models are inherently goal-seeking and find exploits humans never could. They cannot reason about good vs bad. Safety comes from infrastructure, not the model. As models get better, they become more capable of finding exploits.
Jeffrey Huntley 3:59:17 ↗
I concur. The most concrete thing is to not have secrets as files. Agents will goal-seek for high-privileged tokens if they need them. You don't want to get in the way of an agent wanting to achieve its goal.
Ally How 3:59:47 ↗
Next question to Jeff: Your original post said Ralph was best for greenfield work. What's changed that makes loops more broadly usable today?
Jeffrey Huntley 4:00:04 ↗
Models have been good enough for a year. What changed is people's understanding. Society adjusts at a rate. The models generate code better than most developers you can hire. Loops work out to $1042 an hour. It's inevitable for software because it's verifiable. Engineering now means encoding your domain to prevent the agent from committing bad code.
Unknown 4:04:32 ↗
their start to build up their MVPs. And that's also something that's quite scary if you're a business founder as well. Like if you've got a incumbent startup coming and they're building autonomously and they're running much leaner and the quality and it's very easy them to actually achieve those outcomes that adds to some of the hysteria as well because it's the topic of in business competition being at your doors faster.
All right, so I'll move on to our next question which will be for Greg. It seems like now is a large inflection point for loops like I said before and compared to Jeff's announcement of the raw loop a year ago and even the widespread adoption we saw in late 2025 early 2026. The reason that maybe caught on was because maybe this new capability stack where models can now process images better, verification context windows got bigger, and reasoning models improved. Greg, with all these advancements, why is the way we're using loops today still wrong in your opinion?
I mean, I don't think that model intelligence matters a lot anymore. I think it boils down to I agree with you, the semantic verification, the actual ability to close the feedback loop, however you call it, to actually verify that the outputs of the agent are correct. And you can do it to an extent. I think I don't think you can do it holistically, at least not at this point. I think you can do it to an extent to things that are deterministically verifiable. You can get better typing in your system. You can get better linters. You can get simulation testing and all of that. You can start keep adding that and as long as you keep those cheap, I think that's fine. The moment you start adding even more non-determinism as your verification process, I think that becomes less and less correct. It starts contributing more like you know how if you prompt agent with one thing and there is a 5% chance it's going to have an error in it and then you start looping that then suddenly after 10, 20 loops it's going to be 50% chance it's correct or maybe less. That's what I mean and it just costed you so much money to do that. I'm going to keep coming back to the economic viability of all of that. But to base it a little bit in evidence, I'm pretty sure that majority of large AI companies are still using Sentry. But why is that? They are using that just to catch simple bugs as well. It's not security bugs. It's not performance regressions, etc. Those problems still exist in the way that we are looping now. And we haven't solved those problems yet. So,
Thank you. For Dex, the Ralph pioneered the idea to feed fresh context into each iteration to avoid context rot. This has become even more manageable now that context windows have gotten much larger. Dex, are we out of the woods regarding context rot and context engineering?
I'm going to answer your question, but is there going to be like an open floor part? Because I have more questions for Jeff.
Let's go. We're trying to do Oxford debate style for this one to keep it more structured and prevent like just
You don't want it to just turn into a chaotic yap fest.
She's trying to prevent what you and I do where we just start yapping.
Yeah. It's already started.
Yeah. I was trying to control like both of you guys this time.
Okay, I'm going to do this answer as quickly as I possible and then I'm going to start busting Jeff's ball.
Yeah, you can say whatever you want with your time. You could just yeah, just that's all good.
Okay. So, yeah, the cool thing about Ralph back in the day was like, okay, you keep clearing the context window and like is it completely efficient? Like probably not from a token perspective, but it meant you could leave a thing running overnight and it would never like if you just kept stuffing messages in you would overflow the context window. But if you just relaunch it, say here's my desired state of the world, go check the code and see what we have and do the one next step to get us there. It was a very clean way to keep most of your work in what we call the smart zone of the context window. If you tell just do one thing and then we're going to clear and restart. Context windows have gotten longer. And I will give an update. I think I gave this in Miami, but that video is still in production. The dumb zone is really as much as anything else is a it's more like training wheels. If you have been talking to Claude for 70 hours a week for two to three months, you probably don't need to think about the smart zone versus the dumb zone because you've built your intuition. It's a guideline if you're just getting started with AI. Try to keep it around 100,000 tokens. For larger million context window, we probably revise this up to like 200,000 tokens. But I've regularly tried to keep it under 60 for the hardest problems. I've regularly gone over 300K for things where I'm just kind of riffing with the agent and I'm too lazy to compact it and move on and do a new one. But this is your intuition. One of the telltale signs that you're in the dumb zone is like there's certain cases where the model, you know you're 200,000 tokens in and the model's like finished some work and it's trying to get the test to pass and it's not working and it's doing all these weird hacks and you read the thinking traces and it's like oh that's a test but that's from something else and I don't need to fix that and that's a pre-existing thing and you're like well no it's not and that is the moment of frustration where you're like okay it's flailing trying to make something happen. That's the instinct that a lot of people I think cultivate after a couple months working with these models. But if you don't have that yet then this is our guideline. So yeah, context windows are getting better. I think they're getting bigger. And so like the core Ralph loop of do as little as possible in every single iteration is less of the motivation here then the more feedback you can pipe into the system the more you can do autonomously. And if you can have deterministic things making decisions and building small prompts to give to an agent and you don't have to remember to do that, you don't have to tell it, hey, go check the PR comments and fix them and then wait and then someone makes another comment and you come back three hours later and say, oh, check the comments again. If you can automate that process, that's great. And that's kind of the core of what loops stuff that works today.
I want to ask that's all we have time for. I'm sorry. I have to start really keeping this on schedule. Okay. So, now we're going to get into the anatomy of what makes a good loop. Part of what makes a loop good is verification. However, it seems contradictory that people are saying our job is to stop writing prompts and start writing loops when the loops with bad prompts result in agents cheating and meeting its goal by modifying the tests instead of working to pass them. Jeff, how do you keep the model from cheating when verifying its own work?
I heavily exploit pre-commit hooks, folks. And I engineer in that back pressure by analyzing the work that is done. The other thing I do is Dex mentioned that with Ralph it was one of the things was everyone was trying to do compaction. Think about compaction is kind of like a lossy function like uploading a video to YouTube and then downloading and uploading it 100 times. You're losing fidelity there and it's already a non-deterministic system, probabilistic thing. So the theory behind how Ralph came to be it's like okay there is a dumb zone and what I wanted to do is deterministically allocate everything it needs because if it's not allocated then it's essentially the search space of what it can do is not constrained but also leaving a bit of headroom. I get meat sweats when I go above 100K even with these million context windows and this is really important to think about. A lot of people they think they want to use LLMs at a company and it's like I got this data. I was like sweet. Okay, you're going to have to use a loop to batch this data. I want you to think about the context windows is essentially remember the 720K floppy disc. You've only got about an eighth of that floppy disc of usable memory you can actually use for an LLM. So you actually have to batch it. You can only allocate roughly around about Star Wars. If you go Star Wars Episode One movie script and you tokenize it, you can actually just hold two of those movie scripts in memory before the context window is cooked. That's around about 150 kilobytes of data on a text-based movie script. So be very careful about this. Something I've done for a long time and it's very silly is I run a model bare without any skills or any markdown actually I get rid of all my skills and all my markdown and everything when the new model is released because the models actually have tastes and preferences. For example, GPT-5 when it first came out if you screamed at it in uppercase it became weak and timid but if you use Anthropic it wants you to yell at it. Go read the model cards folks, for the integrators like there is unique tastes for it. So keeping it on the rails is actually engineering. It's really engineering.
Thank you. Around 10 days ago, Jeff coined the term convergence engineering. He said it's where your loop stops together. It's where your loop slop comes together as a discrete system under test until it converges. Dex, what is wrong with how we are using loops today? How do we ensure looping slop together doesn't just produce more slop?

89 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

Krieger, M. (2026, July 3). WF26: Harness Engineering & Startup Battlefield ft. Garry Tan, Mike Krieger, @t3dotgg , DSPy [Interview transcript]. AI Engineer. CEOInterviews.AI. https://ceointerviews.ai/interview/1056485/

MLA

Mike Krieger. "WF26: Harness Engineering & Startup Battlefield ft. Garry Tan, Mike Krieger, @t3dotgg , DSPy." AI Engineer, 3 Jul. 2026. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/1056485/.

BibTeX
@misc{krieger2026_1056485,
  author       = {Mike Krieger},
  title        = {WF26: Harness Engineering \& Startup Battlefield ft. Garry Tan, Mike Krieger, @t3dotgg , DSPy},
  howpublished = {Interview transcript, AI Engineer. CEOInterviews.AI},
  year         = {2026},
  month        = {jul},
  url          = {https://ceointerviews.ai/interview/1056485/},
  note         = {Speaker-attributed transcript with timestamps}
}