Models, runtime, and silicon, all working together to create a meaningful experience. We're working with the GitHub team to bring our vision for hybrid intelligence to life on Windows. Earlier this year, GitHub launched Hydra Fusion, which routes each task to the right model in the cloud. Today, we're extending intelligence to Windows, letting Hydra Fusion tap into models running locally.
Now let's talk about these models. At Build, we introduced MAI-1 and Flash, a 130 billion parameter coding model built for real-world developer workloads. Now we're bringing that same intelligence to your device. This is where full-stack engineering comes in. We have taken that model to three bits of precision, reducing its size by nearly 80% while preserving coding quality. That is an incredible feat of engineering. But really, shrinking the model to the device is only half the story. MAI Code can work at an impressive 256K context window right on your device.
Now, our goal is simple: let AI run locally, stay responsive, while leaving room for everything else on your PC. And there are many more powerful models coming to Windows. NVIDIA is working on a new Neotron model. At over 70 billion parameters, it manages to quantize to 2-bit and uses just over 20 gigs of memory, making it ideal to run on RTX Spark.
And why stop there? With advanced quantization methods, it's possible to bring more of the frontier to your PC. At 1.66 bits, DeepSeek Flash V4, a 284 billion parameter model, is running locally on device, fitting in just — that's right, it's quite amazing — fitting in just 60 gigs of memory. This model outperforms GPT-5 in coding and reasoning benchmarks. A year ago, such capabilities were exclusive to frontier cloud models. Now you can have these capabilities running on your PC.
And this is just the beginning. To power this growing ecosystem of highly capable models, you need a performant runtime. And this is where Windows ML comes in. Windows ML is a scalable runtime to deploy models across GPUs, NPUs, and CPUs. And we're continuously evolving it. We're excited to announce we're bringing Llama.cpp to Windows ML.
This gives developers more choice: open-source models on day one, the ability to experiment quickly with emerging models as they land. With intelligent routing, highly capable models, and a performant runtime, a clear recipe emerges for what hybrid intelligence looks like. That's the repeatable pattern we'll apply to light up experiences across coding, security, creation, and work. Kayla is here from the GitHub team to show you this in action. Come on over, Kayla.
Thank you, Pavan. I'm excited to show you how I'm going to use the best of cloud and local models to do my coding workflows in GitHub Copilot on Windows. So, to start, let's just paste in a simple prompt. And I have the Windows Terminal codebase cloned onto my machine. So what I'm asking is I want our agent to detect unused code, imports, packages, and obsolete feature flags, and then give me an actionable summary. So this is just basic code cleanup that I'd like our agent to do. And then here at the bottom, you might notice in the model picker, I have Auto Efficiency enabled. So this is going to select my model depending on the complexity of the task. And now with our new version of Auto, it will automatically pick a local model depending on if that is a more simple task. So here we could see on the left-hand side our GPU has ticked up, and that's because we're leveraging our local model here. And I did run this same prompt before on a different codebase, and we could see that it leveraged MAI Code 1.1 Flash Local. Now, since my code is cloned to my desktop and my model is local, we could just put our whole machine into airplane mode, and this process will still continue to turn because everything is right here on the device. So this is just a simple coding task right to the local model. But what if we asked something a little bit more complex? So in this session, I asked our agent to look at the issues in the Windows Terminal repo that have the needs-repro labels. So these are issues that our team hasn't been able to reproduce. And then I wanted to do some investigation and then hand off to some local sub-agents to do further building and testing and see if we can get more information here. So this prompt actually ended up leveraging GPT-6 Luna. So this is a cloud session here. And then we've got three local sub-sessions happening on the left-hand side leveraging MAI Code 1.1.1 Flash Local. So this is a really great example of hybrid intelligence. If I hop into one of these local sessions, we can also see all of the tokens that I got for free. And on this one in particular, we've got 1.66 million being sent to the model and then 10 and a half thousand coming back from the model. And all of this was free, and this was just one of the sub-sessions that I ran. So I have a lot of sessions going. I like to take a look at what I've got running on my machine and what I've run before. So we have a canvas that we created that is a really great visualizer of my sessions. So a canvas is a custom UI that works inside the Copilot app where both you and your agent can work back and forth and update the UI accordingly. So here in our visualizer, all of my purple sessions are local sessions that I've run with our little leaf, and all of my blue sessions with the cloud are cloud sessions that I've run. And you might notice I've got the same session happening a few times. I'm doing issue triage a lot, and I'm doing it locally. And that's because issue triage is one of the tasks that takes a while, and it's something that we have to do pretty frequently. So I've actually set this up as an automation. So automations can run on a cadence or on a trigger. And I've set my issue triage one to leverage the local model explicitly. So I know for a fact that it's going to use just the model on my machine, and it's not going to spend any credits without me knowing. So I can have this running overnight, and I'm still getting all of those tokens for free even while I'm asleep. So these are just a few scenarios that I've showed today. Imagine how much more you can save doing these workflows always on, even when you step away. Thank you so much. I'm going to hand it back to Pavan.
Thank you, Kayla. That was awesome. Hybrid intelligence Windows powered by Hydra Fusion will be coming to the GitHub Copilot app, CLI, and VS Code starting October 15th. You've seen how Windows is the most secure platform for hybrid intelligence. Now, to make AI a part of everyday life, we need to bring it to places people use the most, into the flow of human ambition. We're moving toward a future where everyone can simply express what they want, and their Windows PC, with their permission, can take action on their behalf. For hundreds of millions of people each day, search on the taskbar is where intent begins. So, it's only natural our story starts there. At its core, the new Windows search experience is quicker, more efficient, so you can find your apps, settings, and files faster.
Starting this fall, on Windows 11 PCs, you can take thousands of quick actions directly from search on the taskbar. You can switch to dark mode. You can turn on Do Not Disturb to help you stay focused. You can turn up your mic volume and then use the inline slider to get to the exact level you want. All of this as fast as you can type. And actions can do more than just settings. You can take a screenshot directly, arrange your windows, all from the taskbar. And my personal favorite, you can send a text message to a friend or colleague just letting them know you're running a little late without having to pull your phone out of your pocket. A quick, intuitive way to move from intent to action.
And of course, right from the taskbar, you'll be able to connect to the new Copilot app when your request calls for it. Ask a quick question, and Copilot will bring the answer right there to you, in line in search. If you want to have a richer conversation, you can jump directly from search to the full Copilot app. To tell you more about the Copilot experience on Windows, please welcome Jacob.
Thanks, Pavan. Recently we introduced the new Copilot with three core experiences: Home, Code, and Autopilot. Today I'm really excited to talk to you about how Windows hybrid intelligence supercharges everything that we've just brought to market. First, let's turn it on. Copilot now has access to three new things. First, local context. Local context means that Copilot can now ground over all of the files on your PC the exact same way that Work IQ helps businesses ground over all of their business context in the cloud. Second, local actions. Local actions means that Copilot can actually take action directly on my machine. It can move files. It can change a setting. It has the same level of the ability to control and change things that I do as a user, but always with permission. Third, local models. Copilot will still use the cloud for the hardest tasks, but for times when cost or privacy matter more, it can delegate down to local models that run directly on your computer. So, let's take a look at how this transforms the rest of Copilot. Let's first talk about Home. Home is the new starting point for Copilot. We started by bringing together chat and coworker. We also introduced Office in Copilot, which brings full, rich document editing and collaboration from Office directly into the Copilot app. And today, Home gets even better with hybrid intelligence. Home now knows what's on my PC. And so the exact same local files that I've been working on locally myself can now be turned into collaboration-ready artifacts directly inside of Copilot.
Next, let's talk about Code. Code is the easiest way for anyone to build and share an app. No developer experience required. Hybrid intelligence makes Code better in three key ways. First, you can now build a native Windows app from a single prompt. This means unlimited apps and widgets, all native on Windows, directly from the Copilot app. Building software for Windows has never been this easy. Next, we introduced MXC support. MXC support means that your code, even locally, is running in the sandbox that we heard about earlier. This means for the first time, not just developers, but anyone can write and run code on their local machine with confidence. And lastly, local models. Copilot is now able to hand work down to a local model as it's coding. This means that projects move faster and cost less, just like we saw in GitHub Copilot. These three things turn your PC into your personal software factory. We really can't wait to see what you build.
Lastly, I want to talk about Autopilot. Autopilot, enterprise-grade, long-running autonomous agent. Let's see how hybrid intelligence makes Autopilot even better. We'll flip over to a quick demo. So, over the last couple of weeks, maybe a little bit late, I was working with Autopilot to pull together a bunch of stuff to file my taxes. And so, I'll go ahead and just kick off this simple prompt. And the first thing that Autopilot will do is it'll go into my inbox and it'll find the email from my accountant and it'll look for the specific list of files that we have to pull together. Now, because of hybrid intelligence, Autopilot can now ground against my entire PC. And the documents I need to file for my taxes are all over the place. They're in Downloads. They're in Documents. Half are organized, half aren't. That doesn't matter. Autopilot's able to use hybrid intelligence to ground over absolutely everything with local context. With my permission, it starts to organize. It moves everything into a new folder, pulls things together, and actually starts to rename files. Now, all of this organization and renaming is actually happening with local models. My tax details are not going to the cloud. It creates some subfolders here to give a little bit more structure, and then it finishes by zipping everything that it found into a single zip file that I can easily send to my accountant. Now, Autopilot of course always goes one step further, and it also actually goes ahead and composes an email for me to send directly back to my accountant with the zip file that I just created attached. But of course, it pauses here and it makes sure that I'm good to go. I review everything. This looks great. We can go ahead and send the email. And just like that, all of these files are now sent off and in my accountant's hands, and we're good to go.
So that's Windows hybrid. Copilot now truly understands your PC. It's able to take action on your computer. It's able to keep sensitive work on your device. And at every single step, you're in full control. Really, it's Copilot built natively for Windows. Copilot features that are powered by hybrid intelligence are coming to Copilot over the next couple of months. And with that, I'll hand things back to Pavan.