About Raj Yavatkar
Raj Yavatkar, Chief Technology Officer at Juniper Networks, has been discussing the company's innovation strategy and technology focus areas in several interviews and events. He described Juniper's "Beyond Labs" initiative, launched in 2024, as an effort to organize innovation activities around pioneering research and experimental technology development, with four pillars: artificial intelligence/machine learning, sustainability, quantum networking, and 5G. Yavatkar stated that the company's thesis for AI/ML is that "self-driving networks will go to the point where they're aware of end users and end applications," which he called "application-aware assurance." He noted that Juniper has been evolving machine learning models to collect data from across the network and applications, and cited a Fortune 1 customer that publicly stated 90% of their troubleshooting tickets are self-resolved by AI.
Yavatkar also addressed developments in Open RAN, private 5G, and sustainability. He said Juniper completed a commercial field trial with Vodafone using Open RAN that met or exceeded traditional KPIs, and that the company is augmenting its cloud-managed AIOps-driven Wi-Fi solutions with private 5G to allow both technologies to coexist under the same management model. On sustainability, Yavatkar said the networking industry lacks benchmarks, so Juniper is defining a "green quant" to measure efficiency and is looking at the picojoules cost per bit as a baseline. He added that Juniper's custom chips have reduced power consumption by about 30% from one generation to another. Regarding quantum networking, Yavatkar stated that Juniper has a product for IP VPN tunnels using quantum key distribution that is shipping, and that the company is working with the UK's Department of Defence to trial the technology.
Source: AI-verified profile updated from Raj Yavatkar's recent appearances.
Browse all interviews →
Transcript (31 segments)
J
John Furrier0:05
Welcome back everyone to the Cube Studio here in Palo Alto. I'm John Furrier, host of the Cube. I'm here with Dave Vellante and a great lineup of leaders. This is the Silicon Valley AI Infrastructure Leaders Program hosted by the Cube and the NYSE Wired Community. Raj Yavatkar is here, he is the CTO of Juniper Networks. He is a leader, I've interviewed him many times, and Juniper Networks is obviously setting the agenda for gen AI and certainly AI native networking, which is part of now all the integrated systems. Raj, great to see you and have you on this inaugural program.
R
Raj Yavatkar0:40
Thank you. Thank you, John. Thanks for having me.
J
John Furrier0:42
It's a really cool community vibe. We've got a lot of leaders coming together, sharing what's going on in the industry, not about just the companies they're involved in, but what they're doing and what they're seeing, sharing their perspective, inspiring, and there's a lot of problems to solve. So if you're an engineer, it's a great time to be in this market. If you're an entrepreneur, there's tons of opportunities, Raj. And this is a special moment in Silicon Valley and the industry because there is so much going on and it's in the fun part of the network. It's the network, it's the infrastructure, it's the chips. And all the action is in areas that felt, I won't say dormant, I mean go back 10 years, the VC market wasn't that robust with semiconductors. Now, it's all the rage, but it's not just chips, it is what's around it, and you guys have been leading the charge. So what's your perspective as you look at the market? What's the exciting part? What do you see?
R
Raj Yavatkar1:34
So I think if you're a networking geek, it is an exciting place to be because everybody is building huge GPU clusters. All of those clusters need to be fed data from storage networks. Then once you have data in, you need to have data transferred. So networking is needed no matter what. Very high performance, a high throughput, low latency kind of network. So that's being built. More exciting part is that a lot of the infrastructure is not just being built in hyperscalers, it's also being built by enterprises on-prem, which is a big shift in the market. From what we are seeing, everything shifting to cloud kind of a market. And that is for a variety of reasons, cost, data privacy, data sovereignty. In terms of cost, people have shown, the ACG research recently showed that cost of doing machine learning training on-prem is about half of cost in the cloud. So there are good reasons to build this infrastructure on-prem. And as a networking company we were fortunate to have been thinking about AI for the last four or five years. So we built what we call AI native networking platform. Our infrastructure, we're building 800 gig switches and all, are complimented by our intent-based automation, we call it Apstra-based automation to support congestion management, flow control and so on, so that we can automate the operations of the network. Either it's front-end network or backend network, it does not matter.
J
John Furrier3:04
The industry has seen this next generation cloud, you kind of pointed it out. When the SaaS wave hit, well, first of all, the .com happened, iPhone happened, apps came out. We saw that right out of the gate. When Facebook came out, people were throwing sheep, its application, FarmVille. And then it settled in, the apps got better and more hardened. Gen AI has got that same kind of feel. The apps, they're not as silly as some of the social networks or some of the gaming apps on the iPhone, but they're kind of experimental. They're fast to stand up. Some are more durable than others, but you start to see more and more use cases of new kinds of applications. And so what's happening is we're seeing a wave one of apps that can sit on the network. They don't need a lot of power because they don't really have a large scale, but now you're starting to see scale become a factor. That's going to separate the wheat from the chaff, so to speak. And that's kind of what you and I were talking about last time, is that enterprises aren't just some server handling a thousand users, they've got a lot going on. Scale matters. Can you share your thoughts on why the infrastructure really needs to be ready for the scale?
R
Raj Yavatkar4:16
I think scale matters because people are building GPU cluster sizes of anywhere from 64 GPUs to 1,000, 10,000 GPUs. When you build that kind of a scale, you need to have networking that scales too. So for one of the first things we did, we were the first ones to introduce 800 gig switches. You need that kind of capacity. You need to build leaf-spine networks, super-spine networks, and for that sort of scale, you need to keep in mind that the clusters of the scale, the tail latency, for example, cannot exceed at the mean latency by more than a factor of two or something like that. Which means now you are providing a large scale networking not just for high throughput, but also congestion-free. So that's why providing non-blocking congestion management, load balancing flow control is important. And that's what we're doing with it.
J
John Furrier5:07
Raj, I want you to share with the audience who are watching, and we talked a little bit about this, that you just had an event on the next-gen AIs coming, networking. Storage networking and compute, and now you have GPUs, XP, and all kinds of devices to process. The roles are changing because of generative AI, can you share your vision and specifics on how you see storage, networking, and compute change based upon the conditions and the situations that generative AI will have? For example, training might require more reads than writes or write more writes than reads. Inference does things differently. You're seeing characteristics that sound a lot like policy-based stuff, but it's not. It's real-time generative. This is kind of putting pressure on what those table stakes will be in the enterprise, the key building blocks.
R
Raj Yavatkar5:58
That's an excellent point. Multiple factors here. First is, as you said, training workloads, the traffic characterization is different. Then the inferencing workload, which is more latency sensitive, so it's not so much as traffic, but the latency matters. Then traditionally, the computer storage networking network as separate networks. Storage used the InfiniBand as networking technology. Computer used general purpose ethernet. What's happening with GPU based clusters, these things are coming together. And the standard ethernet, we are being able to show it's not just good enough, it performs as well or better than InfiniBand-based networks. And that is being recognized by industry now by creating this ultra consortium of all to evolve ethernet standard to add some new capabilities to match the training workloads and inferencing workloads.
J
John Furrier6:52
I put out a research report, and this came with a lot of the work that you guys are doing, as well as Broadcom and some other leaders, the shift from InfiniBand to ethernet and networking, parity of performance, enterprise adoption, scalability and flexibility and cloud integration. These were the drivers, but the challenges in the data center is, how do you feed all these accelerators? What's the data security posture, data management at scale, and then the pitfalls that people fall into just on misconfigurating things? These are big challenges and if you can crack the code on them, there's huge opportunities. What's your reaction to that?
R
Raj Yavatkar7:28
So I think you said it well, right? Misconfiguration is a big problem, especially as you build scalable clusters. And also, understanding not just configuration but the operational on it. Like you said earlier, you need to change the characterization or characteristic of network based on workloads. So what we have done is, we introduced what you call auto-tuning application. What the application does is, it's like applying artificial intelligence to the operation of the AI network itself. So you constantly monitor the network, observe what the configurations are. For example, in a switch, if you have set the packet buffer thresholds for marking, say congestion notifications and you set it too high and you wait too long, then the intelligence will, by observing this across the entire fabric, we change the configurations on the fly. We change the thresholds, parameters on the fly. That's what we mean by auto-tuning application. So that's an example of applying AI for the operation of AI networking and that's what we mean by...
J
John Furrier8:34
Auto-tuning, as you call it, is that human involved? Is that done by the software?
R
Raj Yavatkar8:40
No, completely automated using intent-based automation that we have in our Apstra Networks product. But there's enough data you keep on collecting at the time, series data, so you can apply it continuously.
J
John Furrier8:52
Yeah. This is what I love about this marker right now, because you go back to old school networking, obviously you've been there, done that, you've built the companies there. Policy-based stuff was super important in networking because of the impact of wrong routes, or wrong path, costs. I mean, it was simpler then compared to what it is now, but it still matters. And now you take that at scale, it's not policy, it's intelligence. So how does networking become more intelligent and then how does that intelligence work with the other subsystems, the other key building blocks like compute and storage and even high bandwidth memory? Because you have a lot going on around this new system architecture.
R
Raj Yavatkar9:33
That's right. So I think one good news is that you have a lot of telemetry available out of GPUs, out of compute, out of storage, and networking also. If you start taking all of the telemetry and start building models, machine learning models that know how to correlate across these different components, so we call it application-aware assurance. You start collecting telemetry from application, compute, GPU, networking, operating systems, and you start correlating using machine learning model. If you do that, now you can start finding out where the problem is because you can find the anomalies, you find the switch buffers are running out of space, you find the packet loss has increased, latency has spiked. Based on that, you can point out where the problem is and then you can do automated root cause diagnostics. Once you do that, the next step is automated remediation to the extent possible or to the extent network operators are comfortable somebody doing it. Because many times you tell them, 'This is the problem,' and they can fix it or we can fix it ourselves.
J
John Furrier10:35
Root cause analysis. I mean, a lot of companies are talking about that in the news lately, and this is disruption. This is bad disruption, it's not good disruption. Disruptive in a good way is like, change the game for the better, but having a network that's down means you're offline, right? You've got to have that ability. I also want to get your thoughts on the enterprise architecture. We have two use cases we've been pushing out in the Cube in terms of end user Uber, which we've featured on the Cube, their engineering team has come in and talked with us, shared their environment. They built that from scratch because they had to and they had to deal with the first ever data lake that was built and their system is awesome. Then you've got a company we just interviewed here, TransUnion, their legacy company, which rewrote everything. And the two things in common is, one was kind of built from scratch, from the ground up, so it was engineered with all the engineers. The other one was a legacy company that was re-engineered by engineers. So again, keyword engineers. They have end-to-end workloads, and the benefit that they're seeing now with generative AI across the board is that they do have the end-to-end workloads. What is your thoughts on this? Because we're seeing this as a major trend. If you have end-to-end visibility into the workloads, you can then have better execution up and down the stack. What's your feeling on this?
R
Raj Yavatkar11:52
So I like to use the term mean time to innocence. See, anytime the workloads are not doing well in any enterprise network, they blame the network. Network is slow, network is not working, right? So now what you talked about is end visibility across all these layers in the stack, starting from application, compute, GPU, networking, operating system. Now you can start pointing out where the real problem is. So first thing from networking perspective, we want to be able to point out whether network is a problem or somewhere is the problem. Somewhere else is the problem, then we want to be able to also show where the problem is. Is it in the virtualization layer? Is it in application layer? I think that's powerful. Company like Uber, you mentioned they build their own distributed system infrastructure from scratch. It's easier for them to do because they do health checks, they do monitoring. But I think the traditional companies have to re-engineer or transform their way of thinking about it. They cannot operate in silos of networking, storage, compute, GPU. Now you have to bring them together and that's what I mean by this end-to-end application-aware assurance. You want to monitor, set service-level expectations, measure your performance continuously against those expectations. The moment performance deviates from those expectations, you immediately flag it, point out where the problem is and also provide suggestions for fixing those.
J
John Furrier13:17
Raj, that's great insights, like a masterclass. I want to get your thoughts on some of the key things in the intelligence application layer. Obviously a lot of stuff is going on in the plumbing. I call it plumbing, but a lot of infrastructure and then you've got the data piece. If you look at the companies right now, most of them are really looking at their strategic intent. 'Who do I work with?' They're making a generational decision about who they work with, technologies that they choose. There's a lot of foundational work going on at the infrastructure.
R
Raj Yavatkar13:46
That's right.
J
John Furrier13:47
What is your advice? Because there's a lot of decisions being made. In some cases, we just did a survey and shipped it two days ago, and generative AI on our super cloud seven that showed that in the Snowflake kind of data area, 96% of the decisions being made are being made by platform engineers.
R
Raj Yavatkar14:03
Engineers. That's right.
J
John Furrier14:05
Not data scientists or business workload, line of business. So you're seeing a shift to platformization. People have to make a generational bet. What's your advice to the industry on how to do that? What are some of the criteria that you would go through?
R
Raj Yavatkar14:18
So the most important bet to make is in an open ecosystem. It's very easy to find a short path and get a proprietary vertically integrated system where everything comes from one vendor. It makes it easier, it's a single neck to choke. But that's the mistake, because things are evolving fast. We don't have single large language model, multiple LLMs are coming. Open source models like LLaMA2, LLaMA3 are as good as core systems. So you need to look at all of the stack and how you can get the best of breed components from any vendor by investing in the open ecosystem based on the open standards. So networking part of view, that's why we push ethernet. Ethernet has outlasted any other technology in the last 30 years, 40 years.
J
John Furrier15:09
Ethernet wins.
R
Raj Yavatkar15:10
And it will continue to evolve. Same thing, you go next layer. If you go to the PyTorch-like frameworks, you want to use those frameworks which are open so that you are not locked into a single vendor's ecosystem. That's the most important advice I'll give.
J
John Furrier15:24
Okay, Raj, final question. This is a fun one for us. You and I are dorm roommates in college and we're 23 years old, we're going to graduate, but we're super smart. We know what we know now, what are we building? We're a startup, what opportunity would we look at and go after and start a company? Because there's so much opportunity recognition going on right now and the capture equation shifted. Obviously, if you're in your 20s, you had a little more free time, not a lot of responsibility, what would we go after? I mean, what would be a sweet spot? Where's the big white space?
R
Raj Yavatkar15:57
I think the big white space is outside the infrastructure. It's in applying generative AI for vertical use cases like HR, legal, finance, go-to-market, accounting. All of these functions have lots of workflows which are manually driven or they're outsourced to low-cost countries like India. Those can be now performed by generative AI. So you have white space there to build these vertical use cases as applications, SaaS applications, or on-prem applications that will increase the productivity of these functions by 20, 30%. That's a big opportunity and that's the next wave.
J
John Furrier16:39
Raj, I'll get the AI to write our PowerPoint demo, actually code the prototype, and we'll... I know some VCs that will fund us, I think.
R
Raj Yavatkar16:46
I hope.
J
John Furrier16:48
Raj, thank you so much, and we'll see you soon. And thank you for spending the time out of your busy schedule to be part of this inaugural Silicon Valley AI Infrastructure Leaders Program. Great group of great people, experts sharing their advice and opinion, of course, talking about what they're working on. And thank you so much for taking the time.
R
Raj Yavatkar17:07
No, thanks for the opportunity. It's...
J
John Furrier17:09
Okay. Raj, CTO at Juniper Networks. They're doing some killer work there on some next-gen networking, AI for networking, networking for AI. The network has always been that last area that we've been waiting to innovate. It's happening and continuing to evolve and being a key part of the new chip design systems, clustered systems, compute, networking, storage, all magically working together with all kinds of stuff around it. Juniper Networks leading the way. I'm John Furrier, here in the Cube. Silicon Valley's AI infrastructure leaders are here, we'll be back with more after this.