Back
Sjur Jensen
Senior Vice President, Head of Business Area Markets, Vattenfall AB

OpenInfra Days Nordics 2019 | Keynote | Christoph Buchli & Sjur Jensen

🎥 Oct 16, 2019 📺 Cleura ⏱ 28m
Vattenfall is building sustainable and efficient compute infrastructure leveraging renewables, heat reuse and demand response.
Watch on YouTube
Transcript (7 segments)
C
Christopher Buckley0:03
All right, hi everyone, welcome to our presentation about the real world sustainable compute. My name is Christopher Buckley. I am from Zurich, Switzerland, and I would like to introduce myself with a little story about when I was a little kid, maybe 12, 13 years of age. At that time, an uncle of mine used to work at the water power plant, and he told me something very peculiar. He told me that sometimes they would use electricity to pump water up the hill into a lake, and later on they would let that water flow back down into the valley and generate electricity again. And since I was that little kid, pretty naive, I was speechless. I thought, well, these grown-ups are playing water fountain with our electricity. Is it really a reasonable thing to do? But I wasn't only a naive kid, maybe a stupid kid, but I was also very curious, and that's why I set out to investigate this matter. I interviewed a few people, I even did a school project about that. If I remember correctly, the title of the project was 'Wells: Is it reasonable?' And if you ask to what degree of efficiency is it reasonable to do these pumped storage hydropower plants, it after some investigation turned out it's actually quite reasonable and a very sustainable thing to do, because it's a very efficient and green means of storing electricity. Years later, I joined the compute industry. I did the apprenticeship, I went to school, started, and then I joined the compute industry. And in the compute industry, I learned, like that little kid when I first heard about actually people pumping water up the hill, I learned a few interesting things about sustainability of computing. One thing that I learned is a number: it's the number 14. 14% of the computing infrastructure of the services and everything that we have today are actually being used, which means basically only one in seven servers we actually need on average. Another number that I learned is the number 40. 40% of the electricity that the server uses is used just to keep it idling. So when a server is just up and running but doesn't do any work at all, it already uses 40% of the electricity that it uses. Also I learned that a lot of data centers they spend really a lot of money for cooling and for redundancy and for emergency prevention measures such as double lines for electricity and everything, although I mean the electricity grid is so stable, why do we still do that today? And just like that little kid I was 15 years ago, I am also, I'm still very curious, and I set out to investigate also this matter. And today I am here to tell you that we found a lot of very promising and inspiring things that we can actually do about sustainability of data centers. And to help me tell you this story, I have joined forces with a lot of very inspiring and visionary minds. And one of these visionary minds is the vice president of Vattenfall, a vice president at Vattenfall who I wish I had known when I was that little kid, because for the last 25 years he has worked to run the biggest hydropower infrastructure in Europe. Please welcome Sjur Jensen from Vattenfall.
S
Sjur Jensen3:57
Thank you, Christopher, a very nice introduction. Very nice being here, and I'm glad also that we had sustainability already on the stage, which is great. At Vattenfall, our purpose is to make energy fossil-free within one generation. And actually we don't stop at just making our production fossil-free, which is what we are doing already. We also want to make our customers fossil-free. And the question: how do we do that? Well, one large example that you might have heard of is a product called HYBRIT. This was actually presented at the UN Climate Summit in New York just two weeks ago. This is an initiative where we want to replace coal with green hydrogen from electricity in the steelmaking process. This is actually the biggest emitter of CO2 in Sweden next to transport. So by actually producing hydrogen, we can replace coal in the steelmaking process, and we team up with the big steel maker called SSAB and the miner LKAB in this effort. But there are more sectors which emit CO2, of course. Data centres is why we're here today, and I'll get back to that. But there are also examples like concrete, cement making, or chemical industry, or even brewing. Like the place where today this used to be a brewery, and of course when they boiled beer in the old times it was with coal or natural gas. Today they use electricity. Looking at compute, Christopher showed us some nice curves. Your consumption, you are consuming a hell of a lot of energy, and this is just growing. A good start is that you use electricity. Electricity can actually be produced by renewable means like wind, PV, and hydro. So of course that's what we want to do. We want to try it with green energy. But it doesn't stop there, because the number 14 stuck really to me: one seventh of the compute is used, and six sevenths are just standing there. This is a waste of resources. And it doesn't stop at the computer side. Also another number which I find very interesting is the five nines. This is what data centers ask for in availability requirements on power supply: that's five minutes downtime in one year. Yet only one seventh of the time the compute is used. This doesn't match. I mean 99.999% availability and one seventh of the computers used, this cannot be the same way. So that's one sort of equation to get back to. Also one thing is that data centers are many today. In big data centers maybe up north in Sweden or some in Iceland, but there's also computing needed in cities with growing use of compute in self-driving cars and other edge applications. We will need data centers in the cities, which is today very difficult because the electricity grids are actually very congested. In Stockholm it's hard even to connect a toaster these days. It's all congested, the grid. So if you want to establish a data center in Stockholm, you might have to wait for years because the grid is not strong enough. And why is that? Well, things are changing: different use of electricity, different sources of supply, wind and PV instead of thermal production. So what do we do about that? Actually, if you look at the grid, it's only a few hours per year that is really constrained, maybe a hundred, maximum 200 hours when there's really a constraint. But when you construct the grid, you need to take into account the highest use case because electricity is used for hospitals and used for very sensitive topics. But of course with data centers, if you could find a way where you could actually allow us to reduce the consumption a hundred hours or two hundred per year, we could allow you to connect in the city. So this is time to market for you guys. Another thing is that when you consume energy, it all becomes heat. And what do we do with the heat today? You just emit it to the air. That's a waste of energy. Even if the input is green, we should still use the heat. If you get data centers into cities, there is heat demand. I would actually bet that this building is heated by district heating, and maybe some of that heat is actually coming from a data center. But we can use more of that, especially with the growth. That's our product. So we have some challenges both in data and in energy. I will say there's three main challenges that we face which are similar. It's you have to generate energy. I think in its what actually you don't generate energy, you transform energy. So you have maybe water which is high up, you let it down and it's transformed from gravitation to electricity. Electricity is great to transport, very fast if you have the grid. Yeah, you know, there's a great issue. Back to the example you talked about, it's not only storing energy, it's also transformation of energy. And storage of energy is one of the biggest challenges, especially with wind and with PV. You all know that the wind cannot be steered. So if you have a large share of renewable production, there's a large share you can't actually steer, which means you need to store energy. Storing energy, especially electricity, is very difficult. Batteries are expensive and have a high climate impact. Storing water is cheaper and better, but we will have a shortage of storage. Data is quite cheap to store, isn't it? It's a lot cheaper to store data than to transform data, which you normally call compute. For the move back to this one over seven: why is the compute not used? So we would like to see a cooperation between energy and the data industry to find a middle way between the 99.999% and the 14%. And if there would be a way to queue up compute, if there would be a way to safely turn down compute when there's less electricity in the grid, that would be beautiful. Then we could have data centers in cities, we could use the heat to heat buildings like this. And I think the same technology can be used to also increase these 14% to maybe 70, 80 percent if you don't force it. So Christopher, what do you think? Would that be possible?
C
Christopher Buckley10:53
Yeah, I think that will be possible. And you see, Vattenfall is a perfect partner and inspiration for us to actually tackle these challenges. And for that, we identified three steps that we are going to take. The first step, we already heard that because it's not the first step because it necessarily is, but because it's the easiest: the first step is to reuse heat. Today there is a lot of initiatives around reusing the heat of data centers, but the heat of data centers is actually quite difficult to use because it's 35 degrees warm water and air. Water would be already better, but you cannot really do much with 35 degrees of air. That's why liquid cooling is actually a very important step towards sustainable compute, because without liquid cooling, you really can't do much with the heat. You can boil beer with 35 degrees of air? Obviously not. The second initiative or the thing that we need to work on is to optimize the utilization. We already heard it, because I'd like to remind you of the number 40: 40% of electricity does a server use when it just runs idle. So why don't we just shut it down? Because a server that's not running at all doesn't use those 40% of electricity. And when we spin up a server, then we should use it. We should achieve 70, 80 percent maybe of utilization. And Sjur already told us the way we are going to do that is called native technologies, it's Kubernetes, because with Kubernetes we have APIs to ask: hey, how much of the server actually can we shut it down? Can we merge the workloads of two servers? And third of all, we have a thing called demand response. Because at some point, Sjur or any other electricity company would want to tell us: in 30 minutes you will have to shut down half of your consumption because otherwise the grid is going down. And by that, if we are able to react on the circumstances of the network of the grid, we will be able to help the electricity companies to become much more sustainable in their productions as well. Because it's actually pretty simple: no more diesel generators anymore if our software can take care of actually shutting down a server. And these generators and batteries obviously have a pretty bad environmental impact, and if you can just not put them in the data center anymore, we have already achieved a lot. So these are the three steps that we identified. And as I said, the first one is actually a no-brainer, it's just technology that we have to implement and we have to build the servers in a way that they can be cooled by water. The second one is a little bit harder to do, because like who of you, Sjur mentioned these hundred hours a year that you would need actually to shut down your data center. Who of you would buy servers that, I mean 100 hours, this is 98.8% availability. Who would buy servers at 90% availability these days? Would Solando host their web shop on servers that are there for 98% of the year? It's called spot instances, exactly. But you need a clever mechanism of actually handling spot instances, because as yet the web shop cannot go down for a hundred hours a year. But there is actually we identified a lot of workloads that actually can be done that way. Spot instances are very convenient because they save you a lot of money, and because Amazon internally they do exactly what we are advocating here: they increase the utilization. But why is increasing utilization reserved to the hyper scalers? Why can every other data center, everyone else in the world do that? Because I mean Google they achieve utilization of 60% in their data centers, which is already pretty good, but still the industry only achieves 14%. So like everyone else needs to do something about that as well. We identified a lot of workloads actually. If you think about it in your daily work, I don't know what you do in your daily work, but there is a lot of workloads that are actually offloadable in terms of when you actually execute them. For example, your full backup, some data analysis, refinement, simulations, machine learning trainings, deep learning trainings even better because you can do second-wise snapshots and store them away and continue once the grid is able to handle your workload again. So the more you think about it, actually the more workloads are really not that time critical. Obviously an online shop is time critical, but yeah, doing your data refinement that can wait for half an hour usually. So we already heard of the five nines that Sjur told us was a bit difficult to actually achieve. And also if you think about if you want to achieve it, of course there is a lot of companies that actually get this done, but it's very expensive and there is a lot of capacity actually. There is not five nines 99.999% available and you run your backup jobs on this infrastructure. Why would you do that? Because remember, we are going to sell the waste heat that these data centers emit, so the cooling infrastructure actually pays for itself. Also, we are not because there is no idle time on these servers anymore, we are not paying for idle, which was the very first value proposition of cloud in the first place. And it's fully automated, so you once the job is done you get your result back and you don't have to worry about it. So if you identify proper use cases, I bet you would love to use such an infrastructure because your CFO would love to use such an infrastructure. It's gonna be much cheaper. So who buys this? I we can sum up what I just talked about. There is a lot of important workloads that don't have uptime requirements. And when you don't have uptime requirements, you can shift them a little bit in time. When Sjur calls me and tells me, 'Sorry, for the next hour you can't compute anything in your data center because we have no wind or because you're in London and the sun isn't shining,' then it's probably gonna be more than one hour, but when he calls me and tells me, 'You need to make a break,' and we have a lot of those important workloads without uptime requirement queued up in our backlog of work that we are going to do, then we can actually do that. We can snapshot and store everything and properly shut down the data center, and then an hour later because Sjur is very happy about us, he will call us again and tell us that yeah you can boot up your data center again. And remember, we are also selling the heat to for example to a brewery or to district heating. So all in all, since we use the hardware that we actually have there, all in all the sustainability of such a data center would be much better. Now when we summarize everything that we just talked about, what we really need is we need a global kind of means of coordinating our workloads. Because when we try to do that just inside our data center and we have one power line and when Sjur calls me and I just shift the workloads from one server to another inside the same data center, it's really not going to do much because we're gonna use the same amount of electricity. Maybe some of the workloads we can pause, but some of them yeah probably we shouldn't. Therefore we need kind of like a platform that globally coordinates all the resources that are available at a certain time in a certain place. And then this platform would need to align and prioritize the workloads that are there, then schedule them, execute them when the platform seems fit, and then at the end deploy the result of those workloads back to you. And this is exactly the kind of platform that we at Helia are building. We are building it... it's a little bit long phrase, I'm sorry, and I'll just pause until you have finished reading it. We are building a global platform where these remember the workloads are important but not uptime critical. So all the workloads you can think of, and I bet many of you can think of a lot of kind of workloads in your daily work that would fit such a platform. And I will love to hear about them after over lunch. Yeah, exactly, for example Netflix would be a perfect example. These workloads we coordinate, we deploy them onto the perfect hardware with the availability plans that Sjur is telling us, and we can even prioritize data centers that actually make something useful with the heat, because obviously we would favor a data center that sells their heat because at the end of the day it's cheaper when we can sell the heat. Obviously electricity's almost for free depending on the price that we get for the heat. And we would deploy your workloads and report the result back to you as soon as we are done. And this way, well, we call it a sustainable virtual data center because we don't buy even more hardware just to make it one in eight servers that is actually used, but we use the capacity of those six servers that are currently not used and use them to do your computations. So our virtual data center consists of every ideally in the end every idle compute resource in the world. Basically what we're building is just the AWS spot instance market, but we build it for every computer in the world. So the way we build it, obviously that's why we are here today, because we truly strongly believe that such an infrastructure can only be built on open infrastructure and open source code. Because if you cannot really if you're not used to use these technologies, you will never use such an open platform. Therefore obviously we bet on cloud native technologies, we bet on open source software, because only as Jonathan said this morning, only an open community can achieve such an effort. But the beauty of it is at the end lies truly sustainable computing, and I think we can all agree that that is a good vision. So that's it from me. There is a lot of means of actually reaching us. I just want to thank Sjur and Vattenfall. We have been a great exchange over the past almost two years where we had like sparring partner for all our ideas around sustainable computing. We even founded an alliance, the Sustainable Digital Infrastructure Alliance. Please go and have a look at that website. This is where we try to consolidate all our efforts, obviously, and we're also member there. And I yeah, then for me that's it, and I think we should go and have a delicious lunch because yeah we cannot really build sustainable compute if we don't sustain our bodies with healthy food. Thank you very much. Thank you, thank you. Is there any question that we can you want to ask them? Otherwise, yes, yes, time. Yeah, we I mean I'm from Switzerland, we like to start lunch breaks on time in Switzerland, so if you have questions obviously we still have 15 minutes, but otherwise we were also happy to talk to you over lunch. Okay, heading.
A
Audience Member23:39
So can you give an update on the current state of the technology? I quickly had a look on the Alliance page but there's not so much like changeable things.
C
Christopher Buckley23:55
Yeah, yeah. I mean the Alliance is there to coordinate everything and they are really not a technology, they really don't build anything. But for the technology that we have at Helia, we have a prototype that is working and we are already starting, well for almost half a year, we started to exactly this utilization of idling capacity with our pilot customers. And we expect to have a public beta of this resource management platform by Q1 2020. So yeah, on the website heyleo.exchange there is a form you can sign up for the newsletter and then you will get an invitation to our public beta which is expected to come in first very at the beginning of 2020. But the prototype is working and we are actually currently today using idling resources of that. I mean they still smaller data centers that are attached to our platform and we are just proving the concept that it works. And on our side, of course we do supply green energy already to data centers. What we're looking into now is the liquid cooling. So we have a partnership with a German entity which is installing on-chip water cooling, which means it really gets the cooling efficiency up and we get like 65 degree hot water instead of 35 degrees air, which makes a huge difference. And we're actually looking to deploy that in a small scale as a pilot within also the same time frames like you won this year. Two more questions?
A
Audience Member25:30
About the deep tech what you have, so can you tell us something about the efficiency of it? So your current utilization rate with the POC and some words about the software stack what you're using there.
C
Christopher Buckley25:47
Yeah, of course. And first, one really funny thing about efficiency: I just want to mention that the efficiency of power to heat that a well-equipped server today achieves is almost 100%. So you get almost the same amount of heat out of a server that you put electricity in, which Sjur loves to hear. But that's just a side note. From the energy perspective, from what we achieve with our pilot, we are currently only testing because the crucial thing is how much workloads do we get on that platform. Of a single server that we are currently running, we are using the normal traditional virtualization technologies. Currently we are in one proof of concept we are actually using Kubernetes because with Kubernetes you achieve the highest utilization of the system because every other virtualization technology is not there yet to optimize as far as Kubernetes does. But if you're interested in the number, just drop me an email, my email address is there, and I can plot you the graph on a dashboard, so Cordy according to Grafana our dashboards. But it really depends on the success of this whole venture depends on how much and how fitting workloads we get onto it. And our pilot customers are obviously not filling up our queue enough yet so that we could achieve the utilization that we are needing. And for the technologies, for the technology stack, we actually because we support we want to support any kind of data center, we use whatever they have. So we need to get access to their OpenStack or to their any actually any kind of hypervisor that runs there. We are thinking one partner uses a new Tenex, I don't know if you ever heard of that, and we are thinking about how we can implement that hypervisor. We have obviously partners that are already like giving Kubernetes workloads to us. So we want to be really not too invasive into what you're currently running in your data center. So we want to support everything. And on top of it, what we do is we install the Kubernetes cluster and then just yeah, we have a Kubernetes cluster that is customly managed by us here, and also the API is a custom development. I hope that answers your question. Any more question? Nope, okay, so we go for lunch now. The breakout session starts at 1:00 o'clock, so you're gonna have plenty of time for mingling. And yeah, thank you very much.