Volts
Volts
Making data centers flexible so they can serve the grid rather than stress it out
Sponsored
0:00
-53:24

Making data centers flexible so they can serve the grid rather than stress it out

A conversation with Varun Sivaram of Emerald AI.

Typically, AI data centers are large, inflexible loads that the grid has to build around, which is one reason utilities take so long to connect them. Emerald AI has designed a “digital brain” that can ramp down, move, or delay computing jobs in a data center on demand, making a data center a flexible asset to the grid. I talk with CEO Varun Sivaram about how the software works, what flexibility costs the compute, why it beats just installing batteries, whether utilities can enforce it, and what it would mean for clean energy.

📌 Instructions to add paid episodes to your podcast app via mobile / desktop

As attentive Volts listeners know, the US grid is built for its peak loads, which only materialize for a handful of hours a year. The rest of the time, an enormous amount of grid capacity just sits there.

When a large load like a data center asks to hook up to the grid, conventionally what it’s asking for is continuous high power, including during those peak hours — meaning it raises the peak. And if you raise the peak, you have to build new grid capacity to serve it. That’s a big part of why it’s taking so long to get data centers connected these days: utilities study them, discover that they require a bunch of grid upgrades, and then everyone starts arguing about who pays.

But what if data centers could ramp down their demand during those hours, so they didn’t raise the peak at all? A 2025 study out of Duke’s Nicholas Institute, which has since become somewhat famous in the industry, found that if new loads could curtail demand for just half a percent of their uptime, something like 100 gigawatts of new load could be added to the existing grid without the need for new generating capacity.

Varun Sivaram
Varun Sivaram

So how do you actually make a data center flexible like that? That’s the problem Varun Sivaram set out to solve when he founded Emerald AI in late 2024. The company sells software that orchestrates computing loads — throttling, deferring, and shifting them as the grid demands.

Share

Emerald has raised a lot of money fast, partly on the back of Sivaram’s credibility. He has an almost comically packed résumé: a PhD in physics, CTO of India’s largest clean energy company, chief strategy and innovation officer at Ørsted, managing director for clean energy at the State Department, author of a book, Taming the Sun. I could go on, but we need to get on with the podcast.

We’re going to get into what data center load flexibility actually means, whether it’s reliable or just a pinky promise, and how the grid benefits if it works.

Chapters

  • 00:00 – Introduction

  • 03:05 – Temporal, spatial, and resource flexibility

  • 06:26 – What Emerald Conductor touches on site

  • 08:35 – Who decides which workloads can flex

  • 10:55 – Who is liable when a job slows down

  • 13:45 – Who actually signs the contract

  • 14:49 – Larger and faster grid connections

  • 18:02 – Enforcing the flexibility promise

  • 19:05 – What broke in the demos, from bad nodes to slow telemetry

  • 26:41 – What flexing costs the compute jobs

  • 29:21 – Why not just use batteries, and the demand merit order

  • 35:44 – Training, inference, and substation-scale data centers

  • 40:23 – PJM, ERCOT, and legal enforceability

  • 44:50 – Renewables, gas, and consumer bills

Resources

People & Organizations

Company & Industry News

Books & Articles Discussed

Related Volts Episodes

Transcript

David Roberts: So all right then, with no further ado, Varun Sivaram, welcome to Volts. Thank you so much for coming.

Varun Sivaram: Dave, thanks so much for having me. Been a longtime fan.

David Roberts: Yeah, it’s… This is long overdue. This is a long time coming. I’m excited to talk. I think the place to start is just the literal mechanics here. So you sell software called Conductor. It sits in between the grid and the data center. When the grid comes to Conductor and says, “Hey, I need you to drop 20 megawatts for two hours in some sort of peak or emergency event,” what does Conductor, what is it touching? What is it doing?

Varun Sivaram: Yeah, great question. And to answer this, let’s just back up a moment and say in order for a data center to be flexible, responsive to the grid, there are kind of three ways that you could do that, right? I call them temporal flexibility, spatial flexibility, and resource flexibility. The first way, temporal flexibility, is you might have some things that are running in the data center that don’t need to run that fast at that very moment.

David Roberts: Mmhmm.

Varun Sivaram: And so you might be able to slow down some of your AI computations or pause some of your jobs, reschedule jobs, and you can do this in many ways, by the way. You can downclock some of the individual GPUs or nodes, you can reallocate the number of GPUs working on a particular task. You can force a checkpoint and pause a run.

There are many, like, different ways to do this, but temporal flexibility is in time, can you flexibly change when work happens or the rate at which it happens?

The second is called spatial flexibility, and that’s in space, can you change where the work happens? The amazing thing about data centers is that they’re connected at the speed of light with one another. And so if you’re able to move an AI job from one location to another, it could be separated by 100 or 500 or thousands of miles-

David Roberts: Mm.

Varun Sivaram: and remain within the latency constraints of your customers. The customers might not even notice a difference while you have provided real grid relief in one location by moving work to another one. And then finally, there’s what I call resource flexibility, and this is where you step outside of the data center and you say, “Can we put some other stuff like batteries and generators near the data center that help that data center to reduce its net grid withdrawal or net impact on the grid for some temporary period of time?”

The correct way to solve this problem is to orchestrate all three of these alternatives together. They’re not alternatives, they’re complements, and I use orchestrate intentionally. Emerald Conductor is literally conducting an orchestra of available interventions, temporal, spatial, and resource flexibility.

And so you asked, Dave, you know, what does Emerald Conductor do? It kind of does all the stuff. It is ensuring that the GPUs and the scheduler and the AI jobs are temporally and spatially flexible in time and space, and it’s also making sure to coordinate with a local resource, let’s say a battery that is on site, so that we can use what’s available in the battery without running into its constraints on state of charge or discharge rate or durability. All of this is a very complex problem to solve, and we have developed an AI technology ourselves, which is underlying the Emerald Conductor platform. We are an AI that runs AI, and-

David Roberts: Mm-hmm…

Varun Sivaram: we have AI agents that are inside the data centers that are orchestrating all of these autonomous decisions.

David Roberts: Yeah, that was actually my second question. So you’re not just conductors, not just controlling the GPU, the computing. You also are in touch with on-site batteries or on-site backup generation, or, like, the cooling system? How far do the tendrils go? I mean, is everything on-site in one way or another? Are you touching it? Like, you can control all of it, basically?

Varun Sivaram: That’s exactly right, and we do a lot of this with our partners and our investors that are large Fortune 500 companies that are real experts in this.

We’ve announced our partnership with Siemens, for example, where Siemens has software that controls a lot of this technology on site, whether it’s the battery, the microgrid. We have other partnerships with folks like Eaton. They have these UPS systems, GE Vernova. We have a range of partners, and what we make sure to do is the Emerald Conductor serves as this brain that says, “Hey, what’s all the stuff on site that can provide me some flexibility within its constraints, and how do I fuse that or orchestrate it alongside all of the IT or computational flexibility, temporally and spatially, and put it all together?”

So yes, you can think of Emerald as the brain that turns the data center — or what we like to call AI factories, following NVIDIA’s lead — into a controllable power user. And I think this is a very profound point. Traditionally, you look at a data center or any other load, a car factory, as an unpredictable user of power that the grid has to provision for.

Like you said, David, in the intro, you said, “Hey, we build these massive energy systems and keep building our peak capacity,” but most of the time, transmission and distribution capacity is underutilized. Generation is underutilized. And so we’re really using the grid at 50% of what it could be doing. Well, if the data center itself or these AI factories become precise, controllable energy users, I believe that makes them energy technologies. That’s counterintuitive. Like, you think of nuclear power plants or batteries as energy technologies. But what if data centers were energy technologies?

David Roberts: Here at Volts, we’re very familiar with controllable load and all its many glorious benefits. So we’re gonna get into some of that. But on the computing front, I have some questions about just the computing front. So one of the things your framework has is these tiers of computing workloads. There are jobs that can’t be touched, that have to run all the time. There are jobs that can tolerate 10% throughput loss, 25%, 50%, there’s these tiers. So one of my questions is, who decides what’s in what tier? Is that something that the owners of the compute tell you, or is that something that Conductor analyzes the compute and makes its own judgments about what is and isn’t movable and controllable? Where do the tiers come from?

Varun Sivaram: That’s a great question, Dave. It’s the heart of what we’re doing, which is serving both sides at the same time, the energy grid’s needs and the AI customer’s needs.

Look, at the end of the day, AI customers are doing the most economically important work in America, right? Generating valuable tokens of artificial intelligence, and you don’t wanna mess with that in a way that’s going to degrade their expected quality of service. And so yes, today, for our commercial scale deployments, we receive tags from the customers, the users of AI compute, tagging which of their jobs has which flexibility parameter.

And it’s more than just, you know, this can be flexed and this cannot be. Some jobs, a benchmarking run, for example, you really can’t downclock the clock frequency of the GPU ’cause it’ll mess with the benchmarking run. Some jobs, a fine-tuning run, they can be deferred. They could also be downclocked. But some particular workloads, like a serving inference job, really can’t be downclocked or paused because the user expects an answer within a latency bound, and perhaps that needs to be geo-shifted.

So within different priority tiers, you know, low, medium, high, or on a scale of 0 to 100, as well as across different characteristics, preemptibility, throttle ability, the ability to be paused and checkpointed, or the ability to be moved within residency requirements. It’s a complicated set of constraints, and we feed this in through what we call the normalizer to turn it into actions that we can take, and then our AI agents take the actions.

David Roberts: Well, I’m sort of-- Well, I mean, one of the things I wonder is if somebody gets that wrong, who’s responsible? Like, who’s holding the bag at the end of the day? Like, if you could downclock the power, a job slows down, and for some reason that someone didn’t anticipate, that breaks a job or doesn’t get the result it wants, or there’s an unanticipated result, is that just on the compute load for mislabeling its job, or who’s responsible if something goes wrong with the jobs?

Varun Sivaram: Yeah, great question, Dave, and I’ll use this to mention, you know, this is high stakes, high reward in kind of three different dimensions. Number one, it’s massively high reward because if we make this all work, we can fit 100 gigawatts, as you mentioned from the Tyler Norris paper, we can fit 100 gigawatts of AI data centers onto the grid that we didn’t think we could fit, and we’ll walk through the benefits of that.

So massive reward. High stakes though is on both sides. If you don’t do this right, the grid could face reliability issues. On the other side, if you don’t do this right, the AI customers might face failed expectations, again, for the most economically valuable work in America. And the third dimension, of course, is a security dimension.

AI data centers becoming energy technologies and this fusion of AI and the energy system where we sit at the interface is a potential security threat, right? These are two massive multi-trillion dollar networks that are critical infrastructure for America’s economy. It’s the backbone of our future economy. And so we gotta make sure that we’re doing all of this in a very safe way that protects the reliability of the grid, is resilient to intruders or threats from abroad and domestically, and protects the integrity of the AI workloads.

Your specific question, Dave, is who pays if there’s a screw-up? And the answer to that specific question lies in the contracting, right? There will be, depending on the context, the grid, if they’re offering you a faster connection, they may say, “You know what? If you don’t curtail your power when we ask you to, there may be a financial penalty, or there may even be a technical penalty. I’ll flip your breaker.”

David Roberts: Mm.

Varun Sivaram: And so there’s a contract related to that between the data center and the grid, but the data center will say, “Look, you’re offering me power I otherwise wouldn’t get for 12 years. I’d love to get this power, and let’s find a way to mitigate this risk.” And Emerald Conductor exists in order to make sure that with 100% reliability, we are doing the thing the grid needs in order to fulfill your terms of getting that earlier power connection. And then, of course, there’s the question of, well, what if the customer said they didn’t want this job to be slowed down, and it got slowed down?

David Roberts: Mm.

Varun Sivaram: That’s also something that the reason you bring Emerald Conductor on is because it’s a production-hardened, production-grade software platform that ensures that you protect the workloads that you want it to protect. So I hope the risks are clear. The liabilities lie in contracting, but these are solvable problems because the reward is so massive.

David Roberts: On the contracting, like, who cuts you a check? I guess that’s a basic question. Like, who are you working for? Is it the utility? The, is it the data center operator itself? Is it the operator of the compute, which is sometimes different than who owns the data center? Like, literally, who are you working for?

Varun Sivaram: Great question. We are lucky, Dave, that we work for basically everybody in this ecosystem, and our objective is to make everybody else so much value that we’re sort of barely noticed. You know, utilities get a ton of value, and we should talk about this, Dave.

David Roberts: This gets to my next question, but first, I just wanna know, like, literally who comes to you and says, “Let’s sign a contract”? Who is that entity? Who is the counterparty?

Varun Sivaram: Yeah. To answer it bluntly, it is electric utilities, data center operators, leading AI companies that are using compute, cloud companies. It is everybody in the value chain is our counterparty because we provide services and software across the value chain.

David Roberts: Okay. And so then the next question is then, like, what’s actually being sold here? What’s the value? So the obvious source of value here is that an inflexible data center is gonna have to wait a long time to connect to the grid, whereas a flexible data center might be able to jump the queue and get hooked up earlier. I’m assuming that from the data center’s point of view, that’s the golden ticket there. Like, that’s the main source of value. Is that accurate?

Varun Sivaram: Yeah, roughly right. For data centers, you want larger and faster power connections than you would otherwise get. So you might be an existing data center, and suddenly, instead of 50 megawatts, you get 75 megawatts from the grid. That additional 25 was, you were gonna wait 10 years for, and now you get it immediately. And the reason is because you’re willing to be flexible with that new allocation of power, and the utility said, “Wow, I can just grant it to you because I don’t have to build out any grid infrastructure to provide it to you.”

So that’s larger connections. Faster connections is you’re building a new data center. Like you said, Dave, you’re waiting in a queue for 12 years, and instead, you now get connected much faster because, again, the utility can-- wants to connect you. You provide a valuable service, and you’re not requiring them to build out the grid and get permitting, et cetera.

David Roberts: Can I ask a couple of questions about that? One is, if I’m a built data center, I’m already up and running, and I have no immediate expansion plans, do I get any value from you or… Do you know what I mean? Is there a reason to do this other than wanting more power from the grid?

Varun Sivaram: Yeah, absolutely. Great question, Dave. So if I walk through what’s the value that we give to data centers, the number one value is larger and faster power connections. So if you are an existing data center and you wanna upgrade to the latest liquid cooling and latest NVIDIA GPU, you probably need more power, and so we provide that for you.

But your question, Dave, is what if you’re a happy data center that has no desire to do any upgrades and you just want to hang out? Would you ever work with Emerald? And the answer is you could, ’cause the second value stream is by making you a flexible, grid-friendly, grid-responsive participant, you can be paid. You can be paid by flexibility markets, by utility demand response programs. Now, to be clear, these are far less lucrative than the value of getting access to power than you otherwise wouldn’t have.

David Roberts: Yeah. That’s what I sort of wonder, like, that is a revenue stream. I mean, compared to the value of the compute, you know what I mean? Compared to, like, the value of the data center itself, it’s pretty small beans. I’m just wondering if any data center is gonna take risks with its compute for that prize, ’cause that prize is relatively small.

Varun Sivaram: You know, going forward, as we continue to, as the American economy increasingly uses AI inference, I expect that cost pressures on inference will grow, and power is the variable cost of inference, and so reducing your power cost will become more and more important. But you’re absolutely right, Dave. Today, the key driver of value in the minds of data centers and clouds and AI companies is, “Can I get the GPUs connected to power at any cost?” And as a result, optimizing the power cost is probably not top of mind right now.

David Roberts: Yeah. So another question about that is, if I am promising flexibility to the grid and the grid says, “Okay then, you can have the power,” and then I get hooked up to the power. What’s to stop me from just saying, “You know, upon review, it turns out all my computing is incredibly important and can’t be interrupted”? And what’s to enforce you actually being flexible once you’re hooked up?

Varun Sivaram: The contract upon which you connect. And so, again, that can be a contractual thing. Hey, you have the power so long as you’re complying, and you fail to comply, you know, you get a, maybe it’s a three-strikes policy or two strikes, and you fail to comply, you lose that connection. Or there’s actual physical breaker control. We flip your breaker and you don’t have access to the grid again. I wanna be clear, you said pinky promise early on, Dave. You said, “Hey, will demand response in the form of flexible data centers actually show up?” And my answer is yes, it will actually show up, and we need to make sure that it is enforceable and reliable for utilities.

David Roberts: Right. I wanna get into that a little bit later, that specific question a little bit later, but a couple more questions just about the mechanics of this working. So I read this, you know, you have this Hillsboro test case that you’ve written up. It’s, I think, 250 kilowatts, so it’s relatively tiny compared to what an actual data center’s gonna be.

And there’s a bit in there about what went wrong, where the errors were, and there’s a couple of cases where some of the computing nodes were either unreachable or unmeasurable or uncontrollable for some reason. And if that happens, I think just to have a buffer, you have to assume that that node is running at full power just to make sure you’re coming in under the load you need.

And then there was some issues with sort of control network getting congested when there’s a big event that requires a big power drawdown, sending messages out to all those nodes at once, you get congestion on the power network. All of this is one thing at 250 kilowatts, but then you scale up to, you know, whatever, 96 megawatts, which I think is the Manassas test you’re doing.

I wonder how those errors scale. I guess I’m just wondering, like, what are you worried ab- what are the sort of failure modes that you’re worried about? How many nodes turn out to be unreachable? What are you doing about this congestion issue? Like, what are the sort of failure modes that are keeping you awake at night at when you get up to scale?

Varun Sivaram: Yeah, such a great question, Dave. Let me back up and just give some background about what you read and then where we are. So you read the results of our fifth demonstration on GB300s, Grace Blackwell 300s, the latest NVIDIA chip generation in Hillsboro, Oregon with Portland General Electric. And you’re right.

That was a demo that was less than a megawatt. Now, we’ve done five demos all over the world, including in London with NVIDIA and with other partners like EPRI and their DC Flex program. And I’m excited that we’ve just exited the demonstration phase, and we’ve entered the commercial deployment phase. We’re now actually deployed at multi-megawatt scale at a whole data center, flexing it.

In fact, two nights ago, it was my personal dream to be able to do this, and I’ve thought about this for the last decade, we flexed at the peak of the duck curve. Dave, you and I have followed this duck curve forever. You may know me as the guy who first put a picture of a duck on the duck curve.

David Roberts: Were you, are you the one who drew that crude duck? I must have used that in, like, 100 posts.

Varun Sivaram: Well, no, no. Somebody else drew the crude duck. But in my book, Taming the Sun, I put a very sophisticated duck picture, a rubber ducky. But in any event, so we’re a level of order of magnitude or two orders of magnitude above that. And you’re right, later this year, the 96-megawatt Vera Rubin Test Center with NVIDIA and Digital Realty comes online, where Emerald will make it the world’s first, you know, 100-megawatt scale power flexible facility.

So your question is such a good one, which is what do you get scared is gonna break as you scale from kilowatts to megawatts to hundreds of megawatts, right? And the answer is you don’t know what you don’t know. There’s a lot of unknown unknowns, you can discover at each level. That’s why we did the demos. Our company believes deeply in open sourcing our research and our results, which is why we publicly published a report that said, “Here are the things that went wrong.”

David Roberts: Right.

Varun Sivaram: We’ve published, you know, peer-reviewed research in a Nature journal, scientific journal [NPJ Climate Action] saying, “Here was the code we used to make this all work.” We deeply believe in an ecosystem where all data centers can do this, and regulators can say, “I know how this thing works, and I trust it.” And they’ve done the hard work to go through demos, prove what doesn’t work, improve upon it, and actually make it work.

And so I think about when we first deployed in this multi-megawatt facility, whole data center, we’re flexing it, and we’re all sitting in that conference room around a bunch of laptops doing the first ever event where the utility says, “Hey, we need you to curtail.” The atmosphere, one engineer told me, felt like a SpaceX launch. I think that was our head of product, Mansi Shah. She said, “Hey, it feels like a launch.” You know, we all cheered. The curtailment happened. At some point midway through, the data center power briefly breached or went right back above the limit of the curtailment that the utility had asked.

David Roberts: Mm.

Varun Sivaram: And it was kind of like thinking, you know, the SpaceX rocket, it went all the way up, it came all the way down, and then it nicely landed, but then it kind of tipped over at the end. And we’re like, “Oh, what went wrong? What can we fix here?” And so we fixed it, and we continue to be deployed, and it’s been a fascinating deployment where we’ve learned a lot. So that’s how I think about this.

David Roberts: Well, like the nodes that you couldn’t reach or couldn’t measure, I mean, one of the things that I wonder about, is that an error in the construction or the design or the manufacturing of the node that is out of your control, you’re not making those. Or was it something about your measuring system that you can improve and then measure and test it better the next time? Like, you’re gonna get some bum nodes, I feel like, just the law of large numbers, you know.

Varun Sivaram: Great question, and I didn’t actually answer your question about, like, what keeps me up at night. So the way the Conductor platform is architected uses a principle called virtualization from the IT world, or abstraction.

David Roberts: Mm-hmm.

Varun Sivaram: The goal is no matter what the underlying hardware is and what the underlying telemetry, or the sensors that are reading to us the power data, no matter what it is, we should be able to make this thing work, right?

And we are creating an abstraction or virtualization layer above it so that if anybody looks at a data center, if you’re a grid operator, that data center presents itself to you as if it were a battery or a controllable generator. It looks to you like the same kind of thing, and you don’t have to worry about whether they’re using this measuring equipment or that GPU, but it does become my problem. And so Emerald’s problem that we’ve gotta solve is there will be problems with the equipment.

David Roberts: Right.

Varun Sivaram: You know, one of our demos, we ran into the problem where the telemetry, the stream of data-

David Roberts: Mm-hmm…

Varun Sivaram: only happened every 15 minutes, right? That’s crazy, right? You expect this to come every second or every 30 seconds so you can at least make some quick decisions, but suddenly we’re getting it every 15 minutes. And so our AI technology needs to account for this, and then make decisions understanding that it’s like you’re talking to somebody on Mars and the communication is delayed 15 minutes. That’s one example of a thing we have to engineer around. Ultimately, all of these are engineering challenges. You know, the one you mentioned, Dave, was what if nodes go down? GPU nodes go down all the time.

David Roberts: Yeah.

Varun Sivaram: GPU nodes sometimes don’t report correctly.

David Roberts: Yeah.

Varun Sivaram: This, by the way, this is not an NVIDIA issue. This is often it can be an installation issue, it can be a problems with other parts of the power equipment in the facility. Like, and you have to be able to be fault tolerant or error correcting.

David Roberts: Right.

Varun Sivaram: And so a lot of what we do is build in the buffers and the intelligence to take care of these faults.

So there’s a lot of engineering, but fundamentally my vision is… AI factory should be the most exquisitely controllable energy asset on the grid we’ve ever seen. They do things that nobody else can do, right? You can’t move your load from one location to another at the speed of light. Like, you can’t move an electric vehicle from Indiana to California in a couple hours. It-- let alone milliseconds, which is how fast we move AI workloads. You can’t control a big physical load in the same way you can control a data center from a laptop in a very graceful way that the customer barely notices.

So that’s why I’m so excited. That’s why Emerald does only one thing, AI factory flexibility. Whereas a bunch of other companies do a bunch of other things, which is great, we do one thing, ’cause we think AI is going to account for most of the American economy this century, most of the electrical load, and we’re going to make it a very community-friendly, grid-friendly asset.

David Roberts: In terms of what it costs the compute jobs to be capped or slowed down like this, you know, you have, in your test, NVIDIA attested that the performance was acceptable, that the performance of the jobs was acceptable. But you also acknowledge in this one white paper, and I don’t know if this has changed since then, that no one has actually done the actual testing of the throughput or the latency or the job completion, you know, sort of deltas. So I’m wondering, is it really costless to the jobs themselves, and how tightly measured has that been?

Varun Sivaram: I’m so glad you brought this up, Dave. The central question is, can we make sure it’s not costless, but as, call it inexpensive or as hassle-free for customers, for us to every so often gracefully reduce their power?

David Roberts: Because again, like these jobs are worth way more to these companies than power is, right? Like the cost to-

Varun Sivaram: And again, to the American economy. I mean, this is the lifeblood of the American future economy. We cannot screw this up. And so yes, Dave, to your question, we have in very great and meticulous detail measured this.

Our chief scientist, Professor Ayse Coskun, has done academic peer-reviewed work on ensuring the service level of critical jobs. We ourselves have published papers, for example, with Nebius, NVIDIA, and National Grid in London and EPRI, where we showcased not only that you could precisely meet over 200 grid events, but you could do so while protecting the throughput of… the token throughput, the other performance metrics of critical AI workloads.

There’s no such thing as a free lunch, Dave, and so therefore, there will be some jobs for which there is a reduction in performance or a delay, but we ensure that we are meeting the customer constraints so that if they designate a job that’s non-interruptible or non-preemptible or super high priority, that job culminates in the time that the customer expects it to culminate in.

So we have very detailed metrics tracking. We’ve published it. We’ve also done metrics tracking with Oracle, where we moved AI inference, serving inference workloads. Like Dave, you’re talking to ChatGPT. We’ve moved queries from Virginia to Chicago at the speed of light and kept it within the latency bounds that you would expect when you’re having that conversation with a chatbot. So Dave, you ask, you know, “Hey, is Varun actually informed about what he’s talking about?” And you want that answer to come to you before I give my next answer, right? And it’s gonna come to you on time, no matter if it, you know, got processed here or there within certain bounds. So yes, we have deeply studied it.

David Roberts: Got it. Okay, and here’s a question. This is actually the very first question that occurred to me when I heard about what you’re doing last year, and I’ve been wanting to ask it to you for ages, which is like, as you mentioned, a lot of these facilities have batteries on site, and one way to get flexibility is just to draw from the battery instead of from the grid, right?

I mean, this is the time-tested way of operating grid flexibility. And I’m just sort of wondering, like- isn’t that simpler and easier? And doesn’t it require much less engineering and all this sort of, like, very complicated real-time orchestration? You know, a battery, and it’s meterable. You can just flop over to the battery and flop back to the grid when you need to. Why not just do that? What is the incremental value add of this detailed load orchestration versus just a battery?

Varun Sivaram: Yes, yes, and yes, Dave, is the answer to your question intuitively, right? And then I’m gonna tell you why counterintuitively I have a different answer. But the intuitive answer is yes, yes, and yes. Of course, you should just treat this data center as a black box. Who the heck knows how much power it needs? But we’ll just stack a ton of batteries and generators right next to that data center, and so whatever the grid needs that facility to do, we’ll just take care of it with the batteries and the generators.

And that should work, right? And there are a few reasons to understand why that’s probably not the right solution.

The first is that it’s expensive. If the grid tells you, “Look, in order to connect you, I may need to curtail you. Now, it’s once in a blue moon, but I may need to curtail you for two hours or four hours or eight hours or 16 hours,” or with our multi-megawatt deployment, we actually just did a 24-hour event to prove that if the local substation went down and they had to repair for 24 hours, “could you curtail?” Good luck trying to size the super expensive battery to provide all of this flexibility.

David Roberts: Enough batteries to run your facility for 24 hours would be cost-prohibitive, I think.

Varun Sivaram: And even less than that is pretty cost-prohibitive. But, you know, wouldn’t it be nice if you only had to put a two-hour battery on your facility for part of your consumption, and the rest of it was done through this nuanced, sophisticated- computational orchestration, well, then you’re saving on CapEx, and it’s a better together approach where the battery does some of the work and the computational flexibility does more of the work. So yes, it’s a lot of complexity that goes into our Emerald Conductor platform, but it makes for an ultimately cheaper, better functioning system.

David Roberts: Real quick, just on the duration question, ’cause you brought it up and I wanted to ask about this. So I see the logic where a battery can do two hours, but if you have an event that lasts longer than two hours, orchestrating the computing doesn’t get any more expensive in hour three or hour five or hour 24, right? So, like, on a long duration event, I totally get the advantage here. But in practice, most of these events are not gonna be long, right? I mean, in practice, like, if you look at the bell curve, like, the bulk of the events are gonna be two hours and under, aren’t they? So is the orchestration of the computing just for kind of-

Varun Sivaram: It’s not-

David Roberts: the edges of that bell curve?

Varun Sivaram: It’s not for the edge cases. Now, I will say there is a journey that we’re gonna take the whole industry on, which is we’ll start by working with you on the edge cases to say, “Look, the load orchestration of the compute, the AI flexibility, we’ll do that for edge cases, but your traditional way of operating for the last two decades is never to touch the compute. We’ll largely leave you alone.” But where I wanna bring everybody is, in the long run, AI should really be flexing a lot more often.

David Roberts: Hmm. Not just during emergency events, you mean, like on a routine-

Varun Sivaram: That’s exactly right. In a way, of course, that’s respectful of the true performance constraints and requirements that customers need, but this is the most elegant way to provide flexibility to the grid at scale. And let me just walk you through what this means. Shayle Kann, who led our recent strategic expansion round from Energy Impact Partners, you know, he tells me the reason he wanted to invest was he had this aha moment. He said, “You know, in the electricity space, we have something called the supply curve or the merit order.” And we say, “You know, some things are really cheap to dispatch, some things are really expensive to dispatch, and you, like, arrange them in order, and then you do the cheap things first and the expensive things later.”

And he’s like, “I can now draw a similar curve, but for the things that you could dispatch at the data center.” And in many cases, there will be some AI jobs that show up as very short vertical bars on that merit order dispatch curve. Like, there will be things that are inherently flexible. The customer’s really not gonna mind if that batch inference job or fine-tuning job got slowed down for a bit. It literally was not important to them that it happened at that moment. And so …

David Roberts: So it’ll come before the battery on the merit curve of demand response, basically.

Varun Sivaram: … it comes before the battery, right? That’s exactly right, Dave. So you visualized the same thing that Shayle did. In the long run, Dave, again, I founded this company because I’m pretty deep AI maximalist. I think AI is going to be transformative in a good way. But if we have a ton of AI and it becomes a 50% or more user of the American power grid, we simply won’t be able to build our grid and generation fast enough. We need to build as much as humanly possible. We just won’t be able to build as fast in order to serve it all.

And I see a future, mid-century, 2040, when half the work is done by the underlying physical power grid, the power lines, the electrons, and half the work is done by the AI digital optical grid, the photons and the processing that’s happening at each node ramping up and down and across nodes when work is being transferred from one location to another via fiber optic cables. And that combination, that complementary physical grid and what I call the meta grid, the AI grid working together, that’s the thing that makes electricity affordable and reliable for communities and gets us all of the AI benefits that we want for our economy.

David Roberts: So you’re talking about moving from this is an emergency or peak shaving, you know, edge case to this demand flexibility is providing 40, 50% of the total picture.

Varun Sivaram: It’s the way the grid works in 25 years, and hopefully much earlier than 25 years, hopefully in 10 years. That’s the way the grid works.

David Roberts: Well, let me ask about this then, ’cause one of the things that’s happening in AI, I think listeners may be familiar with the sort of difference between training, inference. Training is when you’re training up these models and you’re doing… It’s a big burst, highly intensive. You need really sophisticated GPUs, et cetera, et cetera.

And moving to inference, which is actually using them in the field, it seems like when training is dominating, those are sort of batchable jobs. You can kind of defer them. Almost by definition, you can defer them ’cause they’re not as maybe time sensitive. But over time, it seems like on a sort of macro level, we’re moving more and more toward inference. More and more of the compute is inference, which again, is much more time sensitive. So is that shrinking the total pool of kind of controllable, movable compute load, is the move into inference tightening the parameters, I guess, is the question?

Varun Sivaram: Yeah, great question. And my answer to this is, it’s probably an oversimplification to separate the world of AI into a bucket called training and a bucket called inference. There’s actually a lot of small, small buckets. You have model fine-tuning. I think this will become a large user of compute as, you know, we start to… Enterprises fine-tune open source models to meet their needs. There’s inference that is extremely latency sensitive, sort of latency sensitive, batch inference that is not very latency sensitive at all. There’s a lot of different ways that we use compute.

I also, by the way, happen to think that we will continue to have large frontier training model runs, and counterintuitively, those are probably the ones that are least amenable to the kind of flexibility that the Emerald Conductor does. Because, look, if you’re training, Opus 5 or GPT-6, it’s a massive, highly synchronized operation across, in many cases, lots and lots of data centers. That is a complex use case that we probably wanna not work on right this second.

David Roberts: Can I just cut in and ask a question? ’Cause this occurred to me actually. Like, if you have Conductor running, say, at five different data centers, and there’s a compute job like, say, the training thing you’re referring to that is itself spread across those five data centers, can Conductor, can it conduct all five at once? Like, are the Conductors communicating with one another such that they can do sort of like macro conducting of multiple data centers at once? Does that make sense?

Varun Sivaram: Yeah, great question. It’s a project we’re doing together with EPRI and others like NVIDIA and Prologis. It’s called the Distributed Data Center… Substation Level Distributed Data Center project, where at the substation level you often have like some capacity, but not a whole lot, I don’t know, 5 to 20 megawatts. But you could build these distributed data centers at multiple substations, each of which is flexible. So every so often the substation’s gonna have 2 megawatts of capacity or otherwise 18 megawatts. And orchestrated together, those five or 10 or 15 distribution level data centers act as a virtual aggregated single data center-

David Roberts: Right. They can sort of move the compute to the substation that’s got slack at the moment, right, is the idea.

Varun Sivaram: Is the idea, or they can slow down or downclock compute that’s happening at a substation that’s constrained. So temporal and geographic or spatial flexibility, you put them together, and then, you know, as you correctly said, multiple smaller data centers can be aggregated into one larger virtual data center. And-

David Roberts: Mm.

Varun Sivaram: Look, we’re gonna be looking for power everywhere. And-

David Roberts: Right…

Varun Sivaram: the lower levels of our network, the distribution level, the substation level, there’s land, there’s power. It’s just, you know, not as convenient as getting a nice large one gigawatt chunk that’s firm. And as we prospect for power, these cool ideas are all gonna utilize this one fundamental concept, power flexibility.

That’s why we think of ourselves as a frontier research lab. We do frontier research on how to make AI power flexible, and it can be used in all these cool cases. Another cool example is you could theoretically put AI on a naval ship, for example, or in a forward deployed theater where there’s an imperfect power grid, and flexible AI compute enables you to handle fluctuating power envelopes. So you can really think outside the box. This is a fundamental attribute of compute, is its power flexibility.

David Roberts: So this is like multiple data centers aggregating, acting as one. This is not happening yet. That’s out on the edge of the sort of research and, you know, forward-looking type of thing.

Varun Sivaram: I think it’ll happen pretty soon. We have a project underway, but I do love the pun you made, Dave. It is happening on the edge. That is right.

David Roberts: Edge computing. Okay, let’s get back to the enforceability issue. So the market monitor in PJM recently called data center flexibility, quote unquote, “a regulatory fiction”, which means that PJM, it can ask for flexibility, but can’t really enforce it, can’t statutorily sort of expect it, and so refuses to kind of plan around it. So I’m wondering, like, would you support making flexibility, if someone says they have flexibility, making that legally enforceable, basically such that entities like PJM… You know, ’cause with a battery, it’s measurable. There’s a meter, you know what I mean? Like, you can measure exactly what happened.

You can enforce it. You can, in some sense, plan around it. This sort of aggregated stuff, I think PJM is not quite confident yet. So is there a way of making it mandatory and legally enforceable that you would support? Or how do you think about that? How do you think about that issue?

Varun Sivaram: Yeah. The answer is yes. Absolutely. If we’re gonna make this work, if data centers are gonna get faster and larger grid connections in return for promising flexibility, that promise needs to be legally enforceable and technically enforceable. And so I 100% believe that we need to prove to regulators, to governors, to utilities that data centers will 100% show up when they’re asked to show up if the reward they got was larger and faster connections.

And the market monitor is right to say, “Hey, historically, the performance just hasn’t been great of demand response.” But let me tell you, Dave, this is a new category, a new class of flexible power users. It’s an energy technology in its own right, more so than anything that’s come before it, and I believe it can be legally enforceable.

So the last thing I’ll say is I think ERCOT, Texas, ERCOT has done a great job of developing a framework, a legal framework for how this could work. They say a flexible data center is welcome to connect It’ll be a, what we call a provisional controllable load resource. It’ll actually be dispatched via the dispatch system. It can provide its willingness to curtail, and the willingness is an economic number, $2,000 a megawatt hours. You really don’t wanna curtail, but-

David Roberts: Mm-hmm…

Varun Sivaram: you know, hey, if we reach that condition, then we will dispatch you. And it’s a, you know, legally binding thing that you have to do, and a condition of being connected is that. That makes the data center the same sort of entity as a generator or a battery or anything else that is dispatched. It is legally equivalent, and I think we have the technology to make that possible.

David Roberts: Interesting. Well, on that note, you know, we were discussing earlier, like, if you’re just a fat and happy data center running along, like, why would you wanna do this? Once it is demonstrated that it is possible for data centers to be flexible, I wonder if you ever foresee a day where it simply becomes mandatory for all data centers to be flexible. In other words, right now I think there’s, like, data centers get grandfathered in if they don’t have this stuff installed, but theoretically, you could come along and retrofit an existing data center with this and make an existing data center flexible. And if that becomes possible, do you think that’s ever gonna become simply mandatory for all data centers?

Varun Sivaram: Well, look, I think it’s wrong to paint all data centers with the same brush. There are a lot of existing or legacy data centers that do a bunch of heterogeneous cloud computing tasks that probably aren’t great candidates for flexibility, and it’s really the AI factories going forward.

David Roberts: So AI specifically, not just all compute, but it’s AI specifically that you think is flexible.

Varun Sivaram: Well, that’s where we’ve chosen to focus. We believe that there’s technical room to run here, that there is an economic value driver. They really want access to larger and faster power. And we believe that the right way to do this is to provide incentives, the carrot approach, not the stick approach. The carrot approach is-

David Roberts: Hmm…

Varun Sivaram: if you want the power, here’s how you come get it, and, you know, if you wanna do the traditional route, go ahead and do the traditional route. It just might take forever. And I think that’s a workable solution that, you know, the AI community can get really excited about.

David Roberts: Interesting. Okay, and Volts is about decarbonization ultimately, so one of the things in that Duke study that everybody's constantly citing — one of their conclusions is that flexibility, all things being equal, favors renewables over gas. I think they found that 20% flexible demand avoids a fifth of new gas construction, and that sort of scales, like 50% avoids half the gas construction. But I’m curious about the mechanism of that, whether you believe in that and whether that’s part of your motivation here.

Varun Sivaram: Dave, it’s why I started the company. I sincerely, strongly believe that by making data centers flexible, again, as AI data centers become the largest user of electricity on the grid. Their flexibility, ceteris paribus, pulls new clean energy from intermittent renewables onto the grid.

David Roberts: Hmm.

Varun Sivaram: You know, I’ve spent my career in the renewable industry. Man, would I have loved to have an offtaker that’s a flexible data center. It’s massive. It is responsive. It helps me use my resources, my electrons when they’re produced. Like, this is a massive win for making the grid more amenable to connecting clean energy resources.

Dave, I’m sure you saw this report that came out in the last couple of days that said the net impact of data centers, published in Nature, a peer-reviewed journal, is going to be bad for climate because it’s going to increase the efficacy of oil and gas extraction, I think that was the primary driver of increased emissions, even as it, in a less pronounced way, advances renewable energy. I think that it has it flipped.

David Roberts: Well, also, and just also, it’s visibly right in front of us pulling a crapload of gas onto the grid. Like, whatever it might do in the future, right now it looks like hyperscalers are herding to gas. Like, the gas pipeline is jammed now, and gas companies, you know, their stock is rising, and they’re building on-site gas plants. So, like, when, if AI is gonna benefit renewables, when?

Varun Sivaram: As soon as we demonstrate demand flexibility and get it scaled up, which is what we’re doing. I mean, I just wanna be clear here, Dave, I think demand flexibility is something both parties can get behind, whether or not you care about climate.

If you do care about climate, you get a lot of clean energy on the grid. If you don’t care about climate, you get a lot of other benefits. Consumers pay lower power bills. Ceteris paribus, all else equal, a flexible data center will serve to reduce cost and affordability pressure on local bills because you don’t have to build out exorbitant infrastructure to serve higher peaks.

It tends to democratize AI by getting more GPUs onto the grid so everybody can get access to life-saving AI innovations. It improves grid reliability because flexible AI data centers can respond to the grid, be good grid citizens, you know, responding to lightning strikes, as we demonstrated in our demos in London. And all of these community-focused benefits happen if AI data centers become flexible. It’s why, you know, I strongly support Republican Senator Dave McCormick’s covenant for how you get data centers to be community-friendly, and plank one of that is just make sure that they have positive grid impacts and don’t raise rates, ceteris paribus, and flexible AI is the only way to do that.

There is no way to do it otherwise. You can’t say, “I’m just gonna go off the grid,” because then you’ll have an implicit cost increase by using equipment that can’t then be used on the public power grid. That’s also, by the way, a death spiral effect for utilities. You really don’t want data centers to just entirely go off grid or even worse, go into space.

So I really think that flex- and I’m not saying that out of jest. I actually, I really admire what Elon Musk and Starcloud and others are doing to bring data centers to space. It’s a remarkable achievement that they’re gonna be able to do that, but the best thing we can do right now is to use every last watt of electricity we’ve got on our existing system for AI, and the best way to do that is through flexible demand.

David Roberts: So you think that Nature paper is wrong?

Varun Sivaram: I think that Nature paper did not assume that Emerald AI and others, like Google by the way, Google is an absolute pioneer here. A set of actors doesn’t make AI flexible, and therefore a fantastic grid participant and citizen.

David Roberts: Mm-hmm. And so talk about the logic or the mechanics of how this lowers consumer bills, ’cause it’s not intuitive.

Varun Sivaram: Yeah, absolutely, and I think this is one of the most important parts of Emerald, right? Again, going back to Senator McCormick’s covenant, if you’re a data center and you wanna be welcomed by the community, don’t raise their power bills, all else equal, and be a good participant on the grid and don’t cause blackouts.

And this is possible today, and the reason it’s possible is if a data center shows up to town, by the way, chances are if a data center shows up and it’s a highly utilized data center, just that new load, the math probably already works out for that data center to support affordability. But to guarantee it, a flexible data center that’s willing to ramp down during peak demand and when the grid needs it, a hot summer day when everyone’s running their air conditioner, for example, that flexible data center can be connected utilizing the existing grid infrastructure.

David Roberts: Right. You don’t have to build new stuff. That, that’s the basic here is you don’t have to build new stuff. You’re using existing stuff.

Varun Sivaram: Well, to be clear, we will have to build a lot of new stuff, right? But Dave, what I would like is for the rise of AI demand to grow even faster than the new stuff we have to build.

Because that way, every new megawatt hour that a data center pays for on power bills, some portion of that goes to pay for the existing network- that it’s utilizing more highly, and that helps to, all else equal, reduce your bill. Now, there’s a bunch of other stuff that we’re doing where, like, weatherizing grids for fire risk, et cetera, and so your bill may end up rising nonetheless, Dave. But all else equal, the net effect of the data center coming, if it’s flexible, is to control that rate pressure. All else equal, it’s to reduce your rates, Dave.

David Roberts: And just it spreads the existing costs, the existing per unit costs of the existing infrastructure over a larger load, and thus kind of the per customer cost reduces.

Varun Sivaram: Nailed it.

David Roberts: This is… Did a whole pod on this dynamic. All right, all right. Well, final question. How long until it is standard for a data center to be flexible like you say? Five years, 10 years?

Varun Sivaram: In the year 2030, you can’t add a gigawatt of data center load to the American grid without it being flexible.

David Roberts: All right then, we will wrap up there. Thanks so much for coming on and talking us through this fascinating stuff you’re doing.

Varun Sivaram: You’re the best, Dave. Thanks so much for the time. It’s so great to do this with you.

David Roberts: Thank you.

Discussion about this episode

User's avatar

Ready for more?