CEOInterviews.AI
Start App
- John
Chief Technology Officer, Cloudflare

Cloudflare's AI Strategy with John Graham-Cumming

📅 Jul 10, 2025 TWiT Tech Podcast Network 36 MIN 457 VIEWS 73 SEGMENTS · 3 SPEAKERS
This week on Intelligent Machines, we spoke to former CTO, and current board member of Cloudflare, John Graham-Cumming, about how they deal with AI scrapers. For the full episode, go to: https://twit.tv/im/826 You can find more about TWiT and subscribe to our full shows at https://podcasts.twit.tv Subscribe: https://twit.tv/subscribe Products we recommend: https://www.amazon.com/shop/twitnetca... TWiT may earn commissions on certain products. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free shows, a members-only Discord, and behind-the-scenes access. Join today: h...

Questions asked in this interview

12
  1. 1:13I have to point out before the show that John, I was saying John are you on Tatooine?
  2. 2:10You have a habit of creating websites, I think, right?
  3. 2:29Well, tell first of all, what's the genesis of that phrase low background steel?
  4. 4:18That's why did you choose that date 2022?
  5. 8:55Cloudflare is interesting. Does Cloudflare hate AI?
  6. 12:13Is it in any way illegal to ignore robots.txt or is it purely a matter of honor?
  7. 17:34How does this vary from what Creative Commons also announced at the same time of a new kind of structure for identifying your conditions for crawling?
  8. 20:21Are there other mechanisms to make data available in a less intrusive way to the companies that exist?
  9. 27:12Something started as a blog and now is a gigantic company and your largest customer?
  10. 29:09I'm glad that O'Reilly kept it in print because it really, I mean the good thing about these is they're eternal, right?
  11. 29:27What would you like to do?
  12. 33:33What were you going to do if the whole world?
Leo Laporte 0:00 ↗
This is Twit. Now I want to introduce our guest. Old friend, dear friend. In fact, we were just reminiscing about the first time we met, which I think goes back to TechTV days more than 25 years ago. I checked and he's been on Twit a number of times because he's a very interesting fellow. He's done so many interesting things. Paris studied his work in school. Apparently one of the first Bayesian spam filters. Popfile. Yep. That was a collegiate era study. You might know him from his Geek Atlas, a guide book we interviewed him for some years ago that has 128 of course geek local. He also was instrumental in engineering an apology from the British government to Alan Turing's family for harassing and ultimately causing Alan Turing's death. But more recently the CTO at one of the best companies on the internet, Cloudflare, now on the board. He's retired somewhat. CTO emeritus I guess we'll call you. John Graham-Cumming. Welcome.
John Graham-Cumming 1:08 ↗
So great to see. Good to be back.
Leo Laporte 1:13 ↗
I have to point out before the show that John, I was saying John are you on Tatooine? Yes I am. This is actually a photo I took in 2009 of the set of Star Wars in Tunisia which is really near to a little town called Fume Tatooine. So George Lucas literally took the name of the local town for Tatooine.
John Graham-Cumming 1:36 ↗
That's fascinating. Were there tourists all about or was this deserted?
Leo Laporte 1:40 ↗
Actually there was almost nobody. I have pictures of just me and one other person wandering around.
John Graham-Cumming 1:48 ↗
That is so cool. Any graffiti? Any people any tributes people have left?
Leo Laporte 1:53 ↗
I didn't see anything. Honestly, at the time I think it was really a little bit remote and not many people were going. And actually some parts of the set were being reclaimed by the desert. The desert was literally just moving over it. So yeah, it looks like it's falling apart, but that was of course exactly how Tatooine looked. So I guess it wasn't. You can't tell.
We originally wanted to get you on because of a new website. You have a habit of creating websites, I think, right? Like it's just a fun thing for you to do.
John Graham-Cumming 2:19 ↗
Yeah. I mean, sometimes it's just like, oh, I have an idea. Worst case, I'll make a website and nobody will look at it and then, you know, who cares?
Leo Laporte 2:29 ↗
But yes, Leo was so happy when he read your latest post. He was about low background. What I just learned? Yeah. Oh, yeah. Yeah. Well, tell first of all, what's the genesis of that phrase low background steel?
John Graham-Cumming 2:40 ↗
Okay. So, the idea is that for some nuclear detection situations where they have to have very very sensitive sensors, you can't have any stray radiation. And in particular, if you have metals that are contaminated with radioactivity, they cause false readings. And it turns out that most steel in the world is contaminated with radioactivity because we set off a whole load of nuclear bombs in 1945 and beyond. And we contaminated everything because the way steel is made is by pushing air into the steel process, the Bessemer process. And so you end up with this slightly radioactive steel. But there is a bunch of steel that is not radioactive and that is steel from ships that sunk prior to the Trinity test. And so there is actually a business of going and pulling up wrecks from the bottom of the sea and getting the steel out of them because it's sort of purer. So the idea was of this low background steel was that there was a defining point at which the world changed, in this case Trinity test nuclear weapons contaminating everything, and we sort of have something similar with AI which is when ChatGPT launched there was a sudden explosion of AI generated content and it's not really possible in general to know did a human write this or not. So I had this idea to use this as a metaphor for it'd be nice to make a website that was kind of a collection of we know humans made this. So stuff that was created before the end of 2022 basically.
Leo Laporte 4:18 ↗
That's why did you choose that date 2022?
John Graham-Cumming 4:20 ↗
Well just because that's roughly when the real explosion happened once ChatGPT was launched. I mean, you really had this incredible Cambrian explosion or even Cambrian explosion if you're into Cambrian explosion of was spam human made or was there a version of machine made in a more mechanical way?
Leo Laporte 4:42 ↗
It's a very good question.
John Graham-Cumming 4:48 ↗
So, yes, there absolutely were. And in fact, one of the things that happened way back when I was doing spam filtering was a lot of humans tried a lot of tweaks to their messages and they tried to get through spam filters and actually that made it easier to filter them because humans were not very good at filtering the mission. Exactly. Yeah. Six. Yeah. And it turns out that in some ways the more the spammers tried to look like they were doing, the more obvious it was and actually it gave even more signal so it became easier in a funny way. But there were people using essentially the techniques that we used in spam filters to write spam, right, and try and get them through. And I actually a long time ago, I think about 2003, did a talk at MIT about this, about how you could take one AI machine learning system and use it against another, which of course people actually do in real life and you could actually fool the spam filters. So yes, there were things that were done. But for the most part spammers were pretty just spray and pray. You know, if someone's got a spam filter, well too bad. We'll get to the people who haven't got a spam filter.
Leo Laporte 6:01 ↗
Actually, it's an interesting question because you wrote Popfile in 2001. I have to think that spam is not any less prevalent in 2025.
John Graham-Cumming 6:16 ↗
No, the big change of course is if you can remember that far back, Leo, which I can just about remember, is that people were downloading their email onto their computers. Oh, yeah. Right. And they weren't using web-based email. And so the problem was the spam was actually coming into their machine. And so if you could, it was doubly annoying. You were wasting bandwidth on it. You were wasting bandwidth and time and maybe even pay per second internet access or whatever. And so it was a bigger problem. It's still there. If you go into your spam folder, I mean mine in Gmail is absolutely full of stuff, but it's all getting filtered out. So it's sort of irrelevant. You know where spam's not getting filtered out for me now? My Twitter DMs or my X DMs. I don't know if you guys have experienced this, but it is the one area where I find spam just pernicious and annoying in a way that I've never experienced. Like I probably have to keep my DMs open because I'm a journalist and sometimes I get tips, but I probably get like 20 to 30 spam DMs a day.
Leo Laporte 7:17 ↗
Really? I don't want to say anything, but I don't. Nobody wants you, Jeff. No, they don't. They don't. Your DMs are not open, Jeff, are they?
Jeff Jarvis 7:28 ↗
Well, good for you because it's quite annoying. Not only are my DMs not open, I haven't logged on to Twitter in more than a year. So, hey, that's healthy and wise.
Leo Laporte 7:39 ↗
So, John, what is on low background steel AI? By the way, the website is lowbackgroundsteel.ai, right?
John Graham-Cumming 7:44 ↗
Yes. What is on lowbackgroundsteel.ai? Actually, there's a bunch of resources that are we know are pre the Cambrian explosion. So there's a dump of Wikipedia. Pre-AI, right? And there's stuff from the Internet Archive, of course, because they've got a bunch of stuff like, you know, you can think about all of the media that was created before that. The Library of Congress has a photo archive. Project Gutenberg, you know, it's got a lot of things, pictures of actual people as opposed to generated people, books actually written by humans. In fact, there is a problem in publishing these days. A number of authors are saying, 'Please don't publish AI written books. We didn't use AI. We really didn't.' But there's no way to verify this. So Project Gutenberg makes sense. That's a good place to go. There's also the Arctic Code Vault that GitHub did. That was in 2020. That was every active repository on GitHub at that point. So it's stuff like that. And this website is just a Tumblr. So anyone can submit. If someone's got an idea, you can go to Tumblr and you can submit something.
Leo Laporte 8:55 ↗
Submit and reshare. I love it. Exactly. So we booked you because I wanted to talk about this but also because I'm a fan and I just love talking to you. But this morning I saw there was a fairly big announcement from Cloudflare. Cloudflare is interesting. Does Cloudflare hate AI?
John Graham-Cumming 9:14 ↗
No, actually we use it quite a lot. And if you look at Matthew Prince, the CEO, put out a blog post about this kind of stuff. It's not really about hating AI. In fact, we have a whole bunch of AI products, right? People can run things like Llama on our network. We deploy GPUs everywhere. Now what we were seeing was that from some of our customers who are publishers, they were getting worried about a fundamental change in the way in which the web has worked in a certain fashion. For a long time you had this agreement in a way between search engines and everybody else that it was okay for a search engine to come to your website and make available information about your website because fundamentally someone was going to click through and go to your website. And what you see with AI happening quite often is that there's no click happening. The website is scraped and an answer is given to someone searching on a search engine using the content that was scraped and there's no click. And that changes fundamentally the business model of the web because the business model was, yeah Google, you're allowed to make a lot of money from us because you send us the clicks. And that's changed and that is the concern that's come from big publishers.
Leo Laporte 10:35 ↗
And it's interesting because this kind of shift in the relationship between publishers and crawlers or search engines started before the AI boom in a way with the advent of Google's knowledge panel. I remember there was a whole big to-do some years ago about Google's use, I think through the acquisition of Genius for lyrics, how Google was scraping lyrics from other lyrics sites. It became a huge sticking point because companies were saying, 'Hey, we've done the work putting out the lyrics here and you're taking them and giving us no clicks or revenue whatsoever.' But it seems like AI has really accelerated that.
John Graham-Cumming 11:17 ↗
Yeah, because you get this longer response, right, which is sort of bringing together maybe multiple sources. And it's not about Google, right? There are many different examples of this with AI companies that are taking in content and giving you a response and some of them are good at linking back to the original content and some of them are not so good. So the idea from Cloudflare's perspective is because we sit in the middle between you the visitor, whether it's a human or a crawler, and the person putting something online, there's an opportunity to say what sort of control do you want to have. And we've done that for quite a while actually. So one thing we did was there's a thing called robots.txt right which is meant to tell robots, i.e. crawlers, how to behave. Some crawlers don't follow that. They will ignore it and so we're actually able to enforce it. So we'll say, 'Oh, wait a minute. We told you not to crawl that. Actually, we're going to block that. We're going to stop you from doing it.'
Leo Laporte 12:13 ↗
So it's historically difficult to do. I remember Steve Huffman, the CEO of Reddit, saying, 'We try to block, you know, they actually ended up making a deal to offer their content to OpenAI, but he said, 'We try to block this.' And these guys go around robots.txt all the time. So, can I ask you a fundamental question there? Is it in any way illegal to ignore robots.txt or is it purely a matter of honor?
John Graham-Cumming 12:41 ↗
I think it's a matter of honor. I mean, I'm not a lawyer, so IANAL as they say. I think the robots.txt is an RFC. It's an agreement on, you know, this is what you should not do, this is how you should behave. And again part of that sort of tacit agreement, right? 'Oh, you won't behave like this. If you don't want this thing crawled, we'll link back to you. Someone will click through when they've come to something that you have crawled.' So Cloudflare happens to be in a position where it's easy to help enforce robots.txt if you want. And then how can you tell it's a crawler? There's a way to technically know that. There's a lot of work. It depends what it is. Obviously some crawlers just announce themselves with a user agent. They say, 'I'm agent.' Or they publish an IP range. 'We'll crawl from this particular IP range.' The larger, very legit crawlers will give you information about it.
Leo Laporte 13:44 ↗
But those are all the people who adhere to robots.txt, no doubt in general.
John Graham-Cumming 13:46 ↗
Yes. Yes. So then you're going to start looking for other signals, right? So, one of the things that Cloudflare and other companies have had to do for a long time is figure out is something a bot or not. We because some people want to block it or whatever. Often when you get that check box that says 'Are you a human?' that's coming through Cloudflare. The thing called Turnstile that Cloudflare created is doing that. And what that's doing is actually looking at your browser and trying to determine if you are actually human. Are you behaving like a human? Is your browser lying about what it is? Because it's amazing actually when I was working on DDoS mitigation at Cloudflare, how many HTTP DDoS requests came from IE6. It hasn't been around in a long time. And what it was was the people doing the DDoS were just copying and pasting the same piece of code all over the place and it was one of these things where it was like, 'I don't think we're getting a thousand requests per second to this website from IE6.' So that's definitely bad behavior. So you're looking for anomalies like that. And then the other thing is at the scale that Cloudflare operates at, you have a lot of data about what normal behavior looks like on the internet. So you can use AI machine learning to make a determination of if this behavior looks like it's correct. So there's a tab on your Cloudflare dashboard that has AI. It says 'Manage AI Crawlers' and there's an AI audit and you can say, 'Hey, I don't want these guys. I don't want these guys.'
Leo Laporte 15:22 ↗
So this is better than robots.txt because you're actually looking for misbehaving.
John Graham-Cumming 15:27 ↗
Yes, correct. We can say, 'Oh, this bot is doing something which is misbehaving in this way, so make it go away for me or block it.' And ultimately what Cloudflare announced was use the HTTP 402 error code, which is payment required. To say to a caller, 'Yeah, I'll give you the content and here's the price.'
Leo Laporte 15:50 ↗
So it's almost blockchainy.
John Graham-Cumming 15:53 ↗
No, no, if you say blockchain, I have to leave.
Leo Laporte 15:57 ↗
Oh, the hives breaking out.
So this is called the new Pay-per-Crawl feature. Now it's not available yet to everyone. It's a closed experiment, closed beta. It is right now. There's a whole bunch of big publishers who signed up for it right from the beginning. Condé Nast etc have joined in this process. And to be clear, AI companies also signed up and said, 'Yes, we want to participate in this way. We want there to be clarity about what's allowed and what's not.'
John Graham-Cumming 16:30 ↗
Is there a separate negotiation? So Condé Nast says, 'We're going to charge you $10 for every make' and you say, 'That's ridiculous.' No. We go and negotiate a separate deal with Condé Nast. I can then inform you that I already I'm okay. You should ask the current CTO of Cloudflare how that works. But yeah, I mean absolutely obviously people can. I mean Reddit is one of the people who said they were happy about what Cloudflare was doing. They're in the press release and obviously Reddit also has signed deals with other people so clearly we're not going to usurp deals they have or whatever.
Leo Laporte 17:04 ↗
So if I'm a nasty AI company and I clearly am guessing they would do this and I say, 'Screw your deals. I'm going to find my way around.'

36 more exchanges in this transcript

Sign in free to read the rest of this interview. No card required.

Sign in to read the full transcript

Cite this transcript

APA, MLA, BibTeX
APA

John, -. (2025, July 10). Cloudflare's AI Strategy with John Graham-Cumming [Interview transcript]. TWiT Tech Podcast Network. CEOInterviews.AI. https://ceointerviews.ai/interview/638723/

MLA

- John. "Cloudflare's AI Strategy with John Graham-Cumming." TWiT Tech Podcast Network, 10 Jul. 2025. Transcript, CEOInterviews.AI, https://ceointerviews.ai/interview/638723/.

BibTeX
@misc{john2025_638723,
  author       = {- John},
  title        = {Cloudflare's AI Strategy with John Graham-Cumming},
  howpublished = {Interview transcript, TWiT Tech Podcast Network. CEOInterviews.AI},
  year         = {2025},
  month        = {jul},
  url          = {https://ceointerviews.ai/interview/638723/},
  note         = {Speaker-attributed transcript with timestamps}
}