Back
Simon Mcdougall
Chief Compliance Officer, ZOOMINFO TECHNO INC

Generative AI, Privacy and Data Protection | Oxford Generative AI Summit 2024

🎥 Dec 13, 2024 📺 Oxford Generative AI Summit ⏱ 58m 👁 78 views
Oxford Generative AI Summit 2024: oxgensummit.org Speakers: Elena Simperl, Director of Research, ODI & Professor of ...
Watch on YouTube

About Simon Mcdougall

Simon McDougall, Chief Compliance Officer at ZoomInfo, has spoken publicly about the intersection of generative AI, privacy, and data protection. At the Oxford Generative AI Summit 2024, he argued that if personal data used in training a model is not recreated in its outputs, he does not see a meaningful privacy risk, but he expressed greater concern about intellectual property, copyright, and the future of work. He also noted that supply chain risks in AI create a "fog" that traditional data protection frameworks are not well equipped to handle, and that the GDPR can "creak at the seams" when applied to generative AI. McDougall stated that ZoomInfo engaged with Anthropic early to obtain assurances about its AI models. McDougall has also described ZoomInfo's data collection practices, stating that the company collects business contact information such as job titles and contact details, not sensitive personal data. He said ZoomInfo sends notification emails to individuals when it creates a record, regardless of local legal requirements, and provides multiple opt-out methods. McDougall previously served at the UK Information Commissioner's Office (ICO), where he focused on technology and innovation, and has discussed the importance of reconciling privacy regulation with competition policy, citing the UK's Digital Regulation Cooperation Forum as a forum for exploring synergies and tensions between those regimes.

Source: AI-verified profile updated from Simon Mcdougall's recent appearances. Browse all interviews →

Transcript (38 segments)
M
Moderator0:02
Thank you very much for coming back after lunch for this particular panel, where we're focused on the topic of generative AI, privacy, and data protection. Another important area of regulation for generative AI. We've got another important set of questions about the intersection between this fast and bold innovation and how data protection law, which has been around for a long time—you can see Simon and I have been around for a long time because we had hair when we started off in this profession over 20 years ago—is now facing the challenge of how to apply this long-standing set of rights and principles to this new emerging context of the technology, and how we can reap the benefits of the technology but still keep those rights and principles in data protection intact, which the public still value and think are important. In the panel today, we're going to look at some of the questions about risks, safeguards, and mitigations, and how they can be balanced. We're also going to think about data protection regulation in the context of other forms of AI regulation, and how data protection law is going to remain relevant and how it may need to intersect with new forms of AI regulation, e.g., the EU AI Act. We've got a great panel here today. I'm going to get them all to introduce themselves in a minute. We've got perspectives from business, from the regulator, and from an academic and someone who's working in a think tank as well. Really looking forward to the discussion. Also, we actually had a panel here last year for those of you who attended the 2023 Summit on generative AI and data protection that I hosted as well. I think what's going to be interesting to reflect on is what's moved on in the last year, how sharp some of the questions have got, or have we got some solutions to some of the things we were discussing last year as well, because I had the pleasure of chairing the panel last year. To open this up, I'm going to go to each of the panelists to ask them this first question, and then they'll introduce themselves and talk about their current work and role in this topic. The first question for the panel is that data protection really intersects with a number of different parts of the life cycle for generative AI in terms of the inputs, the enhancements, and the outputs, and in development and deployment. So from your different perspectives, where do you see the most significant data protection risks, and how does data protection governance need to evolve to address those risks? I think we'll go to Sophia first from the ICO, from the regulator, to kick us off.
S
Sophia2:59
Thank you. Not a difficult question at all. Thanks for inviting us. Just to briefly introduce myself, I'm the Group Manager for AI Policy at the ICO, the data protection regulator. I manage our team of policy and technology experts who are tasked with developing the ICO's emerging positions in terms of AI. We monitor the technology as it evolves and it surfaces new issues. We are updating our guidance, we lead our cross-regulatory cooperation, so we represent the ICO at the Digital Regulation Cooperation Forum, which I think was one of the legacies of Simon's work at the ICO as well, and our cooperation with regulators in the UK and internationally, and a variety of other stakeholders. We're producing guidance, for example, at the moment with the EHRC, the Equality and Human Rights Commission in the UK, and we continue to engage with counterparts internationally at EU level as well. Going back to the question around what are the most significant risks that generative AI surfaces in terms of data protection, I would say it's one of those iterations of AI that has presented most challenges for data protection because there are things that are implicit in how it's being developed that are contradicting or creating tensions with requirements under the GDPR. More specifically, the issue around information rights. Under data protection, you need to be able to fulfill information rights where data subjects are making requests such as exercising the right to erasure. Someone may want to ask you to remove their data from your model or from the training data, but in the context of a model that takes days to train, millions in terms of computing capacity, obviously that comes with market forces that make it practically really difficult to respond to information rights requests on time. You also have accountability issues. There is a lot of ambiguity around who is the controller, who is the processor. What I'm trying to do right now is nudging the conversation towards our generative AI consultation. Back in January this year, we launched a consultation on that specific point—the intersection of data protection and generative AI—and we chose to focus on those five areas where we see the most risk because there is a lot of regulatory uncertainty at the moment. Those areas are: the information rights piece that I just mentioned; accountability—who is accountable for what; purpose limitation—the fact that under data protection you need to be specific in terms of why you process data, but what we're seeing in practice is a lot of ambiguity or confusion from stakeholders around what is a clearly defined purpose, in a way that is also intentional with the notion of general purpose AI; you also have the fundamentals of the technology itself, like the fact that you consume a lot of data and that creates tensions with data minimization; our consultation also touched on accuracy—what our expectations are in terms of accurate outputs and inputs; and I think there is also a question around what are the downstream implications of this technology, because the more we move into an area where those models are being used to power AI agents, we will have to start to think about what are the implications for Article 22. So I can go on forever, to be honest, so I'll just pause here.
M
Moderator7:02
Thanks very much, Sophia. I think that gives you a really good insight in terms of what the regulator is considering in terms of those key risks and areas of regulatory uncertainty, and also obviously, if you're new to some of these topics, go on the ICO's website—there's lots of information there about the ICO's current positions that you can start to look at as well. I'll turn to you now, Elena, to give your perspective on these data protection risks and some of the changes to governance and organizations.
E
Elena Simperl7:29
Thank you. My name is Elena Simperl. I'm the Director of Research at the Open Data Institute. I'm here with my colleague Emma Thwaites, who just took a picture of me. She is, among others, our Director of Policy, so possibly more of an expert to respond to some of these questions. I am also, for my sins, a Professor of Computer Science at King's College London, and so I've been working in AI. I'm going to talk a little bit about the work we've been doing at the Open Data Institute as part of the Data-Centric AI program. This is a program of work we launched at the end of last year. For those interested, there's a 15-point plan of things that need to happen to create a healthy, trustworthy data ecosystem for AI, and that plan is not just for the ODI and shouldn't be just for the ODI to tackle, but it will be a multistakeholder effort. First of all, let me just congratulate you on the question, because often when a very challenging question is asked, it is imposed in such a nuanced way. Yesterday, Tuesday, we published a first version of what we call our AI Data Taxonomy. In that taxonomy, the aim is to unpack what we mean by the 'no data, no AI' headline that many of us have embraced over the last year. You are already making some of these distinctions: inputs, enhancements, outputs, also looking across the whole AI lifecycle. All these different types of data sets and the way data is collected, managed, assured, governed in all these different scenarios is different and tackles important challenges. Whatever we do, we will need to genuinely put an effort into understanding the characteristics and the challenges associated with—I'm just going to give some examples—data that is in the public domain and has been used to pre-train a foundational model; synthetic data that has been created because the actual data can or should not be shared; fine-tuning data that possibly a different organization has created and used on top of an existing foundational model to tailor it for a particular purpose; prompting data that a consumer is adding to a consumer-facing tool like ChatGPT to ensure it responds to whatever needs it has. All these different types of data come with different challenges. The processes, technical and otherwise, to create, maintain, and assure these data sets are different. So without making your job more complicated, because you will need to look into all these different scenarios individually, this is also an opportunity for the providers of privacy-enhancing technology, data protection technology, to come up with really interesting innovative solutions to make a difference. I will just say a few things now with my scientist hat on. In terms of data protection and inputs, I think there's a lot of research that needs to happen. Techniques like model unlearning, for instance, are supposed to help with some of those challenges you mentioned around managing to make those models more mindful of privacy without having to actually bend the weights and spend another hundred million on retraining. On enhancements, we've published a public policy intervention in June around rights, which includes people's rights but also includes rights of those people involved in the supply chain of a data set. In particular, if you're thinking about safety, everyone's talking about AI safety. Safety testing includes a lot of contributions from gig workers that work on digital platforms that, for instance, are supposed to make judgments on questions around harms like toxicity, racism. They do that using their knowledge and experience, so they add their own data to these data sets that are then part of fine-tuning and safety tailoring these models. On outputs, I think there is an opportunity and a risk. I don't know how much the general public understands that at the moment we are putting our knowledge, our data, and our experience in the form of prompts into a handful of proprietary digital platforms. As an academic and a tech optimist, I would like us collectively not to make the same mistakes we've made with previous digital platforms, because we're just prompting into ChatGPT—hundreds of millions of users, several times a day—we're putting that data into these systems, and that data is not available in the same form across the ecosystem.
M
Moderator13:24
Thank you very much, Elena. Just to pick up on some of your points there around potential technology solutions and what we call privacy-enhancing technologies to address some of the concerns that we've got, particularly how can you technically erase personal data connected with a generative AI model because of the way it's constructed in its tokenized form? How confident are you, as a computer scientist, that we may be able to find some of these solutions which can still give effect to people's data protection rights? Where do you think we are on that road?
E
Elena Simperl14:04
Okay, so let me start with a disclaimer. I'm not working in privacy-aware AI; I'm working in symbolic AI, that's the old-fashioned AI that didn't seem to work but now everyone needs. But I think it's quite challenging, and it will very much depend on how that model was trained and what sort of safety testing was applied before it was released. In some cases, depending on the type of data we're talking about, it is going to be hard to almost impossible. Correcting the models is a huge challenge, both from the degree to which the end user or even the developer is able to intervene in a predictable, deterministic way, but also because it is quite hard to locate in those models, in those weights, the exact point where this happens. I'm speaking to an expert audience, but just in case I can't make my point clear: with deep learning, imagine you have a fruit salad and a smoothie. Google search, the AI deep learning AI that is in Google searches, is like a fruit salad. If you want to take a piece of apple, you will take the piece of apple; bits of it will have spread in the salad, but you still can take most of it out. In the smoothie, which is the generative AI part, good luck with that. You will be able to find some bits of the skin, but it's much, much harder. Still, I think there are some ways, depending on the training data and depending on where else that data is available, to isolate those effects.
M
Moderator15:51
Thank you, great. We may have to come back to this next year. I think you heard it here first anyway: regulating generative AI and data protection is like regulating smoothies. Trademark that one. Thanks, it was really great. It's not mine, by the way. Great insight and follow-up to that. I'm now going to turn to Simon, who I used to work with when I was the Deputy Commissioner at the ICO before I came out to work in the private sector. Simon is now working in the private sector for a very data-focused company. You're really dealing with some of these AI governance issues at the coalface in terms of products and services you're launching, Simon. So look forward to hearing your perspective.
S
Simon Mcdougall16:41
Thanks, D. I'm going to try and find some food-related metaphors to throw in as we go along. Yes, so I spent 20 years in privacy, then three years at the ICO, set up the technology team there, the innovation team there, and what I think was the first specialist AI team in a regulator in the world. If anyone can find one before 2018, then let me know. I now work at ZoomInfo, which is a B2B data company, $1.3 billion of revenue, listed on NASDAQ. No one in the UK has ever heard of it because it's B2B data. So, I think the first thing I'd say is that underlying all of this is the fact that privacy and data protection is old now, and the GDPR is old now. That has strengths and weaknesses. That means there's a lot of guidance out there, there's a lot of case law now, and there's a lot of people who have worked in it for a long time. The downside is that I think generative AI in particular makes some bits of GDPR at best creak at the seams, at worst break a little bit. So that's something I'll come back to as we go through these various areas. Most of the best points here, as the last person on the panel speaking, have been taken by everyone else because everyone else has made great points. So I will touch on three things. Firstly, in terms of significant risks, risks around inferences. Again, it's a variation on the smoothie point. There are some privacy regulators now who are saying that there's actually no personal data in a model—not an AI system, but just in the model—because you have training data and you have outputs, but not in the model. This is really hard to grasp if you're used to the old way of thinking around privacy where you had some data and you put it in a database and then you had some outputs, and all the way through it was chunks of apple. Now you can have situations where you could have systems producing data with personal data—defined as data that relates to a living person—almost out of the ether. That might be data which obviously has risks to individuals, risks of discrimination, all those kinds of things. So inferences is hard for privacy to accommodate. That's a big risk: how we manage that in generative AI. Secondly, automated decision-making. We've already touched on this in some points. The thing I'd add to that really is if we think back to Nigel discussing this morning around co-intelligence and different kinds of decision-making, privacy regulation leans a lot on humans in the loop, and that gets harder when you have really great assistants that exacerbate automation bias. If you have co-intelligence and co-decisioning, at what point is the human in the loop or not in the loop, and at what point are they one foot in, one foot out, doing the okey-dokey? So that's going to be a really interesting area to watch over the next few years. There are some not-quite-regulated safe harbors: if you have a human in the loop for an automated decision, that kind of takes you out of some regulation, gives you some protection. But is that really going to stand up? I think we're going to see some breaches where bad things are happening, and normally there's a human in the loop, and then when you actually look at the bad thing that's happened, you go, the human was pretty useless. Thirdly, supply chain risks, again touched on a little bit. Privacy works about data controllers, data processors. It's fairly set ways of working out liability. You have joint controllership; it's kind of we know how that works. I think for lots of organizations, especially in the private sector but also the public sector, what you've suddenly had is a whole new world of supply chains pulled upon you. Sometimes imbalances of power in who's working with who, and murkiness at both ends. You don't know what's happened with the training data upstream, and if you're actually building your own generative AI, you don't actually know what the deployers are going to do with it downstream. So there's this kind of fog around here which you wouldn't get in a normal supply chain, and again that's hard because the data protection world thinks in more simple terms about this. So risks are coming through, but I'm not quite sure whether the GDPR is built to handle all this.
M
Moderator21:14
Thanks so much, Simon. Just to drill down into that supply chain issue right now: how do you think companies are grappling with it in terms of the contracts that they're having between the controllers and the processors, or this concept of joint controllership which the ICO has perhaps put forward as a concept that may need to have more prominence in the generative AI space?
S
Simon Mcdougall21:36
I think joint controllership is context-specific, but it can be useful. Even in the last year, we've seen a lot of technology platforms' contractual clauses start to stabilize and standardize a bit. I've got to say, in the rush we had after ChatGPT 3.5, where suddenly everybody in the company wanted to do stuff with generative AI, there was lots of commercial pressure to engage with folk. Now, we were lucky at ZoomInfo; we engaged with Anthropic early on and sufficiently early, but we were big enough to actually have a meaningful conversation. But I know totally lots of other people just—the generative AI companies were selling ice cream in sunshine, and if you didn't want to take their—to use a food metaphor—pressure, and effectively if you didn't like what they were saying around their information security or the assurances they gave around what data it was trained on, there were five other companies you could go and deal with straight away, so they weren't getting straight answers back. I think that's kind of leveling out over time a little bit. So I think we're getting better on supply chain, but I don't think we're anywhere near sorted on that. And when you get on to downstream and deploying, I think that's even harder. So if you are, as people are building their own agents and adding their own flavors to the open source and closed source models, understanding what that might be used for downstream is especially depressing.
M
Moderator23:22
Okay, thanks so much, Simon. I'm going to move on to a different topic now, which I think a number of you would have been hearing about, and this is actually focused on the inputs and the training phase for generative AI models. Obviously, given the scope of data that goes into the training process, personal data will be harvested during that process. The companies obviously don't really target personal data in building a number of these large language models, but personal data is involved because of the sweep of the data. So we have a question under the GDPR, under Article 6, in terms of what is the lawful basis for that process of AI training. It's an interesting question because obviously for those of you who know a little bit about GDPR, there are a number of options in finding a lawful basis in Article 6. One is consent, another one is necessary for the performance of a contract, and another one is whether it's necessary for a legitimate interest pursued by the controller, e.g., the company, or in the interests of other third parties. But it also needs to be subject to a balancing test against the rights and freedoms of the individual as well. The discussion and the debate that has started to emerge has really looked at what lawful basis is feasible. Can you feasibly get consent for large-scale training of AI data? Obviously, that could be challenging. It could be challenging as well whether it will always be necessary for the performance of a contract—a contract won't be in place between the individuals and the company undertaking the AI training. So the focus has actually fallen on this legitimate interest condition in the law. That's why you've seen some big news stories recently in relation to Meta of PA, the training relating to their models in the EU, because of the challenge they've got currently engaging with the European data protection regulators in terms of whether that is actually meeting the legitimate interest condition in the GDPR. So we're seeing quite an important moment now in the trajectory of generative AI about how it's interacting with data protection law. So my question to the panel really is: what views and thoughts they've got about this tricky question of organizations relying on legitimate interests, and what they might need to consider? And of course, it may be in different use cases and context as well. So Simon, could you start us off on this one and talk about your perspective on it?
S
Simon Mcdougall26:09
Yeah, absolutely. I think from a—I'll say the sensible thing, then I'll say the controversial thing. The sensible thing is that the good news is that really good regulators have given this lots of thought and produced really thoughtful analysis of this. In terms of legitimate interest in generative AI and lawful basis overall, both the ICO's guidance is very good on this, and also the CNIL, the French privacy regulator, produced a series of documents called 'How to'—I don't know why they're called 'How to' documents, but that's what they're called. So if you look for CNIL and 'How to', you'll find them. Both have the same kind of thrust: you could use legitimate interests for processing personal data and training data, with a bunch of sensible provisos and whatnot. And very recently, there's another draft document out from the European Data Protection Board which talks about legitimate interests overall. So smart people have done lots of work in this already, and if you're looking to find things on there, you can get there. They worked really hard to square the circle where, if you had a minimalist reading of the GDPR, you could easily argue the other way and say that you couldn't use legitimate interests for training data, at which point there's no generative AI in Europe. And in terms of economic growth and just in general how much it's being used, everyone's between a rock and a hard place. So I think that's regulators being smart and proportionate and trying to make the best of what they've got. The controversial point is that I don't really care. What I mean by that is, if my personal take is, if somebody uses personal data in a training data set to produce a model which doesn't have that personal data in it—and this is a massive 'if'—and the outputs don't recreate that personal data somewhere, then in terms of the privacy infringement there, I can't get very excited about that. I can't actually see meaningful risk of harm to the individual there. If I'm a coder or a poet and the training data included all my code and poems, and now I'm out of a job because of that, then there's real harm. So in terms of IP, copyright, and future of work, I think there are big issues with use of that data. But purely from a personal data point of view, I'm like, why are we talking so much about the law? I get it's an important part of the regulation, but I'm not actually seeing who we are protecting from that. Anyway, I'll stop there.
M
Moderator28:58
Thanks, you kicked off really well. In terms of that risk point, is part of your message there that we really must consider each use case on its merits as well? Because I suppose the concerns which have risen particularly about the social media companies, the particular maybe heightened risks about using social media data for training as opposed to other data—would you differentiate between any the context?
S
Simon Mcdougall29:20
I personally only in terms of the risk to the outputs. I'll be very hypothetical here because in reality you can't completely divorce those two, but I'm just saying that within our little data protection bubble, you get lots of very animated discussions that are very legalistic about this, and I'd rather we focus on risk of harm because there's plenty of risks out there: personal risks, societal risks, environmental risks. So I feel we take up too much airtime on the legalistic side. It doesn't put the privacy community in a good light in the bigger picture.
M
Moderator30:01
Okay, thanks. Sophia, over to you. The ICO has put some guidance out on legitimate interest as Simon talked about, so it'd be great to hear about the ICO's approach to this question.
S
Sophia30:08
Yeah, I think this is a really interesting discussion. Legitimate interests in terms of generative AI is an area we looked at as part of the consultation, but we looked at that as a lawful basis when you web-scrape data. So there's a difference in our approach when it comes to first-party data—for example, user data that Meta may want to use to develop their models—and a difference between that and the context in which data brokers or developers themselves are basically web-scraping social media or other unidentified sources, because they rarely do actually disclose what they actually scrape. So starting with the first-party data context, in terms of legitimate interest when you're using that basis, you still have to be able to fulfill people's information rights, but those rights are not absolute to the extent that they are, for example, when you're using consent as your lawful basis. When you use consent, you have to ensure that consent is revocable, which we do say is really difficult to actually enact in a generative context. That's one of the reasons we turned toward legitimate interest as the most likely available basis to train generative AI. I think we have been public about the fact that we engaged with Meta on their use of first-party data, and our experience has been that it was a really productive dialogue. In none of our engagements do we end in a place where we can actually provide some kind of certification for the lawfulness of the processing in perpetuity, but we did manage to improve aspects of transparency towards data subjects, and that's why people were alerted to the fact that Meta was planning to do that. I think that's a direction that regulators should be heading towards: do a lot of upstream engagement so you kind of nudge a company into compliance before bad things happen. So that's in terms of first-party data. When it comes to web-scraped data, I think the harm we see is the loss of control of someone's personal data. I think we have to be realistic here: when someone scrapes your data to develop a model, these institutions most of the time will not only use your data for that specific model. There was a press release by one of the big players in the space who was saying that they have trained millions of models in the last couple of years. What we're trying to do as regulators is to translate the legislation in a context in a way that somehow mitigates the risk of too many gateways that enable abuse of the legislation, if that makes sense. We don't—perfection is a really tricky thing to achieve in data protection, but we're trying the best we can.
M
Moderator33:21
Thank you very much, Sophia. I think that's really useful. Flowing on from what Simon said about how you can differentiate the use, and actually learning from how the UK regulator has approached this question of how Meta have used social media data in their model and the improvements that the ICO got from that process, I think that's interesting to learn and contrast from what's happening in the EU as well. Elena, I'll turn to you. I think you were touching on some issues about inputs when you spoke previously, but have you got any particular views on how this GDPR question could be approached from your distinct perspective?
E
Elena Simperl33:59
Yes, following from what you were saying about the fact that lots of these conversations are very technical and we're focusing a lot on the law, I think I would agree with that. I think even if there is this guidance, even if this dialogue with the relevant technology providers is happening and is encouraging, public perception and public trust are different. So it's nice that a company is now more transparent about the fact that they've started to acknowledge that they're using public data to train models, and maybe strictly speaking from a legal point of view they are within their rights to do that, but that does impact the trust that the public has, including myself as a user of some of these platforms, not just in the provider of the service but also in the regulatory and political actors who are supposed to protect my rights. So maybe some of that could be improved with more awareness campaigns, with more digital literacy and data literacy. But if we're thinking about harms down the line, misinformation is a huge issue, and these models are perpetuating and exacerbating misinformation. A lot has been said about the role of these models in elections, in how we engage in the democratic process. Some of that can or should be regulated—I'm not going to comment on that—but I would encourage us to think beyond the law, because there's what the law says and what the regulator can and does enforce to show that regulation is working in practice, and then there's the public perception and trust. As a scientist, I think there could also be a huge opportunity to think a little bit more disruptively. We are talking about a time when the way in which data is used to innovate with technology is changing. Lots of the assumptions are changing because of how the technology works. Isn't that an opportunity to actually think a little bit more about the whole architecture and technology stack of these platforms? Why do we have to have the sort of data-scraping behavior and centralized data controls that these platforms have been built on? Could this not be an opportunity to actually think about different architectures that give people from the very beginning more agency and control over their data? The two things aren't directly related, but there is a relationship, and with everything in technology, timing is everything. If everything has changed and a lot has, this is an opportunity to actually rethink the whole how personal data management is implemented in these platforms as well.
S
Sophia37:44
Can I come in? I probably should have clarified more about the part of our engagement with Meta. One of the things we achieved is basically, apart from the fact that Meta contacted their customers to tell them they're planning to do that, you also had an easy way to object to that processing. So I personally have objected to that processing. So in a way, we empower data subjects to just stay out of it basically. And in terms of, I really support the more kind of disruptive and innovative way of looking at the problems we're facing, because for a lot of what we're seeing at the ICO, we're aware of the fact that law may not be the actual practical solution. It may be the case that we need to think more creatively around new institutions when we want to introduce, for example, data trusts or data intermediaries, or what kind of innovation do we want to promote and tell developers that we actually need, so we don't impede their experimentation with this technology.
S
Simon Mcdougall38:50
I do think building on that point, I do think there's a risk that when we talk about regulation here, we can over-index onto the large technology platforms. Now, they matter because they're massive, they matter because the precedent you're setting through, say, the engagement with Meta sets a precedent, so I'm not saying they're irrelevant—far from it. But again, we're looking at technologies which can be reused and redeployed, fine-tuned or used for a whole range of purposes, and there's going to be lots of other companies that are using them which right now probably have far less sophistication as to how they're deploying these things than the large tech platforms, which are very well resourced and have had this kind of technology in their lab for years, as opposed to a few days, a few weeks, maybe a few months at most. So I think that's one of the challenges in this space. And again, to Sophia's point, it's not really about the GDPR regulation. You could say, well, all those data controllers have to form a legitimate interest assessment—that is a technically true statement. What you really want them to do is say, 'Oh, we need to think about why we're doing this, if this is the right way to go about using this technology, and what are the interests to us and the risk of harm to the people.' I just think think it through. So I think the next stage is making sure that these good messages get distributed.
M
Moderator40:12
Yeah, thanks. I think it's really an interesting discussion about where this nuanced position currently is in the middle, where we can see the technology can continue to grow and evolve, but we are actually able to build in practical safeguards. Sophia actually referenced one there in terms of the right to object to that processing by Meta. So I think all of those points are really interesting. And then going to the bigger picture, I think it's a great wider discussion about also the policy aspect of maybe we don't longer term have to accept the status quo. We've got discussions about, you know, Nigel was mentioning this morning about the utility of smaller models which may therefore use less personal data, whether data protection regulators can actually enforce technically their way through the law. But I think data protection regulators can have a voice in this conversation about it being a sort of data protection by design conversation. So I think some really great points there on the panel. We want a bit of time for questions, so I'm going to do a reasonably quick final question to finish off to the panel, which is just to reflect on the fact that there are other forms of regulation emerging. We've got the EU AI Act obviously coming towards implementation, and that's focused particularly on safety risks, particularly on high-risk models, etc. So it'd be good to get the panel to reflect on what's the position: do you think data protection should take against other forms of regulation which are also going to be important related to generative AI? Elena, could you just give us your thoughts on that quickly?
E
Elena Simperl41:51
So I'm going to focus on a very specific point of the EU AI Act that I know a little bit about, and that is the transparency requirements, and in particular transparency requirements about data. Now, with foundational models, one area that we have seen flourishing where there's a lot of innovation and a lot of potential is synthetic data. Synthetic data is used in situations where the actual data can or should not be shared, but it is also extensively used in creating fine-tuning and other types of data that would be too costly to create. Well, in some of those situations, actually collecting data, including personal data, genuine data, is quite vital. So we need to distinguish between these different use cases, but more importantly, we need to be clear about what things like a sufficiently detailed summary of that synthetic data looks like. So yes, there is a law. I mean, everyone says that the AI Act tries to regulate AI as if it were a microwave or a product safety law at the end of the day. But still, it does create more awareness on the needs for transparency across the AI lifecycle, and we need more guidance on what's appropriate and useful for the different types of data sets, in particular synthetic data sets, which are playing a huge role because at the moment that's the only way or the cheapest way to deal with the fact that model providers are running out of real data.
M
Moderator43:48
Okay, thank you. So that's quite a good point: the EU AI Act is going to ask about transparency which maybe GDPR won't ask, and that can be complementary. Sophia, what are your views? Working in a data protection regulator, I'm sure you have a view on the relevance of data protection, but how do you see these different forms of regulations sitting with data protection?
S
Sophia44:07
I mean, the AI Act itself says basically that it complements and works alongside the GDPR. So in a way, there is no tension there. Transparency is already a legal requirement in data protection, but the kind of transparency that has data subjects as its focus—data protection regulators will never tell you to release all your training data for everyone to see because that is disclosure of someone's personal information. I think a really useful addition that the EU AI Act provides is this notion of systemic risk, which is something that we have discussed—the systemic risks that AI gives rise to in the context of the DRCF, for example—but data protection is not really the best vehicle to address that level of risk. I think the EU AI Act has the ability to really empower data protection authorities in the EU because it really makes explicit requirements that DPAs should have anyway when it comes to the processing of personal data in the context of AI development and use. One advantage of the GDPR compared to the EU AI Act is that point around being a principles-based legislation rather than a product safety one that is more rigid and not really amenable to change or interpretation. And the fact, I would contest, that data protection as a framework has the capacity to look across the supply chain. So we're not interested just at the moment that you introduce a model into the market; we're interested in everything that comes before that.
M
Moderator45:47
Thank you. I think that's a really good overview of how the two can complement each other and what each of the pieces of legislation brings. What's your view on this, Simon?
S
Simon Mcdougall45:54
So I'll mention one other area where I think the current wave of AI stress tests the GDPR, which is around what we'd call in privacy world sensitive data, special categories of data. In the old days, the simple answer was don't collect it—don't collect information around race or sexuality or health—and then you'll be all right. Now, obviously, it's very possible and often in unexpected ways to infer those qualities without collecting that data. And if you haven't collected that data in the first place, there's no way to test the outputs of your generative AI to know whether the output is biased. So you are then stuck in a little bit of a whirlpool: do I collect this data or not, because I've been told not to, but actually my AI might be producing biased outputs which is worse. So that's an issue, and again it's an issue which regulators have grappled with. The DRCF, the group Sophia mentioned before, has worked on this, and European regulators. But it's another example where I think the GDPR is a bit broken. I think the GDPR will begin to be rewritten next year—new regulation in the UK or the EU or both. I should be very clear that I'm talking about the EU right now. If you are terribly dull like me, you read the mission letters. Ursula von der Leyen writes a mission letter to each commissioner saying this is what I want you to do. A theme throughout these letters is that Europe is overregulated and needs to deregulate for growth. That comes up again and again, and the GDPR is explicitly mentioned along the way as being something to—it says things in a very European way: 'consider how GDPR can best regulate, yada yada.' It's on the table. I'm not saying it's going to be a complete rewrite from start to finish by any means at all, but I think if something's going to give in the European regulatory framework, it's not going to be the AI Act because that's too new, it's not going to be all these other data acts, it's going to be the GDPR.
M
Moderator48:09
Okay, well that's an interesting note to finish on. Let's throw out and see if anybody's got any questions for the last few minutes.
A
Audience Member48:19
I've really liked your point about being able to use engineering as another tool in addition to law and standards. To what degree can we as engineers get ahead of this cycle and insist upon a higher standard? I know Karl Friston and his work on active inference is one way they're showing great results—90% less data for the same result. How can we as engineers help?
E
Elena Simperl48:52
I'm happy to try. I was told that I'm not supposed to touch the microphone because then something would change, but anyway. So I think we can and we should. What I observe in my own practice or when I teach is that unfortunately at the moment there's still a bit of a gap between two communities: the engineers or the AI scientists who sit typically in a computer science department, and researchers or practitioners in responsible or trustworthy AI. In that latter area, there are a number of tools like impact assessments and the likes that have been around since 2017, 2018, so they're quite established already. There are things like data cards and model cards, a number of transparency and accountability tools, ways to check and analyze your data before it goes into a model. But those frameworks aren't very much used by engineers, and there's lots of research that shows that even when they are used, they're not very well understood. I'm going to give one example, which is explanations. There's a lot of research that shows that it's actually the engineer, the expert, who tends to over-rely on those explanations even when they're basically random. So a lot of work to be done. I think the way forward would be to take these responsible AI frameworks and tools and test them in practice. I think companies and managers of engineering teams should make an effort to see how these toolkits apply to them and what small measures they could put in place now. I'm going to give one example very quickly. Nigel mentioned small models. Now, with small models, even if they're deep learning, you have a different control, you have different remedies that you can put in place. And as most companies will in time use these sorts of models or retrieval-augmented generation methods, where accuracy and control over the data and the provenance are different, if we manage to take this responsible AI toolkit and really apply them in practice, I think a lot of these issues will not disappear but will have fewer impacts down the road.
S
Simon Mcdougall51:47
Does anybody else want—I'll make a—I'm terrified of the mic. I'll make one small practical point. I think the gap between engineers and the compliance, risk, and legal people has closed a lot in the last 10 years. I can remember when some of the big tech firms started having privacy engineers—that was revolutionary, and that was only, I'd say, 2012 or so. But I still think there's a lot of good operational stuff to be done with engaging between different groups in corporates. So within the risk management compliance functions of ZoomInfo, we send people to the offsites of the other teams, we have people just sit on the Slack channels. When ChatGPT 3.5 got going, most of the activity flowed through three or four Slack channels within our organization, and so we just put people on the Slack channels and they chipped in along the way, and that's how we generated the first generation of guidance on this stuff. So yes to everything that Elena has mentioned, I think there's also just a matter of the interpersonal side being key. So just getting in there.
S
Sophia52:54
Okay, I'm not going to go too close to the microphone. I agree with Simon. I think basically getting this cross-disciplinary conversation to work in a productive and well-informed way is a really complicated exercise, and we're actually learning it as we go along. A lot of the time when we're discussing new policy positions, we will have our technical advisers in the room with legal advisers. I think as long as there is goodwill from all the teams involved and we are humble and accept that there are certain things we do not understand, I think we can move forward with this, but it is going to be a learning process.
M
Moderator53:42
Okay, thank you very much. Do we have a question over there?
A
Audience Member53:57
Thank you. You sort of touched on this a little bit towards the end, but I essentially wanted to ask if you think that the current explosion in generative AI that we've seen in recent years could have happened in the current GDPR landscape that we're seeing here in Europe, or whether, as you sort of alluded to, that's stifling the ability for such new explosive innovation to occur.
E
Elena Simperl54:36
That's directly to you. First stab.
S
Simon Mcdougall54:45
Got it. So my counterpoint to that would be that Google has been sitting on this technology for many years, and the reason why they haven't deployed it but OpenAI did wasn't necessarily because they were operating in Europe; it was about more general considerations and business motives that they had not to release it. So yes, there is an interesting tension between more regulation and less innovation, and the Draghi report certainly emphasizes that strongly enough, and as Simon was saying, apparently the message has been received. Let's see how it's going to be implemented. But it took a company like OpenAI with the structure it has, with the investors it has, to put that technology out there. Lots of the incumbents already had it and didn't do it for many, many years for all sorts of reasons. So what I'm trying to say is that the relationship between innovation and regulation is much more complicated than that.
S
Sophia55:53
As the regulator, I will basically reference something that I think Nigel said earlier today: that necessity is the mother of invention, and regulation basically articulates what is necessary to be done. Responsible innovation, I would say, is innovation that operates within the bounds of the law. I wouldn't call the EU an environment that's overly regulated, but obviously I work for a regulator, I'm biased. I think one pattern we see in general is that a lot of stakeholders who are not founded in the EU or the UK start processing data without actually having EU legislation at front of mind, to be honest, and they do move fast. If they break things or they don't break things, sometimes we may learn about it in a couple of years rather than now. So wait and see, I would say.
S
Simon Mcdougall56:49
I'm very grateful I'm coming last on this because I've actually had time to formulate a point. So yeah, I think on one hand, probably not. On the other hand, we had a company like DeepMind, and because we have issues with our funding and ability to scale great companies in the UK—and that's partly cultural and partly financial—Google went and bought it. So it's a really hard point. I could go on for hours, so if you call me later, I'll happily go into my thinking on this. I think actually the issues we have around our ability to scale great new businesses that spring out of places like this is actually a bigger issue than the regulatory framework, in my take.
M
Moderator57:52
Okay, well I think we're just about at time, so I can't take any more questions. I think we've really delved into some of the key questions in data protection in this panel. Hopefully, I think you've provided some different perspectives and really tried to illustrate how data protection is practically interacting with generative AI. It's a reality, it's here, but equally we can see some pathways forward. We look forward to seeing how different solutions emerge in the coming years. It may be that there are solutions on a range of areas: some of it can be addressed by regulators through helpful guidance, engineering solutions, and Simon's positing that maybe eventually we might see some law reform as well. But I think we can see data protection will play an important role in this area because ultimately it's about protecting people's data and it's about trust and confidence, and that's always going to shape how we use the technology. But thank you very much to my excellent panelists.