Elena Simperl7:29
Thank you. My name is Elena Simperl. I'm the Director of Research at the Open Data Institute. I'm here with my colleague Emma Thwaites, who just took a picture of me. She is, among others, our Director of Policy, so possibly more of an expert to respond to some of these questions. I am also, for my sins, a Professor of Computer Science at King's College London, and so I've been working in AI. I'm going to talk a little bit about the work we've been doing at the Open Data Institute as part of the Data-Centric AI program. This is a program of work we launched at the end of last year. For those interested, there's a 15-point plan of things that need to happen to create a healthy, trustworthy data ecosystem for AI, and that plan is not just for the ODI and shouldn't be just for the ODI to tackle, but it will be a multistakeholder effort. First of all, let me just congratulate you on the question, because often when a very challenging question is asked, it is imposed in such a nuanced way. Yesterday, Tuesday, we published a first version of what we call our AI Data Taxonomy. In that taxonomy, the aim is to unpack what we mean by the 'no data, no AI' headline that many of us have embraced over the last year. You are already making some of these distinctions: inputs, enhancements, outputs, also looking across the whole AI lifecycle. All these different types of data sets and the way data is collected, managed, assured, governed in all these different scenarios is different and tackles important challenges. Whatever we do, we will need to genuinely put an effort into understanding the characteristics and the challenges associated with—I'm just going to give some examples—data that is in the public domain and has been used to pre-train a foundational model; synthetic data that has been created because the actual data can or should not be shared; fine-tuning data that possibly a different organization has created and used on top of an existing foundational model to tailor it for a particular purpose; prompting data that a consumer is adding to a consumer-facing tool like ChatGPT to ensure it responds to whatever needs it has. All these different types of data come with different challenges. The processes, technical and otherwise, to create, maintain, and assure these data sets are different. So without making your job more complicated, because you will need to look into all these different scenarios individually, this is also an opportunity for the providers of privacy-enhancing technology, data protection technology, to come up with really interesting innovative solutions to make a difference. I will just say a few things now with my scientist hat on. In terms of data protection and inputs, I think there's a lot of research that needs to happen. Techniques like model unlearning, for instance, are supposed to help with some of those challenges you mentioned around managing to make those models more mindful of privacy without having to actually bend the weights and spend another hundred million on retraining. On enhancements, we've published a public policy intervention in June around rights, which includes people's rights but also includes rights of those people involved in the supply chain of a data set. In particular, if you're thinking about safety, everyone's talking about AI safety. Safety testing includes a lot of contributions from gig workers that work on digital platforms that, for instance, are supposed to make judgments on questions around harms like toxicity, racism. They do that using their knowledge and experience, so they add their own data to these data sets that are then part of fine-tuning and safety tailoring these models. On outputs, I think there is an opportunity and a risk. I don't know how much the general public understands that at the moment we are putting our knowledge, our data, and our experience in the form of prompts into a handful of proprietary digital platforms. As an academic and a tech optimist, I would like us collectively not to make the same mistakes we've made with previous digital platforms, because we're just prompting into ChatGPT—hundreds of millions of users, several times a day—we're putting that data into these systems, and that data is not available in the same form across the ecosystem.