Nishit Sahay17:37
Yeah, actually, that is a mountain of a problem statement that we realized that we had to solve. So Marvell sits on a lot of data. Even though we are not a B2C company, B2C company tends to have a lot more data than we do, but we have a lot of technology data that's sitting in front of us. Now for companies like us, again, if we have to really take advantage of the AI revolution, the real meat is when we bring our context, our knowledge into it. Now the question is, how do we bring it? And that's where the unstructured data part of it came into being. So it's not so much about reporting, but it's more to do to feed in these AI pipelines that we are developing. Now, coming up with where we are with it, we are in a 'what should we do' mode, because when you start to drill down into this problem statement, you're looking at it from various different angles. Companies like us are very good at external security, this is an example. We do excellent source system security, but as soon as you move the data across, how would you track if you're doing things right? So that's one example on the security angle. But now if you compound that with the compliance need that most of the geographies that we actually operate in, whether it's US, Europe, and lately in India and Vietnam and China, China of course has been very popular in that area in terms of their strict requirements and compliance. When you bring this content in and feed an AI pipeline, all of a sudden your data protection, your data compliance requirements just becomes multifold. Now what we are doing is we are trying to understand what we have. So there are two things. One is like we're building these AI pipelines, we're building these AI systems, and as we feed them, we actually look at the data and make sure that everything is clean. But now we realize if you had to scale it up, right now most companies are in a build state, you're piloting things for 100 users, 200 users, you're curating data for those engines. But if you want to scale up, if you want to feed in large amount of data for the AI system to really move the needle, you have to look at data foundations. So think about, most companies have data all over the company. Most companies don't think about data quality when they're feeding data because it didn't matter to them. If you're writing a technical spec, let's take that as an example. Everybody knows this is the final version of the technical spec which has to be used by developer and the product is built, then who cares about the technical spec till a revision is required? Now you might have 100 versions of the technical spec and one is the one that you really want to use to feed in your pipeline. Which one? So you have the redundancy issue that I just brought in, in the data itself or in the spec itself. You might start putting in people's information, like a requirement coming from XYZ, and all of a sudden you have personal information in the spec that you never thought of. It's not, again in this example, it's not HR data, you're talking about technical data. And then you might have some customer information embedded in it, because you have a customer requirement coming in, and you have customer's data in it. All unstructured, all in this format. Looking at the tag or the metadata of the document, you can figure out that this problem actually exists. Then you're starting to feed it into these knowledge graphs where everything just comes together in a vector format, and now you try to figure out controls on top of it. Now that's very difficult. So for us, understanding our data, understanding the ownership, understanding the redundancy of it, and trying to figure out at domain level how do we control the quality, how do we control the compliance, how do we tag the data, how do we classify the data, that is the problem statement that we have started to work on. We're building teams around it. The first thing you have to do is you have to hire a data officer, which is what we did. So now we're building teams around it, we're working with the security team, working with the compliance team, and more importantly, working with the domains, the business unit and the functional leads, to have a program structure in place, and then we'll prioritize certain things too so that those things are coming. But this is where we are. We have recognized the problem statement.