Back
Goran Garevski
Chief Technology Officer, HYCU, Inc.

Modern Data Protection Starts with the Right Architectural Foundation with HYCU

🎥 May 31, 2023 📺 Tech Field Day ⏱ 28m 👁 121 views
To be able to address the challenges raised by HCI and multi-cloud adoption, data protection solutions need to be architected right from the ground up. Goran shares more on how HYCU was developed with solving not only the immediate needs of multi-cloud data protection but also effectively handling the emergence of as-a-Service and SaaS application use.  Goran Garevski, CTO and co-founder of HYCU, discusses the challenges of data protection in the era of numerous data silos and the company's approach to simplifying and streamlining the process. The goal is to abstract the problem across differ...
Watch on YouTube
Transcript (49 segments)
G
Goran Garevski0:09
My name is Goran Garevski, CTO and co-founder of HYCU. So as Simon actually opened up, what are the challenges you try to solve, and how we approach this era of thousands of data silos, how to simplify it and make efficient data protection for the customers. One of the biggest challenges when you want to conquer the world is how to abstract, how to do it in an efficient way. 17,000 SaaS applications with a different type of data, different type of metadata, different type of configuration, different type of high availability and error functionality built by the vendor itself. How to abstract it and make it easy to use so that customers actually have a zero learning curve — this was the goal of our cloud platform.
What we learned actually in the past with our previous experiences is that we wanted to avoid the trap of the first architecture definition. I usually say for every company there's a DNA. Either they're coming from a storage background, from a monitoring background, from a developer background — usually the first DNA is visible in the architecture and everything they do later. You can see very few cases in the whole industry, back 30, 50 years, that people redesigned the whole thing and fine-tuned it to the new environment.
So the key for us were two goals. First is how to abstract the problem across different types of sources — classical data center file systems, containers, SaaS applications, databases, platform as a service. The abstraction means that what is, at least in our definition, nirvana in data protection: the same policy applied on a database as a service, on a SaaS service, on a VM, or on a file server. That's the abstraction we are looking for in order to address this problem, including 17,000 plus applications. The key here is the central intelligence for cross-cloud data protection management and standard functionality. No one has abstraction on the policy level so that you apply it to different sources.
I
Interviewer3:00
Assuming, Goran, that this is originally based on the secret sauce of HYCU, which of course is application awareness — which you're way ahead of where others are — do your policies also adapt based on the type of underlay? Like you said, there we can have one abstraction, but there are policy elements that wouldn't exist for one data type versus another.
G
Goran Garevski3:25
Great question. You're abstracting all of the components of the policy. From the original design — and Simon mentioned the penicillin moment — they're very simple in each case, as they should be: RPO, RTO goal, retention, copies, archives, and then potentially a fine-tuning on the policy itself. If you look at the definition of the policy, you can pretty much apply it to anything. In case it's not applicable, of course users should be aware, but in 99% of cases you apply those simple properties of a policy to any type of source. This is our approach from the very beginning. If you recall, the application awareness for VM-hosted applications, file systems, and so on — that was the famous penicillin moment. Of course we started with infrastructure as a service, the classical things we need to cover: public clouds, the data center, and VM-hosted applications.
But the key was to add the future here, as Simon mentioned — tons of SaaS applications, different in nature, different in types of data, different in metadata. So for this, we actually extended our application or SaaS awareness model to also cover platform as a service, database as a service, and SaaS. But in order to address the 17,000 plus problem — which will become 20,000 or 50,000 in five years — you need the ecosystem. You cannot do it yourself. The original SaaS vendors, if they want to participate, or ISVs, or channel partners specialized in a specific SaaS application — the key is there must be a legal entity behind it to join the development program. There's a defined process: how we work together, how we certify, how we check for vulnerability and security, and how we publish it later in the marketplace.
So the key here is the ability to extend the platform, to abstract the functionality — how you protect not just SaaS but potentially any workload — and to apply it in an efficient way on top of it. We make it known ourselves by auto-discovering things. Originally in the VM-hosted world, but now even more in the much more complex world, which is the SaaS application and PaaS applications. Again, we stepped back and looked at who are usually the key sources for this kind of thing. We will discuss the discovery process later in more detail, but this is Okta and Azure AD, at the end of the day. Of course, as we go forward we will try to bring in more additional sources, but currently this is pretty much a stable situation in terms of having efficient discovery of the SaaS, PaaS, and IaaS workloads.
The key for us, and one of the key advantages — at least what we strongly believe against the competition — is that we do not lock your data. It's your storage, it's your data. We are providing a service that will protect your data, and we will store them on storage that you define, that you have, that you own at the end of the day. This is applicable to the platform itself. On top of it, clear to us is also the ability to automate, to orchestrate anything which is happening in the platform. This is the REST API, which is used also by our own user interface. Anything we do in the user interface can be done through the REST APIs too.
And the key addition here is the HYCU Marketplace. In order to properly address the problem, you can see HYCU has its own marketplace with all of the functionality required for a typical marketplace — content, tracking the consumption, publishing, updating, billing — either through us or through different partner marketplaces on top of it. Let's dive deeper into a couple of key components of our cloud platform and its extensibility.
But first, when you look at 17,000 plus SaaS sources, what is the challenge? This is not the same problem as we had in the past, where you can directly access the infrastructure, the storage, the hardware, and anything behind the application. We are talking about a black box. SaaS and PaaS are pretty much a black box, so you cannot do many miracles. You have an API, which usually limits access to the data. On top of it, any SaaS company has its own API throttling, network throttling, and limited bulk data APIs. When I say bulk data, I mean bulk data transfer in order to perform efficient backup and recovery operations. Not to mention the granular recovery, which is the most key feature for end users. The most wow effect we got from end users on our cloud is with discovery and granular recovery.
Why? Because every application has different types of objects, and this is a nightmare for a backup vendor to protect and recover. Not to mention more complex applications that have their own plugins and custom ISVs that add metadata and other things on top of what exists within the SaaS service. And of course configuration recovery, because there are always data, metadata, but also the configuration of a service or a sub-component of the service. So these are the challenges we faced when we approached the problem.
A
Audience Member10:10
So if we're talking about storing data, like let's say I'm in Azure, AWS, GCP, whatever, I'm still storing that data. I'm still ultimately responsible for storing that data. So where would HYCU come into play there?
G
Goran Garevski10:24
Yes, for the policy enforcement piece. Data movement, data protection, recovery — everything on top of it. We are providing pure data protection as a service on top.
A
Audience Member10:37
Okay, got it. And then one more question — from a multi-cloud perspective, let's say I'm using Azure storage accounts and I'm using S3. I have some data in S3 and I want to replicate that to Azure. Is that capability available?
G
Goran Garevski10:52
That's available through backup and recovery functionality. Got it. We implement DR through the basic backup and recovery principles.
A
Audience Member11:01
How do you avoid the problem of a customer doing something really dumb, like backing up their on-prem infrastructure to their on-prem infrastructure? Do you notice that?
G
Goran Garevski11:17
I can say the customers are very creative. So we proactively monitor behind the SaaS service but also through our telemetry in terms of what is happening, and try to guide them. We have a continued success — we are very fanatic about customer success. There's a proactive service that we monitor, but also daily continuous synchronization between the customer success team and the customer. Can we catch everything? No. But I think the NPS score tells itself.
I
Interviewer11:57
One of the things I'd love to see — as we look at the policy, RPO, RTO — general Highland policies are super cool. We need that at the application layer: this application is protected in this way via this policy. Presuming you have a proprietary — we will have a demo or whatever — but also how deep into the elements and the artifacts of that? Because Salesforce data versus O365, you know, at a record level...
G
Goran Garevski12:28
You'll address the theoretical part in the following slides, and then you will see a demo for a couple of SaaS services.
I
Interviewer12:35
Oh, and also one quick one — I feel like they've been paired. I've spent a lot of time in Justia for the last week finding what you guys are doing. You got an incredible patent attorney, he writes a lot. I've studied way too much. One thing about air-gapped environments — I know you have federal customers. Presumably you've got some situations where you may have to do with air-gapped.
G
Goran Garevski13:00
So we support that, yes.
I
Interviewer13:07
All of these different data types, services have very different data profiles. So you're building hundreds, thousands of some data movers?
G
Goran Garevski13:19
That's a hell of a challenge. Just having data movers — not so business challenge — the goal is to abstract the platform so that multiple partners can curve out efficient modules that will support different sources.
I
Interviewer13:36
Yes. If I understand correctly, you're not building the data movers — you have an API to which they will go in the following slots, if you don't mind.
G
Goran Garevski13:43
Right. Let's go first on the discovery process, because this is one of the wow effects. Simon was talking about the graph and the value there. Let me explain a couple of key challenges there. It's not so easy to discover a SaaS source or a PaaS source. Why? Because there's no central registry of all SaaS sources. Usually the CIOs don't even know about 20, 30 of the SaaS services within the company. The closest approximation or hint about SaaS services is from the IDP vendors — Okta, Azure AD. What we do there is collect data about how these services are configured, but that's far from enough. We have pre-built intelligence and information about SaaS applications and their built-in SaaS protection. It's very important to understand what are the HA or operational backup capabilities behind the service itself, because data protection is a multi-layered approach. You have operational stuff behind the SaaS service, then compliance stuff, and so on. The key is to inform the people, because this information is not available instantly. Sometimes you dig into different sources.
Key for us is to give it to the user so they understand what is the protection behind the SaaS service without any backup vendor approaching. Then we do advanced logic to identify, because there's a massive gap between the data provided by IDP vendors and the actual SaaS instance you need to protect or what the user has. There's advanced identification, filtering, and auto-mapping logic to remove the noise, because there are a lot of irrelevant things within the IDP event or catalog. And at the end, as you filter that, we are very proud also of our classification. We help you identify by grouping specific SaaS services into industry categories. For example, if you have Jira, it will be mapped into the engineering subsection of the graph or subtree. If you have HR, you'll have BambooHR part of the subtree there. We're trying to group and automate that discovery process so the CIOs or IT managers can see and have the quality information to understand what is happening in the environment.
And of course, ideally — because our graph is not just for services to be clear — we add the classical workloads, VMs, database as a service. This is a visualization of the whole cloud data environment, to propagate and provide the state of the whole environment about data protection and compliance. So this is ideally the single picture you need to see to understand if you're safe or not.
I
Interviewer17:28
You talk about a RESTful API. You provide also access to the graph or focus?
G
Goran Garevski17:34
Yes. I see more and more people are beginning to understand the graph has different ways in which you can present that data. Absolutely. A lot of people also want to add value on the graph, potentially additional information. So a lot of interesting stuff in front of us. What I wanted to give you a clue also is what part of the typical backup application is done by the platform and what part is done by the module itself. Our goal was to minimize the effort the partner — or potentially the customer, if they do their own modules — needs to do in order to have an efficient data protection solution for a specific SaaS source. We managed to minimize the effort, putting the module to a bare-bones effort. Anything the development partner needs to do is tightly tied to the SaaS application itself — no additional layers, no architectural services. The beauty of the abstraction: policy management, storage management, managing the lifecycle of backup copies, retention — all of those tough and ugly things are done by the platform. The module itself focuses purely on the interaction with the SaaS service.
A
Audience Member19:03
I do have a question. You're doing a very good job managing literally all the data across all these different platforms. What if I wanted to use this to migrate from service X to service Y? I could see something like Dropbox to OneDrive would be straightforward, but what if you wanted to do Microsoft Dynamics CRM to Salesforce? Is that possible?
G
Goran Garevski19:27
We can talk for the next five days on this topic. The key for us was to protect. Each platform has its own format of the data. It's extremely tough to migrate between one SaaS service and another, because there are a lot of things around. But for us, the first goal was to provide protection and give you the data. In the future, everything is possible. But first we need to protect the data and give you control of the data itself.
S
Simon20:07
Thank you, that was a great question, David. One of the things we've been hearing from CIOs a lot this year, in light of the macroeconomic environment, is the word 'negotiability' — which I wasn't sure is a real word, but I've heard it from three different CIOs. What they're all saying is it's really becoming a problem because all of their budgets are cut and they have the same SaaS vendors they've got to negotiate with. The hardest thing about negotiating with SaaS vendors is they've got all your data. The idea that they actually have a white slate, a canvas where they can put all of that data so it can then be migrated somewhere else increases their negotiating position. Simply by virtue of the fact that data is now in their control — and we're seeing that as a big value add for a lot of customers.
I think a lot of these questions, especially data portability between systems — we intellectually like to think that we could do it, but then in practicality, do you actually see people moving data between these systems, or is it really more in your view or anecdotal experience that we just moved it in and out of that same platform, but we need the policy level? SaaS is such a wide place, it's such a big ocean that we're swimming in. Let's take Jira for example. It's not necessarily that people are moving off of Jira, but there is a massive move from Jira on-prem to Jira in cloud. When you look at things like that, I think that's where we're seeing that shift, and I think that's where HYCU can really help support it — on a use-case basis, for sure. I like the way that you approach that, because you're right, there are some things that people want to see. There's no one red button you can hit to move everything from here to here. But in areas where we are really doubling down, like Atlassian, there's a tremendous amount of value we can provide in that process.
A
Audience Member22:46
Thank you. I think to summarize some of the questions I have that are kind of related to these — like all data protection presentations, you're talking a lot about backup. But my snarky statement is: no one needs backups, we only need recovery. And therefore I want to hear more about recovery, or migration, or all the words that go around that. Backing up data — great, my emotional security is through the roof, I'm happy. But recovery is why I sleep at night.
S
Simon23:00
I absolutely love that, and I hope — I'm glad that's on tape somewhere, because I may try to use it. One of their patents — and again, I do brag about patents probably too much — I just looked up your first one. It's in UI recovery. I think this is one of the most critical aspects of our cloud. It is one thing, as you rightly say, just to be able to backup data and then recover it to a CSV file — it's completely nonsense, it's useless. What you really want to do is be able to recover your data in a simple format that looks and feels like the platform you're using. You want it to feel as close to an in-app recovery experience as possible, and that's what we've done. It's customizable so that the verbiage is associated to each individual SaaS app. If the SaaS app uses 'events' as a category, you're going to see 'events' when you recover. I think that's the kind of customization our cloud provides that is really, really game-changing.
A
Audience Member23:57
Yeah, like show us the recovery — that's what I'm saying.
G
Goran Garevski24:00
So we definitely have demos after this. Just to add on top of Simon's comments — each module, our cloud has a mechanism so each module self-describes the types of data, what can be recovered, and where it can be recovered. Which means it's not burnt into the platform but it's logic brought by the module itself, by the integration. If the module supports 100 types of recovery use cases, the platform will show them. That's the beauty of the extensibility.
D
David24:39
How do I suppose — so if the SaaS vendors are making the modules and they define where it can be backed up and what can be backed up and recovered, if I'm in AKS, for example, I want to move my application to DKE, I don't want the AKS module to say I can't restore to pre-paid.
G
Goran Garevski24:58
I understand. The container services are actually managed by HYCU. We provide cross-cloud support. We are talking here about SaaS applications in nature. The computing platforms, container services, file systems — the core infrastructure as a service is covered by HYCU, and we provide multi-cloud operations there. Anything SaaS-related, anything PaaS, which is also very specific — it's done by the external partner.
D
David25:21
So like, for example, Microsoft isn't creating the integration for AKS — HYCU's creating the integration?
G
Goran Garevski25:29
Yes. And then anything SaaS-related, anything PaaS, which is also very specific, is done by the external partner.
D
David25:36
Okay, that makes sense. But as long as the coding of the application is up to the SaaS vendor, they're not going to build a workflow to restore it to their competitor. It's just not going to happen. You say you have the data, you have control over the API, but the module defines what can be restored and where it can be restored.
G
Goran Garevski25:58
If the module is done by the independent software vendor, they can define their own use cases. If the module is done by the original SaaS vendor, they define their use cases.
D
David26:13
So you mentioned a vetting process for when they submit the module. Is any of that vetting process covering the restore points or how they're storing the data to make sure that customers have full access?
G
Goran Garevski26:31
We review if this is a product that addresses the quality level and the restore use cases for a typical application. There must be minimal functionality — not just backup without restore. We have a review process where we review the use cases. There must be core value, minimal value, and then we go from there.
D
David26:53
Are you trying to sell more to software-as-a-service vendors or to end customers?
G
Goran Garevski27:00
I think there are 17,000 applications to be covered. Anyone is welcome. It's absolutely end customers. The SaaS vendors are simply a means to expedite the source-by-source integration and development process. Our customer base are the end customers of those services.
D
David27:27
Is any of this open source, or is the code available? In a sense of, let's say you're using one vendor and you want to restore to another vendor, but the integration that vendor A created doesn't have that process. Can I go in and write the code to do that, or is this all closed?
G
Goran Garevski27:44
In order to write a module, you need to be part of our development partner program and go through all of the checks. But you can dump the data if you want. You can use the REST API to customize additional workflows on top.