Eric Xu1:20
Ladies and gentlemen, friends from the press, analysts, and friends, good morning, good afternoon, because we have more than 100 friends from the press joining us online. They come from Europe, so good morning to you. Many of you may have attended the Huawei Connect conference last year in Shanghai. At that event, I officially announced Huawei's AI strategy as well as our full-stack, all-scenario AI portfolio. I also outlined major changes that we need to make to help make AI more pervasive and accessible. We hope all industry players will work together to drive changes and close the gap between AI-related reality and AI-related potential. And over the fall, we have been working hard on several of those areas. So today, the launch event is about an update of our latest progress. First of all, please allow me to brief you on our AI strategy, which has five pillars that I introduced last year. First, we invest in AI research. We set up a Noah's Ark Lab. We develop fundamental ML capabilities in computer vision, natural language processing, and decision inference, focusing on data and power efficiency, namely using less data, computing, and energy. Security and trustworthiness is another area of focus. Automation and autonomy is the third one. So that's the responsibility of our Noah's Ark Laboratory. The second part of our strategy is to build a full-stack AI portfolio adaptive to all scenarios, including both standalone and cooperative scenarios between cloud, edge, and the device. We also aspire to provide abundant and affordable computing power and an efficient, easier-to-use AI platform with all four pipeline services. The third pillar is to cultivate a challenging and open ecosystem, working extensively with global academia, industries, and partners. Fourthly, we will strengthen the existing portfolio by bringing AI into all of Huawei's products and solutions to create greater value and make them more competitive. The fifth part of that is to drive operational efficiency by using AI to automate high-volume, repetitive tasks for better efficiency and quality.
Last year, I also launched our full-stack, all-scenario AI portfolio. By taking today's opportunity, I'd like to reiterate and take you through this full-stack, all-scenario portfolio. The portfolio covers all development scenarios, including public cloud, private cloud, edge computing, IoT, industry devices, and consumer devices. That's what all-scenario is about. And the full-stack is from a technical or functionality perspective, that includes chips, chip enablement, training and inference frameworks, as well as application enablement. There are a couple of areas for full-stack. One is the Ascend IP and the chip series, which is based on a unified, scalable architecture. In this series, we have Ascend 910, Max, Mini, Lite, Tiny, and Nano. The other part is the CANN chip operators library and highly automated operators developing toolkit. The third part is MindSpore, a unified training and inference framework for device, edge, and cloud. The other part is application enablement, which is ModelArts, including full pipeline services, hierarchical APIs, and integrated solutions. So last year, we launched our AI strategy as well as the full-stack, all-scenario AI portfolio. Also last year, we launched our very first AI processor, the Ascend 310, primarily used for edge computing and inference. Today, Ascend 310 has been widely adopted in a range of edge scenarios. There are three primary series based on Ascend 310: MDC, Mobile DC; the other is Atlas, including as well as computing AI accelerated modules and servers; the other is cloud services. Today, these solutions are widely used in our partnership with carmakers to address intelligence and automation needs. The Atlas cards and servers are widely used by our partners, and they use them for smart transportation, smart power. It's also used widely for heating and wastewater processing that is closely linked with people's livelihood. Cloud services based on Ascend 310 are for video analysis. There are more than 50 APIs related to this. Based on Ascend 310, providing cloud services, the daily API calls has exceeded 100 million, and this figure is expected to hit 300 million by the end of this year. At the same time, we also launched ModelArts, a full-pipeline model production service. It provides model developing services spanning the full pipeline from data collection, model development, to model training and deployment. Today, there are more than 4,000 training tasks per day for a total of 32,000 training hours. 85% are related to visual processing, 10% audio processing, and 5% are related to machine learning. And today, there are more than 30,000 developers using our ModelArts. Today, what I'm going to show you is Ascend 910, the industry's most powerful AI processor.
In October last year, we actually disclosed the tech specs of Ascend 910. Today, I will show you more about how it actually performs in tests. Test results show that the Ascend 910 processor delivers on its performance goals. For high-precision floating-point operators, Ascend 910 delivers 256 teraflops, and for integer precision calculations, it delivers 512 teraops. More importantly, its max power consumption is only 310 watts, much lower than specified. Under the same conditions, its computing power is twice that of the mainstream benchmark. Ascend 910 performs much better than we expected and better than our design specs. It is now already used for full AI model training. In a typical training session based on ResNet-50, the combination of Ascend 910 and MindSpore trains models about two times faster than other mainstream training cards. Using TensorFlow, Ascend 910 can train 1,803 images per second, while the existing benchmark is 965 images per second. As for the key technologies, we have a video clip. This is a 7nm Da Vinci architecture-based AI core for the compute engine. In addition to scalar and vector units, the AI core integrates a 3D cube computing engine. It completes 4,096 MAC ops per cycle, more than two orders of magnitude larger than what CPUs and GPUs can deliver. Ascend 910 has 32 cubes inside, providing 256 teraflops. Not just a powerful AI coprocessor, Ascend 910 is a highly integrated SoC comprised of CPUs, DVPP, and task scheduler. As well, Ascend 910 has self-managing capability that can minimize data interchange with the host. By eliminating this overhead, the unprecedented computing power can be fully utilized. Highly efficient communication mechanism is another key for a training system. Ascend 910's self-developed HCCS can deliver 240 gigabits per second for each port. Latest PCIe guarantees twice the throughput compared with previous generations. Lastly, on-chip 100G RoCE enables direct data exchange between multiple nodes, which improves training cluster performance and flexibility. High computing power, high level of integration, and ultra-fast interconnection together build Ascend 910, the world's most powerful AI processor.
Ascend 910 is only the start. Moving forward, we continue investing in AI processors to meet different needs of a broad range of scenarios. There are a couple of Ascend series for edge computing based on Ascend 310. Our plan is to launch Ascend 310 in 2021. The existing MDC is based on Ascend 310, currently used for autonomous driving. In the future, for large-scale commercial use, we'll be launching Ascend 1616 and Ascend 620. For AI training, today we officially launched Ascend 910. In the future, we are going to launch Ascend 920, so that we can come up with a whole series of AI processors for different scenarios, so as to provide powerful and affordable computing resources for research and industrial adoption. Today, I would also like to announce the release of MindSpore, our all-scenario AI computing framework. We know AI computing frameworks are critical to making AI application development easier, making AI applications more pervasive and accessible, and ensuring privacy protection. That's where an AI application development framework comes in. For that purpose, at Huawei Connect last year, we announced three development goals for our AI framework: easy deployment, which substantially reduces training time and costs; efficient execution, which uses the least amount of resources; and most importantly, it needs to be adaptable to all scenarios, including device, edge, and clouds. MindSpore marks significant progress towards all three dimensions. It is adaptable to all scenarios across all devices, edge, and clouds, and provides on-demand cooperation between them. Its 'AI algorithm as code' design concept allows developers to develop advanced AI applications. That means people working on algorithms do not necessarily have to have strong coding capabilities, so that would bring much ease of use and train models much more quickly. Through framework innovation as well as co-optimization of MindSpore and Ascend processors, our solution can ensure stronger performance and more efficient execution. MindSpore does not only support Ascend 910, it also supports other GPUs and CPUs or other processors available in the industry. Many people have asked me the question: now that we have TensorFlow or PyTorch, what's the point of MindSpore that Huawei works on? I've been telling these people that right now, none of the existing frameworks can support all scenarios. Well, you can look at our portfolio, we cover devices, edge, and clouds. Furthermore, data privacy protection has become a more important topic than ever, and the support for all scenarios is essential for enabling secure, pervasive AI. This is also a key component in our MindSpore framework. Resource budgeting environments can be big or small as needed. MindSpore also helps ensure user privacy because it only deals with the gradient and model information that has already been processed. It doesn't process the data itself, so private user data can be effectively protected even in cross-scenario environments. In addition, MindSpore has built-in model protection technology to ensure the models are secure and trustworthy. In addition, MindSpore is built on a concept called 'AI algorithm as code.' This design concept allows developers to develop advanced AI applications with ease and train their models more quickly. In a typical neural network for natural language processing, MindSpore has 20% fewer lines of code than existing frameworks on the market, and it helps developers raise their efficiency by at least 50%. Through framework innovation as well as co-optimization of MindSpore and Ascend processors, our solution can help developers more effectively address complex AI computing challenges and the need for a diverse range of computing power. This results in stronger performance and more efficient execution. In addition to Ascend processors, MindSpore also supports other processors as well. And we also make MindSpore developer-friendly. Adopting an 'AI algorithm as code' concept, MindSpore enables a host of new technologies, including source-to-source automatic differentiation, which vastly outperforms graph and operator overload. MindSpore can achieve differential expression and compiler optimization for any operator and automatic generation of inverse operators. This makes model development as easy as ABC. As data sets and model scales grow ever larger, parallel processing is inevitable. Regardless of how time-consuming and difficult it is to implement, MindSpore is here to help. By simply defining a standalone model, MindSpore can automatically implement multi-machine hybrid parallel operations. MindSpore also supports both static and dynamic graphs, allowing for seamless switching with only one statement, making model debugging easy and efficient. MindSpore provides Ascend-native runtime efficient technology to maximize the computing power of Ascend chips. In the host-device model, interaction between the CPU and GPU creates large memory and data overhead. MindSpore controls and executes the entire neural network model training on the chip, reducing interaction with the host CPU and improving training speeds. Existing distributed training models introduce central controls to define gradient synchronization points. MindSpore instead uses point-to-point distributed gradient aggregation and completely eliminates control overhead. Software and hardware are optimized to map different types of operators to their optimal computing unit and data layout, achieving exceptional performance. With MindSpore, mathematicians and researchers are being given an invaluable tool that makes theoretical exploration and innovation easier and more efficient. To drive AI application, in order to encourage the entire industry to build MindSpore into a real all-scenario computing framework, MindSpore will go open-source in the first quarter of next year, so as to help every developer and encourage every developer to participate. With this launch of Ascend 910 and MindSpore, Huawei has unveiled all the key components of our full-stack, all-scenario AI portfolio. As you might be aware, Kirin Lite is already used in Kirin 980, used in Honor and Nova phones. And Ascend Tiny, we're going to have Kirin 990, which Richard Yu will be launching at IFA this year. Kirin 990 is based on a neural processing unit, Ascend Lite. So last year, we promised a full-stack, all-scenario AI portfolio, and today we delivered. This is a new milestone for Huawei, it's also a new start. We look forward to working together more deeply with partners around the world to build pervasive AI to benefit every individual, every home, every organization. We also hope all partners can work together with Huawei to promote the development of AI so that AI can benefit society. Thank you.