The Chinese AI developer and telecoms giant are building programming infrastructure for Huawei’s Ascend chips, shifting the contest from raw silicon toward the software layer that has long reinforced Nvidia’s advantage.

DeepSeek and Huawei have moved their collaboration into one of the most strategically important layers of artificial intelligence computing: the software that lets developers extract performance from chips. On September 30, DeepSeek said it had worked with Huawei to develop programming infrastructure optimised for Huawei’s Ascend processors and was releasing parts of that work as open source, according to Reuters.
The initiative includes compute and communications libraries for the Ascend platform and work on a supernode configuration built from 128 Ascend 950 chips. DeepSeek also highlighted TileLang, a high-level language designed to make chip programming easier while preserving low-level performance. The development is significant because Nvidia’s competitive strength is not only in the performance of its processors. It is also in CUDA, the software environment that developers have used for years to build, optimise and deploy accelerated workloads.
The battle is moving above the chip
Semiconductor competition is often described as a race to build faster processors. That is only part of the contest. A chip becomes useful at scale when developers can program it efficiently, move data between devices, debug workloads and connect it to established AI frameworks. Software therefore determines how quickly a new architecture can become practical for researchers and companies.
Nvidia’s CUDA ecosystem is powerful precisely because it reduces friction. Years of developer tools, libraries, documentation and optimised code have made it difficult for rivals to compete solely on hardware specifications. DeepSeek and Huawei are attempting to narrow that advantage by building an alternative stack in which the compiler, communications layer and high-level programming model are designed together around Ascend hardware.
Why TileLang matters
DeepSeek said TileLang can provide a simpler programming model while still reaching high hardware performance. The open-source TileLang-Ascend project describes an Ascend-oriented implementation designed for high-performance AI kernels using Python-like syntax and compiler infrastructure.
That kind of abstraction matters because AI workloads increasingly depend on specialised kernels for matrix multiplication, attention and communication between accelerators. If developers have to rewrite large volumes of low-level code for every new processor, switching costs remain high. A higher-level language can reduce that burden if it delivers reliable performance across models and hardware configurations.
Huawei is trying to build a system, not just a processor
Huawei unveiled its next generation of AI processors and supernode systems earlier in September and said it expected them to be used widely for model training next year. The company’s strategy increasingly resembles a systems approach: individual chips are combined with networking, memory, software libraries and large-scale node designs that can operate as a unified computing platform.
That approach reflects a wider shift in the AI industry. The performance of a single processor matters less when large models require hundreds or thousands of accelerators to work together. Communication efficiency, memory movement and software scheduling become just as important as the raw throughput of each device.
DeepSeek brings a demanding workload to the partnership
DeepSeek’s relevance goes beyond its public profile as an AI model developer. A company training and serving large models has direct experience with the bottlenecks that appear when software, networking and accelerators interact at scale. That gives the partnership a practical testing environment: the tools are being developed around workloads that need high utilisation and efficient communication.
For Huawei, that can help close the gap between a hardware specification and a usable developer experience. For DeepSeek, closer access to the Ascend stack reduces dependence on imported accelerator ecosystems at a time when China’s technology sector faces continuing restrictions on access to advanced U.S. chips.
Open source is part technical strategy, part adoption strategy
Releasing parts of the programming infrastructure as open source lowers the barrier for outside developers to test, modify and extend it. It can also accelerate debugging because performance problems are exposed to a broader community. That does not automatically create an ecosystem comparable with CUDA, but it changes the adoption model from a closed vendor relationship to a more collaborative one.
The choice is also strategic. Developers are more likely to experiment with unfamiliar hardware when the tooling is transparent and accessible. If enough useful libraries accumulate, the software itself can become a reason to consider the hardware. Ecosystems are therefore built through repetition: each successful tool or framework makes the next developer’s decision slightly easier.
Nvidia’s moat is still deep
None of this means Nvidia’s software advantage has disappeared. CUDA benefits from a long development history, extensive documentation, optimised libraries and a global developer base. Many AI systems were designed around Nvidia hardware from the beginning, and changing platforms can require testing, retraining engineers and rewriting production workflows.
The real question is therefore not whether Huawei and DeepSeek can reproduce every part of CUDA immediately. It is whether they can make enough high-value workloads run efficiently enough that switching becomes economically acceptable. In markets where access to Nvidia hardware is constrained, the threshold for adoption may be different from markets where customers can choose freely among suppliers.
The supernode is an answer to scale
DeepSeek said it and Huawei jointly advanced a supernode solution based on 128 Ascend 950 chips, optimising both computation and communication. That architecture points to another important competitive dimension: clustering. Modern AI training increasingly depends on how effectively large numbers of processors can act as one system.
A poorly connected cluster can waste expensive compute capacity while devices wait for data. A well-designed interconnect and software stack can raise utilisation without changing the chip itself. The economics are significant because AI infrastructure is capital intensive; small improvements in utilisation can change the cost of training and serving models at scale.
China’s AI strategy is becoming more vertically integrated
The partnership also illustrates how Chinese technology companies are responding to external constraints by linking model developers, chip makers and software projects more closely. The goal is not simply import substitution. It is to create a domestic stack in which each layer can be improved with knowledge of the others.
That model has advantages and risks. Tight integration can speed optimisation, but it can also create fragmentation if multiple incompatible domestic platforms emerge. The value of open standards and shared programming tools will therefore depend on whether different companies can build on the same foundations rather than recreating separate ecosystems.
A software race may decide the hardware race
The most important consequence of the DeepSeek-Huawei announcement is that it reframes the semiconductor contest. Hardware performance remains essential, but the winning platform will also need tools that make developers productive, clusters efficient and models portable.
For Nvidia, that means the challenge from China is becoming more sophisticated. For Huawei, the test is whether the new stack can move from demonstrations and open-source releases into stable production use. And for DeepSeek, the project offers a way to shape the computing environment on which its future models may depend. The next phase of AI competition may be decided as much in compilers and libraries as in fabrication plants.




