Deepseek Ships Open-Source Tools for Huawei Ascend Chips

Deepseek Ships Open-Source Tools for Huawei Ascend Chips

Deepseek is putting its software weight behind Huawei. The Chinese AI developer has built a set of programming tools for Huawei's Ascend AI chips and is releasing all of it as open source. The goal is simple to state and hard to reach: make domestic Chinese hardware as easy to program as Nvidia's.

The announcement came through Deepseek's official WeChat channel, with further details reported by Reuters and The New York Times.

What Deepseek released

The package includes libraries for running computations on Ascend chips and for moving data between chips. According to Deepseek, Huawei "fully supported" the effort. The two companies also worked together to optimize a supernode, which is a cluster of 128 Ascend 950 chips linked to act as one large system.

The main piece is TileLang, an open-source programming language for AI chips. Researchers at Peking University originally developed it, and Deepseek has used it for roughly a year. The company first tried it on older Nvidia chips. According to the NYT, TileLang is now Deepseek's primary tool for its work on artificial general intelligence (AGI).

Deepseek's reasoning is straightforward. Any effort to build an independent software ecosystem for AI chips needs a universal language first. That language has to be easy to write, and it still has to get full performance out of the hardware. Deepseek says TileLang offers a simpler programming model than CUDA, Nvidia's software platform.

Why software is the real bottleneck

This partnership targets one of the biggest weak points in China's AI industry. Domestic chips exist, but they need software that can extract their full performance.

Nvidia's lead was never only about chip design. It also rests on an estimated four million developers worldwide who write code for CUDA. That base of developers is what keeps rivals out. Even when competitors like AMD shipped hardware that looked just as strong on paper, they could not get past it.

According to the NYT, Chinese model makers such as Z.ai and Moonshot AI have so far moved faster than the country's chipmakers. Huawei is trying to catch up. Two weeks before Deepseek's announcement, it presented new AI processors and supernode systems and said they would be widely used for model training next year.

Huawei also says it cannot meet demand at home, so it plans to sell fewer chips abroad. Eric Xu, Huawei's current rotating chairman (the company rotates its top leadership role among several senior executives), pointed to US export controls. He said Huawei cannot accept a future that depends on whether others are willing to sell chips to China.

Is the CUDA moat still holding?

The research firm SemiAnalysis has been measuring how much of Nvidia's software advantage remains. After testing Jalapeño, OpenAI's inference chip, its analysts described the CUDA moat as "potentially dead." Their reason was how quickly OpenAI gets new models running on its own hardware.

In most of the tested scenarios, Jalapeño beat Nvidia's Blackwell on performance per watt. SemiAnalysis also noted that OpenAI models helped design the chip, and those models themselves run on Nvidia GPUs.

The analysts added caveats of their own:

  1. Easy workloads only. They tested scenarios that are relatively simple to optimize, with about 8,000 input tokens and 1,000 output tokens.
  2. No agent benchmark yet. They have not yet run AgentX, a benchmark that measures how AI agents handle multistep tasks.

That second point matters. In August, SemiAnalysis found Nvidia well ahead on exactly this kind of work. With AMD's current software stack, Nvidia would still be cheaper per token even if AMD gave its hardware away. In the analysts' view, Nvidia's lasting advantage is not the silicon. It is the software that links many chips into one working system.

Huawei's chips were not part of the AgentX comparison. In an earlier review of DeepSeek V4, however, SemiAnalysis noted that Huawei's CANN software stack was the only one besides CUDA to support that model on day one.

The Bigger Picture

The pattern is familiar: hardware alone does not break Nvidia's grip, and software is where the real work is. Deepseek's move suggests China's AI industry has taken that lesson seriously. A leading model lab is now investing in the programming layer for domestic chips, rather than waiting for chipmakers to build it.

For developers and companies following the chip race, the open-source release is the part to watch. A language that is easier than CUDA and not tied to one vendor could lower the cost of switching hardware, but only if enough people adopt it. Four million CUDA developers is a large base to compete with.

SemiAnalysis's findings add a useful limit. Beating Nvidia on simple inference tests is one thing. Matching it on multistep agent workloads, where software that links many chips decides the outcome, is another. It is worth watching whether Huawei's supernodes and the new Deepseek tools show up in independent tests like AgentX, and whether Huawei's training systems reach wide use next year as the company says they will.