Chinese AI Models Echo State Doctrine, Aleph Alpha Finds
Ask a Chinese language model about Tiananmen, Taiwan or Xinjiang, and there is a good chance you will get the official line, a polite dodge or no answer at all. That is the main finding of a new study from Aleph Alpha, which built its own benchmark to measure how models from China handle politically sensitive questions.
The study deserves a careful read, and so does the source. Aleph Alpha, like Cohere, sells itself as a provider of "sovereign AI" for governments. In practice, that means models that public institutions can run without depending on foreign suppliers. The company therefore has a clear commercial reason to show that Chinese alternatives fall short.
What the benchmark tested
Aleph Alpha put models from three Chinese developers through its test: Alibaba's Qwen, DeepSeek, and Moonshot AI's Kimi. The benchmark covers 967 hand-picked topics that are taboo in China, including Tiananmen, Taiwan and Xinjiang.
The responses were graded by Aleph Alpha's own AI scoring system. It judged only 17 to 41 percent of answers as balanced. The remaining answers fell into three groups:
- repetition of state doctrine
- deflection away from the question
- outright refusal
None of this comes as a shock. China's AI rules require public-facing models to reflect "socialist core values," and the results match both earlier audits and a steady flow of anecdotal reports from users.
The bias does not stay on topic
The more interesting finding is that the slant can show up where nobody asked about China. When Qwen 3.6 was asked about censorship in the United States, its answer opened in a fairly even-handed way. It then ended by defending China's position on global internet governance: "Many countries, including China, also manage information to ensure social stability and national security."
The Central European Institute of Asian Studies (CEIAS), a research organisation focused on Asia, documented the same spillover in an earlier study. When prompts touched on human rights, opposition or surveillance, models often fell back on familiar Beijing phrasing. Examples included the "principle of non-interference in internal affairs" and a "community with a shared future for mankind."
How the values travel into other models
Aleph Alpha also points at a direct rival. Nvidia's Nemotron Cascade 2 showed party-line patterns in 17 percent of its responses. Aleph Alpha traces this to about 3,500 training examples, out of 9.3 million in total, that were generated with DeepSeek and Qwen.
One example stands out. Asked to write a speech in favour of recognising Taiwan, the Nvidia model refused. It produced a patriotic text defending Beijing's One-China principle instead. The context matters here: Nvidia is pushing its own models harder into government and enterprise markets, the same ground Aleph Alpha and Cohere want to win.
The broader point is not limited to China. Language models absorb cultural and political values because some viewpoints are overrepresented in their training data, and because data can be selected on purpose. Researchers warn that billions of people seeing the same kind of AI output again and again could shape how they think and express themselves.
The US angle
Ideological tuning is happening in the United States too. Elon Musk has repeatedly had Grok modified to give more right-leaning answers. Studies still suggest that models in general lean left, possibly because their answers draw more heavily on scientific evidence. For the EU, this leaves an awkward choice between two foreign value systems, unless European models can match them on performance and gain wider adoption.
Our Take
The headline numbers are striking, but readers should weigh them against who produced them. Aleph Alpha designed the benchmark, picked the topics and used its own AI system to score the answers. That does not make the findings wrong, especially since they line up with CEIAS and other audits. It does mean independent replication would carry more weight than a vendor study.
The most useful lesson for developers is the Nemotron result. If roughly 3,500 synthetic examples in a set of 9.3 million can leave visible traces, then distilling data from other models is not neutral. Teams that build on open weights or synthetic data may inherit values they never chose. This suggests that provenance checks on training data could become as important as performance benchmarks.
For public buyers, the timing fits a wider trend, as US labs also court governments with offerings such as Claude for Government. It is worth watching whether "value alignment" becomes a formal procurement criterion, and whether independent auditors, rather than competing vendors, end up running these tests.
