The headline reads clean. China expects to train frontier AI models on domestic hardware by 2028. Clean, declarative, and almost entirely useless without a layer of forensic unpacking. Hype dies. Data breathes. And the data here points to a conclusion that cuts against the popular narrative: the single-chip gap is closing, but the real battle is being fought in cluster interconnect, software entropy, and a supply chain that no government decree can bend.
The plan, as reported, is thin on specifics. No baseline model definition. No cluster scale. No mention of what happens if the United States extends export controls to high-bandwidth memory or advanced packaging equipment. This is not a technical roadmap. It is a strategic marker, placed deliberately at a time window that aligns with the end of a policy cycle and the next iteration of domestic chip roadmaps. The year 2028 is not arbitrary. It is engineered.
The first thing to isolate is the technical trajectory. Based on my experience auditing hardware ecosystems during the 2020 DeFi cycle, I learned that advertised specs are a starting point, not a conclusion. The same discipline applies here. Huawei's Ascend 910B delivers roughly 320 TFLOPS in FP16, close to the A100's 312. The 910C is expected to reach 70 to 80 percent of the H100's capability. Cambricon's Siyuan 590 approaches the A100 in energy efficiency for training workloads. On paper, the gap is narrowing faster than most Western observers projected.
But paper is not production. Single-card performance is a node in a much larger system. The dominant constraint is interconnect. NVIDIA's NVLink and NVSwitch paired with InfiniBand create a cluster-level fabric that scales with near-linear efficiency. The Chinese ecosystem relies on HCCS for chip-to-chip and a proprietary RoCE network for node-to-node. Industry estimates place the linear scaling efficiency of a 10,000-card domestic cluster at 70 to 85 percent of NVIDIA's equivalent. The 2028 target implicitly demands at least 90 percent. That gap is the difference between a flagship announcement and a functional training environment. This is not a problem that marketing solves. This is an engineering problem, measured in latency, packet loss, and checkpoint recovery times.
Your emotion is not my edge. Neither is the emotion of a government planner who believes policy directives can accelerate thermodynamics. The software ecosystem is the more insidious bottleneck. Compute without a mature toolchain is a paperweight. CUDA is not just a library set; it is a gravitational field that holds the entire machine learning community in orbit. PyTorch, TensorFlow, Megatron-DeepSpeed, FSDP — these frameworks are optimized for NVIDIA architectures through years of developer iteration. The Chinese alternatives, CANN and MindSpore, are improving. Huawei reports over 2 million registered developers in its Ascend community. But raw developer counts do not equal optimization depth. Operator libraries remain thinner. Distributed training profiles remain less tuned. Debugging a stack that has not been battle-tested at 10,000-card scale is a latency tax that no hardware spec sheet can offset.
There is also the physical constraint that no amount of national resolve can bypass: advanced process nodes. US export controls restrict access to sub-7nm fabrication. The Chinese response is tactical and partially effective: chiplet packaging and architectural optimization on mature nodes. The strategy trades area for performance. That trade comes due in the form of higher power draw and higher unit costs. A domestic chip may require 30 to 50 percent more energy to deliver the same useful FLOPs as an equivalent NVIDIA part. Now multiply that inefficiency across 10,000 cards. A cluster that size draws 50 to 100 megawatts. That is a small city's electricity consumption. The "East Data West Computing" project, which routes compute workloads to western provinces with abundant energy, partially relieves the constraint, but it introduces new variables: network latency, operational distance, and harder logistics for hardware maintenance. Simplicity scales. Complexity collapses. And a sprawling, multi-region training architecture is complexity by design.
The most consequential metric that is almost entirely absent from public discourse is Model FLOPs Utilization. MFU is the ratio of achieved compute throughput to theoretical peak. Industry estimates put domestic clusters at 30 to 40 percent MFU. NVIDIA's reference clusters reach 50 to 60 percent. That 20-point delta is not a rounding error. It is a 30 to 40 percent effective loss of installed compute. To train frontier models, you do not just need more cards. You need more efficiency per card. The 2028 target, therefore, is not a question of whether Chinese chips can reach parity. It is a question of whether the entire system — interconnect, software, power, cooling, fault tolerance — can be engineered to approach parity. That is a significantly harder problem.
The commercial logic of this campaign is well understood by anyone who watches capital allocation. Policy procurement is the foundational layer. Government agencies and state-linked entities are directed to prioritize domestic compute. The market share of domestic AI chips in China is currently estimated at 15 to 20 percent, with projections reaching 40 to 50 percent by 2028. The "Xinchuang" framework is the umbrella under which this procurement happens. The incentives are structural, not market-driven. Cloud providers such as Alibaba, Tencent, and Huawei Cloud are integrating domestic accelerators into their service catalogs. Open-source model support for Llama, Qwen, and DeepSeek on domestic hardware is expanding. The pieces are moving. But the litmus test is not procurement volume. It is whether private, profit-driven enterprises with no political mandate choose domestic compute voluntarily. On that front, the data is still inconclusive.
The geostrategic implications are real. NVIDIA's revenue exposure to China has historically been significant. Domestic substitution will erode that share, pushing NVIDIA to cultivate markets in the Middle East, Southeast Asia, and Europe. A dual-compute-system world is not a hypothetical. It is a trajectory. And "compute sovereignty" is becoming a policy concept that will shape global technology governance, much like data sovereignty did in the previous decade. But there is a mirror-image risk that the planners discount: a parallel ecosystem that is isolated, smaller, and less battle-tested. The cost of migration off CUDA is a multi-year tax. Developers do not abandon a matured ecosystem because a policy paper tells them to. They move when the tools are better or the access is otherwise impossible. Right now, the latter condition is doing the heavy lifting.
The 2028 outcome, in my estimation as someone who has been burned by narrative-driven investments and who survived the Terra collapse by auditing reserves rather than trusting promises, is a "usable but not optimal" result. The realistic output is a domestic compute stack capable of training models that are close to the global frontier, with clear yardsticks but not dominant leadership. The actual measure of success will not be the model that China trains. It will be three specific data points: the mass production yield and real-world performance of the Ascend 910C, the pace of domestic high-bandwidth memory industrialization, and the achieved MFU of a 10,000-card cluster in sustained production. Watch those variables. They will tell you more than any announcement from a planning meeting.
The question is not whether China can build the hardware. The question is whether the Chinese ecosystem can make the hardware compute with near-zero waste. That is where the race is won. And the clock is already ticking.


