Hook: The Hardware Spec Sheet Is a Lie
Huawei's Ascend 910B delivers 320 TFLOPS in FP16. NVIDIA's A100 delivers 312. On paper, the gap is gone. On paper, China can train frontier models on domestic silicon by 2028.
That paper is worthless.
I've spent the last decade auditing infrastructure, not marketing decks. And every single time a spec sheet catches up to the incumbent, the real fight moves elsewhere. The fight isn't the chip. The fight is what happens when you string ten thousand of those chips together and ask them to train a model that doesn't collapse into NaN losses.
China's 2028 target—training frontier AI models exclusively on domestic hardware—isn't a silicon problem. It's a systems problem. And systems are where empires go to die.
Context: The Infrastructure Reality Check
Let's establish the baseline. The Chinese government's plan, reported by Crypto Briefing, is a single data point: by 2028, train frontier AI models using domestic hardware. No specifics on cluster size, model benchmarks, or acceptable performance thresholds. Just the goal.

That vagueness is strategic. But my job isn't to parse policy rhetoric. My job is to assess whether the infrastructure can bear the weight of the ambition.
Current state of play: Huawei's Ascend 910C is expected to hit 70-80% of H100 performance. Cambricon's Siyuan 590 approaches A100 efficiency. Single-card metrics are respectable. But the moment you scale beyond a single node, the picture degrades. NVLink and InfiniBand give NVIDIA clusters 900GB/s+ of interconnect bandwidth. Huawei's HCCS plus RoCE networks deliver roughly 400-500GB/s. That's a 50% gap in the backbone that carries gradient updates during distributed training.
And then there's the software stack. CUDA isn't just a library—it's a gravitational field. PyTorch, TensorFlow, Megatron-DeepSpeed, FSDP—all optimized for NVIDIA hardware. Huawei's CANN platform and MindSpore framework are improving, but developer inertia is a real tax. Every migration costs time, performance, and sanity.
Add the physical constraint: advanced process nodes below 7nm are off-limits due to export controls. China's chipmakers compensate with chiplets and clever packaging, but that trades performance for power and cost. A 30-50% power penalty per unit of compute is the price of doing business under sanctions.
Core: The Cluster Is the Product
The real metric that matters isn't TFLOPS. It's Model FLOPs Utilization—MFU. This measures what percentage of a cluster's theoretical compute is actually used during training. NVIDIA clusters achieve 50-60% MFU. Chinese clusters, by industry estimates, sit at 30-40%.
That gap is the entire ballgame.
A 10,000-card cluster at 35% MFU delivers the effective compute of a 3,500-card cluster at 60% MFU. You're not just behind on hardware. You're behind on the entire engineering stack that turns silicon into intelligence.
Scaling to 10,000 cards introduces problems that don't exist at 1,000 cards. Network congestion. Fault tolerance. Checkpoint frequency. Thermal management. Power delivery. A single failed GPU in a 10,000-card cluster can stall an entire training run for hours. NVIDIA has spent years refining these systems. China's experience at this scale is measured in months.
Based on my audit experience, I've seen this pattern before. In 2017, I built arbitrage bots between Binance and Poloniex. The hardware was fine. The infrastructure was the bottleneck—API rate limits, exchange downtime, network latency. Code is law, but infrastructure is reality. The same principle applies to AI training clusters.

There's also the HBM problem. High-bandwidth memory is the lifeblood of AI accelerators. Huawei's chips rely on HBM2E and HBM3 from Samsung and SK Hynix—both subject to US export controls. Domestic HBM production, led by ChangXin Memory Technologies, is in early stages. If HBM supply tightens further, the entire 2028 timeline gets compressed.

Contrarian: The Hidden Assumption Everyone Misses
Here's the counter-intuitive angle: the 2028 goal might not require China to beat NVIDIA. It might only require China to achieve "good enough" compute to train models that are competitive, not necessarily frontier-leading.
The definition of "frontier" is elastic. If the benchmark is "match GPT-4 capabilities," that's already achievable with domestic hardware at scale. If the benchmark is "match whatever OpenAI releases in 2028," that's a completely different game.
Smart money understands this ambiguity. The Chinese government isn't announcing a technical roadmap. It's announcing a strategic direction. The flexibility is intentional—it allows for face-saving adjustments while the real work proceeds quietly.
But here's the part the crypto media misses: this isn't just about AI. This is about creating an alternative compute ecosystem that can survive sanctions. The Chinese plan is a hedge against total technological decoupling. Even if the 2028 goal is only partially met, the effort builds a domestic supply chain for chips, memory, networking, and software that reduces strategic vulnerability.
I didn't fully grasp this until I saw the infrastructure play unfold with Bitcoin ETFs in 2024. The real money wasn't in the ETFs themselves—it was in the custody and compliance plumbing. Similarly, the real strategic value of China's 2028 plan isn't the models. It's the ecosystem that gets built along the way.
Takeaway: The Real Trade
The infrastructure gap is real, but it's closing. The question isn't whether China will have domestic AI compute in 2028. It will. The question is whether that compute will be efficient enough to train models that matter.
Watch the MFU numbers. Watch the HBM supply chain. Watch the developer adoption of CANN versus CUDA. These are the signals that will determine whether 2028 is a milestone or a mirage.
I'd rather track those metrics than any spec sheet. Spec sheets don't train models. Systems do. And systems are where the real battle is being fought.