Tianshu Zhixin's new GPU beats Nvidia's Hopper on AI benchmarks — and it's ready to ship

China’s GPU race just got a new contender. At the World Artificial Intelligence Conference in Shanghai on Sunday, Tianshu Zhixin unveiled the Tian Gai 300 — a general-purpose GPU that the company says beats Nvidia’s Hopper architecture on several key AI workload metrics.

The chip is built on a SIMT (Single Instruction, Multiple Threads) general computing architecture and handles scalar, vector, and tensor computation types. But the interesting part is how it performs on the workloads that actually matter for today’s large language models.

Tianshu Zhixin claims the Tian Gai 300 achieves over 90% Attention efficiency. On 64k-context Attention tasks — the kind needed for long-document analysis and code generation — it outperforms Nvidia’s Hopper-based solutions by 10%. That means it uses GPU compute resources more effectively, especially during long-sequence training and inference.

The MoE (Mixture of Experts) performance is similarly impressive. The Tian Gai 300 delivers about 10% higher average MoE efficiency compared to Hopper, which translates to more inference tasks completed with the same raw compute budget. Given how many production AI systems are moving toward MoE architectures, that’s a meaningful advantage.

Latency numbers also look strong. First-token latency is roughly 20% lower than Hopper-based systems, and average communication latency drops by about 13%. Tianshu Zhixin attributes the communication improvements to lower protocol overhead and optimized data routing — both critical during the decode phase when models shuttle small data packets across multiple GPUs. Decode-stage efficiency comes in 10% higher than Hopper.

Perhaps more important than the raw benchmarks: the company says the Tian Gai 300 is already past the prototype phase. It’s gone through deep compatibility testing with domestic cloud providers, server manufacturers, interconnect ecosystems, and supernode systems. Deployment is possible at both single-machine and large-scale cluster levels.

Tianshu Zhixin is one of a growing number of Chinese GPU startups trying to build alternatives to Nvidia’s hardware ecosystem. The Tian Gai 300 won’t unseat Hopper or its successor Blackwell overnight — software ecosystems and CUDA compatibility remain massive moats. But the performance claims, if they hold up in third-party testing, suggest the gap is narrowing faster than many expected.