H200 monthly rental exceeds 100,000 yuan yet still unavailable! Domestic AI chips are stuck on the encoding hurdle: high-value tokens can only rely on Nvidia.

According to reports, as AI shifts from model development to large-scale deployment, the demand for inference computing power is expanding at an unprecedented rate. However, domestic AI chips still need improvement in handling complex inference tasks such as coding, forcing domestic AI companies to allocate computing power sparingly from a limited supply of Nvidia chips.
Guan Jiawei, vice president of Qujing Technology, recently told the media that the inference market has shown clear polarization. Demand for high-quality tokens far exceeds supply, and high-end tasks impose extremely stringent performance requirements that domestic processors cannot yet consistently meet.
Especially in scenarios like coding, users are willing to pay a premium, and these high-value tasks remain heavily dependent on Nvidia chips.
It is precisely this supply-demand imbalance that has directly driven up rental prices for Nvidia chips. Data from Qujing Technology shows that the monthly rental price of Nvidia H200 chips has surged from 50,000 to 80,000 RMB at the end of the third quarter of 2025 to over 100,000 RMB earlier this year.
The explosive growth in token usage is the direct driver on the demand side. The latest statistics show that China's daily token call volume surged to nearly 175 trillion in June 2026, an increase of more than a thousandfold compared to early 2024.
Facing severe supply pressure, Chinese AI companies have been forced to seek technological breakthroughs. Moonshot AI has designed a data transfer library for heterogeneous inference scenarios, enabling efficient data exchange between Nvidia H20 and Huawei Ascend 910B chips.
Qujing Technology, meanwhile, has adopted a "heterogeneous P/D separation" technology, where the Ascend 910B handles the initial phase of requests and the Nvidia H20 is responsible for final response generation. Qujing Technology's inference engine KTransformers has successfully run the full DeepSeek model on a single RTX 4090, reducing costs by 10 to 20 times.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 4 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 5 hours ago


