In August's text-to-image ranking, a domestic model takes the global top spot, surpassing GPT!

SuperCLUE Image released its August Chinese text-to-image benchmark rankings.
This round tested 14 mainstream models from both domestic and international developers, adding a new evaluation dimension for complex information layout to specifically assess models' capabilities in creating infographics, storyboards, and UI interfaces.
The biggest highlight of this ranking: SenseTime's SenseNova U1 Pro scored 93.81 points, edging out OpenAI's GPT Image 2 by a slim 0.58-point margin to claim the top spot globally.
ByteDance's Doubao Seedream 5.0 Pro and Alibaba's Qwen Image 3.0 Pro followed closely behind. Among the top five overall, domestic models secured three slots.
Looking at the scores, many models have already reached maturity in image fidelity and creative generation, with both categories averaging close to 90 points.
The real differentiator lies in text-image consistency and the newly added complex information layout, where averages barely exceed 72 points. The complex information layout category shows the most extreme disparity, with a 74.2-point gap between the highest and lowest scores.
In short, there are plenty of large models that can generate visually appealing images, but very few can pack large amounts of text, tables, and flowcharts into a single image.
The rankings also reveal clear polarization. The top two models—SenseTime and OpenAI—rank in the first tier across all six evaluation dimensions, making them all-round performers.
However, many mid-to-lower-tier models show severe specialization, excelling in one capability while lagging badly in others. A single standout feature is no longer enough to climb the rankings.
Domestic models don't dominate every category. They perform better in Chinese character generation and realistic reproduction, but in text-image consistency, they trail overseas models by a significant 16.56 points.
Overall, leading domestic text-to-image models have reached the global top spot, but they don't lead across the board. The ability to handle complex text-image information has become the core benchmark for judging a text-to-image model's usefulness.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 3 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 4 hours ago


