Who's Really Using Them: Chinese Models Take Over Half of Overseas Image and Video Generation Usage

From August 24 to October 5, Overchat AI logged 1.09 million users and 31.6 million model interactions. When you lay out this anonymous dataset, the most striking thing isn’t who scored a point higher on some leaderboard, but who people actually voted for when they got to work: over half of image generation and video generation—the two most compute-intensive and expensive tasks—was completed by Chinese models.
These aren’t benchmark scores; they’re usage numbers. Scores can be gamed with one inflated test run, but usage can only be accumulated through repeated clicks, one after another.
Let’s first be clear about the methodology of this data, so it isn’t mistaken for a market census. It comes from a platform that puts multiple models into the same mini-app, and it counts 31.6 million interactions across 1.2 million sessions from 1.09 million users between August 24 and October 5. Share is calculated as “each time a user uses a given model, it counts once,” rather than accumulating by each API call. The sample is naturally skewed toward users willing to try multiple models at once, so it’s more like a snapshot of “who is being repeatedly opened” than a revenue ledger for the entire industry. But precisely because all models sit behind the same entry point and face the same group of people, this comparison is more persuasive than the benchmark scores each vendor publishes on its own.
Text Is Still America’s; Visuals Are Changing Hands
Start with what hasn’t changed. In text interactions, Gemini, GPT, Claude, and Grok together account for 86.2%, and U.S. vendors have almost no rivals in this area. What has changed is visuals: Chinese models’ image generation share rose from 54.0% in the previous period to 59.5%, and video generation from 43.5% to 53.6%. Over the same period, Chinese models accounted for 36.9% of all interactions.
Zoom in on specific families, and the division of labor becomes clearer. On the image side, ByteDance’s Seedream holds first place at 33.5%, Google’s Nano Banana at 26.1%, and Alibaba’s Qwen Image at 26.0%—the latter nearly double its 12.3% from the previous period. On the video side, xAI’s Grok Imagine remains the top individual model (30.6%, though down from 39.6% in the previous period), ByteDance’s Seedance jumped from 10.8% to 25.2%, and Kling held at 21.9%.
What really deserves a pause is the two-week window: from September 22 to October 5, Qwen Image became the most-used model family across all categories at 17.2%, beating GPT text at 10.9% and Gemini at 10.8%. A family born from open weights and known for image editing was opened more often than all text models combined in an entry point that aggregates models from every vendor.

Cheap Doesn’t Mean Used; the Qwen Numbers Say Something Else
The easiest explanation is price. Overchat AI specifically added: Qwen Image and GPT Image 2.5 have the same unit price on the platform, and both are included in the free tier. Prices are equal, yet usage differs by about four times. So what’s at work here isn’t who’s cheaper, but something else: the usable output rate, how well edits preserve the original image, and whether users are willing to try a second time after their first attempt.
A supporting data point explains this stickiness well. Among new users’ model choices, OpenAI and Anthropic together account for 23.8%; among returning users, that figure drops to 19.8%. At the same time, Google rose from 20.9% to 28.0%, and Chinese vendors rose from 36.3% to 39.8%. The first choice is driven by familiarity; after that, it’s actual experience. People do switch, and they switch toward visual capabilities.
Usage Is Changing, and So Are Devices
Regional differences are bigger than expected. In Japan, 74.4% of AI usage comes from Gemini, making it almost a one-player market; Spain and Indonesia clearly lean toward Qwen Image. The U.S. is an exception: it’s the only major market where video generation is the top category, at 36.8%. Brazil leans most toward text, at 49.3%.
Mobile accounts for about 74% of sessions, and text, image, and video are almost evenly split, at roughly 33% each. Desktop clearly skews toward text, at 51.3%, with GPT, Gemini, and Claude as the dominant families. This contrast reveals one thing: image and video generation is becoming something people do “on the fly”—users try it on their phones as soon as they think of it, rather than saving it up for serious work at a computer. The preferred families on mobile confirm this too: Seedream, Grok Imagine, Seedance.
Don’t Mistake Usage for Quality
The places that require restraint are equally clear. This is a single-platform sample, and its user base skews toward people willing to try new models, so it can’t be directly extrapolated to the entire market. High usage only means it gets opened repeatedly; it doesn’t mean every image is better. Seedance costs about three times as much as Grok Imagine 1.5, yet its share still more than tripled, which shows that price has never been users’ only concern.
But put this sample together with other signals, and the direction is no longer blurry. During the same period, Alibaba opened up the weights for Qwen-Image (though the license changed to research-only), ByteDance introduced a mode for Seedance that first produces a 480P draft and generates the final video only after you’re satisfied, and Shengshu Technology pushed the launch price of its flagship video model down to 0.09 yuan per second. These moves all point to the same goal: making “trying a few more times” affordable.
Broaden the scope a bit more, and this data also explains something more macro: the adoption of text-to-video and text-to-image is trickling down from professional studios to ordinary people. Mobile accounts for more than 70% of sessions, and the three content types are nearly evenly split, which means most people aren’t building workflows—they’re just trying things out. This group can’t determine the upper bound of frontier models’ benchmark scores, but it does determine how often models get used and how many people they get recommended to. Whoever is more consistent at “one try yields a usable result” is more likely to stay on the first screen of someone’s phone.
In the end, competition in image and video isn’t about the ceiling of a single finished piece, but about how many times you can try, how many times you can revise, and how many you’re willing to keep on the same budget. The usage data simply writes out this simple truth in 30 million clicks.
Related articles
AI Short Drama Hit the Hot Search, and Live-Action Shoots Fell 70%
Only one of the top 20 titles on a major Chinese drama chart was made with real actors. AI costs about a tenth as much to produce, and it is rewriting the whole pipeline.
Decagon's PACT Protocol Wants Consent to Be a Standard
Decagon open-sourced a protocol for verifying a personal agent's identity and the permissions a customer granted it. It is plumbing, and it decides whether the agent economy works.
The AI Reunion Wave China Can't Decide How to Feel About
AI tribute films brought departed public figures back to Chinese screens and pulled in hundreds of thousands of likes. Then the backlash arrived, and it was about consent.
Apple Published an Open Multimodal Model and Barely Told Anyone
Apple's research-first release slipped past the mainstream, but the model's fine-grained visual grounding says a lot about where its AI stack is heading.