← Back to blog
AiAbout 6 min read

Bilibili Open-Sourced a 150-Language Translation Family Under Apache-2.0

Published Oct 2, 2026
Bilibili Open-Sourced a 150-Language Translation Family Under Apache-2.0

Bilibili's Index LLM team released Index-Translate on September 30, a family of multilingual translation models built on Alibaba's Qwen3.5. The text models cover 150 languages, including Chinese and English, and are published under the Apache-2.0 license on Hugging Face and ModelScope in 2B, 9B and 35B-A3B sizes, the last of those still in preview, together with a technical report, code on GitHub and an online demo.

A translation model family from a video platform is worth a second look, because it says something about what the platform actually needs. Bilibili is a Chinese video community with a heavy subtitle culture and a growing appetite for content that crosses language boundaries. For a company like that, translation is infrastructure rather than a side project.

What makes it more than a translation model

The straightforward part is the language count, which is high but not unprecedented. What distinguishes Index-Translate is how much it is engineered around the parts of translation that are not translation.

The models follow instructions about terminology, formatting, and content that must be left untouched. In an example from the project page, the 9B model translated a JSON game-maintenance notice into Korean while preserving its structure, its star symbols and a hashtag the user had asked to keep. That sounds modest. Anyone who has run a real localization pipeline knows it is not. Markup, placeholders, product names, legal boilerplate and inline formatting are where machine translation most often breaks things, and where the cleaning-up afterwards eats the time the automation was supposed to save.

Handling those constraints is what turns a model from a demo into something you can wire into production. A translator that mangles a JSON structure is not usable no matter how fluent the Korean sentence is.

The three branches, and why they matter

The family extends the same multilingual base into three specialized directions, and those branches are where the release gets genuinely interesting.

Index-Echo produces translated subtitles or dubbed speech that keeps the source speaker's voice characteristics. Its packaged speech-to-speech release covers Chinese to English, Spanish and Japanese, and the reverse. Voice-preserving dubbing is a hard problem that sits at the intersection of speech recognition, translation, voice cloning and speech synthesis, and errors compound across every stage. Getting a voice to sound like the same person in another language is the difference between a dub that viewers accept and one they switch off.

Index-Homura adjusts a translation toward a target syllable count. This is a dubbing and subtitling constraint that most translation systems ignore entirely. When you are fitting words to a moving mouth or a fixed caption window, a perfectly accurate translation that is 40 percent too long is a failed translation. Constraining length while preserving meaning is exactly the kind of narrow, practical problem that separates a research model from a tool a production team can use.

Index-NativeLong, published under model IDs named Index-Nailong, translates whole documents and keeps references consistent across them. Document-level consistency addresses another common failure: a term translated one way on page two and another way on page nine, because each sentence was translated in isolation. For manuals, contracts and long-form content, that inconsistency is a correctness problem, not a style preference.

The open-license choice, and what it signals

Publishing under Apache-2.0 rather than a research-only or non-commercial license is a significant decision. It means anyone can build on these models commercially, including for products that compete with the platform that made them. That is a broadly permissive stance, and it is increasingly the default among Chinese labs releasing models in this cycle.

The pattern has been visible all month. Several Chinese teams shipped open-weight models in September across text, image generation, video and translation, and the common thread was permissiveness: MIT here, Apache-2.0 there, weights downloadable rather than rented. The strategic logic is consistent. Open weights build ecosystem adoption faster than an API alone, and adoption creates the kind of developer gravity that pays off in the next generation of products.

Building on Qwen3.5 rather than training from scratch is also telling. Qwen has become something close to a shared foundation for a number of Chinese model releases, the way Llama once was in the West. When multiple independent teams converge on one base model, that base becomes a de facto standard, and the work that differentiates them moves up the stack into specialized branches like the ones Index-Translate includes.

Localization is one of the quiet cost centers of video work, and it has been hard to automate without quality falling apart. Subtitling and dubbing each impose their own constraints, and a pipeline that ignores them produces output a human has to repair line by line. By packaging voice-preserving dubbing and syllable-constrained translation alongside a general text model, the release targets that repair step directly. Whether it removes the step or merely shrinks it is something only a real localization workflow will reveal, but the intent is legible: make the machine's output close enough to usable that the human's job becomes review rather than rewrite.

What it means outside China

For teams outside the Chinese ecosystem, the practical consequence is access. A permissively licensed, 150-language translation family with speech and document branches, downloadable and runnable locally, is a genuinely useful asset for anyone building localization into a product, especially where sending source text to a hosted API is not acceptable for privacy or cost reasons.

The speech branches deserve particular attention from anyone working on video. Voice-preserving dubbing and syllable-constrained translation are the two capabilities that decide whether automated localization produces something watchable or something that needs a human to fix every line. A model that handles both, under an Apache-2.0 license, lowers the barrier for small teams to localize video content that would previously have required a dubbing budget.

The caveats are the usual ones for a fresh release. Released benchmarks come from the team that built the models, and independent evaluation takes time. A 35B-A3B preview is not the same as a finished model. And permissive licensing does not automatically make a model the best choice for a given language pair or domain. That still has to be tested against your own content.

What is clear is the direction. Translation is being rebuilt as a set of constrained, instruction-following tasks rather than a single open-ended one, and the models that win will be the ones that respect the awkward realities of subtitles, dubbing and document structure. Index-Translate is a well-aimed shot at exactly those realities, and it is free to build on.

Related articles