← Back to blog
AiAbout 6 min read

VibeVoice Only Open-Sourced Half: The Missing Half Is the Lesson

Published Oct 11, 2026
VibeVoice Only Open-Sourced Half: The Missing Half Is the Lesson

Microsoft's VibeVoice repository carries an MIT license, about 55,000 stars, and a description that reads like a promise: open-source frontier voice AI. What it actually contains is a lesson about what "open source" means once a model becomes powerful enough to worry about.

The story runs in two directions. In one direction, the repository keeps growing, and the parts that remain are genuinely useful. In the other, the most capable piece disappeared. In September 2025, Microsoft removed the inference code for VibeVoice-TTS from the repository after finding people using it in ways the company said were inconsistent with the project's intent. The weights for that model are still on Hugging Face. The code that runs them is not. Anyone who wants that specific model today is on their own.

That decision shapes everything about how the repository is used in 2026, and it is more instructive than any benchmark.

What is actually in the box

The centerpiece today is VibeVoice-ASR, a 7-billion-parameter speech recognition model that transcribes up to 60 minutes of audio in a single pass. It does not just write down words. It returns who spoke, when they spoke, and what they said, producing structured output with speaker labels and timestamps. It handles more than 50 languages without a language flag, and it accepts custom hotwords, so you can hand it the names and technical terms you expect and let it stop mangling them.

Around that core sit several trimmed variants. There is a streaming ASR version that emits text chunk by chunk while the audio is still arriving, useful when you cannot wait for the full recording. There is a BitNet build that runs on a CPU with no GPU at all, compressed to about 1.58GB, which is remarkable for a model of this size. And there is VibeVoice-Realtime-0.5B, a small streaming text-to-speech model that reaches its first audio in roughly 200 to 300 milliseconds, with preset voices and, notably, no voice cloning.

A waveform splitting into a solid bright half and a faded, locked half

The 7.5 Hz trick

The technical idea that ties the family together is a speech tokenizer that runs at an unusually low frame rate: 7.5 frames per second. Most speech tokenizers run at 50 Hz or higher. The lower rate matters because it fits far more audio into the same context window. That is what allows a single pass to cover an entire hour of conversation, because fewer tokens per second means more seconds fit in the model's memory.

Low frame rates usually cost fidelity, and VibeVoice handles that with a two-stage design. A language model handles the semantics, deciding what is being said and who is saying it, and a diffusion head generates the fine acoustic detail. The tokenizer only needs to preserve enough audio quality for the diffusion head to do its part. It is a clean division of labor, and it is the reason the model can be both long-window and detailed at the same time.

What you can and cannot build

For anyone working with speech, the usable half of VibeVoice is substantial. Meeting transcription with speaker labels, podcast diarization, interview workflows that need to know who said what and when, streaming recognition for live applications, and a lightweight voice for an assistant that runs on modest hardware: all of these are within reach today. The CPU build in particular opens the door to running recognition on machines without a discrete GPU, which matters for cost and for deployment in places where GPUs are scarce.

The unusable half is the part that made VibeVoice famous. The original VibeVoice-TTS could synthesize up to 90 minutes of speech with up to four distinct speakers, a genuine research advance that was accepted as an oral paper at ICLR 2026. That capability is exactly what you cannot run from the repository today. The famous version was, in practice, the withdrawn one.

Open source as a spectrum

The VibeVoice case complicates a tidy narrative that the industry likes to tell. On one side are models said to be open, meaning the weights are downloadable. On the other are models said to be closed. VibeVoice sits in between. The weights are open, the license is permissive, and the code needed to actually use the most promising model is gone. Releasing weights without the inference code is a middle path, and it may be where more capable models land as the stakes rise.

There are practical wrinkles too. The repository describes itself as a research framework rather than a finished product, so expectations should match. There is no official Python package you can install from PyPI under that name, and the package that shares the name is not Microsoft's project. Microsoft also advises that the models are for research and not recommended for commercial use without further testing, even though the license itself is permissive, which is the kind of gap that surprises teams who read the license and skip the README.

None of this makes the release a failure. The ASR model is strong, the streaming variant is useful, and the CPU build is genuinely clever. But it is a reminder that "open source" now covers a wide range of arrangements, and the specific arrangement matters more than the label. A team that needs the TTS capability should know before it plans a project, not after.

The question worth asking

The honest question VibeVoice raises is whether removing code is the right response to misuse. The weights are still out there, so anyone determined to run the model can find a way. Meanwhile, the people who might have used the code responsibly, under the license they were given, lost access to it. That is the tension every lab faces as capable models spread: restrict the tool and you inconvenience the careful users more than the careless ones, and leave open the tool and you accept that some misuse will happen.

Microsoft chose the middle. It kept the useful speech recognition work open and withdrew the synthesis code that carried more risk. Other labs will make different calls, and users will keep having to read the fine print instead of trusting the headline. VibeVoice is worth studying not because it is unusual, but because it is a preview of how more releases will be structured as the technology matures.

The lesson for anyone who builds on open models is to verify before committing. Check that the weights you want come with code that runs them. Read the license and the README, because they can disagree. Confirm that the version you plan to depend on will still be maintainable in a year. None of that is exciting, and all of it is cheaper than discovering six months into a project that the half of the model you needed was never released. VibeVoice did the field a favor by making the distinction between "open weights" and "usable" impossible to miss.

Related articles