1 MediaTek Research2 Internship at MediaTek Research3 National Taiwan University4 NVAITC, NVIDIA, Taiwan5 NVAITC, NVIDIA, Santa Clara, USA* Equal contribution
Start with VoiceBot to run the system; Checkpoints hosts the Merge and Direct dialogue models.
Abstract
A full-duplex system keeps listening while it talks: it decides when to start and stops when interrupted. Building one means processing speech as it arrives without losing the linguistic competence the model inherited from text pretraining, and those goals pull against each other. Most spoken language models turn speech into extra discrete tokens and interleave them with the text, so a short reply becomes a long sequence and the patterns learned during text pretraining get diluted. TASTE (Text-Aligned Speech Tokenization and Embedding) instead attaches one speech representation to each text token. We present TASTE2, which rebuilds that utterance-level method into an incremental dialogue stack.
A shared text-token vocabulary removes word-level averaging across the Speech Tokenizer, Spoken LM, and Speech Detokenizer. Dialogue training predicts one continuous audio latent for each text token rather than interleaving two kinds of token, and an incremental Speech Detokenizer enables streaming synthesis through CosyVoice2. The sequence therefore stays exactly as long as the text. On LLaMA-Questions, a spoken question-answering benchmark, TASTE2 (Merge) answers 56.3% correctly where the same questions given as text to Qwen2.5-7B Instruct yield 57.3%, a 98.2% semantic retention.
TASTE2 VoiceBot is a research system that consumes user speech incrementally, streams synthesized audio, and halts generation when the user interrupts. On Full-Duplex-Bench v1.0, TASTE2 (Merge) stops on every interruption and reacts in 0.060 s while keeping its continuation coherent. The deployed VoiceBot is slower: 2.701 s from the end of user speech to its first audio. We also characterize explicit paralinguistic control in a TASTE-based model: speaking-rate control survives under both training strategies, emotion control turns out to be strategy dependent, and the remaining attributes stay weak.
Method
The system takes speech in and puts speech out. An ASR front end supplies the text tokens, the Spoken LM predicts each text token together with its audio latent, and the Detokenizer streams the waveform, so every reply exists as text and as audio at once. The original TASTE aligns one speech representation with each text token this way, but it was built for complete utterances. Its Whisper-based tokenizer and LLaMA backbone disagree on vocabulary, so representations are averaged over word boundaries, and its synthesis stack waits for the whole utterance before producing audio. TASTE2 rebuilds all three components, Tokenizer, Spoken LM, and Detokenizer, to remove both limits.
Each text token carries one continuous audio latent.
Linguistic content stays in the text stream; the sequence never grows longer than the text.
Architecture
Shared vocabulary
Speech Tokenizer, Spoken LM, and Speech Detokenizer all use the Spoken LM's tokenizer. Word-level averaging and its language-dependent segmentation disappear, and a standard chat template needs no conversion between token boundaries.
Frozen Distil-Whisper encoder; the decoder's embedding layer is swapped for that vocabulary and its output goes to a Residual VQ instead of an LM head.
Aligned prediction
Each position is a weighted sum of a text-token embedding and its aligned audio latent, not an extra element in the sequence. The audio target is shifted by one position, so the model commits to a word before deciding how to say it.
Qwen2.5-7B backbone with LoRA. The model regresses the audio vector directly instead of picking from a codebook of discrete speech tokens.
Streaming synthesis
The Speech Detokenizer converts text-latent pairs into the discrete speech units that CosyVoice2's vocoder turns into a waveform, emitting 15 of them for every 5 input pairs, so audio starts playing before the reply is finished.
Adapted from CosyVoice2's Qwen2.5-0.5B text-to-S3 model; the flow-matching model and vocoder downstream stay frozen.
Training
Stage 1
Representation
The Speech Tokenizer and Speech Detokenizer are trained jointly, one learning to extract audio latents and the other to reconstruct audio from text-token and latent pairs. This stage fixes the latent space everything else depends on.
Stage 2
Spoken LM pretraining
A pretrained text LM is adapted with LoRA to predict aligned text-latent pairs over roughly 40,600 hours of speech from Emilia and LibriTTS, both English. Every corpus and benchmark reported here is English. How well an acoustic attribute is represented here turns out to bound what any later fine-tuning can control.
Stage SFT
Dialogue fine-tuning
143,507 dialogues (781,443 turns, 1,493 hours) teach response generation, turn boundaries, and interruption behavior. Of those, 466 hours come from the DeepDialogue corpus and 1,027 hours are synthesized, because natural corpora rarely mark interruptions: 778 hours of general conversation, part rewritten to contain interruptions and seeded with disfluencies, and 249 hours carrying explicit style requests.
Design notes
Two ways to get instruction following
Merge starts Stage 2 from the base Qwen2.5-7B and then adds instruction-following ability by weight arithmetic, taking the difference between the instruct and base checkpoints and merging it in, which costs no additional training. Direct starts from Qwen2.5-7B-Instruct instead. Everything else is identical, which makes the pair a controlled comparison, and the two variants diverge in ways the results below depend on.
The deployed stack is this 7B Spoken LM plus the tokenizer and a 0.5B Detokenizer, 8B in total, which is how the released checkpoints are named. The Whisper encoder, flow-matching model, and vocoder are frozen; the system runs at batch size 1 on two RTX A6000.
Teaching a model to be interrupted
Turns are laid out in the ChatML template Qwen2.5 already uses, with audio attached to each text position instead of getting special tokens of its own. For an interrupted assistant turn, the loss weight on the closing <|im_end|> is set to zero, because the turn ended but not by the model's own choice. The resulting <|im_end|> probability then serves directly as the turn-decision gate at deployment, with no separate classifier to train.
Results
The evaluations support a specific claim rather than a general one. Text-aligned tokenization keeps the text model's competence nearly intact and delivers the full-duplex behavior that depends most directly on that representation; latency and paralinguistic breadth are where it still falls short.
Evaluation
TASTE2 result
Finding
Semantic retention (LLaMA-Questions)
98.2%
56.3% accuracy against the 57.3% Qwen2.5-7B Instruct reference. The remaining gap comes mostly from how instruction tuning is introduced, not from speech modeling: TASTE2 (Direct) retains 92.4% under the same tokenizer and data. Interleaved spoken LMs retain between 29.7% and 83.1% of their own text reference, roughly in proportion to how much of the sequence their acoustic tokens claim, though each is scored against a different backbone and protocol. Llama-Omni's 93.9% comes from emitting text alone and delegating acoustics downstream.
Interruption response (stop latency)
0.060 s
The model stopped on every interruption and reacted fastest among the systems compared, with GPT-4o rating the continuation 3.275 out of 5. Because the history is text, it can be truncated exactly where the user cut in, so the model resumes from what was actually heard.
Pause handling (false-start rate)
14.6%
How often the model wrongly starts speaking while the user pauses mid-thought: 14.6% on synthetic cases, the best of the compared systems. On CANDOR, a corpus of real recorded conversations, it rises to 38.9%. That is behind Gemini Live at 31.0% and behind TASTE2 (Direct) at 31.9%, the one axis where Merge is not the stronger variant.
Smooth turn taking (latency)
1.207 s
On natural CANDOR conversations, turn ends are identified correctly 90.8% of the time, but the decision is slow: the fastest system compared answers in 0.265 s. This is the design's direct cost.
Mean time to first audio
2.701 s
22% faster with TensorRT (down from 3.469 s), ahead of Gemini Live at 4.032 s but behind Grok at 2.310 s and ChatGPT Voice at 1.370 s. The bottleneck sits in speech generation, not the Spoken LM. Commercial figures were measured through the official mobile apps in March 2026 on unknown hardware.
Numbers are for TASTE2 (Merge). The first row is the LLaMA-Questions spoken QA benchmark. Rows two through four measure the Spoken LM on Full-Duplex-Bench v1.0, which covers four interactive behaviors, three of which TASTE2 is evaluated on. The last row measures the deployed VoiceBot end to end.
TASTE2 VoiceBot
The VoiceBot puts TASTE2 behind a live audio loop: voice activity detection notices when the user stops speaking, the model's own end-of-turn probability decides whether to answer or keep listening, generated audio streams out as it is produced, and playback stops when the user speaks over it. Because that gate is only consulted after speech end, the deployed system never speaks while the user is speaking: it produces no backchannels, the "mm-hm" a listener offers mid-sentence. Active overlap behaviors are not yet deployed.
TASTE2 VoiceBot deployment architecture.
A recorded interactive session with the deployed VoiceBot is available here.
Paralinguistic control
If an audio latent accompanies every text token, the model has somewhere to put acoustic intent, but only if training gives it something to put there. We test instruction-conditioned control on a 41-item subset of VStyle's acoustic attributes, scored 1 to 5 by three human raters. Speaking rate is the one attribute family that clears the threshold under both training strategies, though slow only just does so under Merge.
Feature
Stage 2 continuation
TASTE2 (Merge)
TASTE2 (Direct)
fast
4.87
4.15
3.53
slow
4.27
3.00
3.77
happy
3.80
1.67
1.67
sad
3.56
2.00
1.89
surprise
3.33
2.42
3.42
whisper
3.27
1.00
1.00
fear
2.40
1.17
1.33
angry
2.00
1.33
3.33
Stage 2 behaves like a ceiling: an attribute with no solid representation there is not rescued by fine-tuning, and one that only marginally clears it there, such as whisper at 3.27, may not survive it either. Emotion is the least settled case: angry and surprise clear the threshold under Direct and collapse under Merge. Both models return valid audio for 39 of the 41 items, so what fails is control, not generation.
Mean human rating on a 1 to 5 scale, where 3 is the threshold for usable control. The Stage 2 column is the model before dialogue fine-tuning, asked only to continue an unfinished utterance carrying the attribute; the last two are the fine-tuned models asked to produce it on instruction. Stage 2 means average over five openings per attribute excluding invalid outputs, so angry rests on 2 and sad on 3. Coverage differs from the full VStyle protocol, so these scores are not comparable with published VStyle numbers.
Samples
Instruction
Response
Model
Rating
Fast speaking rate
"In a very fast tempo, recite the seven days of the week."
TASTE2 (Direct)
5.0
Slow speaking rate
"With a very slow pace, count the months from January to June."
TASTE2 (Direct)
5.0
Two of the highest-rated generations for speaking rate. Both come from TASTE2 (Direct), the stronger variant on slow.
Conclusion
Text-aligned speech modeling supports full-duplex voice interaction without an interleaved token stream. Pairing one continuous audio latent with each text token retains 98.2% of the text reference's accuracy while the model emits acoustic information at every step. It also delivers the behavior that depends most directly on that representation: the model stops and resumes correctly when the user cuts in, and it identifies turn ends without a dedicated classifier, at the cost of over a second of turn latency. The VoiceBot shows the same stack running as a live system.
Natural conversations are still harder than synthetic ones, deployed synthesis is slow enough to be felt, and paralinguistic control works for speaking rate while most other attributes do not follow. Because an attribute absent from Stage 2 is never recovered by fine-tuning, widening what Stage 2 represents is the prerequisite for the rest.
Acknowledgments
The authors gratefully acknowledge the NVIDIA AI Technology Center (NVAITC) for providing access to the Taipei-1 supercomputer and the computational resources used in this research.
Citation
@article{chen2026taste2,
title = {TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction},
author = {Chen, Yi-Chang and Chen, Chun Wei and Wu, Dien-Ruei and Lin, Jie and
Fu, Yu-Kuan and Lin, Yang-Hsien and Huang, Eddie TC and See, Simon and
Lee, Hung-yi and Shiu, Da-Shan},
journal = {arXiv preprint arXiv:2609.08956},
year = {2026}
}