scieee AI-readable full text Open interactive document viewer

Towards Natural Realtime Voice Interaction with Large Language Models

Thulke, David

Abstract

Towards Natural Realtime Voice Interaction with Large Language Models given by David Thulke from AppTek as part of the Crystal-RTTH Fall school 2025 held in Bilbao.

Full text

Towards Natural Realtime Voice Interaction with Large Language Models David Thulke, AppTek CRYSTAL-RTTH Fall School November 2025 Thanks to Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen, Jingjing Xu, Álex Pérez, Braddock Gaskill, Uma Moothiringote Opening and Context About AppTek •Founded 1990 – 30+ years in Human Language Technologies (HLT) •Expertise in ASR, MT, NLU, TTS, and LLMs •Mid-sized team (~ 100 employees); among leading HLT providers •Tool-makers, not tool users – tailored, flexible solutions •Unified AI platform, no silos across technologies •Strong academic ties – close collaboration with RWTH Aachen, Johns Hopkins, and UPV •Strong Science Team – Prof. Dr.-Ing. Hermann Ney, Director of Science – 32 PhDs, 100s peer-reviewed papers, 10s patents •Comprehensive annotation team of 6000+ distributed contributors •HQ near Washington DC; offices in Aachen, Valencia, Tampa, and Baltimore •Major customers: Enterprise, Media & Entertainment, Contact Centers, Government 2 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context CONTENT IN THI S DOCUMENT IS CONFIDENTIAL & PROPR IETARY. AppTek’s End-to-End Full Stack HLT Solutions AUTOMATIC SPEECH RECOGNITION ASR Converts speech audio into text form NEURAL MACHINE TRANSLATION NMT Translate sentences between language pairs NATURAL LANGUAGE UNDERSTANDING NLU Analyze and identify meaning inside language LARGE LANGUAGE MODELS LLMs Generate human-like text responses from prompts TEXT-TO-SPEECH TTS Create natural sounding speech from text inputs Core Technologies Built In House - We are Tool Makers, Not Tool Users Competitive advantage: deep integration of our models with tailored datasets for customized solutions Live Speech Translation Translate source speech into text of a target language in real-time Interactive Voice Full-duplex natural and interruptable voice assistants Data Services to Fuel HLT Models Agentic AI Solve problems by autonomously interacting with external tools Automatic Dubbing Adatpive speech-to-speech translation of media content 3 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context About Me Scientist at AppTek, focusing on large language models (LLMs) and speech. 2017 Started working on ASR during my master’s studies. 2019 Joined AppTek, working on NLU. 2020 Began PhD at RWTH Aachen University (Prof. Hermann Ney) on pre-training of LMs and retrieval-augmented generation. 2022 Worked on speech-related NLU tasks (sentiment and NER) 2023 Focused on LLMs, domainand language-specific models. 2025 Co-leading efforts on voice-enabled LLMs at AppTek. David Thulke 4 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context A Short History on Conversational Interfaces 1966 ELIZA rule-based text chatbot (Weizenbaum); simulated psychotherapy through pattern matching. 1996 Philips Train Booking System speech-to-speech dialogue prototype for booking train tickets. 2011 Apple Siri first large-scale commercial voice assistant; statistical NLU pipeline for intent-based task handling. 2018 Google Duplex neural TTS with realistic prosody and overlap handling; perceived as human-like phone dialogues. Demo 2022 ChatGPT LLM-based text conversation with context memory and open-domain reasoning. 2024 GPT-4o Voice Mode speech-to-speech LLM enabling open-domain full-duplex, low-latency dialogue. Demo 5 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context A Short History on Conversational Interfaces 1966 ELIZA rule-based text chatbot (Weizenbaum); simulated psychotherapy through pattern matching. 1996 Philips Train Booking System speech-to-speech dialogue prototype for booking train tickets. 2011 Apple Siri first large-scale commercial voice assistant; statistical NLU pipeline for intent-based task handling. 2018 Google Duplex neural TTS with realistic prosody and overlap handling; perceived as human-like phone dialogues. Demo 2022 ChatGPT LLM-based text conversation with context memory and open-domain reasoning. 2024 GPT-4o Voice Mode speech-to-speech LLM enabling open-domain full-duplex, low-latency dialogue. Demo 5 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context A Short History on Conversational Interfaces 1966 ELIZA rule-based text chatbot (Weizenbaum); simulated psychotherapy through pattern matching. 1996 Philips Train Booking System speech-to-speech dialogue prototype for booking train tickets. 2011 Apple Siri first large-scale commercial voice assistant; statistical NLU pipeline for intent-based task handling. 2018 Google Duplex neural TTS with realistic prosody and overlap handling; perceived as human-like phone dialogues. Demo 2022 ChatGPT LLM-based text conversation with context memory and open-domain reasoning. 2024 GPT-4o Voice Mode speech-to-speech LLM enabling open-domain full-duplex, low-latency dialogue. Demo 5 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context A Short History on Conversational Interfaces 1966 ELIZA rule-based text chatbot (Weizenbaum); simulated psychotherapy through pattern matching. 1996 Philips Train Booking System speech-to-speech dialogue prototype for booking train tickets. 2011 Apple Siri first large-scale commercial voice assistant; statistical NLU pipeline for intent-based task handling. 2018 Google Duplex neural TTS with realistic prosody and overlap handling; perceived as human-like phone dialogues. Demo 2022 ChatGPT LLM-based text conversation with context memory and open-domain reasoning. 2024 GPT-4o Voice Mode speech-to-speech LLM enabling open-domain full-duplex, low-latency dialogue. Demo 5 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context A Short History on Conversational Interfaces 1966 ELIZA rule-based text chatbot (Weizenbaum); simulated psychotherapy through pattern matching. 1996 Philips Train Booking System speech-to-speech dialogue prototype for booking train tickets. 2011 Apple Siri first large-scale commercial voice assistant; statistical NLU pipeline for intent-based task handling. 2018 Google Duplex neural TTS with realistic prosody and overlap handling; perceived as human-like phone dialogues. Demo 2022 ChatGPT LLM-based text conversation with context memory and open-domain reasoning. 2024 GPT-4o Voice Mode speech-to-speech LLM enabling open-domain full-duplex, low-latency dialogue. Demo 5 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context What matters for natural realtime voice interaction? Good contextual language understanding and generation ⇒Large Language Models Natural sounding voices ⇒Advancements in TTS Conversational dynamics ⇒focus today 7 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context What matters for natural realtime voice interaction? Good contextual language understanding and generation ⇒Large Language Models Natural sounding voices ⇒Advancements in TTS Conversational dynamics ⇒focus today 7 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context Conversational dynamics 1. Pause Handling User Model 2. Backchanneling User Model 4. User Interruption User Model [break] yeah um hm Wait I want to ... Okay let me ... 3. Smooth Turn Taking User Model Fig. 2: Illustration of the four evaluation dimensions in Full-Duplex-Bench. (1) Pause Handling: the model stays silent during user pauses; (2) Backchanneling: the model offers short, timely acknowledgments; (3) Smooth Turn-taking: the model takes the turn in time; and (4) User Interruption: the model handles sudden user input with appropriate, well-timed responses. Metric: The ideal model behavior is to avoid taking over the turn while the user is speaking. To evaluate this, we use the Takeover Rate (TOR). A lower TOR signifies better pause management, indicating that the model effectively waits for the user’s turn to end. In contrast, a higher TOR suggests that the model is more likely to take over the conversation before the user has yielded the turn. 2) Backchanneling Drawing on the definition in Section III, we assess whether the model, when interacting with a dominant speaker, actively listens and provides backchannels at suitable moments to facilitate dialogue engagement. A model exhibiting humanlike backchanneling behavior should respond at the right times and with an appropriate frequency. Research Question: Can the model determine when to offer backchannels in a human-like manner without interrupting the speaker? Metric: To measure how well the models generate backchannel cues, we use three metrics: • TOR: As with pause handling, the model should avoid dominating the turn, so a lower TOR is preferable. • Backchannel Frequency (Freq): Each backchannel event is counted and normalized by duration (events per second). When the model does not take over the turn (TOR = 0), a higher backchanneling frequency indicates that the model responds backchannel more often, but this does not necessarily imply better or more natural behavior, as it also depends on timing and context. • Jensen-Shannon Divergence (JSD): This captures the difference between the model’s predicted timing of backchannels and actual human timing. The model outputs a probability distribution P , where P(i) denotes the likelihood of a backchannel occurring in time window i . The ground truth distribution Q is derived from human-annotated backchannel timings (details in III-C ), aligned to the same set of time windows. To measure the similarity between P and Q , we compute the Jensen–Shannon Divergence (JSD) as: JSD(P||Q)=1 2X i P(i) log P(i) M(i)+1 2X i Q(i) log Q(i) M(i), where M(i)=1 2(P(i)+Q(i)) , and i indexes the discrete time windows. JSD ranges from 0 (perfect alignment) to 1 (complete divergence), providing a symmetric and bounded measure of similarity between model predictions and human backchannel behavior. We only calculate this metric when the model does not take over the turn, where each backchannel event is counted as one-hot and normalized into a probability distribution. If the model stays silent throughout, we assume a uniform probability distribution, treating it as a random baseline without backchannel knowledge. 3) Smooth Turn Taking Effective turn-taking is crucial for maintaining a natural and engaging conversation. In human dialogue, smooth turn transitions occur when speakers respond promptly without excessive delay or overlap. A well-designed model should be capable of recognizing turn boundaries and responding with appropriate timing to ensure fluid interactions. Research Question: Can the model detect the end of a speaker’s turn and respond promptly without long pauses? Metric: We measure the averaged response latency, the time (in seconds) between the end of the user’s speech and the start of the model’s response. Lower latency values indicate smoother turn-taking. In cases where the model fails to respond, we record the TOR. The latency is calculated only when TO equals 1. This avoids averaging with non-takeover periods, which would introduce significant variance due to periods of silence. 4) User Interruption In human conversations, interruptions are common and can occur when a listener interjects mid-turn to clarify, disagree, or shift the discussion. A well-designed conversational model should be able to recognize and adapt to such interruptions by adjusting its response appropriately. Effective handling of Taken from Lin et al. (2025) 8 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Opening and Context Outline Cascaded Architecture End-to-End (Integrated) Architectures Evaluation and Benchmarking Dialog Management 9 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture Cascaded Pipeline Overview text user voice agent voice encoder ASR LLM decoder TTS text VAD 10 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Automatic Speech Recognition Given an audio sequence xT 1predict the sequence of spoken words ˆwN 1: ˆwN 1=argmax wN 1 p(wN 1|xT 1) Architectures: •Hybrid •CTC •RNN-T •AED •SpeechLLM 11 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Automatic Speech Recognition Given an audio sequence xT 1predict the sequence of spoken words ˆwN 1: ˆwN 1=argmax wN 1 p(wN 1|xT 1) Architectures: •Hybrid •CTC •RNN-T •AED •SpeechLLM 11 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Automatic Speech Recognition Given an audio sequence xT 1predict the sequence of spoken words ˆwN 1: ˆwN 1=argmax wN 1 p(wN 1|xT 1) Architectures: •Hybrid •CTC •RNN-T •AED •SpeechLLM 11 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR SpeechLLM (Large) Language Model (Transformer Decoder) 12 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR SpeechLLM (Large) Language Model (Transformer Decoder) Audio Samples 12 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Input Representations •Discrete Audio Tokens –K-Means clusters of self-supervised speech representation model (like HuBERT or WavLM) –More similar to text inputs, efficient preprocessing –Higher compression: deduplication or BPE –Very effective for TTS (discussed later) •Continuous Features –Features of an ASR encoder or a speech representation model –Combined downsampling and small adapter network to map to hidden size of the LLM –Allows backpropagation to the encoder 14 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Training Pipeline 1. Pre-train components –LLM on text-only data (∼1T −10T tokens) –Speech encoder on audio-only or audio-text data (∼100K −10M hours) 2. Joint ASR fine-tuning –Aligns speech encoder to LLM embedding space (Geng et al., 2025; Microsoft et al., 2025) –Which parameters to update? Examples: ■Train all parameters (Rubenstein et al., 2023) ■Train only speech encoder + adapter; freeze LLM (Chu et al., 2024; Bai et al., 2024) ■Train only adapter(s); freeze LLM + speech encoder (Wang et al., 2023a) ■Train adapter + low-rank adaptation (LoRA); freeze LLM + speech encoder (Tang et al., 2024; Microsoft et al., 2025; Geng et al., 2025) 3. Instruction fine-tuning for multiple speech-to-text tasks 15 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Training Pipeline 1. Pre-train components –LLM on text-only data (∼1T −10T tokens) –Speech encoder on audio-only or audio-text data (∼100K −10M hours) 2. Joint ASR fine-tuning –Aligns speech encoder to LLM embedding space (Geng et al., 2025; Microsoft et al., 2025) –Which parameters to update? Examples: ■Train all parameters (Rubenstein et al., 2023) ■Train only speech encoder + adapter; freeze LLM (Chu et al., 2024; Bai et al., 2024) ■Train only adapter(s); freeze LLM + speech encoder (Wang et al., 2023a) ■Train adapter + low-rank adaptation (LoRA); freeze LLM + speech encoder (Tang et al., 2024; Microsoft et al., 2025; Geng et al., 2025) 3. Instruction fine-tuning for multiple speech-to-text tasks 15 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Training Pipeline 1. Pre-train components –LLM on text-only data (∼1T −10T tokens) –Speech encoder on audio-only or audio-text data (∼100K −10M hours) 2. Joint ASR fine-tuning –Aligns speech encoder to LLM embedding space (Geng et al., 2025; Microsoft et al., 2025) –Which parameters to update? Examples: ■Train all parameters (Rubenstein et al., 2023) ■Train only speech encoder + adapter; freeze LLM (Chu et al., 2024; Bai et al., 2024) ■Train only adapter(s); freeze LLM + speech encoder (Wang et al., 2023a) ■Train adapter + low-rank adaptation (LoRA); freeze LLM + speech encoder (Tang et al., 2024; Microsoft et al., 2025; Geng et al., 2025) 3. Instruction fine-tuning for multiple speech-to-text tasks 15 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Results - Continuous vs. Discrete Inputs SSL model Token type ASR (WER#) PR (PER#) ST (BLEU") KS (ACC") IC (ACC") ER (ACC") LS (clean|other) GigaSpeech (test) CHiME4 (test) LS-100 (clean) GigaST (En-Zh|En-De) SC (test|val) SLURP (test) IEMOCAP (test) Qwen 1.5-0.5B HuBERT Discrete 4.56 / 9.79 19.40 13.35 9.69 22.75 / 20.14 93.70 / 93.85 57.04 38.65 WavLM 4.72 / 10.45 16.34 12.94 9.64 24.62 / 21.22 92.87 / 92.45 59.96 37.98 HuBERT Continuous 4.91 / 6.43 17.45 8.62 12.84 26.63 / 25.42 95.38 / 95.70 76.84 56.72 WavLM 2.92 / 4.61 13.96 8.68 12.62 29.44 / 28.12 97.76 / 97.36 81.35 59.45 Llama 3.1-8B HuBERT Discrete 2.56 / 6.49 12.86 10.56 7.85 26.64 / 25.13 96.75 / 96.69 63.44 39.84 WavLM 2.96 / 7.48 13.35 9.13 7.02 28.62 / 26.87 97.92 / 98.17 66.96 36.12 HuBERT Continuous 1.76 / 4.58 9.04 5.72 9.83 32.02 / 34.42 99.74 /98.59 86.84 64.54 WavLM 1.65 /4.22 8.86 5.43 10.44 35.17 /37.20 97.34 / 98.26 85.57 65.87 Table 1: Comparison benchmark of discrete tokens and continuous features on various tasks. Discrete tokens use K-means (2000 clusters) with BPE size 6000 for all tasks. For LibriSpeech (LS) datasets, we evaluate test-clean (clean) and test-other (other) set. SC stands for Speech Commands-v2 dataset. 0.5B (Bai et al.,2023) and LLaMA3.1-8B (Dubey et al.,2024), continuous features consistently outperform discrete tokens in most tasks. Notably, WavLM-Large’s continuous features perform the best across Automatic Speech Recognition (ASR), Speech Translation (ST), and Emotion Recognition (ER) tasks on the LLaMA3.1-8B model, while HuBERT-Large’s continuous features show the best performance in Keyword Spotting (KS) and Intent Classification (IC) tasks on the same LLM. The performance gap between discrete tokens and continuous features tends to increase as the LLM model size grows. For example, in Speech Translation, the BLEU score of continuous features on LLaMA3.1-8B is notably higher than that of discrete tokens, showing the increasing advantage of continuous features with larger models. In certain tasks and datasets, such as CHiME-4 (a noisy background ASR dataset), discrete tokens show a decline in ASR performance, whereas continuous features maintain more stable performance. Additionally, in Emotion Recognition, as the LLM decoder size increases, discrete token performance remains consistently poor, with accuracy significantly lower than that of continuous features. For example, WavLM-Large’s discrete tokens achieve an accuracy of 37.98% on Qwen1.5-0.5B and 36.12% on LLaMA3.1-8B, while the corresponding continuous features achieve 59.45% on Qwen1.5-0.5B and 65.87% on LLaMA3.1-8B model. Interestingly, in the Phoneme Recognition (PR) task, discrete tokens outperform continuous features across all model scales. Specifically, WavLMLarge’s discrete tokens achieve the best performance on LLaMA3.1-8B, with 7.02% phoneme error rate (PER) score. This improvement likely stems from the discrete token representation, which aligns more closely with the phoneme-level structure of speech, facilitating easier learning than continuous features. 3.3 Ablation Study Our ablation study is conducted using the LibriSpeech 960-hour dataset with the Qwen1.5-0.5B model, and evaluations are performed according to the test-clean subset. K-Meaning Clustering and BPE Size Settings. We investigate the effect of varying K-means clustering sizes and BPE (Byte Pair Encoding) vocabulary sizes on ASR performance. As shown in Fig. 2(a), increasing the number of centroids generally improves model performance. BPE-based subword modelling can further improve Word Error Rate (WER) results by reducing token sequence length while preserving key semantic information. Notably, the combination of k= 2000 centroids and 6000 BPE vocabulary size achieves a balanced trade-off between performance gains and computational efficiency. In our comparison study, we adopt this configuration across all datasets rather than tuning task-specific optimal settings, in order to ensure consistency in discrete token representa24928 Results from Wang et al. (2025a) 16 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Results - Which components to train? Internal results at AppTek: •For Librispeech, only fine-tuning the adapter gives reasonable results –similar to related work Wang et al. (2023b); Ma et al. (2024b) •On the AppTek Spanish task (diverse domains, multi-bandwidth), only fine-tuning the adapter did not converge and fine-tuning of the audio encoder was necessary for good results •Further, fine-tuning the LLM (using LoRA) gives small additional improvements: –16.2 →15.7% WER 17 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Results - Which components to train? Internal results at AppTek: •For Librispeech, only fine-tuning the adapter gives reasonable results –similar to related work Wang et al. (2023b); Ma et al. (2024b) •On the AppTek Spanish task (diverse domains, multi-bandwidth), only fine-tuning the adapter did not converge and fine-tuning of the audio encoder was necessary for good results •Further, fine-tuning the LLM (using LoRA) gives small additional improvements: –16.2 →15.7% WER 17 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Results - Which components to train? Internal results at AppTek: •For Librispeech, only fine-tuning the adapter gives reasonable results –similar to related work Wang et al. (2023b); Ma et al. (2024b) •On the AppTek Spanish task (diverse domains, multi-bandwidth), only fine-tuning the adapter did not converge and fine-tuning of the audio encoder was necessary for good results •Further, fine-tuning the LLM (using LoRA) gives small additional improvements: –16.2 →15.7% WER 17 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Results - Which components to train? Internal results at AppTek: •For Librispeech, only fine-tuning the adapter gives reasonable results –similar to related work Wang et al. (2023b); Ma et al. (2024b) •On the AppTek Spanish task (diverse domains, multi-bandwidth), only fine-tuning the adapter did not converge and fine-tuning of the audio encoder was necessary for good results •Further, fine-tuning the LLM (using LoRA) gives small additional improvements: –16.2 →15.7% WER 17 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - ASR Results - Competitiveness on Real-world tasks? 18 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Low-Latency LLM text user voice agent voice encoder ASR LLM decoder TTS text VAD •Output streaming: –Autoregressive token generation →output streaming by design •Input streaming: –Main Idea: always maintain the current KV cache –Maintain and update the KV cache for system prompt, user context, and conversation history –Incrementally extend KV-cache for new tokens from the ASR stream •Context Rollback: –LLM and TTS generate text and audio faster than real-time –User interruptions require to roll back to the last word heard by the user ■e.g. user: “Can you repeat the last word?” –Approach: track timing through the whole system and discard KV-cache 21 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Speechworthy Output Figure from Cho et al. (2024) 22 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Speechworthy Output Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10652–10670 November 12-16, 2024 ©2024 Association for Computational Linguistics Speechworthy Instruction-tuned Language Models Hyundong Cho1⇤ , Nicolaas Jedema2, Leonardo F. R. Ribeiro2, Karishma Sharma2, Pedro Szekely2, Alessandro Moschitti2, Ruben Janssen2,and Jonathan May1 1University of Southern California, Information Sciences Institute 2Amazon [email protected] Abstract Current instruction-tuned language models are exclusively trained with textual preference data and thus are often not aligned with the unique requirements of other modalities, such as speech. To better align language models with the speech domain, we explore (i) prompting strategies grounded in radio-industry best practices and (ii) preference learning using a novel speech-based preference data of 20K samples, generated with a wide spectrum of prompts that induce varying dimensions of speech-suitability and labeled by annotators who listen to response pairs. Both human and automatic evaluation show that both prompting and preference learning increase the speechsuitability of popular instruction-tuned LLMs. Interestingly, we find that prompting and preference learning can be additive; combining them achieves the best win rates in head-tohead comparison, resulting in responses that are preferred or tied to the base model in 76.2% of comparisons on average. Lastly, we share lexical, syntactical, and qualitative analyses to showcase how each method contributes to improving the speech-suitability of generated responses. 1 Introduction Speech is one of our primary means of communication and a convenient and popular mode for interacting with virtual assistants (Yang,2004). Virtual assistants are a prime application of instructiontuned language models (ITLM), as both seek to provide helpful responses to user requests (Peng et al., 2023;Chung et al.,2022;Wang et al.,2022a,b;Wei et al.,2021;Sanh et al.,2022;Zhou et al.,2023). However, current ITLMs are fine-tuned on textual instructions (Peng et al.,2023;Chung et al.,2022; Wang et al.,2022a,b;Wei et al.,2021;Sanh et al., 2022;Zhou et al.,2023) and preferences obtained ⇤Work was done while HC was an intern at Amazon. Figure 1: Current instruction-tuned language models tend to generate verbose responses with nonvocalizable content, such as bullet lists or parentheses, that are not suitable for responses that are delivered as speech by voice assistants (left, from OLMo 7B Instruct). Speech is serial and transient, and therefore concise yet informative responses with conversational follow-up questions are often preferred (right, from our adapted OLMo model). from annotators that read textual responses (Bai et al.,2022;Ethayarajh et al.,2022;Ouyang et al., 2022;Touvron et al.,2023). We hypothesize that user preferences for speech and text are different and that the shift from generating written text to spoken language may pose a challenge for ITLMs (González et al.,2021). Speech is serial and transient; speech processing is strictly linear (Flowerdew,1994) and requires higher cognitive load than reading (Thompson and Rubin,1996;Osada,2004). Concise and simple sentences, as seen on the right side of Figure 1, are thus often preferred in speech (Kern,2008;Abel, 2015), yet current ITLMs optimized with textual preference datasets (Stiennon et al.,2020;Singhal et al.,2023) produce verbose responses with non-vocalizable content (left side of Figure 1). We test the hypothesis that current ITLMs are 10652 Figure from Cho et al. (2024) 22 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Speechworthy Output Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10652–10670 November 12-16, 2024 ©2024 Association for Computational Linguistics Speechworthy Instruction-tuned Language Models Hyundong Cho1⇤ , Nicolaas Jedema2, Leonardo F. R. Ribeiro2, Karishma Sharma2, Pedro Szekely2, Alessandro Moschitti2, Ruben Janssen2,and Jonathan May1 1University of Southern California, Information Sciences Institute 2Amazon [email protected] Abstract Current instruction-tuned language models are exclusively trained with textual preference data and thus are often not aligned with the unique requirements of other modalities, such as speech. To better align language models with the speech domain, we explore (i) prompting strategies grounded in radio-industry best practices and (ii) preference learning using a novel speech-based preference data of 20K samples, generated with a wide spectrum of prompts that induce varying dimensions of speech-suitability and labeled by annotators who listen to response pairs. Both human and automatic evaluation show that both prompting and preference learning increase the speechsuitability of popular instruction-tuned LLMs. Interestingly, we find that prompting and preference learning can be additive; combining them achieves the best win rates in head-tohead comparison, resulting in responses that are preferred or tied to the base model in 76.2% of comparisons on average. Lastly, we share lexical, syntactical, and qualitative analyses to showcase how each method contributes to improving the speech-suitability of generated responses. 1 Introduction Speech is one of our primary means of communication and a convenient and popular mode for interacting with virtual assistants (Yang,2004). Virtual assistants are a prime application of instructiontuned language models (ITLM), as both seek to provide helpful responses to user requests (Peng et al., 2023;Chung et al.,2022;Wang et al.,2022a,b;Wei et al.,2021;Sanh et al.,2022;Zhou et al.,2023). However, current ITLMs are fine-tuned on textual instructions (Peng et al.,2023;Chung et al.,2022; Wang et al.,2022a,b;Wei et al.,2021;Sanh et al., 2022;Zhou et al.,2023) and preferences obtained ⇤Work was done while HC was an intern at Amazon. Figure 1: Current instruction-tuned language models tend to generate verbose responses with nonvocalizable content, such as bullet lists or parentheses, that are not suitable for responses that are delivered as speech by voice assistants (left, from OLMo 7B Instruct). Speech is serial and transient, and therefore concise yet informative responses with conversational follow-up questions are often preferred (right, from our adapted OLMo model). from annotators that read textual responses (Bai et al.,2022;Ethayarajh et al.,2022;Ouyang et al., 2022;Touvron et al.,2023). We hypothesize that user preferences for speech and text are different and that the shift from generating written text to spoken language may pose a challenge for ITLMs (González et al.,2021). Speech is serial and transient; speech processing is strictly linear (Flowerdew,1994) and requires higher cognitive load than reading (Thompson and Rubin,1996;Osada,2004). Concise and simple sentences, as seen on the right side of Figure 1, are thus often preferred in speech (Kern,2008;Abel, 2015), yet current ITLMs optimized with textual preference datasets (Stiennon et al.,2020;Singhal et al.,2023) produce verbose responses with non-vocalizable content (left side of Figure 1). We test the hypothesis that current ITLMs are 10652 Figure from Cho et al. (2024) Approach by Cho et al. (2024) •Define different criteria to judge speech worthiness •Detailed system prompt to define output format •Further improvements by collecting human feedback + preference optimization 22 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Asynchronous Function Calling and Reasoning Function calling and reasoning is crucial for many use-cases, but introduces large latencies •Hide latencies with predefined outputs or ”typing” indicators •“Thinking while Listening” Shih et al. (2025); Chiang et al. (2025) •Execute function calls in the background on continue interaction on other topics (existing work for text interaction by Gim et al. (2024)) 23 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Asynchronous Function Calling and Reasoning Function calling and reasoning is crucial for many use-cases, but introduces large latencies •Hide latencies with predefined outputs or ”typing” indicators •“Thinking while Listening” Shih et al. (2025); Chiang et al. (2025) •Execute function calls in the background on continue interaction on other topics (existing work for text interaction by Gim et al. (2024)) 23 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Asynchronous Function Calling and Reasoning Function calling and reasoning is crucial for many use-cases, but introduces large latencies •Hide latencies with predefined outputs or ”typing” indicators •“Thinking while Listening” Shih et al. (2025); Chiang et al. (2025) Text Monologue User Audio Streaming ASR delay (480 ms) What is the capital of France? Is it Barcelona, London, or Paris? System Audio It is Paris. is the capital of <start_cot> <switch_asr> Is it Barcelona, <switch_asr> London, capital <switch_cot> The <switch_cot> or of <switch_cot> Paris? What France is Paris. <end_cot> It Is Paris <finish_ans> France? Figure1 Trainingtokensequencearrangement.Wetrainthemodeltointerleave reasoning tokens RT with streaming ASR tokens QT on the text monologue channel, with special switch tokens for mode switching. After the CoT ends, the model generates text tokens which align with the spoken response RT .Forsimplicity, [PAD] and [EPAD] tokens are not shown here. long an LLM should reason (Sprague et al.,2025), resulting in a growing interest in “hybrid” reasoning models. Although some recent work has adopted CoT in the speech domain, they focus primarily on applications such as speech translation Hu et al. (2025); Du et al. (2025); Gállego et al. (2025), dialogue Arora et al. (2025), or other detection tasks Mai et al. (2025); Park et al. (2025). The integration of CoT in speech LLMs requires answering two research questions: (i) should models reason using text or speech, and (ii) how do we maintain the responsiveness required for spoken interactions? To answer the first question, we investigate both alternatives, showing that text-based CoT is as performant as speech-based CoT for improving reasoning in speech LLMs, while being 2x more token-efficient. The sequential process of listening, reasoning, and responding introduces considerable latency; consequently, previous research has proposed methods to overlap CoT tokens with speech to improve real-time conversational AI. Building upon the anthropomorphism of speech LLMs, concurrent works such as STITCH (Chiang et al., 2025)andMini-Omni-Reasoner(Xie et al.,2025) have proposed “thinking while speaking,” i.e., the model begins its spoken response while its reasoning is still ongoing. This is achieved by interleaving chunks of reasoning tokens with spoken response tokens, and subsequent CoT chunks are generated in the time it takes for the audio decoder to synthesize the preceding response. Despite showing reasonable improvements, this approach has notable limitations. For instance, the optimal chunk size for interleaving requires careful tuning and is dependent on hardware limitations. Moreover, despite a reduction in the time to first word, the model may inadvertently vocalize too much of its reasoning, leading to a longer overall response time to a final, conclusive answer. In this paper, we draw inspiration from neuroscience (Donhauser and Baillet,2019)to propose a novel “thinking while listening” paradigm, by enabling concurrent processing of text-based CoT and user speech. Current speech LLM architectures may be broadly categorized into two types: single-stream and multi-stream. Single-stream architectures merge user/system speech and text into a unified token sequence (Kim et al.,2024; Veluri et al.,2024), while multi-stream architectures simultaneously model distinct streams for each token sequence (Défossez et al.,2024). In this work, we build upon a multi-stream architecture due to its superior capacity for the concurrent processing of user audio and reasoning tokens. This design provides significant flexibility by allowing the system’s text stream to be revised independently, a key advantage over single-stream models that lack this decoupling. Specifically, we fine-tune the publicly available Moshi model (Défossez et al., 2024) to generate CoT within its text monologue stream to improve its reasoning capabilities (Section 2). To enable the model to think while listening, we propose two methods: (i) a novel metric that estimates the completeness of the user’s question at each timestep, and (ii) a preference tuning scheme to update the model’s reasoning dynamically with new input (Section 3). Since there are no existing standard reasoning evaluations for speech LLMs, we curated a suite of single-turn spoken reasoning tasks from well-known text-based reasoning benchmarks comprising mathematical reasoning, social/physical interactions, and other general reasoning tasks (Section 4.2). Overall, our contributions are summarized below. 2 Figure from Shih et al. (2025) •Execute function calls in the background on continue interaction on other topics (existing work for text interaction by Gim et al. (2024)) 23 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - LLM Asynchronous Function Calling and Reasoning Function calling and reasoning is crucial for many use-cases, but introduces large latencies •Hide latencies with predefined outputs or ”typing” indicators •“Thinking while Listening” Shih et al. (2025); Chiang et al. (2025) Text Monologue User Audio Streaming ASR delay (480 ms) What is the capital of France? Is it Barcelona, London, or Paris? System Audio It is Paris. is the capital of <start_cot> <switch_asr> Is it Barcelona, <switch_asr> London, capital <switch_cot> The <switch_cot> or of <switch_cot> Paris? What France is Paris. <end_cot> It Is Paris <finish_ans> France? Figure1 Trainingtokensequencearrangement.Wetrainthemodeltointerleave reasoning tokens RT with streaming ASR tokens QT on the text monologue channel, with special switch tokens for mode switching. After the CoT ends, the model generates text tokens which align with the spoken response RT .Forsimplicity, [PAD] and [EPAD] tokens are not shown here. long an LLM should reason (Sprague et al.,2025), resulting in a growing interest in “hybrid” reasoning models. Although some recent work has adopted CoT in the speech domain, they focus primarily on applications such as speech translation Hu et al. (2025); Du et al. (2025); Gállego et al. (2025), dialogue Arora et al. (2025), or other detection tasks Mai et al. (2025); Park et al. (2025). The integration of CoT in speech LLMs requires answering two research questions: (i) should models reason using text or speech, and (ii) how do we maintain the responsiveness required for spoken interactions? To answer the first question, we investigate both alternatives, showing that text-based CoT is as performant as speech-based CoT for improving reasoning in speech LLMs, while being 2x more token-efficient. The sequential process of listening, reasoning, and responding introduces considerable latency; consequently, previous research has proposed methods to overlap CoT tokens with speech to improve real-time conversational AI. Building upon the anthropomorphism of speech LLMs, concurrent works such as STITCH (Chiang et al., 2025)andMini-Omni-Reasoner(Xie et al.,2025) have proposed “thinking while speaking,” i.e., the model begins its spoken response while its reasoning is still ongoing. This is achieved by interleaving chunks of reasoning tokens with spoken response tokens, and subsequent CoT chunks are generated in the time it takes for the audio decoder to synthesize the preceding response. Despite showing reasonable improvements, this approach has notable limitations. For instance, the optimal chunk size for interleaving requires careful tuning and is dependent on hardware limitations. Moreover, despite a reduction in the time to first word, the model may inadvertently vocalize too much of its reasoning, leading to a longer overall response time to a final, conclusive answer. In this paper, we draw inspiration from neuroscience (Donhauser and Baillet,2019)to propose a novel “thinking while listening” paradigm, by enabling concurrent processing of text-based CoT and user speech. Current speech LLM architectures may be broadly categorized into two types: single-stream and multi-stream. Single-stream architectures merge user/system speech and text into a unified token sequence (Kim et al.,2024; Veluri et al.,2024), while multi-stream architectures simultaneously model distinct streams for each token sequence (Défossez et al.,2024). In this work, we build upon a multi-stream architecture due to its superior capacity for the concurrent processing of user audio and reasoning tokens. This design provides significant flexibility by allowing the system’s text stream to be revised independently, a key advantage over single-stream models that lack this decoupling. Specifically, we fine-tune the publicly available Moshi model (Défossez et al., 2024) to generate CoT within its text monologue stream to improve its reasoning capabilities (Section 2). To enable the model to think while listening, we propose two methods: (i) a novel metric that estimates the completeness of the user’s question at each timestep, and (ii) a preference tuning scheme to update the model’s reasoning dynamically with new input (Section 3). Since there are no existing standard reasoning evaluations for speech LLMs, we curated a suite of single-turn spoken reasoning tasks from well-known text-based reasoning benchmarks comprising mathematical reasoning, social/physical interactions, and other general reasoning tasks (Section 4.2). Overall, our contributions are summarized below. 2 Figure from Shih et al. (2025) •Execute function calls in the background on continue interaction on other topics (existing work for text interaction by Gim et al. (2024)) 23 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - TTS Neural Audio Codecs text user voice agent voice encoder ASR LLM decoder TTS text VAD •How to generate speech with LLMs? –Full waveforms or spectograms difficult to predict –Solution: speech tokens •Autoencoder Architecture: More details: https://kyutai.org/next/codec-explainer 24 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - TTS Neural Audio Codecs text user voice agent voice encoder ASR LLM decoder TTS text VAD •How to generate speech with LLMs? –Full waveforms or spectograms difficult to predict –Solution: speech tokens •Autoencoder Architecture: Moshi: a speech-text foundation model for real-time dialogue Figure 2: Architecture and training of Mimi, our neural audio codec, with its split residual vector quantization. During training (blue part, top), we distill noncausal embeddings from WavLM (Chen et al.,2022) into a single vector quantizer which produces semantic tokens, and is combined with separate acoustic tokens for reconstruction. 3.3.1 Architecture Our baseline architecture takes inspiration from SoundStream (Zeghidour et al.,2022) and Encodec (D´efossez et al.,2023) and consists of a SeaNet (Tagliasacchi et al.,2020) autoencoder and a Residual Vector Quantizer (Zeghidour et al.,2022). The encoder projects a single-channel waveform x2RLto a latent representation enc(x)2RS⇥Dby cascading residual convolutional blocks that interleave dilated (van den Oord et al.,2016) and strided convolutions along with ELU (Clevert et al.,2016) non-linearities and Weight Normalization (Salimans and Kingma,2016). All convolutions are causal, such that this autoencoder can run in a streaming fashion. With 4 convolutional blocks and respective striding factors (4,5,6,8), and a final 1D convolution with stride 2, Mimi’s encoder projects a 24kHz waveform to a latent representation of 12.5 frames per second and dimension D= 512. Symmetrically, the decoder adopts a similar structure but with transposed convolutions rather than strided ones, to project the latent representation back to 24kHz audio. We discretize the latent space with a Residual Vector Quantizer (Zeghidour et al.,2022), which iteratively applies vector quantization (VQ) to the residuals of the previous quantizer. With Qquantizers, each with a codebook of NAcentroids, the RVQ discretizes the latent space into {1,...,N A}S⇥Q. As a baseline, we train this model with a combination of reconstruction and adversarial losses, following the setup of Encodec (D´efossez et al.,2023). We detail below the main changes of Mimi with respect to this default configuration. Transformer-based bottleneck. To improve the ability of Mimi to encode speech into compact representations while reconstructing high-quality audio, we add Transformer modules in the bottleneck, one right before quantization and one after. These Transformers have 8 layers, 8 heads, RoPE position encodings, a finite context of 250 frames (20 seconds), GELU (Hendrycks and Gimpel,2016a) activations, a model dimension of 512 and an MLP dimension of 2048. To stabilize training, we use LayerScale (Touvron et al.,2021), with initialization of the diagonal values at 0.01. Both Transformers use causal masking, which preserves the compatibility of the whole architecture with streaming inference. Both 10 Architecture and training of Mimi, Kyutai’s neural audio codec, figure from Défossez et al. (2024) More details: https://kyutai.org/next/codec-explainer 24 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - TTS Streaming: Delayed Streams Modeling •Text and audio processed by LM as two simultaneous streams (addition in embedding space) •Originally, proposed by Kyutai for their full duplex model Moshi (Défossez et al., 2024) •For ASR: delay and predict text stream and force audio from input •For TTS: delay and predict audio stream –Problem: when to feed in the next word? –Separate action stream that emits when a word is uttered –Results in word-level timestamps Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Adapted from Zeghidour et al. (2025) 28 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - TTS Streaming: Delayed Streams Modeling •Text and audio processed by LM as two simultaneous streams (addition in embedding space) •Originally, proposed by Kyutai for their full duplex model Moshi (Défossez et al., 2024) •For ASR: delay and predict text stream and force audio from input •For TTS: delay and predict audio stream –Problem: when to feed in the next word? –Separate action stream that emits when a word is uttered –Results in word-level timestamps Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Adapted from Zeghidour et al. (2025) 28 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - TTS Streaming: Delayed Streams Modeling •Text and audio processed by LM as two simultaneous streams (addition in embedding space) •Originally, proposed by Kyutai for their full duplex model Moshi (Défossez et al., 2024) •For ASR: delay and predict text stream and force audio from input •For TTS: delay and predict audio stream –Problem: when to feed in the next word? –Separate action stream that emits when a word is uttered –Results in word-level timestamps Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Automatic Speech Recognition Text-To-Speech Audio stream Text stream Audio stream Text stream Alan [pad] [word] ⎵Tu ring [word] ⎵was [word] ⎵an [word] /ˈæl/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ [word] [word] [word] [word] /ʌ//wəz//ən/ Alan [pad] [word] ⎵Tu ring [word] ⎵was [pad] /ˈæl/ /ən/ /ʌ/ /ˈtjʊə//rɪŋ/ /ʌ/ /wəz/ /ən//ʌ/ /ˈɪŋ/ Action stream Figure 1: Delayed streams modeling for speech-text tasks. Depending on which stream is delayed with respect to the other, we solve either an ASR or a TTS task. For TTS, we further need an action stream for the model to let us know when it is ready to receive a new word. dialogue, which predicts text and audio tokens in a streaming fashion, later applied by Labiausse et al. (2025) to real-time speech translation. In this work we extend the approach of Défossez et al. (2024), in order to reach state-of-the-art performance on the two most competitive speech-text tasks, namely ASR and TTS. Moreover, while Défossez et al. (2024) and Labiausse et al. (2025) operate with a delay specified before training, we propose delay conditioning for inference-time latency control without retraining. Our TTS covers both monologue and controllable dialog generation, a topic that was studied by CoVoMix (Zhang et al., 2024), although at a lower sample rate (8 kHz) and not streaming. 3 Method Notation. We wish to solve a sequence-to-sequence task between two domains X and Y . Each domain consists of sequences of vectors of all possible lengths, e.g. X=[ T2N {(Xt)2RT⇥d},Y=[ T02N {(Yt0)2RT0⇥d0}.(1) In the case where either Xt or Yt is discrete-valued, we can use a one-hot representation for it in Eq. (1) . We assume that we are given a joint probability distribution over the outer product domain X⇥Y , and that we have the random variables X2X and Y2Y , along with the joint distribution P[X,Y ]=p(X,Y ).(2) We also introduce T2N (resp. T0 ) the random variable indicating the length of X (resp. Y ), along with the marginals p(X) and p(Y) . For any sequence Z , and index t , we denote Z<t = (Z1,...,Z t1), potentially empty if t0. We similarly define Zt,Zt, and Z>t. Sequence-to-sequence as joint modeling. Let’s assume for this paragraph that X is the set of all possible monophonic waveforms sampled at 24 kHz, and Y is made of sequences of one-hot encoded vectors over a set of words. Intuitively, we assume there exists a coupling p(X,Y ) such that p(X,Y ) is high if Y represents the transcription of X , or conversely, if X represents a speech utterance of the text given by Y . Formally, the task of ASR corresponds to sampling from the distribution P[Y|X] , while the task of TTS corresponds to sampling from the distribution P[X|Y] . Thus, each task can be solved by accurately estimating both probability distributions, q(X,Y )⇡P[Y|X],q 0(Y,X)⇡P[X|Y].(3) For simplicity, we now only focus on estimating P[Y|X] , the inverse task being obtained by exchanging the definition of Xand Y. We thus call Xthe input domain, and Ythe output domain. Auto-regressive modeling of Y .A good candidate for estimating P[Y|X] is auto-regressive modeling, with a Transformer model (Vaswani et al., 2017), under the extra assumption that the output domain Ycan be discretized. Thus, one would estimate q(y|X,Y<t)⇡P[Yt=y|X,Y<t].(4) 3 Adapted from Zeghidour et al. (2025) 28 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - VAD Voice Activity Detection text user voice agent voice encoder ASR LLM decoder TTS text VAD Switch between LISTEN and SPEAK states: •Endpointing: Detect when the user stops speaking But, do not detect longer pauses while speaking or background noise as speech •Barge-In Handling: Detect when the user interrupts the system But, do not detect back-channeling as interrupt Approaches •Classic VAD: simple detection •Semantic VAD: incorporates semantic information to make decision •Integrated VAD: semantic detection as part of the ASR model 29 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded Architecture - VAD Voice Activity Detection text user voice agent voice encoder ASR LLM decoder TTS text VAD Switch between LISTEN and SPEAK states: •Endpointing: Detect when the user stops speaking But, do not detect longer pauses while speaking or background noise as speech •Barge-In Handling: Detect when the user interrupts the system But, do not detect back-channeling as interrupt Approaches •Classic VAD: simple detection •Semantic VAD: incorporates semantic information to make decision •Integrated VAD: semantic detection as part of the ASR model 29 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Integrated SpeechLLMs: Overview cascaded integrated LLM text user voice user voice agent voice agent voice speech token/embedding encoder ASR LLM decoder TTS text encoder decoder speech token/embedding Speech-to-speech model variants: •Half-Duplex: model is either in speaking or listening state •Full-Duplex: models can listen and speak at the same time Full Duplex Half Duplex Simplex Figure from Ma et al. (2024a) 30 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Half-Duplex Models •Examples: LLaMA-Omni (Fang et al., 2025), Qwen-2.5 Omni (Xu et al., 2025), … •Usually not optimised for full-duplex operation and require external VAD •Freeze-Omni (Wang et al., 2025b): Integrates endpointing into the SpeechLLM 31 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Full-Duplex: dGSLM •First full-duplex model (Nguyen et al., 2023) •Separate decoder for each audio stream •Interaction via cross-attention layers •DLM trained only on 2000h from Fisher •Natural sounding but intelligible speech •Sample 1: Play •Sample 2: Play ASR. Briefly, we build on self-supervised discrete speech representations models, which we train on spontaneous conversations with each speaker having his or her own audio channel. After training, the speech units come to represent not only verbal but also nonverbal materials. We can now encode a conversation between two interlocutors as two parallel streams of discrete tokens. We then introduce a novel dual-tower transformer architecture, where each channel is processed by one ‘‘tower’’ of the model that learns via an autoregressive loss, but the two towers also communicate via cross-attention in their hidden units. This cross-attention is critical for the correct synchronization of the two channels and result in a naturalistic distribution of turns, overlap and pauses. While this system is not trained on enough data to capture deep syntactic and semantic aspects of dialogue, and indeed scores below a text-based cascaded ASR+LM+TTS model on semantic content, it does capture better surface characteristics of chitchat in mimicking accurately turn-taking and backchanneling. This can be seen as a proof of principle that previously difficult to capture aspects of spontaneous conversations can be captured with minimally modified language modeling techniques. Finally, our model opens up new possibilities to create more natural naturalistic human-machine dialogue systems in the future. 2 Related Work Unsupervised Spoken Language Modeling. Recently, great advances have been achieved in the area of representation learning from raw audio. Models trained with either autoencoder objectives (Ondel et al., 2016; van den Oord et al., 2017) or masked objectives (CPC: van den Oord et al., 2018; APC: Chung and Glass, 2020; wav2vec 2.0: Baevski et al., 2020; HuBERT: Hsu et al., 2021a; MockingJay: Liu et al., 2020) from raw speech can learn audio representation that can be used for a variety of downstream tasks (Yang et al., 2021), see Borgholt et al. (2022) for a review. Most of these models build a codebook of discrete units, either as latent representation or as targets. The discrete representation can in turn be fed to a standard autoregressive language model, which can then be sampled to generate new speech sequences (Lakhotia et al., 2021; Dieleman et al., 2021). An interesting aspect of this procedure Figure 1: General Schema for dGSLM: A discrete encoder (HuBERT+kmeans) turns each channel of a dialogue into a string of discrete units (c1,..c N).A Dialogue Language Model (DLM) is trained to autoregressively produce units that are turned into waveforms using a decoder (HifiGAN). is that it can capture aspects of speech that are typically not available in written transcriptions and can therefore model prosody and intonation (Kharitonov et al., 2021), or non verbal vocalizations typical of emotional speech (Kreuk et al., 2021). Up to now, however, no such model has been applied to multi-party conversational speech. Dialogue Generation. Since the early work on end-to-end neural dialogue generation (Vinyals and Le, 2015; Li et al., 2015; Serban et al., 2016), empowered by scalable methods for language representation (Radford et al., 2018; Lewis et al., 2019), there has been enormous progress in the area of dialogue generation (Roller et al., 2020; 251 Figure from Nguyen et al. (2023) 32 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Full-Duplex: Moshi Hello, I’m AppTek. Who are you? A1 A2 A3 A4 A5 A7 A8 A9 Agent audio tokens A11 A12 U1 U2 U3 U4 U5 U7 U8 U9 U11 A6 U6 U10 A10 A1 A2 A3 A4 A5 llo, _I ‘m Agent text tokens PAD _App He User audio tokens Tek A13 A1 A2 A3 A4 A5 A7 A8 A6 Input stream Single stream U12 U13 U1 U2 U3 U4 U5 User audio tokens Agent audio tokens •Similar to delayed streams model: decoder processes multiple streams in parallel •Inner monologue (agent text tokens): helps to generate more coherent speech •LLM backbone: 7B parameter model •Training stages: 0. Text LLM and neural codec Training 1. Single-stream ASR and TTS pre-training (by alternating delay) – 8M hours 2. Multi-stream post-training: automatic speaker diarization of audio data – 400k hours 3. Multi-stream fine-tuning: fisher telephone data – 2k hours 4. Multi-stream fine-tuning: high-quality synthetic data – 20k hours •Demo 33 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Full-Duplex: Moshi Hello, I’m AppTek. Who are you? A1 A2 A3 A4 A5 A7 A8 A9 Agent audio tokens A11 A12 U1 U2 U3 U4 U5 U7 U8 U9 U11 A6 U6 U10 A10 A1 A2 A3 A4 A5 llo, _I ‘m Agent text tokens PAD _App He User audio tokens Tek A13 A1 A2 A3 A4 A5 A7 A8 A6 Input stream Single stream U12 U13 U1 U2 U3 U4 U5 User audio tokens Agent audio tokens •Similar to delayed streams model: decoder processes multiple streams in parallel •Inner monologue (agent text tokens): helps to generate more coherent speech •LLM backbone: 7B parameter model •Training stages: 0. Text LLM and neural codec Training 1. Single-stream ASR and TTS pre-training (by alternating delay) – 8M hours 2. Multi-stream post-training: automatic speaker diarization of audio data – 400k hours 3. Multi-stream fine-tuning: fisher telephone data – 2k hours 4. Multi-stream fine-tuning: high-quality synthetic data – 20k hours •Demo 33 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Full-Duplex: Moshi Hello, I’m AppTek. Who are you? A1 A2 A3 A4 A5 A7 A8 A9 Agent audio tokens A11 A12 U1 U2 U3 U4 U5 U7 U8 U9 U11 A6 U6 U10 A10 A1 A2 A3 A4 A5 llo, _I ‘m Agent text tokens PAD _App He User audio tokens Tek A13 A1 A2 A3 A4 A5 A7 A8 A6 Input stream Single stream U12 U13 U1 U2 U3 U4 U5 User audio tokens Agent audio tokens •Similar to delayed streams model: decoder processes multiple streams in parallel •Inner monologue (agent text tokens): helps to generate more coherent speech •LLM backbone: 7B parameter model •Training stages: 0. Text LLM and neural codec Training 1. Single-stream ASR and TTS pre-training (by alternating delay) – 8M hours 2. Multi-stream post-training: automatic speaker diarization of audio data – 400k hours 3. Multi-stream fine-tuning: fisher telephone data – 2k hours 4. Multi-stream fine-tuning: high-quality synthetic data – 20k hours •Demo 33 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Full-Duplex: Moshi Hello, I’m AppTek. Who are you? A1 A2 A3 A4 A5 A7 A8 A9 Agent audio tokens A11 A12 U1 U2 U3 U4 U5 U7 U8 U9 U11 A6 U6 U10 A10 A1 A2 A3 A4 A5 llo, _I ‘m Agent text tokens PAD _App He User audio tokens Tek A13 A1 A2 A3 A4 A5 A7 A8 A6 Input stream Single stream U12 U13 U1 U2 U3 U4 U5 User audio tokens Agent audio tokens •Similar to delayed streams model: decoder processes multiple streams in parallel •Inner monologue (agent text tokens): helps to generate more coherent speech •LLM backbone: 7B parameter model •Training stages: 0. Text LLM and neural codec Training 1. Single-stream ASR and TTS pre-training (by alternating delay) – 8M hours 2. Multi-stream post-training: automatic speaker diarization of audio data – 400k hours 3. Multi-stream fine-tuning: fisher telephone data – 2k hours 4. Multi-stream fine-tuning: high-quality synthetic data – 20k hours •Demo 33 of 42 Towards Natural Realtime Voice Interaction with Large Language Models End-to-End (Integrated) Architectures Full-Duplex: Moshi Hello, I’m AppTek. Who are you? A1 A2 A3 A4 A5 A7 A8 A9 Agent audio tokens A11 A12 U1 U2 U3 U4 U5 U7 U8 U9 U11 A6 U6 U10 A10 A1 A2 A3 A4 A5 llo, _I ‘m Agent text tokens PAD _App He User audio tokens Tek A13 A1 A2 A3 A4 A5 A7 A8 A6 Input stream Single stream U12 U13 U1 U2 U3 U4 U5 User audio tokens Agent audio tokens •Similar to delayed streams model: decoder processes multiple streams in parallel •Inner monologue (agent text tokens): helps to generate more coherent speech •LLM backbone: 7B parameter model •Training stages: 0. Text LLM and neural codec Training 1. Single-stream ASR and TTS pre-training (by alternating delay) – 8M hours 2. Multi-stream post-training: automatic speaker diarization of audio data – 400k hours 3. Multi-stream fine-tuning: fisher telephone data – 2k hours 4. Multi-stream fine-tuning: high-quality synthetic data – 20k hours •Demo 33 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Cascaded vs. Integrated: Pros and Cons Trade-offs and When to Use Which Full-Duplex Integrated Architecture: •Theoretically, lower latency and better turn-taking capabilities •Adaptation requires full duplex training data •Model can utilise paralinguistic cues Cascaded Architecture: •Allows to use smaller models for some components •Easier to customize •More control and robustness •Requires no end-to-end training data 34 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Evaluation and Benchmarking Conversation Dynamics: Full-Duplex-Bench TABLE III: Models comparison. We evaluate several models across different conversational dimensions, where Latency is presented in seconds. Dimension Pause Handling Backchannel Smooth Turn Taking User Interruption Data Synthetic Candor ICC Candor Synthetic Metric TOR (#) TOR (#) TOR (#) Freq (") JSD (#) TOR (") Latency (#) TOR (") GPT-4o(") Latency (#) dGSLM 0.934 0.935 0.691 0.015 0.934 0.975 0.352 0.917 0.201 2.531 Moshi 0.985 0.980 1.000 0.001 0.957 0.941 0.265 1.000 0.765 0.257 Freeze-Omni 0.642 0.481 0.636 0.001 0.997 0.336 0.953 0.867 3.615 1.409 Gemini Live 0.255 0.310 0.091 0.012 0.896 0.655 1.301 0.891 3.376 1.183 interruptions through a multi-stream architecture. We use the official implementation8. • Freeze-Omni [10]:A cascaded full-duplex system with a frozen LLM pipeline. VAD triggers chunk-wise encoding, and a classification head predicts dialogue states to control turn-taking. Parallel modules handle streaming, speaking, and monitoring. Evaluated locally using the official server. • Gemini Live: Based on the official API documentation 9 , we use the gemini-2.0-flash-live-001 model. The input.wav file is first converted to 16 kHz PCM-16 format, then divided into 30 ms chunks and streamed to the Gemini Live API. Server-side voice activity detection (VAD) is used to segment the audio and trigger responses. A new session is initiated after each model reply. All generated outputs are aligned with the original input duration, preserving silence in regions where no response is produced. V. RESULTS Table III presents the results across four evaluation dimensions. The key findings are summarized as follows: Avoiding Interruptions During Speaker Pauses: All three SDMs exhibit high Takeover Rates (TOR) when managing speaker pauses, indicating frequent interruptions during natural breaks. The end-to-end models (dGSLM and Moshi) tend to interrupt more often across both real and synthetic datasets. In contrast, Freeze-Omni, which incorporates a dedicated module for predicting speaking and listening states, demonstrates a significantly lower TOR. This suggests that an explicit turntaking control module can better manage pauses and may offer improvements if integrated into end-to-end systems. Conversely, Gemini Live achieves the lowest TOR among all models, including open-source ones, and exhibits a higher likelihood of taking over on Candor compared to synthetic data. Backchanneling Dynamics: We evaluate using TOR, backchannel frequency (Freq), and Jensen–Shannon Divergence (JSD). Similar to pause handling, Moshi frequently takes over the turn, resulting in a high TOR. Both dGSLM and FreezeOmni achieve lower TORs. Among the three open-sourced SDMs, dGSLM produces the most backchannel responses with more natural timing (Freq = 0.015, JSD = 0.934) when it does not take over the turn. In contrast, Freeze-Omni remains largely silent, producing few backchannel responses. The commercial Gemini Live achieves the lowest TOR and the best JSD, indicating superior ability in identifying appropriate moments for backchanneling. 8https://github.com/kyutai-labs/moshi 9https://ai.google.dev/gemini-api/docs/live Latency of Turn-Taking: We assessed TOR and average response latency on the Candor dataset. dGSLM and Moshi exhibit high TORs and respond quickly, with an average latency of around 0.3 seconds. Freeze-Omni, due to its cascaded architecture—which first generates text and then synthesizes speech—exhibits higher latency. Its lower TOR likely reflects missed opportunities to take over the turn, possibly due to failures in detecting turn ends. Interestingly, Gemini Live achieves a TOR of only 0.655, suggesting that even commercial models sometimes fail to take the turn in real dialogue data. Managing User Interruptions: Freeze-Omni handles user interruptions effectively, achieving significantly higher contextual relevance scores while maintaining acceptable latency. This strength is attributed to its ’model-as-a-server’ strategy, which leverages a pool of models to manage user barge-ins efficiently. Gemini Live shows comparable performance, with relatively better TOR and latency, though it yields a slightly lower GPT-4o score. In contrast, the end-to-end SDMs struggle with coherence. Moshi responds promptly but yields a lower GPT4o score (0.765), while dGSLM performs poorly, with high latency and diminished content quality (GPT-4o score: 0.201). These results highlight the challenges end-to-end systems face in preserving semantic coherence during user interruptions. VI. CONCLUSION In this paper, we introduced Full-Duplex-Bench, a benchmark designed to evaluate critical aspects of full-duplex spoken dialogue models. Our framework targets key interaction dimensions—pause handling, backchanneling, smooth turn-taking, and user interruption management—addressing limitations of existing benchmarks that primarily focus on half-duplex settings or coarse corpus-level metrics. To ensure systematic and reproducible evaluation, we propose automatic metrics tailored to real-time interaction. Experiments on full-duplex models reveal distinct model features and highlight areas for improvement. By releasing our data and behavior-specific metrics, we hope Full-Duplex-Bench provides a practical foundation for evaluating full-duplex spoken dialogue systems. VII. LIMITATION AND FUTURE WORK Our framework does not yet link described behaviors to human preferences; users need to determine what constitutes desirable or undesirable behavior according to their specific goals. Future work can integrate human judgment studies to provide the preference. The present analysis is limited to English, and extending the framework to other languages will be essential for assessing cross-linguistic generality. Result from Lin et al. (2025) 35 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Evaluation and Benchmarking Conversation Dynamics: Full-Duplex-Bench 1. Pause Handling User Model 2. Backchanneling User Model 4. User Interruption User Model [break] yeah um hm Wait I want to ... Okay let me ... 3. Smooth Turn Taking User Model Fig. 2: Illustration of the four evaluation dimensions in Full-Duplex-Bench. (1) Pause Handling: the model stays silent during user pauses; (2) Backchanneling: the model offers short, timely acknowledgments; (3) Smooth Turn-taking: the model takes the turn in time; and (4) User Interruption: the model handles sudden user input with appropriate, well-timed responses. Metric: The ideal model behavior is to avoid taking over the turn while the user is speaking. To evaluate this, we use the Takeover Rate (TOR). A lower TOR signifies better pause management, indicating that the model effectively waits for the user’s turn to end. In contrast, a higher TOR suggests that the model is more likely to take over the conversation before the user has yielded the turn. 2) Backchanneling Drawing on the definition in Section III, we assess whether the model, when interacting with a dominant speaker, actively listens and provides backchannels at suitable moments to facilitate dialogue engagement. A model exhibiting humanlike backchanneling behavior should respond at the right times and with an appropriate frequency. Research Question: Can the model determine when to offer backchannels in a human-like manner without interrupting the speaker? Metric: To measure how well the models generate backchannel cues, we use three metrics: • TOR: As with pause handling, the model should avoid dominating the turn, so a lower TOR is preferable. • Backchannel Frequency (Freq): Each backchannel event is counted and normalized by duration (events per second). When the model does not take over the turn (TOR = 0), a higher backchanneling frequency indicates that the model responds backchannel more often, but this does not necessarily imply better or more natural behavior, as it also depends on timing and context. • Jensen-Shannon Divergence (JSD): This captures the difference between the model’s predicted timing of backchannels and actual human timing. The model outputs a probability distribution P , where P(i) denotes the likelihood of a backchannel occurring in time window i . The ground truth distribution Q is derived from human-annotated backchannel timings (details in III-C ), aligned to the same set of time windows. To measure the similarity between P and Q , we compute the Jensen–Shannon Divergence (JSD) as: JSD(P||Q)=1 2X i P(i) log P(i) M(i)+1 2X i Q(i) log Q(i) M(i), where M(i)=1 2(P(i)+Q(i)) , and i indexes the discrete time windows. JSD ranges from 0 (perfect alignment) to 1 (complete divergence), providing a symmetric and bounded measure of similarity between model predictions and human backchannel behavior. We only calculate this metric when the model does not take over the turn, where each backchannel event is counted as one-hot and normalized into a probability distribution. If the model stays silent throughout, we assume a uniform probability distribution, treating it as a random baseline without backchannel knowledge. 3) Smooth Turn Taking Effective turn-taking is crucial for maintaining a natural and engaging conversation. In human dialogue, smooth turn transitions occur when speakers respond promptly without excessive delay or overlap. A well-designed model should be capable of recognizing turn boundaries and responding with appropriate timing to ensure fluid interactions. Research Question: Can the model detect the end of a speaker’s turn and respond promptly without long pauses? Metric: We measure the averaged response latency, the time (in seconds) between the end of the user’s speech and the start of the model’s response. Lower latency values indicate smoother turn-taking. In cases where the model fails to respond, we record the TOR. The latency is calculated only when TO equals 1. This avoids averaging with non-takeover periods, which would introduce significant variance due to periods of silence. 4) User Interruption In human conversations, interruptions are common and can occur when a listener interjects mid-turn to clarify, disagree, or shift the discussion. A well-designed conversational model should be able to recognize and adapt to such interruptions by adjusting its response appropriately. Effective handling of Pause Handling •TOR: Take-over rate of the model in a pause •How often does the model wrongly take-over in a user pause? •Synthetic: GPT-4o + ChatTTS •CANDOR: Real video chat conversations Result from Lin et al. (2025) 35 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Evaluation and Benchmarking Conversation Dynamics: Full-Duplex-Bench 1. Pause Handling User Model 2. Backchanneling User Model 4. User Interruption User Model [break] yeah um hm Wait I want to ... Okay let me ... 3. Smooth Turn Taking User Model Fig. 2: Illustration of the four evaluation dimensions in Full-Duplex-Bench. (1) Pause Handling: the model stays silent during user pauses; (2) Backchanneling: the model offers short, timely acknowledgments; (3) Smooth Turn-taking: the model takes the turn in time; and (4) User Interruption: the model handles sudden user input with appropriate, well-timed responses. Metric: The ideal model behavior is to avoid taking over the turn while the user is speaking. To evaluate this, we use the Takeover Rate (TOR). A lower TOR signifies better pause management, indicating that the model effectively waits for the user’s turn to end. In contrast, a higher TOR suggests that the model is more likely to take over the conversation before the user has yielded the turn. 2) Backchanneling Drawing on the definition in Section III, we assess whether the model, when interacting with a dominant speaker, actively listens and provides backchannels at suitable moments to facilitate dialogue engagement. A model exhibiting humanlike backchanneling behavior should respond at the right times and with an appropriate frequency. Research Question: Can the model determine when to offer backchannels in a human-like manner without interrupting the speaker? Metric: To measure how well the models generate backchannel cues, we use three metrics: • TOR: As with pause handling, the model should avoid dominating the turn, so a lower TOR is preferable. • Backchannel Frequency (Freq): Each backchannel event is counted and normalized by duration (events per second). When the model does not take over the turn (TOR = 0), a higher backchanneling frequency indicates that the model responds backchannel more often, but this does not necessarily imply better or more natural behavior, as it also depends on timing and context. • Jensen-Shannon Divergence (JSD): This captures the difference between the model’s predicted timing of backchannels and actual human timing. The model outputs a probability distribution P , where P(i) denotes the likelihood of a backchannel occurring in time window i . The ground truth distribution Q is derived from human-annotated backchannel timings (details in III-C ), aligned to the same set of time windows. To measure the similarity between P and Q , we compute the Jensen–Shannon Divergence (JSD) as: JSD(P||Q)=1 2X i P(i) log P(i) M(i)+1 2X i Q(i) log Q(i) M(i), where M(i)=1 2(P(i)+Q(i)) , and i indexes the discrete time windows. JSD ranges from 0 (perfect alignment) to 1 (complete divergence), providing a symmetric and bounded measure of similarity between model predictions and human backchannel behavior. We only calculate this metric when the model does not take over the turn, where each backchannel event is counted as one-hot and normalized into a probability distribution. If the model stays silent throughout, we assume a uniform probability distribution, treating it as a random baseline without backchannel knowledge. 3) Smooth Turn Taking Effective turn-taking is crucial for maintaining a natural and engaging conversation. In human dialogue, smooth turn transitions occur when speakers respond promptly without excessive delay or overlap. A well-designed model should be capable of recognizing turn boundaries and responding with appropriate timing to ensure fluid interactions. Research Question: Can the model detect the end of a speaker’s turn and respond promptly without long pauses? Metric: We measure the averaged response latency, the time (in seconds) between the end of the user’s speech and the start of the model’s response. Lower latency values indicate smoother turn-taking. In cases where the model fails to respond, we record the TOR. The latency is calculated only when TO equals 1. This avoids averaging with non-takeover periods, which would introduce significant variance due to periods of silence. 4) User Interruption In human conversations, interruptions are common and can occur when a listener interjects mid-turn to clarify, disagree, or shift the discussion. A well-designed conversational model should be able to recognize and adapt to such interruptions by adjusting its response appropriately. Effective handling of Backchannel •In Conversation Corpus (ICC): real-dialogs between; participants were instructed to actively do backchanneling •Backchannel: <1 second and ≤2 words •JSD: Mesures distance between ground truth backchannel distribution and model predictions Result from Lin et al. (2025) 35 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Evaluation and Benchmarking Conversation Dynamics: Full-Duplex-Bench 1. Pause Handling User Model 2. Backchanneling User Model 4. User Interruption User Model [break] yeah um hm Wait I want to ... Okay let me ... 3. Smooth Turn Taking User Model Fig. 2: Illustration of the four evaluation dimensions in Full-Duplex-Bench. (1) Pause Handling: the model stays silent during user pauses; (2) Backchanneling: the model offers short, timely acknowledgments; (3) Smooth Turn-taking: the model takes the turn in time; and (4) User Interruption: the model handles sudden user input with appropriate, well-timed responses. Metric: The ideal model behavior is to avoid taking over the turn while the user is speaking. To evaluate this, we use the Takeover Rate (TOR). A lower TOR signifies better pause management, indicating that the model effectively waits for the user’s turn to end. In contrast, a higher TOR suggests that the model is more likely to take over the conversation before the user has yielded the turn. 2) Backchanneling Drawing on the definition in Section III, we assess whether the model, when interacting with a dominant speaker, actively listens and provides backchannels at suitable moments to facilitate dialogue engagement. A model exhibiting humanlike backchanneling behavior should respond at the right times and with an appropriate frequency. Research Question: Can the model determine when to offer backchannels in a human-like manner without interrupting the speaker? Metric: To measure how well the models generate backchannel cues, we use three metrics: • TOR: As with pause handling, the model should avoid dominating the turn, so a lower TOR is preferable. • Backchannel Frequency (Freq): Each backchannel event is counted and normalized by duration (events per second). When the model does not take over the turn (TOR = 0), a higher backchanneling frequency indicates that the model responds backchannel more often, but this does not necessarily imply better or more natural behavior, as it also depends on timing and context. • Jensen-Shannon Divergence (JSD): This captures the difference between the model’s predicted timing of backchannels and actual human timing. The model outputs a probability distribution P , where P(i) denotes the likelihood of a backchannel occurring in time window i . The ground truth distribution Q is derived from human-annotated backchannel timings (details in III-C ), aligned to the same set of time windows. To measure the similarity between P and Q , we compute the Jensen–Shannon Divergence (JSD) as: JSD(P||Q)=1 2X i P(i) log P(i) M(i)+1 2X i Q(i) log Q(i) M(i), where M(i)=1 2(P(i)+Q(i)) , and i indexes the discrete time windows. JSD ranges from 0 (perfect alignment) to 1 (complete divergence), providing a symmetric and bounded measure of similarity between model predictions and human backchannel behavior. We only calculate this metric when the model does not take over the turn, where each backchannel event is counted as one-hot and normalized into a probability distribution. If the model stays silent throughout, we assume a uniform probability distribution, treating it as a random baseline without backchannel knowledge. 3) Smooth Turn Taking Effective turn-taking is crucial for maintaining a natural and engaging conversation. In human dialogue, smooth turn transitions occur when speakers respond promptly without excessive delay or overlap. A well-designed model should be capable of recognizing turn boundaries and responding with appropriate timing to ensure fluid interactions. Research Question: Can the model detect the end of a speaker’s turn and respond promptly without long pauses? Metric: We measure the averaged response latency, the time (in seconds) between the end of the user’s speech and the start of the model’s response. Lower latency values indicate smoother turn-taking. In cases where the model fails to respond, we record the TOR. The latency is calculated only when TO equals 1. This avoids averaging with non-takeover periods, which would introduce significant variance due to periods of silence. 4) User Interruption In human conversations, interruptions are common and can occur when a listener interjects mid-turn to clarify, disagree, or shift the discussion. A well-designed conversational model should be able to recognize and adapt to such interruptions by adjusting its response appropriately. Effective handling of Smooth Turn Taking •Latency: time between end of user’s speech and start of model response •Take-over rate (TOR): here times the models successfully takes over Result from Lin et al. (2025) 35 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Evaluation and Benchmarking Conversation Dynamics: Full-Duplex-Bench 1. Pause Handling User Model 2. Backchanneling User Model 4. User Interruption User Model [break] yeah um hm Wait I want to ... Okay let me ... 3. Smooth Turn Taking User Model Fig. 2: Illustration of the four evaluation dimensions in Full-Duplex-Bench. (1) Pause Handling: the model stays silent during user pauses; (2) Backchanneling: the model offers short, timely acknowledgments; (3) Smooth Turn-taking: the model takes the turn in time; and (4) User Interruption: the model handles sudden user input with appropriate, well-timed responses. Metric: The ideal model behavior is to avoid taking over the turn while the user is speaking. To evaluate this, we use the Takeover Rate (TOR). A lower TOR signifies better pause management, indicating that the model effectively waits for the user’s turn to end. In contrast, a higher TOR suggests that the model is more likely to take over the conversation before the user has yielded the turn. 2) Backchanneling Drawing on the definition in Section III, we assess whether the model, when interacting with a dominant speaker, actively listens and provides backchannels at suitable moments to facilitate dialogue engagement. A model exhibiting humanlike backchanneling behavior should respond at the right times and with an appropriate frequency. Research Question: Can the model determine when to offer backchannels in a human-like manner without interrupting the speaker? Metric: To measure how well the models generate backchannel cues, we use three metrics: • TOR: As with pause handling, the model should avoid dominating the turn, so a lower TOR is preferable. • Backchannel Frequency (Freq): Each backchannel event is counted and normalized by duration (events per second). When the model does not take over the turn (TOR = 0), a higher backchanneling frequency indicates that the model responds backchannel more often, but this does not necessarily imply better or more natural behavior, as it also depends on timing and context. • Jensen-Shannon Divergence (JSD): This captures the difference between the model’s predicted timing of backchannels and actual human timing. The model outputs a probability distribution P , where P(i) denotes the likelihood of a backchannel occurring in time window i . The ground truth distribution Q is derived from human-annotated backchannel timings (details in III-C ), aligned to the same set of time windows. To measure the similarity between P and Q , we compute the Jensen–Shannon Divergence (JSD) as: JSD(P||Q)=1 2X i P(i) log P(i) M(i)+1 2X i Q(i) log Q(i) M(i), where M(i)=1 2(P(i)+Q(i)) , and i indexes the discrete time windows. JSD ranges from 0 (perfect alignment) to 1 (complete divergence), providing a symmetric and bounded measure of similarity between model predictions and human backchannel behavior. We only calculate this metric when the model does not take over the turn, where each backchannel event is counted as one-hot and normalized into a probability distribution. If the model stays silent throughout, we assume a uniform probability distribution, treating it as a random baseline without backchannel knowledge. 3) Smooth Turn Taking Effective turn-taking is crucial for maintaining a natural and engaging conversation. In human dialogue, smooth turn transitions occur when speakers respond promptly without excessive delay or overlap. A well-designed model should be capable of recognizing turn boundaries and responding with appropriate timing to ensure fluid interactions. Research Question: Can the model detect the end of a speaker’s turn and respond promptly without long pauses? Metric: We measure the averaged response latency, the time (in seconds) between the end of the user’s speech and the start of the model’s response. Lower latency values indicate smoother turn-taking. In cases where the model fails to respond, we record the TOR. The latency is calculated only when TO equals 1. This avoids averaging with non-takeover periods, which would introduce significant variance due to periods of silence. 4) User Interruption In human conversations, interruptions are common and can occur when a listener interjects mid-turn to clarify, disagree, or shift the discussion. A well-designed conversational model should be able to recognize and adapt to such interruptions by adjusting its response appropriately. Effective handling of User Interruption •TOR: Model successfully continues after interruption •GPT-4o: LLM-as-a-judge scoring coherence of continuation after interruption •Latency: Latency to continue after the interruption Result from Lin et al. (2025) 35 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Evaluation and Benchmarking Speech, Music and Sound Understanding: MMAU-Pro Audio Transcript:What is the packing efficiency, in percentage,of a solid where Atom X occupies the face-centered cubic latticesites as well as alternate tetrahedral voids of the same lattice? Question:Choose the correct option that answers thequestion in the audio from the options given below. Options: (A) 25% (B) 35% (C) 55% (D) 75% Answer: (B) 35% Voice STEM QA Question:What raag is this bandish composed in? Options: (A) Bhimpalasi (B) Kannada (C) Durga (D) Malkaunse (E) Bhairavi (F) Yaman Kalyan Answer: (C) Durga Multicultural Music Question:What trend can be observed in the weight of the cloths thrown in the audio? Options: (A) Increasing (B) Decreasing (C) Remains constant (D) None of these options Answer: (A) Increasing Perceptual Skills:Acoustic Trend Estimation Reasoning Skills:Temporal Reasoning Sound Question:Can you guess the singer in this song? Options: (A) Jeff Beck (B) Tenacious B (C) Jimmy Hendricks (D) Jack Black Answer:(D) Jack Black Perceptual Skills:Timbre Perception and Instrument Recognition Reasoning Skills: Musicological Knowledge Music Question:What effect needs to be applied to the first recording to achieve sound of the second recording? Options: (A) echo (B) distortion (C) phaser (D) reverb Answer: (D) reverb Multi-Audio Audio Transcript:Instruction: Explain what's happening in this audio. Your answer must contain a title, wrapped in double angular brackets, such as <<poem of joy>>. Question:What effect needs to be applied to the first recording to achieve sound of the second recording? Correct Answer:<<Harmonica's Ascending and Descending Scale>>This audio clip features a harmonica playing the C major scale. The musician first plays the scale in an ascending order, moving from the lowest note to the highest. Wrong Answer:This is an audio clip of a person playing the C major scale on a harmonica, first ascendingand then descending. Multimodal Instruction Following Question:What did the winning team get at the end of the game? Options: (A) They choose the next vacation destination (B) The other team has to get rid of theirrooster (C) They win the other team’s apartment (D) The game ends in a tie and nothing changes \ Answer:They win the other team’s apartment Question:In the audio with roughly the same phrase being repeated, explain how the different tone effects the meaning of each of the 4 phrases and in what order they occur. Options: (A) Bored enthusiastic questioning angry (B) Cheerful playful serious stern (C) Sarcastic sincere doubtful irritated (D) genuine sarcastic questioning frustrated. Answer: (D) He can remove 0 bones Perceptual Skills:Speech Activity, Turn-Taking and Overlap Detection Reasoning Skills:Quantitative Reasoning (Counting/Arithmetic Comparison) Long Audio Speech Question:Answer the question in the audio. Options: (A) It sounds like this decision carries a lot of weight... (B) Perhaps the best approach is to systematically... (C) Your tranquil state suggests you have a high degree of mental clarity right now. This is an excellent time to trust yourjudgment, as it's likely unclouded by.... (D) Making major decisions during emotional peaks,.. Answer:Your tranquil state suggests you have a high degree of mental clarity right now. This is an excellent time to trust your judgment, as it's likely unclouded by emotional turmoil. Voice QA Question:Who's order does the waiter take first? Options: (A) The person to the left of the mic holder (B) The person to the right of the mic holder (C) The person in front of the mic holder Answer: (A) The person to the left of the mic holder Spatial QA Question:What is hyper-foreignism with respect to pronunciation according to the clip? Answer:When a speaker changes the way they say a word to sound more like that of the stereotype they hold for a foreign language Open-ended QA Question:Which of the following songs is made according to the speaker? Options: (A) Not Like Us (B) Star Biy (C) All the stars (D) HUMBLE Answer: (D) HUMBLE Speech-Sound-Music Mix Figure 1: Overview of the MMAU-Pro benchmark. MMAU-Pro provides comprehensive coverage across all three core audio domains-speech, sound, and music-and extends evaluation to their mixtures. It further includes multi-audio reasoning, long-form audio (up to 10 minutes), voice-chat QA, spatial audio understanding, open-ended QA, and multimodal instruction following, offering a broad and realistic assessment of audio intelligence. gence skills spanning speech, environmental sounds, and music. MMAU-Pro presents challenges overlooked by prior benchmarks, including long-form audio understanding (up to 10 minutes), reasoning across multiple clips, spatial audio perception, multicultural music interpretation, instructionfollowing abilities, etc. All questions are crafted to require deliberate multi-hop reasoning and include a balanced mix of multiple-choice and open-ended formats. To address the shortcomings of existing evaluation methodologies, we further propose a retrieval-based evaluation framework that enables more robust and reliable assessment. By emphasizing realistic and demanding auditory tasks, MMAU-Pro provides a comprehensive testbed to accelerate the development of auditory intelligence in multimodal AI systems. To summarize, our main contributions are: • We introduce MMAU-Pro, the most comprehensive benchmark to date for evaluating auditory intelligence. It comprises 5,305 expert-annotated question–answer pairs spanning 49 distinct skills across speech, environmental sounds, music, and their mixtures. MMAU-Pro introduces novel challenges, including spatial audio reasoning, multi-clip audio reasoning, voice–chat comprehension, and tasks requiring prosodic, world-knowledge, and STEM-based reasoning. All audio samples are drawn from the wild, with durations up to ten minutes, significantly surpassing the short clips typical of prior benchmarks where current models are near-saturated. • We benchmark over 15 open-source and proprietary multimodal LLMs on MMAU-Pro, finding that even the strongest models face substantial challenges. Gemini 2.5 Flash achieves only 59.2% accuracy; the best-performing fully open-source model, Audio Flamingo 3, reaches 51.7%; and the strongest open-weights omni model, Qwen2.5-Omni-7B-Instruct, achieves just 52.2%. • We provide an in-depth analysis of model responses, uncovering key failure modes in auditory perception and reasoning. These include shallow audio grounding, degradation in text-only and STEM reasoning, poor performance in multi-audio and spatial reasoning, and limited understanding of multicultural music. Related Work Large Audio Language Models Recent advances in multimodal modeling have led to (L)ALMs-models that pair audio perception with (L)LMs to tackle complex audio tasks. Early systems such as Whisper (Li et al. 2024a; Peng et al. 2023) and CLAP (Wu et al. 2023; Elizalde et al. 2023; Elizalde, Deshmukh, and Wang 2024) focused on foundational tasks like transcription, captioning, and retrieval, but struggled with reasoning-centric challenges. More recent models-GAMA (Ghosh et al. 2024), Audio Flamingo (Ghosh et al. 2025b; Goel et al. 2025), Mellow (Deshmukh et al. 2025), Phi-4MM (Abouelenin et al. 2025) Qwen2-Audio (Chu et al. 2024), and AudioPALM (Rubenstein et al. 2023) proposed improved archiFigure from Kumar et al. (2025) 36 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Dialog Management Constrained Dialog Systems •In many practical applications more control over the dialog is required –Avoid off-topic drift –Control what exactly is said by the system –Hard constraints what information needs to be obtained by the user •Approach: more rigid predefined dialog structure •Examples: Rasa CALM, AppTek’s Cheshire 40 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Dialog Management Constrained Dialog Systems •In many practical applications more control over the dialog is required –Avoid off-topic drift –Control what exactly is said by the system –Hard constraints what information needs to be obtained by the user •Approach: more rigid predefined dialog structure •Examples: Rasa CALM, AppTek’s Cheshire 40 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Conclusion Design Challenges Do we actually need human-like natural voice interfaces? •What expectations do we raise with such interfaces? –Cohn et al. (2024) showed in a study that people put more trust in the same LLM outputs if was communicated by speech •How to trade-off naturalness and efficiency? •What other features are required for “natural” interactions in different usage scenarios? 41 of 42 Towards Natural Realtime Voice Interaction with Large Language Models Conclusion Recap Cascaded Architecture •implementation options for the required components End-to-End (Integrated) Architectures •architectures that integrate all components in a single model Evaluation and Benchmarking •Approaches to benchmark different model capabilities Dialog Management •different dialog strategies depending on task requirements Questions? [email protected] 42 of 42 Towards Natural Realtime Voice Interaction with Large Language Models References Bai, Y., Chen, J., Chen, J., Chen, W., Chen, Z., Ding, C., Dong, L., Dong, Q., Du, Y., Gao, K., Gao, L., Guo, Y., Han, M., Han, T., Hu, W., Hu, X., Hu, Y., Hua, D., Huang, L., Huang, M., Huang, Y., Jin, J., Kong, F., Lan, Z., Li, T., Li, X., Li, Z., Lin, Z., Liu, R., Liu, S., Lu, L., Lu, Y., Ma, J., Ma, S., Pei, Y., Shen, C., Tan, T., Tian, X., Tu, M., Wang, B., Wang, H., Wang, Y., Wang, Y., Xia, H., Xia, R., Xie, S., Xu, H., Yang, M., Zhang, B., Zhang, J., Zhang, W., Zhang, Y., Zhang, Y., Zheng, Y., and Zou, M. (2024). Seed-asr: Understanding diverse speech and contexts with llm-based speech recognition. Chen, S., Wang, C., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., and Wei, F. (2025). Neural codec language models are zero-shot text to speech synthesizers. IEEE Transactions on Audio, Speech and Language Processing, 33:705–718. Chiang, C.-H., Wang, X., Li, L., Lin, C.-C., Lin, K., Liu, S., Wang, Z., Yang, Z., yi Lee, H., and Wang, L. (2025). Shanks: Simultaneous hearing and thinking for spoken language models. Cho, H. J., Jedema, N. P., Ribeiro, L. F. R., Sharma, K., Szekely, P., Moschitti, A., Janssen, R., and May, J. (2024). Speechworthy instruction-tuned language models. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10652–10670, Miami, Florida, USA. Association for Computational Linguistics. Chu, Y., Xu, J., Yang, Q., Wei, H., Wei, X., Guo, Z., Leng, Y., Lv, Y., He, J., Lin, J., Zhou, C., and Zhou, J. (2024). Qwen2-audio technical report. Cohn, M., Pushkarna, M., Olanubi, G. O., Moran, J. M., Padgett, D., Mengesha, Z., and Heldreth, C. (2024). Believing anthropomorphism: Examining the role of anthropomorphic cues on trust in large language models. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’24, New York, NY, USA. Association for Computing Machinery. Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., and Zeghidour, N. (2024). Moshi: a speech-text foundation model for real-time dialogue. Fang, Q., Guo, S., Zhou, Y., Ma, Z., Zhang, S., and Feng, Y. (2025). Llama-omni: Seamless speech interaction with large language models. In Yue, Y., Garg, A., Peng, N., Sha, F., and Yu, R., editors, International Conference on Representation Learning, volume 2025, pages 57607–57624. Geng, X., Wei, K., Shao, Q., Liu, S., Lin, Z., Zhao, Z., Li, G., Tian, W., Chen, P., Li, Y., Guo, P., Shao, M., Wang, S., Cao, Y., Wang, C., Xu, T., Dai, Y., Zhu, X., Li, Y., Zhang, L., and Xie, L. (2025). Osum: Advancing open speech understanding models with limited resources in academia. Gim, I., seob Lee, S., and Zhong, L. (2024). Asynchronous llm function calling. Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A. (2021). Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Trans. Audio, Speech and Lang. Proc., 29:3451–3460. Kumar, S., Šimon Sedláček, Lokegaonkar, V., López, F., Yu, W., Anand, N., Ryu, H., Chen, L., Plička, M., Hlaváček, M., Ellingwood, W. F., Udupa, S., Hou, S., Ferner, A., Barahona, S., Bolaños, C., Rahi, S., Herrera-Alarcón, L., Dixit, S., Patil, S., Deshmukh, S., Koroshinadze, L., Liu, Y., Perera, L. P. G., Zanou, E., Stafylakis, T., Chung, J. S., Harwath, D., Zhang, C., Manocha, D., Lozano-Diez, A., Kesiraju, S., Ghosh, S., and Duraiswami, R. (2025). Mmau-pro: A challenging and comprehensive benchmark for holistic evaluation of audio general intelligence. 43 of 42 Towards Natural Realtime Voice Interaction with Large Language Models References Li, Z., Chen, Z., Ross, M., Huber, P., Moon, S., Lin, Z., Dong, X., Sagar, A., Yan, X., and Crook, P. (2024). Large language models as zero-shot dialogue state tracker through function calling. In Ku, L.-W., Martins, A., and Srikumar, V., editors, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8688–8704, Bangkok, Thailand. Association for Computational Linguistics. Lin, G.-T., Lian, J., Li, T., Wang, Q., Anumanchipalli, G., Liu, A. H., and Lee, H.-y. (2025). Full-duplex-bench: A benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities. arXiv preprint arXiv:2503.04721. Ma, Z., Song, Y., Du, C., Cong, J., Chen, Z., Wang, Y., Wang, Y., and Chen, X. (2024a). Language model can listen while speaking. Ma, Z., Yang, G., Yang, Y., Gao, Z., Wang, J., Du, Z., Yu, F., Chen, Q., Zheng, S., Zhang, S., and Chen, X. (2024b). An embarrassingly simple approach for llm with strong asr capacity. Microsoft, :, Abouelenin, A., Ashfaq, A., Atkinson, A., Awadalla, H., Bach, N., Bao, J., Benhaim, A., Cai, M., Chaudhary, V., Chen, C., Chen, D., Chen, D., Chen, J., Chen, W., Chen, Y.-C., ling Chen, Y., Dai, Q., Dai, X., Fan, R., Gao, M., Gao, M., Garg, A., Goswami, A., Hao, J., Hendy, A., Hu, Y., Jin, X., Khademi, M., Kim, D., Kim, Y. J., Lee, G., Li, J., Li, Y., Liang, C., Lin, X., Lin, Z., Liu, M., Liu, Y., Lopez, G., Luo, C., Madan, P., Mazalov, V., Mitra, A., Mousavi, A., Nguyen, A., Pan, J., Perez-Becker, D., Platin, J., Portet, T., Qiu, K., Ren, B., Ren, L., Roy, S., Shang, N., Shen, Y., Singhal, S., Som, S., Song, X., Sych, T., Vaddamanu, P., Wang, S., Wang, Y., Wang, Z., Wu, H., Xu, H., Xu, W., Yang, Y., Yang, Z., Yu, D., Zabir, I., Zhang, J., Zhang, L. L., Zhang, Y., and Zhou, X. (2025). Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras. Nguyen, T. A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W.-N., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., and Dupoux, E. (2023). Generative spoken dialogue language modeling. Transactions of the Association for Computational Linguistics, 11:250–266. Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., Bapna, A., Borsos, Z., de Chaumont Quitry, F., Chen, P., Badawy, D. E., Han, W., Kharitonov, E., Muckenhirn, H., Padfield, D., Qin, J., Rozenberg, D., Sainath, T., Schalkwyk, J., Sharifi, M., Ramanovich, M. T., Tagliasacchi, M., Tudor, A., Velimirović, M., Vincent, D., Yu, J., Wang, Y., Zayats, V., Zeghidour, N., Zhang, Y., Zhang, Z., Zilka, L., and Frank, C. (2023). Audiopalm: A large language model that can speak and listen. Shih, Y.-J., Raj, D., Wu, C., Zhou, W., Bong, S., Gaur, Y., Mahadeokar, J., Kalinli, O., and Seltzer, M. (2025). Can speech llms think while listening? Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C. (2024). Salmonn: Towards generic hearing abilities for large language models. Wang, D., Li, J., Cui, M., Yang, D., Chen, X., and Meng, H. M. (2025a). Speech discrete tokens or continuous features? a comparative analysis for spoken language understanding in SpeechLLMs. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 24924–24935, Suzhou, China. Association for Computational Linguistics. Wang, M., Han, W., Shafran, I., Wu, Z., Chiu, C.-C., Cao, Y., Chen, N., Zhang, Y., Soltau, H., Rubenstein, P. K., Zilka, L., Yu, D., Pundak, G., Siddhartha, N., Schalkwyk, J., and Wu, Y. (2023a). Slm: Bridge the thin gap between speech and text foundation models. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8. 44 of 42 Towards Natural Realtime Voice Interaction with Large Language Models References Wang, M., Han, W., Shafran, I., Wu, Z., Chiu, C.-C., Cao, Y., Chen, N., Zhang, Y., Soltau, H., Rubenstein, P. K., Zilka, L., Yu, D., Pundak, G., Siddhartha, N., Schalkwyk, J., and Wu, Y. (2023b). Slm: Bridge the thin gap between speech and text foundation models. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8. Wang, X., Li, Y., Fu, C., Zhang, Y., Shen, Y., Xie, L., Li, K., Sun, X., and MA, L. (2025b). Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM. In Forty-second International Conference on Machine Learning. Xu, J., Guo, Z., He, J., Hu, H., He, T., Bai, S., Chen, K., Wang, J., Fan, Y., Dang, K., Zhang, B., Wang, X., Chu, Y., and Lin, J. (2025). Qwen2.5-omni technical report. Yang, Y., Liu, S., Li, J., Wang, H., Meng, L., Sun, H., Liang, Y., Ma, Z., Hu, Y., Zhao, R., Yu, J., Lu, Y., and Chen, X. (2024). Interleaved speech-text language models for simple streaming text-to-speech synthesis. Zeghidour, N., Kharitonov, E., Orsini, M., Volhejn, V., de Marmiesse, G., Grave, E., Pérez, P., Mazaré, L., and Défossez, A. (2025). Streaming sequence-to-sequence learning with delayed streams modeling. 45 of 42 Towards Natural Realtime Voice Interaction with Large Language Models