Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: HKGAI
All Institutions: HKGAI
YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
The paper proposes a unified architecture (YuE2) that bridges symbolic and audio music generation using an AR-NAR Mixture-of-Transformers (MoT). The core methodological contribution is the "symbolic planning" stage, where the model first generates a readable score (melody, harmony, form) before expanding into semantic tokens and finally acoustic latents via flow matching. This hierarchical approach allows for explicit control over composition. Additionally, the paper introduces two auxiliary models: MERT2 for semantic representation learning and SheetSage2 for symbolic supervision (lead-sheet transcription), which are used to train the main model on unaligned audio data. The integration of these components into a single checkpoint that supports editing and cover generation is a significant architectural advancement.
The evaluation is extensive, utilizing both objective benchmarks (WildSongBench, SongBench, MARBLE) and subjective expert listening tests. The results show YuE2 outperforming public baselines and being competitive with proprietary systems like Suno v4.5 and v5. The ablation study comparing generation with and without symbolic planning demonstrates a clear preference for the planned approach (49.3% vs 34.6%). The introduction of MERT2 and SheetSage2 is supported by state-of-the-art results on their respective benchmarks (MARBLE and transcription metrics). The inclusion of zero-shot cover generation and agentic editing case studies further validates the utility of the unified framework.
As a technical report from an industry lab (HKGAI), the paper likely lacks full open-source code and model weights at the time of release, which limits immediate reproducibility. However, the detailed description of the architecture (AR-NAR MoT, flow matching) and the specific benchmarks used provides a clear roadmap for replication. The reliance on proprietary datasets for training the auxiliary models may pose a barrier for independent researchers.
The paper relies heavily on expert listening tests, which can be subjective and difficult to scale. The comparison with proprietary systems (Suno) is limited to specific versions and may not reflect the full capability of those closed-source models. The "best-of-8" selection strategy, while effective for benchmarking, introduces a computational overhead that may not be practical for real-time applications. Additionally, the paper does not extensively discuss the computational cost of the multi-stage generation process compared to direct audio generation.
This work has significant implications for the music production industry by providing a tool that allows for both high-quality audio generation and explicit compositional control. The ability to edit scores and generate covers zero-shot opens up new possibilities for music creation, education, and collaboration. The unification of symbolic and audio domains could lead to more interpretable and controllable generative music systems, potentially influencing how AI is integrated into professional music workflows. YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
Primary: MIT CSAIL
All Institutions: MIT CSAIL, MIT RLE, KAIST GSCT
The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
The paper introduces a differentiable, GPU-accelerated acoustic simulator for the vocal tract based on the linearized Euler equations. The core technical contribution is the shift from time-domain finite difference methods to a frequency-domain formulation using transmission line matrices (TLM). This allows for parallelization across frequencies, resulting in a 70x speedup and, crucially, smoother gradient landscapes that facilitate stable optimization. The authors also propose a differentiable turbulence model using softplus gating and softmax-based constriction localization to handle consonants, unifying vowel and consonant synthesis. Finally, they leverage neural fields (implicit neural representations) to parameterize the vocal tract geometry, which acts as a regularizer to escape local minima in the non-convex inverse problem.
The experiments are extensive and well-designed. The authors validate the simulator's accuracy by comparing Frequency Domain Synthesis (FDS) against Time Domain Synthesis (TDS), showing significant improvements in SI-SDR, STOI, and PESQ metrics. They demonstrate the utility of the simulator in two novel applications: (1) a self-supervised autoencoder that maps audio to vocal tract area functions across 11 languages, outperforming paired-data baselines in intelligibility and speaker identity preservation; and (2) a generative model that reconstructs MRI videos of the vocal tract from speech alone, without paired data. The emergence of the IPA vowel chart from the latent space of the MRI generator provides strong evidence for the physical grounding of the model.
The paper provides detailed derivations of the physics, including the linearized Euler equations, Webster's equation, and the circuit interpretation. Specific hyperparameters, such as window lengths, hop sizes, and Reynolds number thresholds, are provided. The use of standard architectures like Wav2Vec 2.0 and StyleGAN2 aids reproducibility. However, the specific implementation of the differentiable turbulence model and the exact neural field architectures (RFF vs MFN) would benefit from more code-level details or a public repository link beyond the project page.
The model relies on a 1D acoustic tube approximation, which may not capture all 3D effects of the vocal tract, particularly for complex articulations or nasal sounds (though the nasal tract is noted as a limitation due to MRI data availability). The inverse problem is inherently ill-posed, meaning multiple geometries can produce the same sound, leading to potential ambiguities in reconstruction. The MRI reconstruction quality is limited by the resolution and coverage of the training MRI dataset.
This work has significant implications for linguistics, speech pathology, and medical imaging. By providing a physically grounded link between speech and articulatory geometry, it enables new tools for voice coaching, language acquisition studies, and non-invasive visualization of vocal tract movements. The self-supervised nature of the autoencoder and MRI generator makes these tools scalable and accessible without the need for expensive paired data collection. The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
Primary: The Chinese University of Hong Kong
All Institutions: The Chinese University of Hong Kong, Microsoft Corporation
The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
The paper proposes DuraS2ST, a framework that reformulates duration-aligned speech-to-speech translation (S2ST) as a reasoning problem rather than an acoustic post-processing task. The core methodological contribution is the integration of Chain-of-Thought (CoT) reasoning into the speech generation pipeline. The model first generates an explicit textual rationale planning the target wording and phonetic length, followed by the synthesis of interleaved text-acoustic tokens. This is supported by the construction of DuraSet-440K, a large-scale corpus specifically designed for this purpose. The training paradigm is two-phase: Supervised Fine-Tuning (SFT) on the new corpus, followed by Group Relative Policy Optimization (GRPO). The RL phase introduces two specific technical innovations: the Duration Margin Reward (DMR), which acts as a soft constraint to balance translation quality with duration consistency, and Modality-Aware Reward Attribution (MARA), which prevents cross-modality reward contamination by assigning duration rewards only to acoustic tokens and quality rewards to both text and acoustic spans. This approach is technically sound and addresses a specific gap in current S2ST systems where duration control is often handled via black-box speed tokens or post-hoc time-stretching.
The experiments are conducted on the CVSS-T benchmark, which is disjoint from the training data, providing a valid out-of-domain evaluation. The paper compares DuraS2ST against strong baselines, including commercial models (GPT-4o, Qwen2.5-Omni, Kimi-Audio) and open-source SLMs (Step-Audio-2-mini). The results demonstrate that DuraS2ST achieves a superior balance between translation quality (BLEU, COMET) and duration consistency (SLC-0.2, SLC-0.4, MADE, MRDE). Notably, the method shows substantial gains in a zero-shot RL setting, highlighting the effectiveness of the reward design. The inclusion of both open-source and commercial baselines strengthens the experimental validation.
The paper provides a project page with a GitHub repository link. It details the training setup, including hyperparameters for SFT and GRPO, and describes the construction pipeline for DuraSet-440K in detail. The use of standard metrics (BLEU, COMET, SECS, SLC) and a public benchmark (CVSS-T) enhances reproducibility. However, the specific implementation details of the "duration-controllable TTS model" used for data synthesis are not fully specified in the main text, though likely detailed in the appendix or code.
The method relies on a specific base model (Step-Audio-2-mini-Think), and its generalizability to other SLM architectures is not extensively tested. The construction of DuraSet-440K is resource-intensive, requiring LLM-based translation candidate generation and duration-controllable TTS synthesis. The evaluation is limited to English-Chinese pairs; performance on other language pairs with different phonetic structures remains unknown. Additionally, the "zero-shot RL" claim needs careful interpretation as it likely refers to the RL phase without additional SFT on the target domain, but the model is still SFT'd on DuraSet-440K.
This work has significant implications for video dubbing, simultaneous interpretation, and real-time communication applications where audio-visual synchronization is critical. By treating duration alignment as a reasoning task, it offers a more interpretable and controllable approach compared to traditional acoustic constraints. The release of DuraSet-440K and the code will likely accelerate research in duration-controlled speech generation and multimodal reasoning. The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
Primary: MIT CSAIL
All Institutions: MIT CSAIL, MIT RLE, KAIST GSCT
The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
The paper introduces a differentiable, GPU-accelerated acoustic simulator for the vocal tract based on the linearized Euler equations. The core technical contribution is the shift from time-domain finite difference methods to a frequency-domain formulation using transmission line matrices (TLM). This allows for parallelization across frequencies, resulting in a 70x speedup and, crucially, smoother gradient landscapes that facilitate stable optimization. The authors also propose a differentiable turbulence model using softplus gating and softmax-based constriction localization to handle consonants, unifying vowel and consonant synthesis. Finally, they leverage neural fields (implicit neural representations) to parameterize the vocal tract geometry, which acts as a regularizer to escape local minima in the non-convex inverse problem.
The experiments are extensive and well-designed. The authors validate the simulator's accuracy by comparing Frequency Domain Synthesis (FDS) against Time Domain Synthesis (TDS), showing significant improvements in SI-SDR, STOI, and PESQ metrics. They demonstrate the utility of the simulator in two novel applications: (1) a self-supervised autoencoder that maps audio to vocal tract area functions across 11 languages, outperforming paired-data baselines in intelligibility and speaker identity preservation; and (2) a generative model that reconstructs MRI videos of the vocal tract from speech alone, without paired data. The emergence of the IPA vowel chart from the latent space of the MRI generator provides strong evidence for the physical grounding of the model.
The paper provides detailed derivations of the physics, including the linearized Euler equations, Webster's equation, and the circuit interpretation. Specific hyperparameters, such as window lengths, hop sizes, and Reynolds number thresholds, are provided. The use of standard architectures like Wav2Vec 2.0 and StyleGAN2 aids reproducibility. However, the specific implementation of the differentiable turbulence model and the exact neural field architectures (RFF vs MFN) would benefit from more code-level details or a public repository link beyond the project page.
The model relies on a 1D acoustic tube approximation, which may not capture all 3D effects of the vocal tract, particularly for complex articulations or nasal sounds (though the nasal tract is noted as a limitation due to MRI data availability). The inverse problem is inherently ill-posed, meaning multiple geometries can produce the same sound, leading to potential ambiguities in reconstruction. The MRI reconstruction quality is limited by the resolution and coverage of the training MRI dataset.
This work has significant implications for linguistics, speech pathology, and medical imaging. By providing a physically grounded link between speech and articulatory geometry, it enables new tools for voice coaching, language acquisition studies, and non-invasive visualization of vocal tract movements. The self-supervised nature of the autoencoder and MRI generator makes these tools scalable and accessible without the need for expensive paired data collection. The paper presents a novel differentiable acoustic simulator that enables the reconstruction of vocal tract geometry and MRI videos from speech alone. By combining frequency-domain physics with neural field parameterization, the authors solve a notoriously non-convex inverse problem, offering a robust framework for physically grounded speech analysis and synthesis.
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of $ฯ$-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on $ฯ$-Voice with an acceptable increase in the delegation rate.
Primary: Tsinghua University
All Institutions: Tsinghua University, ByteDance
SALMONN-duo introduces an adaptive dual-system voice agent that balances real-time responsiveness with deep reasoning via knowledge-boundary-aware training and cost-aware reinforcement learning. The paper presents a rigorous methodology for training a full-duplex frontend to delegate tasks to an asynchronous backend only when necessary, demonstrating superior performance-cost trade-offs and improved safety in complex, multi-turn, environment-grounded tasks compared to existing full-duplex voice agents.
The paper proposes SALMONN-duo, a dual-system architecture for full-duplex voice agents. The core methodological contribution is the separation of a lightweight, always-on full-duplex frontend (System 1) from a powerful, asynchronous backend agent (System 2). The novelty lies in the "adaptive delegation" mechanism, where System 1 learns to decide when to answer directly versus when to delegate to System 2. This is achieved through two key training strategies: (1) Knowledge-Boundary-Aware Supervised Fine-Tuning (SFT), which uses the model's own rollout correctness to label delegation targets, and (2) Cost-Aware Group Relative Policy Optimization (GRPO), which introduces a reward function that balances task success, safety (hallucination/policy violation), and the cost of unnecessary delegation. The design is practical, addressing the latency-cost trade-off inherent in real-time voice interactions. The use of a delegation token within the response stream rather than a separate head is a specific architectural choice validated in the paper.
The experimental evaluation is comprehensive, covering single-turn QA, multi-turn daily conversations, and complex environment-grounded tasks (using a customized $\tau$-Voice benchmark). The paper compares against recent full-duplex and half-duplex baselines (Moshi, VoiceChat, etc.). Results show that SALMONN-duo achieves a better performance-cost trade-off, with lower invocation rates on easier tasks and higher accuracy on complex reasoning tasks compared to baselines that either always invoke or rarely invoke tools. The inclusion of safety rewards in the RL phase demonstrates a reduction in hallucinations and policy violations. The evaluation of interruption handling (barge-in) is a strong point, showing the model can maintain conversational context while waiting for backend responses.
The paper provides detailed descriptions of the data generation pipeline, including the roles of user, assistant, and supervisor simulators. It specifies model architectures (Llama-3.1-8B, CosyVoice2), training hyperparameters (learning rates, batch sizes, GPU counts), and reward weights. However, the reliance on proprietary models (GPT-4o, GPT-5.2, OpenAI TTS) for simulation and evaluation limits full open-source reproducibility. The code and specific weights are not explicitly linked in the provided text, though the paper claims to be open-source work.
The system relies on a strong proprietary backend (GPT-5.2) for the "System 2" component, which may not be accessible to all researchers or deployable in privacy-sensitive on-device scenarios. The evaluation assumes ground-truth dialogue history for the backend in some settings, which may overestimate performance in noisy real-world ASR conditions (though ASR sensitivity is analyzed). The "cost" is defined primarily as the number of backend invocations, not necessarily computational cost or latency in a strict sense, though latency is simulated.
This work is significant for the development of practical, low-latency voice assistants that can handle complex tasks without sacrificing conversational fluidity. The dual-process approach offers a scalable path for integrating heavy reasoning capabilities into real-time speech interfaces. It addresses a critical gap in current voice agents: the inability to efficiently manage the trade-off between responsiveness and deep reasoning/tool use. SALMONN-duo introduces an adaptive dual-system voice agent that balances real-time responsiveness with deep reasoning via knowledge-boundary-aware training and cost-aware reinforcement learning. The paper presents a rigorous methodology for training a full-duplex frontend to delegate tasks to an asynchronous backend only when necessary, demonstrating superior performance-cost trade-offs and improved safety in complex, multi-turn, environment-grounded tasks compared to existing full-duplex voice agents.
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the generative trajectory. Our analysis reveals that pretrained TTS models are substantially editable in emotion, but such editability is architecture- and trajectory-dependent and can be disrupted by early flow-matching steps, while cross-speaker emotion transport carries additional acoustic attributes beyond emotion. To address these limitations, we propose SEmoEdit, the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target emotions, enabling robust, flow-based speech emotion editing directly within pretrained TTS models. SEmoEdit unifies three core operations: emotion replacement, emotion erasure, and continuous emotion interpolation, requiring neither parameter updates nor task-specific optimization. To systematically evaluate these capabilities, we introduce SEmoEditBench, a dataset comprising 600 editing cases, and conduct extensive experiments across state-of-the-art (SOTA) models and backbones. Our results show that SEmoEdit is highly effective and broadly applicable, outperforming existing training-based and activation-steering methods. Ultimately, this work reveals that pretrained speech flows possess rich, latent emotion-editing capabilities, providing useful guidance for real applications. Code, benchmark, and Audio samples are available at https://github.com/imxtx/SEmoEdit.
Primary: The Hong Kong University of Science and Technology (Guangzhou)
All Institutions: The Hong Kong University of Science and Technology (Guangzhou)
The paper presents a novel training-free framework for speech emotion editing by leveraging dynamic velocity transport in flow-matching TTS models, supported by a systematic analysis of editability and a new benchmark. The technical contribution is strong, offering a robust alternative to training-based and static steering methods, with comprehensive experiments validating its effectiveness across multiple state-of-the-art backbones.
The paper proposes SEmoEdit, a training-free framework for speech emotion editing based on flow-matching TTS models. The core contribution is formulating emotion editing as dynamic velocity transport. Instead of using static activation steering vectors (which are often unstable and require paired data to derive), the method dynamically computes the velocity difference between source and target conditions at each step of the ODE integration. This approach is theoretically grounded in the properties of Conditional Flow Matching (CFM). The method also introduces "emotion bridging" to handle interpolation artifacts, where intermediate states are not well-formed speech samples. The inclusion of a systematic probing study (Sec. 2) to diagnose *when* and *how* editing works along the trajectory is a strong methodological addition, providing empirical justification for the design choices (e.g., skipping early steps to preserve timing).
The authors introduce SEmoEditBench, a 600-case benchmark, which is a valuable contribution to the field as standardized benchmarks for speech editing are scarce. Experiments are conducted on three SOTA backbones (F5-TTS, CosyVoice 2, IndexTTS 2), demonstrating broad applicability. The evaluation metrics are comprehensive, covering effectiveness (TEP, SES), preservation (WER, S-SIM), quality (UTMOS), and subjective evaluation (MOS). The results show that SEmoEdit outperforms training-based and activation-steering baselines. The ablation studies on trajectory timing and attribute factorization are insightful and support the paper's claims.
The paper provides a GitHub link with code, benchmark, and audio samples. The methodology is described in sufficient detail (equations for velocity transport, noise coupling, and bridging) to allow reproduction. The use of standard open-source TTS models (F5-TTS, CosyVoice 2) further enhances reproducibility.
The method relies on the quality of the underlying TTS model's velocity field; if the base model is poor at emotion generation, the editing will suffer. The "emotion bridging" step for interpolation adds computational overhead and complexity. The benchmark, while useful, is relatively small (600 cases) compared to large-scale speech datasets. The paper focuses on English and Chinese (implied by ESD/IEMOCAP/RAVDESS/CREMA-D usage), and generalization to other languages is not explicitly tested.
This work has significant implications for the controllability of generative speech models. By demonstrating that pretrained models contain latent editing capabilities that can be unlocked via inference-time manipulation, it opens the door to more flexible and efficient speech editing pipelines without the need for fine-tuning. This is particularly relevant for applications in dubbing, accessibility, and personalized voice assistants where emotional nuance is important. The insights into trajectory-dependent editability can guide the design of future controllable TTS systems. The paper presents a novel training-free framework for speech emotion editing by leveraging dynamic velocity transport in flow-matching TTS models, supported by a systematic analysis of editability and a new benchmark. The technical contribution is strong, offering a robust alternative to training-based and static steering methods, with comprehensive experiments validating its effectiveness across multiple state-of-the-art backbones.
Explicit textual Chain-of-Thought (CoT) has improved the reasoning ability of large audio language models (LALMs). However, textual CoTs are often constructed from text captions of audio and provide limited access to the acoustic evidence, which can introduce problems like hallucination. To address this modality-gap issue, we introduce JELAR, a Joint-Embedding Predictive Architecture (JEPA)-based latent reasoning framework that conditions latent reasoning supervision on acoustic representations learned from raw waveforms. During training, a frozen WavJEPA model provides representations learned directly from raw waveforms. A non-causal expert first constructs answer-aware queries, which cross-attend to WavJEPA embeddings to produce latent reasoning targets. The LALM is trained to predict these targets before generating its response. Experimental results show that JELAR improves the Audio-Reasoner baseline by 2.70 and 9.10 absolute percentage points on MMAU-mini and MMAR, respectively, demonstrating the effectiveness of JEPA-conditioned latent reasoning as an alternative to explicit textual CoT supervision.
Primary: Nanyang Technological University
All Institutions: Nanyang Technological University, AI Singapore, Peking University
[One sentence main contribution]. The paper introduces JELAR, a JEPA-conditioned latent reasoning framework that improves audio reasoning accuracy by grounding latent supervision in raw waveform representations, effectively addressing the modality gap inherent in textual Chain-of-Thought approaches.
The paper proposes JELAR, a framework that replaces explicit textual Chain-of-Thought (CoT) with latent reasoning conditioned on acoustic representations. The core methodological contribution is the use of a frozen WavJEPA encoder to generate acoustic embeddings, which are then processed by a non-causal expert (BERT-based) to create answer-aware latent targets. These targets serve as supervision for the LALM during training. The approach is logically sound, addressing the "modality gap" where text-based CoT may hallucinate or miss acoustic details. However, the reliance on a separate, non-causal expert for target generation introduces complexity and a potential train-test mismatch (teacher forcing vs. autoregressive inference), which is a known challenge in latent reasoning. The integration of JEPA is a creative application of predictive coding principles to audio reasoning.
The experiments are conducted on MMAU-mini and MMAR benchmarks, comparing JELAR against the Audio-Reasoner baseline. The reported improvements are significant, particularly on MMAR (+9.10 points), suggesting the method is effective for multi-step reasoning. Ablation studies confirm the contribution of acoustic conditioning and latent reasoning length. However, the evaluation is limited to these two benchmarks, and comparisons with other state-of-the-art audio LLMs (like GPT-4o or Gemini) are not directly made in the main results table for the proposed method, only for context. The lack of qualitative analysis of the latent reasoning process is a minor weakness.
The paper provides sufficient detail on the architecture and training procedure, including the use of frozen WavJEPA and BERT-base. Hyperparameters like K=20 are specified. However, specific details on the training data composition, optimization hyperparameters (learning rates, batch sizes), and the exact implementation of the cross-attention module are not fully detailed in the provided text. The code repository is not linked, which hinders immediate reproducibility.
The method requires a separate non-causal expert during training, which increases computational overhead and complexity. The latent reasoning space is not interpretable, making it difficult to debug errors. The performance gains, while positive, are based on a specific baseline (Audio-Reasoner) and may not generalize to other LALM architectures without further adaptation. The paper does not discuss the computational cost of inference compared to standard CoT or direct answering.
This work contributes to the growing field of latent reasoning in multimodal models, offering an alternative to text-centric CoT. It highlights the utility of JEPA-style representations for audio tasks, potentially inspiring similar approaches in other modalities. The focus on reducing hallucination by grounding reasoning in acoustic evidence is a valuable direction for improving the reliability of audio LLMs. [One sentence main contribution]. The paper introduces JELAR, a JEPA-conditioned latent reasoning framework that improves audio reasoning accuracy by grounding latent supervision in raw waveform representations, effectively addressing the modality gap inherent in textual Chain-of-Thought approaches.
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Primary: HKGAI
All Institutions: HKGAI
YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
The paper proposes a unified architecture (YuE2) that bridges symbolic and audio music generation using an AR-NAR Mixture-of-Transformers (MoT). The core methodological contribution is the "symbolic planning" stage, where the model first generates a readable score (melody, harmony, form) before expanding into semantic tokens and finally acoustic latents via flow matching. This hierarchical approach allows for explicit control over composition. Additionally, the paper introduces two auxiliary models: MERT2 for semantic representation learning and SheetSage2 for symbolic supervision (lead-sheet transcription), which are used to train the main model on unaligned audio data. The integration of these components into a single checkpoint that supports editing and cover generation is a significant architectural advancement.
The evaluation is extensive, utilizing both objective benchmarks (WildSongBench, SongBench, MARBLE) and subjective expert listening tests. The results show YuE2 outperforming public baselines and being competitive with proprietary systems like Suno v4.5 and v5. The ablation study comparing generation with and without symbolic planning demonstrates a clear preference for the planned approach (49.3% vs 34.6%). The introduction of MERT2 and SheetSage2 is supported by state-of-the-art results on their respective benchmarks (MARBLE and transcription metrics). The inclusion of zero-shot cover generation and agentic editing case studies further validates the utility of the unified framework.
As a technical report from an industry lab (HKGAI), the paper likely lacks full open-source code and model weights at the time of release, which limits immediate reproducibility. However, the detailed description of the architecture (AR-NAR MoT, flow matching) and the specific benchmarks used provides a clear roadmap for replication. The reliance on proprietary datasets for training the auxiliary models may pose a barrier for independent researchers.
The paper relies heavily on expert listening tests, which can be subjective and difficult to scale. The comparison with proprietary systems (Suno) is limited to specific versions and may not reflect the full capability of those closed-source models. The "best-of-8" selection strategy, while effective for benchmarking, introduces a computational overhead that may not be practical for real-time applications. Additionally, the paper does not extensively discuss the computational cost of the multi-stage generation process compared to direct audio generation.
This work has significant implications for the music production industry by providing a tool that allows for both high-quality audio generation and explicit compositional control. The ability to edit scores and generate covers zero-shot opens up new possibilities for music creation, education, and collaboration. The unification of symbolic and audio domains could lead to more interpretable and controllable generative music systems, potentially influencing how AI is integrated into professional music workflows. YuE2 introduces a unified AR-NAR Mixture-of-Transformers architecture that unifies symbolic and audio music generation through a symbolic planning stage, achieving frontier quality and enabling editable, controllable music creation. The paper demonstrates that explicitly modeling composition (score) before audio realization improves perceived quality and musicality, supported by strong benchmark results and expert preferences. The introduction of MERT2 and SheetSage2 provides robust semantic and symbolic supervision, setting new standards in music representation learning and transcription. This work represents a significant step towards interpretable and controllable generative music systems, bridging the gap between symbolic composition and audio production.
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
Primary: The Chinese University of Hong Kong
All Institutions: The Chinese University of Hong Kong, Microsoft Corporation
The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
The paper proposes DuraS2ST, a framework that reformulates duration-aligned speech-to-speech translation (S2ST) as a reasoning problem rather than an acoustic post-processing task. The core methodological contribution is the integration of Chain-of-Thought (CoT) reasoning into the speech generation pipeline. The model first generates an explicit textual rationale planning the target wording and phonetic length, followed by the synthesis of interleaved text-acoustic tokens. This is supported by the construction of DuraSet-440K, a large-scale corpus specifically designed for this purpose. The training paradigm is two-phase: Supervised Fine-Tuning (SFT) on the new corpus, followed by Group Relative Policy Optimization (GRPO). The RL phase introduces two specific technical innovations: the Duration Margin Reward (DMR), which acts as a soft constraint to balance translation quality with duration consistency, and Modality-Aware Reward Attribution (MARA), which prevents cross-modality reward contamination by assigning duration rewards only to acoustic tokens and quality rewards to both text and acoustic spans. This approach is technically sound and addresses a specific gap in current S2ST systems where duration control is often handled via black-box speed tokens or post-hoc time-stretching.
The experiments are conducted on the CVSS-T benchmark, which is disjoint from the training data, providing a valid out-of-domain evaluation. The paper compares DuraS2ST against strong baselines, including commercial models (GPT-4o, Qwen2.5-Omni, Kimi-Audio) and open-source SLMs (Step-Audio-2-mini). The results demonstrate that DuraS2ST achieves a superior balance between translation quality (BLEU, COMET) and duration consistency (SLC-0.2, SLC-0.4, MADE, MRDE). Notably, the method shows substantial gains in a zero-shot RL setting, highlighting the effectiveness of the reward design. The inclusion of both open-source and commercial baselines strengthens the experimental validation.
The paper provides a project page with a GitHub repository link. It details the training setup, including hyperparameters for SFT and GRPO, and describes the construction pipeline for DuraSet-440K in detail. The use of standard metrics (BLEU, COMET, SECS, SLC) and a public benchmark (CVSS-T) enhances reproducibility. However, the specific implementation details of the "duration-controllable TTS model" used for data synthesis are not fully specified in the main text, though likely detailed in the appendix or code.
The method relies on a specific base model (Step-Audio-2-mini-Think), and its generalizability to other SLM architectures is not extensively tested. The construction of DuraSet-440K is resource-intensive, requiring LLM-based translation candidate generation and duration-controllable TTS synthesis. The evaluation is limited to English-Chinese pairs; performance on other language pairs with different phonetic structures remains unknown. Additionally, the "zero-shot RL" claim needs careful interpretation as it likely refers to the RL phase without additional SFT on the target domain, but the model is still SFT'd on DuraSet-440K.
This work has significant implications for video dubbing, simultaneous interpretation, and real-time communication applications where audio-visual synchronization is critical. By treating duration alignment as a reasoning task, it offers a more interpretable and controllable approach compared to traditional acoustic constraints. The release of DuraSet-440K and the code will likely accelerate research in duration-controlled speech generation and multimodal reasoning. The paper introduces DuraS2ST, a novel framework that integrates Chain-of-Thought reasoning and reinforcement learning to achieve duration-aligned speech-to-speech translation. By constructing a dedicated corpus (DuraSet-440K) and designing modality-aware reward mechanisms (DMR and MARA), the authors effectively address the challenge of temporal consistency in S2ST, demonstrating superior performance over strong commercial and open-source baselines on the CVSS-T benchmark.
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
Primary: Alibaba Group
All Institutions: Shanghai Jiao Tong University, Alibaba Group, Tsinghua University, University of Cambridge, Nankai University, Chinese University of Hong Kong, SII
Pruned CTC enables memory-efficient training of large-vocabulary ASR models by exploiting the sparsity of CTC alignments, and LLM-CTC successfully adapts pretrained LLMs for fast, non-autoregressive speech recognition with strong performance in both offline and streaming settings.
The paper introduces "Pruned CTC," a memory-efficient implementation of Connectionist Temporal Classification (CTC) that exploits the sparsity of valid alignments. The core insight is that while the softmax normalization requires the full vocabulary, the dynamic programming alignment only involves the target tokens and the blank symbol. By restricting the alignment computation to this small subset while retaining full-vocabulary normalization via a "complement" class, the authors prove exact equivalence in loss and gradients. This is a mathematically sound and clever optimization. The method is further extended to "LLM-CTC," which adapts pretrained Large Language Models (LLMs) for non-autoregressive ASR. This is a significant architectural contribution, as it allows leveraging the linguistic knowledge of LLMs without the latency of autoregressive decoding, and it handles streaming via bounded-history attention masks without requiring chunk-level alignments.
The experiments are robust and cover a wide range of model sizes (0.6B to 32B Qwen3 models) and datasets (GigaSpeech, etc.). The results demonstrate a 5.1x reduction in memory with minimal time overhead. The comparison against LLM-CE (Cross-Entropy) shows that LLM-CTC achieves comparable accuracy (within 7% relative WER) while being 7-10x faster in recognition. The streaming results are particularly strong, showing that the method avoids the complex alignment requirements of previous streaming LLM-ASR methods. The inclusion of both offline and streaming evaluations provides a comprehensive view of the method's utility.
The authors provide a GitHub repository link and state that code and pretrained models will be open-sourced. The paper includes detailed algorithmic descriptions (Algorithm 1) and mathematical proofs, which aids reproducibility. The specific hyperparameters for the alignment beam and chunking are provided in the appendices (referenced).
The method relies on the assumption that the union of target tokens in a batch is small relative to the full vocabulary, which holds for typical ASR tasks but might not for extremely diverse or noisy data. The "finite-beam" approximation introduces a small error, though the paper claims it is negligible ($10^{-11}$). The streaming extension requires careful management of KV caches and attention masks, which can be complex to implement in production systems.
This work has high potential impact on the deployment of LLM-based ASR systems. By enabling non-autoregressive decoding with LLMs, it significantly reduces inference latency and computational cost, making real-time, high-quality speech recognition more accessible. The memory efficiency also allows for training larger models or larger batches on standard hardware. This could accelerate the adoption of LLMs in speech processing pipelines. Pruned CTC enables memory-efficient training of large-vocabulary ASR models by exploiting the sparsity of CTC alignments, and LLM-CTC successfully adapts pretrained LLMs for fast, non-autoregressive speech recognition with strong performance in both offline and streaming settings.
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation. Its main weakness is retokenization: decoding generated speech to a waveform and encoding it again can change token identities and erode the watermark. To make the watermark robust to these changes, we propose Redwing, REtokenization-Durable Watermarking IN Generation. It builds a graph from the token substitutions observed under retokenization, whose Laplacian yields a basis that assigns similar values to tokens likely to substitute for one another. Over this basis, embedding and detection functions are jointly optimized to preserve watermark signal through retokenization while limiting embedding distortion and detector variability on unwatermarked speech. On the Moshi full-duplex system, after eight consecutive passes of Mimi resynthesis, Redwing achieves 80.7% TPR at a calibrated 1% FPR, compared with 8.3% for KGW and at most 7.3% for WMAR. It also has the highest TPR after eight passes through three other neural codecs (77.5-93.0%), and the gains generalize to TTS models at a speech-quality cost close to that of KGW. These results show that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.
Primary: University of Zurich and ETH Zurich
All Institutions: Institute of Neuroinformatics, NAVER Cloud, University of Zurich, ETH Zurich
[One sentence main contribution]. The paper introduces Redwing, a spectral graph-based token watermarking method that exploits the structure of codec retokenization to achieve high robustness against repeated resynthesis attacks, significantly outperforming existing token-level and post-hoc watermarking techniques in speech generation systems.
The paper proposes "Redwing," a token-level watermarking scheme for speech generation that leverages the structural properties of neural codec retokenization. Unlike standard token watermarks (like KGW) that rely on specific token identities, which are often altered during the decode-encode cycle, Redwing constructs a graph where nodes are tokens and edges represent substitution frequencies observed during retokenization. By using the eigenvectors of the graph Laplacian (specifically those with small eigenvalues) as a basis, the method ensures that tokens likely to substitute for one another have similar watermark values. The embedding and detection functions are jointly optimized over this basis to maximize robustness against resynthesis while minimizing distortion to the generated speech. This is a clever application of spectral graph theory to a practical audio security problem, moving beyond treating retokenization as noise to exploiting its deterministic structure.
The experiments are rigorous, testing the watermark on the Moshi full-duplex dialogue system and two TTS models (CosyVoice3, MOSS-TTS). The primary attack simulated is repeated resynthesis through the Mimi codec (8 passes). Redwing achieves an 80.7% True Positive Rate (TPR) at 1% False Positive Rate (FPR), significantly outperforming baselines like KGW (8.3%) and WMAR (7.3%). The paper also demonstrates generalizability to other neural codecs (77.5-93.0% TPR) and maintains speech quality comparable to the KGW baseline. The inclusion of multiple models and codecs strengthens the claim of general applicability.
The method is described as "training-free" in the sense that it does not require retraining the speech model, but it does require building the substitution graph from a corpus and solving for embedding/detection functions. The paper states it applies to released models unchanged, which is a strong reproducibility feature. However, specific details on the optimization process for the embedding/detection functions and the exact corpus used for graph construction are not fully detailed in the truncated text, though the methodology is clearly defined.
The method relies on the stability of the retokenization substitution patterns. If a codec is updated or if the attack involves a different type of processing (e.g., heavy noise addition, pitch shifting) that disrupts the token substitution graph, the watermark's robustness may degrade. Additionally, the construction of the graph requires access to the codec's encoder/decoder to perform the substitution counting, which might not be feasible for all black-box systems.
This work has significant implications for the provenance of AI-generated speech. As speech models become more capable and widespread, robust watermarking is essential for distinguishing synthetic from human speech. By addressing the specific weakness of token-level watermarks in audio (retokenization), this paper provides a viable path for deploying watermarks in real-world speech generation systems without modifying the underlying model architecture. [One sentence main contribution]. The paper introduces Redwing, a spectral graph-based token watermarking method that exploits the structure of codec retokenization to achieve high robustness against repeated resynthesis attacks, significantly outperforming existing token-level and post-hoc watermarking techniques in speech generation systems.
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.
Primary: Sapienza University of Rome
All Institutions: Sapienza University of Rome, Moises Systems, Inc., Paradigma
SAGE introduces a compact, spectrogram-native VAE with semantic distillation that achieves state-of-the-art reconstruction and latent semantics at low inference cost. The paper demonstrates that adapting vision transformers (SwinV2) to audio spectrograms, combined with careful curriculum learning for semantic alignment, can outperform larger, waveform-based models in both fidelity and utility for downstream generative tasks.
The paper proposes SAGE, a 105M parameter Variational Autoencoder (VAE) for audio that operates on the complex Short-Time Fourier Transform (STFT) rather than raw waveforms or magnitude spectrograms. The core architectural innovation is the adaptation of the SwinV2 vision transformer backbone to audio, utilizing rectangular patches and attention windows to handle the time-frequency structure of spectrograms. A key methodological contribution is the "semantic distillation" strategy, where the latent space is aligned with embeddings from a frozen CLAP (Contrastive Language-Audio Pretraining) model. This alignment is introduced via a delayed schedule (detached warm-up) to ensure reconstruction quality is established before semantic constraints are applied, preventing the semantic loss from degrading fidelity. The training process is divided into two phases: pretraining with a full adversarial objective (using a WavTokenizer discriminator) and a fine-tuning phase where the encoder is frozen and the decoder is refined. The use of a sum-and-difference loss for stereo imaging is also a notable technical detail, addressing the common issue of stereo collapse in autoencoders.
The experimental evaluation is rigorous and comprehensive. The authors compare SAGE against five strong baselines (Stable Audio Open, SAME-L, SAME-S, CoDiCodec, Music2Latent) across five different held-out datasets, including in-domain (FMA) and out-of-domain (MoisesDB, MusicCaps, Song Describer) sets. Metrics cover perceptual (CLAP similarity), distributional (Frรฉchet Audio Distance in MERT, PANN, and CLAP spaces), and sample-exact (SI-SDR, STFT distance) categories. Crucially, the paper includes a MUSHRA listening test with 21 raters, showing SAGE is statistically indistinguishable from the much larger SAME-L model. The semantic evaluation is extensive, using 19 probing tasks (genre, artist, instrument, etc.) in the MAEB format, where SAGE achieves state-of-the-art results on all tasks. The ablation studies on the semantic distillation schedule and adversarial components provide strong evidence for the design choices.
The paper offers high reproducibility. The authors provide a link to the GitHub repository containing code, weights, and the evaluation harness. The training data consists of publicly available corpora (FMA, MTG-Jamendo, M4Singer), and the paper details the specific splits and preprocessing steps. The hyperparameters for both training phases are listed in the appendix. The use of standard libraries and the release of the evaluation harness allows other researchers to easily verify the results and compare their own models.
The primary limitation is the reliance on the CLAP model for semantic alignment; if the CLAP embeddings are biased or limited in their musical understanding, SAGE's latent space will inherit these limitations. The model is trained on music, so its performance on non-music audio (speech, sound effects) is not evaluated and likely poor. The inference cost, while competitive, is still higher than simple convolutional codecs, which may be a barrier for real-time mobile applications. The paper focuses on music, so generalization to other audio domains is an open question.
SAGE provides a lightweight, high-fidelity, and semantically rich audio encoder that can serve as a drop-in replacement for existing autoencoders in latent diffusion pipelines for music generation. Its semantic structure could enable better control and conditioning in generative models. The efficient inference cost makes it suitable for interactive applications. The comprehensive benchmarking harness contributes to the standardization of audio autoencoder evaluation. SAGE introduces a compact, spectrogram-native VAE with semantic distillation that achieves state-of-the-art reconstruction and latent semantics at low inference cost. The paper demonstrates that adapting vision transformers (SwinV2) to audio spectrograms, combined with careful curriculum learning for semantic alignment, can outperform larger, waveform-based models in both fidelity and utility for downstream generative tasks.
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%. More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent transfer performance of a two-model TTS-voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.
Primary: Ca' Foscari University of Venice
All Institutions: Ca' Foscari University of Venice, Kandinsky Lab, Sleeping AI, Singapore University of Technology and Design
The paper introduces DEFINE, a framework that decouples speaker identity and accent in zero-shot TTS by using separate audio exemplars and a prototype-anchored encoder, achieving comparable accent transfer to a two-model cascade while better preserving speaker identity. The technical contribution is significant, with a novel supervision mechanism for accent embeddings and a continuous inference-time control for accent strength. The experiments are thorough, including objective metrics and a small-scale listening test, demonstrating the effectiveness of the approach. The work is well-positioned to influence future research on disentangled speech synthesis and fine-grained control in TTS systems.
The paper proposes DEFINE, a framework for disentangling speaker identity and accent in zero-shot TTS. The core technical contribution is the "prototype anchoring" mechanism, which uses a learned table of accent prototypes to supervise an exemplar encoder (based on frozen XLS-R features). This addresses the weak supervision signal for accent representations in flow-matching objectives. The method injects the accent embedding into the F5-TTS backbone via additive shifts to timestep and text embeddings, scaled by RMS to ensure proper magnitude. A key feature is the inference-time guidance weight $w$ that allows continuous control over accent strength without retraining. The use of LoRA for parameter-efficient adaptation is standard but effective. The methodology is sound, though the reliance on a specific backbone (F5-TTS) and the specific injection point (adaLN modulation via embedding shifts) limits generalizability to other architectures.
The experiments are comprehensive, evaluating seen, held-out, and out-of-domain accents. The use of a 15-way logistic regression probe on XLS-R features for accent accuracy is a reasonable objective metric, though it is a proxy for perceptual accent. The comparison against a Seed-VC cascade is strong, showing that DEFINE matches accent transfer performance while improving speaker similarity. The listening test with 11 listeners is a positive addition, confirming that the accent changes are perceptually real and that the guidance weight controls the perceived accent strength. However, the sample size for the listening test (11 listeners, 10 trials) is relatively small, and the statistical significance, while present, is based on a limited number of votes. The WER and UTMOS metrics are standard and reported, showing that intelligibility and quality are maintained.
The paper provides a GitHub link, which is a strong indicator of reproducibility. The training data sources (Common Voice, in-house studio recordings) are described, but the in-house data is not publicly available, which limits full reproducibility. The hyperparameters for LoRA, the prototype table size, and the guidance weight are specified. The use of frozen XLS-R and specific layer selection (layer 15) is detailed. Overall, reproducibility is good, assuming access to the code and the ability to recreate the in-house speaker prompts or use a similar public dataset.
The main limitation is the reliance on a specific TTS backbone (F5-TTS), which may not generalize to other architectures without significant modification. The accent control is limited to the accents seen during training or those that can be represented by the exemplar encoder; truly novel accents not in the training distribution may not be well-captured. The listening test is small-scale, and the objective accent probe may not fully capture perceptual accent nuances. The method requires separate exemplar clips for accent, which adds a data requirement at inference time.
The work has significant implications for personalized TTS systems, allowing users to control not just the voice but also the accent of the synthesized speech. This could be useful for language learning, accessibility, and creative applications. The decoupling of identity and accent is a step towards more fine-grained control over speech synthesis, which is a key goal in the field. The method's ability to generalize to out-of-domain accents is particularly promising for real-world applications where the target accent may not be in the training set. The paper introduces DEFINE, a framework that decouples speaker identity and accent in zero-shot TTS by using separate audio exemplars and a prototype-anchored encoder, achieving comparable accent transfer to a two-model cascade while better preserving speaker identity. The technical contribution is significant, with a novel supervision mechanism for accent embeddings and a continuous inference-time control for accent strength. The experiments are thorough, including objective metrics and a small-scale listening test, demonstrating the effectiveness of the approach. The work is well-positioned to influence future research on disentangled speech synthesis and fine-grained control in TTS systems.
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Primary: Unknown
All Institutions: Unknown
The paper introduces VoxMem, a comprehensive benchmark for evaluating multi-session spoken conversational memory across acoustic evidence types and memory operations. It provides a principled taxonomy and rigorous construction pipeline, revealing significant gaps in current LALMs' ability to retain non-lexical acoustic information over long contexts.
The paper proposes a principled two-dimensional taxonomy for spoken conversational memory, crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal). This is a significant methodological advance over prior benchmarks that treated memory as a single-session, lexical-only problem. The construction pipeline is rigorous, utilizing a three-stage process (planning, dialogue writing, speech synthesis) with specific controls to ensure "acoustic necessity" (i.e., the answer cannot be derived from the transcript alone). The use of Higgs-TTS-3 and VCTK voices, along with mixed environmental sounds from ESC-50, provides a controlled yet realistic audio environment. The inclusion of "haystack" and "filler" sessions to create multi-session histories of varying lengths (8K-64K tokens) is a strong design choice that isolates the effect of context length from question difficulty.
The evaluation covers 15 Large Audio Language Models (LALMs), including both open-weight and proprietary models. The results reveal a clear hierarchy of difficulty: models perform significantly better on speech semantics than on audio-native cues (speaker, paralinguistic, environmental). The finding that no model exceeds 40% accuracy at the 32K token budget highlights a critical gap in current LALM capabilities. The error analysis is particularly valuable, distinguishing between binding failures (wrong speaker) and retention failures (lost cue), which provides actionable insights for future model development. The controlled scaling analysis shows that performance degrades as history length increases, with non-lexical information degrading faster than lexical information.
The paper provides detailed descriptions of the construction pipeline, quality control checks, and evaluation protocol. The use of specific TTS systems, voice datasets, and sound event datasets enhances reproducibility. The release of the benchmark instances, question templates, and judge prompts (mentioned in the text) further supports reproducibility. However, the reliance on proprietary LLMs (Gemini-3.7-Flash, GPT-5.6-Luna) for generation and evaluation introduces a dependency on external services, which may limit full reproducibility for all researchers.
The benchmark relies on synthetic speech generated by TTS systems, which may not fully capture the variability and noise of real-world human speech. The "acoustic necessity" check, while rigorous, uses a specific LLM (Gemini-3.7-Flash) for validation, which could introduce bias if that model has specific strengths or weaknesses in audio understanding. The paper does not extensively discuss the computational cost of evaluating models on 64K token audio histories, which may be prohibitive for some researchers. Additionally, the focus on English speech (implied by the datasets used) limits the generalizability of the findings to other languages.
VoxMem addresses a fundamental challenge in the development of long-term conversational AI systems. By providing a standardized benchmark for multi-session, audio-native memory, it enables fair comparison of LALMs and guides future research toward improving non-lexical acoustic memory. The taxonomy and benchmark are likely to become a standard reference in the field, driving progress in speaker identification, paralinguistic understanding, and environmental sound recognition within conversational contexts. The findings on the distinct failure modes for different evidence types will inform the design of specialized memory modules in LALMs. The paper introduces VoxMem, a comprehensive benchmark for evaluating multi-session spoken conversational memory across acoustic evidence types and memory operations. It provides a principled taxonomy and rigorous construction pipeline, revealing significant gaps in current LALMs' ability to retain non-lexical acoustic information over long contexts.
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
Primary: KAIST
All Institutions: KAIST, VGG, University of Oxford
The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
The paper employs a rigorous mechanistic interpretability framework, combining Representational Similarity Analysis (RSA) and Causal Mediation Analysis (CMA) to dissect the internal workings of Audio-Visual LLMs. The authors successfully identify a three-stage symbolic binding mechanism (Anchor ID Retrieval, Target ID Selection, Feature Retrieval) that relies on modality-specific symbolic variables (Temporal IDs for audio, Position IDs for vision). This is a sophisticated approach that moves beyond black-box evaluation to understand *how* models process cross-modal information. The proposed intervention, using an off-the-shelf Active Speaker Detection (ASD) model to overlay bounding boxes, is a clever, training-free (and lightly fine-tuned) solution that directly targets the identified bottleneck in audio-visual alignment.
The experiments are extensive and well-structured. The authors validate their mechanistic findings across four different AVLLM architectures (video-SALMONN2+, Qwen2.5-Omni, MiniCPM-o-4.5) using both synthetic toy datasets and real-world benchmarks (SocialOmni, AVSpeaker, DailyOmni, etc.). The ablation studies effectively isolate the contribution of the ASD prompting, showing that it outperforms other training-free decoding methods and standard fine-tuning. The generalization to broader audio-visual benchmarks (DAVE, OmniBench, WorldSense) further strengthens the claim of the method's utility.
The paper provides detailed implementation specifics, including LoRA ranks, training steps, and dataset construction methods. The use of standard off-the-shelf models and clear descriptions of the prompting strategy enhances reproducibility. However, the specific synthetic video generation pipeline and the exact ASD model used (referenced as [CITATION]) would need to be clearly specified in the final publication for full reproducibility.
The primary limitation is the reliance on a synthetic toy dataset for the core mechanistic analysis, which may not fully capture the complexity of real-world multi-speaker scenarios, although the authors do validate on real-world data. Additionally, the method depends on the accuracy of the external ASD model; if the ASD model fails, the prompting strategy may degrade performance. The fine-tuning, while lightweight, still requires access to the model's weights and a GPU, which may not be feasible for all users.
This work has significant implications for the development of more robust multimodal AI systems. By identifying specific failure modes in cross-modal binding, it provides actionable insights for improving model architectures and training strategies. The proposed ASD prompting method is a practical, low-cost solution that can be easily integrated into existing AVLLM pipelines, potentially improving performance in applications like video conferencing, accessibility tools, and video understanding. The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.
Primary: Sakana AI
All Institutions: Sakana AI
The paper introduces a symbiotic architecture that decouples audio prefilling from the LLM backbone by directly generating the KV cache via a lightweight injector, thereby reducing computational cost and preventing catastrophic forgetting. The technical contribution is significant in addressing two major bottlenecks in audio-LLM development, though the experimental validation is limited to small-scale models, leaving the full scalability benefits to be verified in future work.
The paper proposes a "symbiotic" architecture where an audio injector module generates the Key-Value (KV) cache for a frozen Large Language Model (LLM), bypassing the need to pass audio embeddings through the LLM's self-attention layers during prefilling. The injector is a CNN-based module (using LConv blocks) that maps audio encoder outputs directly to the LLM's KV space. Key technical contributions include a "KV Scale Matching" strategy to align the injector's output distribution with the LLM's internal KV distribution, and "Noisy RoPE" training to improve robustness to sequence lengths unseen during training. The approach is theoretically sound, leveraging the fact that the KV cache is the interface between context and generation, allowing modality-specific processing to be decoupled from the backbone.
Experiments are conducted on a compact model (Qwen3-0.6B) with WavLM as the audio encoder. The paper evaluates ASR (LibriSpeech), Audio QA (Clotho), and Acoustic Scene Classification (CochlScene). The proposed method outperforms the "Encoder-only" (SLM-style) baseline significantly on non-ASR tasks and approaches the performance of a fine-tuned "Monolithic" model while using fewer active parameters during audio prefilling. Crucially, it preserves text-only performance (WikiText-2, HellaSwag, GSM8K) by construction, avoiding catastrophic forgetting. However, the evaluation is limited to a single small backbone, and the speedup is modest (157s vs 199s) in the tested regime.
The paper provides detailed architectural descriptions, including specific hyperparameters for the injector (kernel size, subsampling stride), training configurations (learning rates, batch size, steps), and specific techniques for stabilization (RMSNorm initialization, noisy RoPE parameters). The use of standard open-source components (WavLM, Qwen3) enhances reproducibility. However, code availability is not explicitly confirmed in the text, and the specific implementation of the KV injection into the inference engine is not detailed.
The primary limitation is the scale of the experiments; using a 0.6B LLM does not fully validate the scalability claims for larger models where the prefilling bottleneck is more severe. The speedup demonstrated is relatively small in the current setup, and the authors acknowledge that the CNN-centric injector may limit performance on complex acoustic tasks compared to attention-based injectors. The method relies on specific internal details of the backbone (e.g., Qwen3's RMSNorm), which may require adaptation for other LLM families.
This work offers a practical solution to the high computational cost of multimodal prefilling and the risk of catastrophic forgetting in LLM fine-tuning. By decoupling audio processing from the backbone, it enables the use of larger, more capable LLMs for audio tasks without proportional increases in inference latency for the audio portion. This could facilitate the deployment of ALMs in edge devices or real-time applications where prefilling latency is critical. The paper introduces a symbiotic architecture that decouples audio prefilling from the LLM backbone by directly generating the KV cache via a lightweight injector, thereby reducing computational cost and preventing catastrophic forgetting. The technical contribution is significant in addressing two major bottlenecks in audio-LLM development, though the experimental validation is limited to small-scale models, leaving the full scalability benefits to be verified in future work.
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.
Primary: Sungkyunkwan University
All Institutions: Sungkyunkwan University
The paper introduces a practical and effective method for reducing KV cache memory in full-duplex speech models by converting acoustic history to text during listening slack, demonstrating significant memory savings with preserved conversational quality.
The paper proposes a novel "acoustic-to-text KV compression" strategy for full-duplex speech models. The core idea is to exploit "listening-time slack"โthe computational gap between audio unit arrivalsโto run a lightweight transcription side channel (implemented via LoRA) that converts incoming speech into text tokens. These text tokens are retained in the KV cache while older acoustic KV states are evicted. The methodology is technically sound, leveraging the existing language model backbone rather than a separate ASR model. A significant methodological strength is the use of knowledge distillation from the frozen original model to preserve native listening/speaking behaviors (turn-taking, interruption) which would otherwise be disrupted by the new transcription objective. The training setup uses word-level forced alignments and a composite loss function including cross-entropy for ASR and KL divergence for behavior preservation.
The experiments are conducted on the MiniCPM-o 4.5 model. The evaluation covers three key areas: 1) Long-form speech understanding (LongSpeech benchmark), showing improved WER and QA accuracy compared to native streaming, and competitive performance with external ASR cascades. 2) Full-duplex interaction (Full-Duplex-Bench), demonstrating that the distillation component is critical for maintaining natural conversational dynamics (pause handling, turn-taking). 3) Efficiency, showing a 64.6% reduction in peak KV cache size with no real-time deadline misses on H100 GPUs. The ablation study effectively isolates the impact of the distillation loss, showing that without it, the model fails at turn-taking. The comparison with external ASR cascades is particularly relevant, showing that the proposed method achieves similar quality with fewer added parameters and integrated memory management.
The paper provides sufficient detail for reproduction, including the specific model (MiniCPM-o 4.5), training data (LibriSpeech train-clean), hyperparameters (LoRA rank 16, learning rate, batch size), and the specific eviction strategy (5-unit retention window). The use of standard benchmarks (LongSpeech, Full-Duplex-Bench) facilitates comparison. However, the code is not explicitly linked in the text provided, and the specific implementation details of the "listening-time slack" scheduling might require access to the source code for precise replication.
The method is evaluated primarily on a single model architecture (MiniCPM-o 4.5), so generalizability to other full-duplex models is not established. The transcription side channel relies on the model's ability to predict the start of the ASR segment; errors in this gating mechanism could lead to missed transcriptions. The method assumes a consistent "listening-time slack" of ~900ms, which may not hold for all hardware configurations or model variants. Additionally, the WER (12.9%) is higher than dedicated streaming ASR models (e.g., FastConformer at 10.3%), suggesting a trade-off between integration and pure transcription accuracy.
This work addresses a critical bottleneck in deploying full-duplex speech agents: memory consumption during long conversations. By converting acoustic history to text, it enables longer context windows without proportional memory increases. This is highly relevant for real-time voice assistants and interactive agents. The approach of using "slack" time for auxiliary tasks (transcription) is a generalizable principle for efficient inference in streaming multimodal models. The paper introduces a practical and effective method for reducing KV cache memory in full-duplex speech models by converting acoustic history to text during listening slack, demonstrating significant memory savings with preserved conversational quality.
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.
Primary: Microsoft
All Institutions: Microsoft, University of Michigan
The paper presents a robust adaptation of a personalized speech enhancement model for online audio-visual target-speaker extraction, demonstrating significant improvements in target recovery and listening quality over existing extractors on both synthetic and real-world meeting data. By leveraging the prior knowledge of a deployed enhancer and integrating visual cues through a gated conditioning mechanism, the method effectively addresses the selection challenge while maintaining high speech quality, supported by extensive objective and subjective evaluations that highlight its practical deployability and the trade-offs involved in such adaptation.
The paper proposes a pragmatic and effective adaptation of a deployed personalized speech enhancement model (PVQE) into an online audio-visual target-speaker extraction system (V+E). The core methodological contribution is the integration of visual cues (mouth motion via AV-HuBERT) into the existing speaker-conditioning path of the enhancer, rather than training a separator from scratch. The use of a gated residual connection to fuse enrollment embeddings with visual features is a sound architectural choice that allows the model to retain its enhancement capabilities while learning selection. The fine-tuning strategy, which involves removing lookahead operations and rearranging decoder filters for causal processing, is well-described and technically rigorous. The approach effectively leverages the prior knowledge of the enhancement model to solve the extraction problem, addressing the "selection" gap in personalized enhancement.
The experimental evaluation is extensive and high-quality. The authors test on both synthetic mixtures (LRS3, VoxCeleb2) and, crucially, recorded meeting corpora (AMI, MCoRec, MTM, UniTalk), which is a significant strength as it moves beyond the standard synthetic benchmark limitations. The inclusion of a large-scale human listening test (170 listeners, P.835 protocol) provides strong evidence for the perceptual quality claims. The results show consistent improvements over the baseline extractor (AVASE) in both objective metrics (SI-SNRi, WAcc) and subjective quality (MOS). The preservation and rejection tests further validate the system's deployability, showing it does not distort the target when no interference is present. The analysis of the UniTalk results, where background noise suppression degrades, is honest and provides valuable insight into the trade-offs of this adaptation strategy.
The paper provides sufficient detail for reproducibility, including model architecture dimensions, training data sources, and fine-tuning procedures. The use of standard public datasets (LRS3, VoxCeleb2, AMI, MCoRec) aids reproducibility, although the internal MTM dataset is not publicly available. The specific hyperparameters and initialization strategies (e.g., zero-initialization of the enrollment projection) are clearly stated. However, the code and model weights are not explicitly linked in the provided text, which is a minor drawback for immediate reproducibility.
The primary limitation is the degradation in background noise suppression on in-the-wild data (UniTalk) compared to the original enhancement model, indicating that the adaptation to extraction comes at a cost to general noise robustness. The system relies on the presence of synchronized video, which may not be available in all deployment scenarios. Additionally, the comparison with PVQE (the starting model) is somewhat confounded because PVQE is not designed for extraction, so the "quality retention" claim is relative to a model that fails at the primary task (selection) in many cases.
This work has significant practical impact for real-time communication systems (e.g., video conferencing). By demonstrating that a deployed enhancement model can be adapted for extraction with minimal architectural changes and high perceptual quality, it offers a viable path for improving user experience in noisy or multi-speaker environments. The focus on low-latency, online processing makes it highly relevant for industrial applications. The rigorous evaluation on recorded meetings sets a new standard for how such systems should be tested, encouraging the field to move beyond synthetic benchmarks. The paper presents a robust adaptation of a personalized speech enhancement model for online audio-visual target-speaker extraction, demonstrating significant improvements in target recovery and listening quality over existing extractors on both synthetic and real-world meeting data. By leveraging the prior knowledge of a deployed enhancer and integrating visual cues through a gated conditioning mechanism, the method effectively addresses the selection challenge while maintaining high speech quality, supported by extensive objective and subjective evaluations that highlight its practical deployability and the trade-offs involved in such adaptation.
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39~dB, compared with 2.49~dB for our strongest baseline. See the demo page at https://benschou.com/notesep.
Primary: Purdue University
All Institutions: Purdue University, Loyola University Chicago, University of Michigan
The paper introduces the first deep learning framework for score-informed note separation, achieving state-of-the-art performance by combining a dual-stream extraction network with a novel joint energy reallocation mechanism. By addressing the specific challenges of harmonic overlap and mixture consistency, NoteSep provides a robust tool for note-level audio manipulation, significantly advancing the field of music source separation beyond instrument-level stems.
The paper proposes NoteSep, a two-stage framework for score-informed note separation. The first stage, NoteGrab, utilizes a dual-stream TFC-TDF U-Net architecture to process harmonic and percussive components (via HPSS) separately, linked by bidirectional cross-attention. A key methodological contribution is the "selective harmonic gating," which conditionally suppresses lower-octave interference based on score alignment, addressing a specific physical challenge in polyphonic music. The second stage, Adaptive Set Ownership (ASO), is a novel joint separation module that reallocates mixture energy among concurrent note estimates to enforce consistency, using a learned gate and log-gain prediction. This approach effectively tackles the "double counting" problem inherent in independent note extraction.
The authors curate a substantial training set (SCNS-Train) and a disjoint evaluation set (SCNS-Eval) covering 16 instruments. The evaluation is rigorous, comparing against both traditional methods (Score-Informed NMF) and commercial software (Melodyne). The results show a significant improvement in SI-SDR (7.39 dB vs 2.49 dB for the strongest baseline). Furthermore, the paper evaluates the system on real-world ensemble recordings (PHENICX-Anechoic, Bach10) by aggregating note estimates into instrument stems, demonstrating competitive performance against specialized instrument separation models. The inclusion of an editing evaluation (SCNS-Edit) further validates the utility of the separated notes for downstream tasks.
The paper commits to releasing code, model weights, and datasets. The architectural details are described with sufficient specificity (e.g., TFC-TDF v3, FiLM conditioning, specific loss functions), and the demo page provides audio examples. The use of standard libraries (librosa) for preprocessing aids reproducibility.
The method relies heavily on accurate score alignment (pitch, onset, offset); errors in the input score will propagate to the separation quality. The computational cost is non-trivial, requiring K passes for K notes plus an ASO pass, which may limit real-time application on consumer hardware. The evaluation on real recordings is indirect (via instrument stem aggregation), so the fidelity of individual notes in complex, noisy real-world scenarios is not directly quantified with the same rigor as the synthetic benchmark.
This work enables precise, note-level audio editing, which has significant applications in music production, education (isolating specific notes for learning), and restoration. It bridges the gap between symbolic music information and audio signal processing, potentially facilitating more granular music information retrieval tasks. The paper introduces the first deep learning framework for score-informed note separation, achieving state-of-the-art performance by combining a dual-stream extraction network with a novel joint energy reallocation mechanism. By addressing the specific challenges of harmonic overlap and mixture consistency, NoteSep provides a robust tool for note-level audio manipulation, significantly advancing the field of music source separation beyond instrument-level stems.
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
Primary: University of Pisa
All Institutions: University of Pisa
The paper demonstrates that Part-of-Speech categories in LLMs are encoded as distributed but compact groups of SAE latents, with significant structural differences between open and closed classes. By combining probing, feature salience, and coverage analysis, the authors provide a detailed map of how morpho-syntactic information is organized in sparse latent spaces, offering valuable insights into the interpretability of modern language models.
The paper employs a rigorous three-step interpretability pipeline to analyze Part-of-Speech (PoS) encoding in Sparse Autoencoder (SAE) latents. First, it establishes recoverability using L1-regularized logistic regression probes on SAE activations from LLaMA-3-8B. Second, it moves beyond binary accuracy to localize information by analyzing feature salience (classifier coefficients) and coverage (the minimal number of latents required to cover 95% of instances for a given PoS tag). Third, it validates these localized groups on held-out data and a controlled synthetic dataset to test stability and additivity. The methodology is sound, moving logically from "is the information there?" to "where is it?" and "is it stable?". The use of a controlled dataset with minimal templates is a strong methodological choice to isolate lexical effects from syntactic ones.
The experiments are comprehensive, utilizing the GUM Treebank for natural data and a custom controlled dataset for validation. The results clearly demonstrate that PoS information is distributed across compact groups of latents rather than single monosemantic units. The distinction between Open-class (nouns, verbs) and Closed-class (determiners, conjunctions) PoS is well-supported, with closed classes showing more compact and stable latent groups. The ablation studies, including random label controls and layer-wise analysis, effectively rule out trivial explanations like lexical memorization or layer-specific artifacts. The finding that a small union of these latents (498 out of 131,072) supports strong multi-class classification is a significant quantitative result.
The authors provide high reproducibility by releasing code, data, and the specific SAE checkpoint used. The GitHub repository contains the pipeline for extracting activations and running the probes. The use of standard libraries (Scikit-learn, Sparsify) and public models (LLaMA-3-8B, EleutherAI SAE) ensures that other researchers can easily replicate the findings. The detailed description of the subword-to-token alignment strategy (leftmost subword anchoring) is crucial for reproducibility in this domain.
The study is limited to a single model (LLaMA-3-8B) and a single SAE variant, which may limit the generalizability of the findings to other architectures or sparsity regimes. The analysis is restricted to English, and it is unclear how these patterns would manifest in morphologically rich languages. The controlled dataset, while useful, is small and covers only a subset of syntactic structures. Additionally, the reliance on linear probes means that non-linear interactions between latents are not captured, potentially underestimating the complexity of the representation.
This work contributes to the broader field of mechanistic interpretability by providing a structured framework for analyzing how linguistic categories are represented in sparse latent spaces. It challenges the assumption of strict monosemanticity for linguistic features, suggesting instead a distributed but localized organization. This has implications for how interpretability tools are designed and how we understand the internal representations of LLMs. The findings could inform the development of more targeted interpretability methods that focus on groups of features rather than individual units. The paper demonstrates that Part-of-Speech categories in LLMs are encoded as distributed but compact groups of SAE latents, with significant structural differences between open and closed classes. By combining probing, feature salience, and coverage analysis, the authors provide a detailed map of how morpho-syntactic information is organized in sparse latent spaces, offering valuable insights into the interpretability of modern language models.
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
Primary: Xiamen University
All Institutions: Xiamen University
[One sentence main contribution]. The paper introduces EditVoice, a variable-length non-autoregressive zero-shot TTS model using Edit Flows and Complementary Prompt Sampling, achieving competitive performance in TTS and speech editing with high inference efficiency.
The paper proposes EditVoice, a non-autoregressive (NAR) zero-shot TTS model based on Edit Flows. The core methodological contribution is the application of Edit Flows to speech token sequences, enabling variable-length generation through insertions, deletions, and substitutions without pre-specifying sequence length. The authors introduce a speech-infilling training objective that unifies TTS and speech editing. A key heuristic contribution is Complementary Prompt Sampling (CPS), which leverages the complementary predictions of prefix and suffix prompt placements to improve generation quality, along with a local consistency rule to prevent artifacts. The method also exploits the model's ability to generalize to editing plausible (non-noise) token sequences for post-generation refinement.
Experiments are conducted on Seed-TTS Eval EN and LibriSpeech-PC for TTS, and RealEdit for speech editing. The model is trained on 10K hours of GigaSpeech. Results show competitive WER and speaker similarity compared to autoregressive baselines like CosyVoice2 and other NAR models. The 16-NFE variant achieves a low RTF of 0.0989. Ablations confirm the effectiveness of CPS and post-generation refinement. Subjective evaluations (NMOS, SMOS) are included.
The paper provides detailed architectural specifications (14-layer LLaMA-style Transformer, Conformer text encoder) and training hyperparameters. It specifies the use of frozen S3Tokenizer2 and CosyVoice2 acoustic decoder. However, the code repository is not explicitly linked in the text provided (only a demo page), which may limit immediate reproducibility compared to papers with open-source code.
The model relies on a specific tokenizer (S3Tokenizer2) and acoustic decoder (CosyVoice2), which may limit generalizability to other speech tokenization schemes. The CPS heuristic, while effective, is a manual rule-based approach that may not scale optimally to all conditions. The training data size (10K hours) is significantly smaller than some state-of-the-art systems (1M+ hours), though the paper argues for efficiency.
This work contributes to the efficiency of NAR TTS by removing the need for length prediction, a common bottleneck. The unification of TTS and editing in a single model is a step towards more flexible speech manipulation tools. The use of Edit Flows for speech is a novel application of this generative framework. [One sentence main contribution]. The paper introduces EditVoice, a variable-length non-autoregressive zero-shot TTS model using Edit Flows and Complementary Prompt Sampling, achieving competitive performance in TTS and speech editing with high inference efficiency.
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-improvement system for audio understanding. To our knowledge, it is the first to evolve the model, waveforms, questions, and difficulty in one closed loop. EvoAudio uses the current model's performance to set the focus and difficulty of the next training data. A library of audio tools then constructs questions whose answers follow from how the audio was made, providing verifiable supervision without new human annotation. Reinforcement learning updates the model, and validation decides whether it enters the next evolution round. Across 13 rounds, EvoAudio improves five models with different audio encoders and language backbones on MMSU, MMAU-Pro, and MMAR. It achieves the highest average for every backbone, raising overall performance by up to 6.3 points. The improvement unfolds over successive rounds, with each stronger model starting the next round.
Primary: The Chinese University of Hong Kong, Shenzhen
All Institutions: The Chinese University of Hong Kong, Shenzhen, Tsinghua University, Tencent Hunyuan, Amphion Technology Co., Ltd.
EvoAudio introduces a recursive self-improvement system that jointly evolves the model, waveforms, questions, and difficulty in a closed loop, achieving significant gains in audio understanding across multiple LALM backbones. The paper presents a rigorous methodology for generating verifiable synthetic data and adapting curricula based on model performance, demonstrating that recursive self-improvement is effective for enhancing perceptual skills in audio language models.
The paper proposes EvoAudio, a recursive self-improvement framework for Large Audio Language Models (LALMs). The core innovation is a closed-loop system where the model's performance on a fixed set of verifiable skills dictates the curriculum for the next training round. Unlike previous methods that use static synthetic data or rely on stronger teacher models for labeling (which introduces bias and error propagation), EvoAudio generates training data using a library of 24 audio tools that construct waveforms with known ground-truth properties (e.g., pitch, tempo, speaker count). This ensures "verifiable supervision" without human annotation. The system uses GRPO (Group Relative Policy Optimization) for reinforcement learning, where rewards are derived from the verifiable answers. A key methodological strength is the adaptive difficulty adjustment: the system monitors the "mixed rate" (variance in model responses) to ensure questions remain in the model's learning zone, avoiding saturation or frustration. The separation of the "verifier" (checking if the answer matches the construction) and the "acoustic check" (ensuring the rendered audio actually contains the intended cue) is a robust design choice that mitigates synthesis errors.
The experiments are extensive and rigorous. The authors evaluate five different LALM backbones (Qwen2.5-Omni, Audio Flamingo 3, Kimi-Audio, MiMo-Audio, MiniCPM-o) across three major benchmarks (MMSU, MMAU-Pro, MMAR). The results show consistent improvements over base models and strong baselines (Static-profile GRPO, Pooled GRPO). The ablation studies are particularly valuable, isolating the impact of adaptive difficulty, skill quotas, and acoustic verification. The finding that removing acoustic verification leads to "reward hacking" (where the model learns to exploit rendering errors) is a significant insight. The improvement of up to 6.3 points on average is substantial in the context of audio understanding, where progress is often incremental. The comparison against "Pooled GRPO" effectively demonstrates the necessity of the recursive, adaptive nature of the curriculum rather than just having a large pool of synthetic data.
The paper provides high-level details on the system architecture, the number of tools (24), question types (47), and training hyperparameters (learning rate, KL coefficient, number of rounds). However, the specific implementation of the "audio tools" library is not fully detailed in the text, and the code repository is not linked in the provided text (only a demo page). The reliance on specific TTS models (Qwen3-TTS) and datasets (LibriSpeech, FSD50K) is clear, but the exact prompts for the LLM proposer and the specific logic for the "acoustic checks" (e.g., how pitch is remeasured) would require the code for full reproduction. The demo page suggests some availability, but a full open-source release of the tool library would be necessary for high reproducibility.
The primary limitation is the scope of the tool library. The system can only improve on skills that can be synthesized and verified by the current 24 tools. Complex semantic or cultural reasoning (tested in MMAR) sees less improvement because these aspects are harder to construct synthetically with verifiable ground truth. The method is also computationally expensive, requiring 13 rounds of training and multiple rollouts per question. Additionally, the "verifiable" nature of the data means the model may not generalize well to open-ended audio questions that do not have a single correct answer derived from construction parameters.
This work has significant implications for the training of multimodal models. The concept of "self-evolving" curricula with verifiable rewards is applicable beyond audio to other modalities where ground truth can be constructed (e.g., code, math, visual geometry). It addresses the bottleneck of high-quality, diverse training data for perceptual tasks. By demonstrating that models can improve their own perceptual abilities without human labels, it offers a scalable path for developing more robust audio understanding systems. The insights into reward variance and curriculum adaptation are also valuable for the broader reinforcement learning community. EvoAudio introduces a recursive self-improvement system that jointly evolves the model, waveforms, questions, and difficulty in a closed loop, achieving significant gains in audio understanding across multiple LALM backbones. The paper presents a rigorous methodology for generating verifiable synthetic data and adapting curricula based on model performance, demonstrating that recursive self-improvement is effective for enhancing perceptual skills in audio language models.
Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited labels for their intent. We present a pipeline for synthesizing intent-labeled, two-channel conversational speech from relational event lists. An LLM authors each event's speaker, text, conversational act, and attachment to an earlier event without predicting absolute timestamps. Events are synthesized independently, aligned with their source text, and placed on a shared clock, so turn-taking landmarks are measured from the rendered signal while silence durations are specified or sampled from turn-taking distributions. The pipeline covers 42 phenomena across eight families in English and Mandarin, derives frame-level system actions from authored intent, and promotes diversity using small, diverse sets of prior examples and batch prompts that request alternatives with self-reported probabilities. Ablations show gains in each targeted diversity dimension. On a four-action label space for taking, holding, releasing, and not holding the conversational floor, a semantic voice-activity detector using only current and past audio reaches start-speaking and start-listening F1 scores of 0.819 and 0.802. When generating its own responses, the full-duplex speech model Moshi takes 0.85 of the reference turns after fine-tuning on the generated corpus, compared with 0.44 before fine-tuning. Its frame-level precision for predicting system-floor occupancy rises from 0.46 to 0.88. With reference context at each step, its frame-level floor F1 rises from 0.893 to 0.962. These results show that controlled synthesis can provide learnable and transferable supervision for full-duplex turn management.
Primary: Tencent Americas
All Institutions: Tencent Americas, School of Electronic Information Wuhan University, Columbia University
The paper presents a novel pipeline for synthesizing intent-labeled, full-duplex conversational speech by decoupling LLM-based semantic authoring from acoustic realization via forced alignment. This approach effectively addresses the supervision gap in existing corpora by providing precise, intent-aware labels for turn-taking events, leading to significant improvements in the performance of full-duplex speech models like Moshi in managing conversational floor occupancy.
The paper proposes a robust pipeline for synthesizing full-duplex conversational data by decoupling semantic authoring from acoustic realization. The core innovation is the "relational event representation," where an LLM defines conversational acts (turn, barge-in, backchannel) and their temporal relationships to previous events without predicting absolute timestamps. This is a significant methodological improvement over previous approaches that attempted to have LLMs predict precise timing, which is unreliable. The pipeline uses forced alignment to measure actual speech boundaries and inserts silence based on turn-taking distributions, ensuring that the acoustic timing is natural while the intent labels remain precise. The use of "diversity hubs" to prevent mode collapse in LLM generation is a clever application of negative memory. The separation of intent (semantic) from activity (acoustic) allows for the creation of frame-level labels that distinguish between a barge-in (which should stop the system) and a backchannel (which should not), a critical distinction for full-duplex models that voice activity detection (VAD) alone cannot provide.
The evaluation is strong and directly addresses the paper's claims. The authors train a semantic voice-activity detector and fine-tune the Moshi full-duplex speech model on the generated corpus. The results show significant improvements: Moshi's ability to take the correct turn increases from 0.44 to 0.85, and frame-level floor occupancy precision rises from 0.46 to 0.88. These metrics are highly relevant to the task of turn-taking management. The ablation studies on diversity mechanisms further validate the pipeline's components. The use of both English and Mandarin data demonstrates the pipeline's cross-lingual applicability.
The paper provides detailed descriptions of the pipeline stages, including the specific models used (DeepSeek-V4-Pro for authoring, Qwen3-TTS for synthesis, Qwen3-ForcedAligner for alignment). The mathematical formulation for timeline assembly is clear. However, the specific hyperparameters for the turn-taking distributions and the exact prompt templates are referenced in appendices that are not fully visible in the provided text, which slightly limits immediate reproducibility. The reliance on specific commercial/proprietary LLMs (DeepSeek, Qwen) may also pose accessibility challenges for some researchers, though these models are increasingly open-source.
The pipeline relies on the quality of the underlying TTS and LLM models; if the TTS produces unnatural prosody, the "naturalistic" claim is weakened. The evaluation is primarily on the Moshi model; generalization to other full-duplex architectures is not tested. The "intent" labels are derived from the LLM's authoring, so any errors in the LLM's understanding of conversational acts will propagate to the labels. The paper does not extensively discuss the computational cost of the LLM-based authoring and validation loop for large-scale corpus generation.
This work addresses a critical bottleneck in training full-duplex dialogue systems: the lack of high-quality, intent-labeled data. By providing a scalable pipeline to generate such data, the paper enables the development of more natural and responsive conversational AI. The focus on distinguishing between different types of overlap (barge-in vs. backchannel) is particularly important for user experience, as incorrect handling of these events leads to frustrating interactions. The methodology could be extended to other modalities or languages with appropriate configuration overlays. The paper presents a novel pipeline for synthesizing intent-labeled, full-duplex conversational speech by decoupling LLM-based semantic authoring from acoustic realization via forced alignment. This approach effectively addresses the supervision gap in existing corpora by providing precise, intent-aware labels for turn-taking events, leading to significant improvements in the performance of full-duplex speech models like Moshi in managing conversational floor occupancy.
Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.
Primary: People Make Things
All Institutions: People Make Things
The paper introduces TACT, a novel intent-conditioned benchmark and scoring framework for full-duplex spoken dialogue models that replaces binary turn-taking metrics with a strictly proper, continuous ranked probability score. By conditioning evaluation on the speaker's latent intent and behavioral memory, the authors demonstrate that current state-of-the-art models exhibit significant rigidity and over-eagerness, failing to adapt to conversational context as humans do, thereby providing a more rigorous and human-aligned standard for assessing dialogue timing.
The paper proposes a theoretically grounded shift from binary, fixed-window turn-taking metrics to a continuous, intent-conditioned scoring framework (TACT). The core methodological contribution is the use of a threshold-weighted Continuous Ranked Probability Score (twCRPS) where the weights are derived from intent-specific human floor-transfer-offset (FTO) distributions. This is a sophisticated approach that correctly identifies that "silence" and "overlap" are not inherently good or bad, but depend on the speaker's latent intent (e.g., rhetorical vs. answer-seeking). The authors provide rigorous proofs for boundedness, strict propriety, and the reduction of their metric to existing binary metrics in degenerate cases, which is a strong theoretical foundation. The integration of a "memory profile" for each speaker to condition the intent posterior is a novel and practical addition that addresses the variability in human conversational styles.
The experimental setup is robust, utilizing a large-scale benchmark (9,728 episodes, 73.2 hours) constructed from five diverse public corpora. The evaluation of eleven frontier systems (including open-source and proprietary APIs) provides a comprehensive landscape of current full-duplex model capabilities. The finding that models exhibit an "over-eagerness pathology" and fail to adapt to speaker-specific memory profiles is a significant empirical contribution. The validity study, showing a Spearman correlation of 0.81 with human judgments compared to 0.46 for binary metrics, strongly supports the utility of the proposed benchmark.
The paper claims that data, kernels, and prompts are available in the Supplementary Material, which is a positive sign for reproducibility. However, the reliance on proprietary APIs (GPT-4o, Gemini) for some evaluations limits full reproducibility for those specific data points. The detailed description of the annotation protocol and the mathematical formulation of the scoring rules enhances the potential for independent verification.
The primary limitation is the discretization of intent into six classes, which may oversimplify the continuous nature of human intent. The reliance on an LLM judge for fusing the intent posterior introduces potential biases or failure modes shared with the underlying LLM. Additionally, the kernels are fitted on English corpora, limiting the immediate generalizability to other languages without refitting. The "memory profile" mechanism, while effective in ablation, adds complexity to the inference pipeline for evaluated systems.
This work has high potential impact on the development of natural spoken dialogue systems. By providing a more nuanced and human-aligned evaluation metric, it can guide the training of models that are not just fast, but contextually appropriate. The identification of the "over-eagerness" problem highlights a critical gap in current model architectures that future research must address. The framework could be extended to other interactive AI domains where timing and intent are crucial. The paper introduces TACT, a novel intent-conditioned benchmark and scoring framework for full-duplex spoken dialogue models that replaces binary turn-taking metrics with a strictly proper, continuous ranked probability score. By conditioning evaluation on the speaker's latent intent and behavioral memory, the authors demonstrate that current state-of-the-art models exhibit significant rigidity and over-eagerness, failing to adapt to conversational context as humans do, thereby providing a more rigorous and human-aligned standard for assessing dialogue timing.
End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipulation. We formalize imperceptible attacks on full-duplex agents as optimization over additive perturbations confined beneath the psychoacoustic masking threshold of the carrier speech, under three goals: targeted semantic hijacking, response suppression, and policy jailbreaking. Against an undefended Moshi-style agent, white-box attacks succeed in up to 91.7% of trials. We then introduce psychoacoustically aligned latent smoothing (PALS), which injects anisotropic Gaussian noise shaped by local codebook covariance at the residual-vector-quantized latent interface, with input noise shaped by the masking threshold constraining the attacker and trained by a Kullback--Leibler consistency objective. Deployed with no inference-time cost, PALS reduces hijack to 8.3%, mute to 11.2%, and jailbreak to 9.1% at clean quality within 2.3%. A Monte Carlo-smoothed variant certifies an ellipsoidal latent radius up to 0.616, a guaranteed floor that the empirical robustness far exceeds.
Primary: People Make Things
All Institutions: People Make Things
The paper introduces a principled, psychoacoustically aligned latent smoothing defense for full-duplex speech agents, significantly reducing adversarial attack success rates while maintaining utility, though its certified guarantees are conservative and its institutional provenance is obscure.
The paper proposes Psychoacoustically Aligned Latent Smoothing (PALS), a defense mechanism for full-duplex speech-to-speech (S2S) models. The core innovation lies in applying randomized smoothing at the latent interface of the Residual Vector Quantization (RVQ) codec, using anisotropic Gaussian noise shaped by local codebook covariance. This is coupled with an input-space augmentation that matches the psychoacoustic masking threshold, derived via a maximum-entropy argument. The training objective incorporates a KL consistency term (similar to TRADES) to ensure robustness without requiring adversarial examples during training. The theoretical contribution includes a certified robustness radius for the smoothed operator, though the authors correctly note that the certificate is a lower bound and the empirical robustness of the deployed model (PALS-Train) is the primary metric of interest. The methodology is sound, leveraging the geometric properties of the quantizer to define the noise distribution, which is a clever and non-trivial design choice.
The experiments are comprehensive, evaluating three distinct attack types (semantic hijacking, response suppression/muting, and jailbreaking) against a Moshi-style agent. The paper reports significant reductions in attack success rates (e.g., hijacking drops from 91.7% to 8.3%) while maintaining clean quality metrics (UTMOS, WER, Turn-Taking F1). Ablations clearly demonstrate the contribution of each component (anisotropic covariance, input-space dual, KL consistency). The inclusion of over-the-air (OTA) tests with reverberation and babble noise adds practical relevance. However, the evaluation relies on a specific "Moshi-style" agent architecture; generalizability to other S2S architectures (e.g., those without RVQ or with different codec designs) is not tested.
The paper provides detailed hyperparameters, training objectives, and attack configurations. It mentions a released `certify.py` harness. However, the code and model checkpoints are not explicitly linked in the provided text (no GitHub URL found), and the "Moshi-style" agent is described rather than provided as a specific open-source artifact. The reliance on specific proprietary or large-scale datasets (CANDOR, Seamless Interaction) may limit immediate reproducibility for smaller groups, though the methodology is clearly specified.
The certified robustness radius is relatively small (0.616 latent units, corresponding to -26 dB input budget), which is far below the attack budgets tested empirically. The authors acknowledge this, positioning the certificate as a "floor" rather than a guarantee against strong attacks. The defense is specific to RVQ-based codecs; its applicability to continuous latent spaces or other quantization schemes is unclear. The "People Make Things" affiliation is obscure, raising questions about the scale of resources behind the work, though the technical depth suggests significant effort.
As full-duplex voice agents become prevalent, securing them against imperceptible audio attacks is critical. This work provides a strong baseline for latent-space defenses in speech, potentially influencing future designs of secure speech interfaces. The psychoacoustic alignment of the defense with the attack constraint is a principled approach that could be adapted to other sensory modalities. The paper introduces a principled, psychoacoustically aligned latent smoothing defense for full-duplex speech agents, significantly reducing adversarial attack success rates while maintaining utility, though its certified guarantees are conservative and its institutional provenance is obscure.
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage $ฮ$ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks $ฮ$ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
Primary: People Make Things
All Institutions: People Make Things
The paper introduces a geometric framework linking subspace entanglement to demographic disparity in speech encoders, validated by a large-scale synthetic audit grid and a lightweight repair adapter. It makes a significant contribution to the theory of fairness in representation learning by providing provable bounds on disparity based on measurable geometric properties, although its reliance on synthetic data for primary validation limits the strength of its causal claims.
The paper proposes a rigorous geometric framework for auditing demographic fairness in speech encoders. The core contribution is the definition of "axis-fidelity" and "subspace leakage" (principal angles between semantic and voice subspaces). The theoretical contribution is a proposition linking aggregate leakage to average probe disparity and a corollary linking peak leakage to worst-case disparity. This is a strong methodological advance in moving fairness from black-box output metrics to white-box representation geometry. The proposed repair mechanism, ORCA, uses a residual adapter with orthogonal contrastive heads, which is a sound architectural choice for disentangling factors without destroying utility.
The experimental setup is extensive, utilizing a custom "TRIAD" grid of 28,800 synthetic utterances to control for confounding variables (content, affect, demographics). The evaluation covers 10 open-weights encoders and 2 closed-source models. The correlation between measured leakage and disparity (r=0.93) is a compelling empirical validation of the theory. The ablation studies effectively isolate the contribution of the orthogonality penalty. However, the reliance on a single TTS system for the grid introduces a significant validity threat, which the authors acknowledge but only partially mitigate with real-speech validation checks.
The paper claims that the grid, estimators, and code are available in the Supplementary Material. However, no specific URLs or repository links are provided in the text. The detailed hyperparameters (AdamW, lr=1e-3, 20k steps) and dataset construction details (Gemini 3.1 Flash TTS) are provided, which aids reproducibility if the supplementary materials are accessible. The use of a specific, potentially proprietary or rapidly changing TTS model (Gemini 3.1) may limit long-term reproducibility.
The primary limitation is the reliance on synthetic data for the main audit grid. While the authors perform validation on real speech (CANDOR, SSSD, etc.), the causal attribution of disparity to demographic attributes is weaker in synthetic data due to potential TTS artifacts. The theory is limited to linear probes; non-linear disparities are not bounded by the proposed leakage metrics. The institution "People Make Things" is obscure, raising questions about the resources and peer-review rigor compared to major academic or industrial labs.
This work provides a principled tool for diagnosing bias in speech models, which is critical as these models are deployed in high-stakes applications (healthcare, legal, customer service). The geometric perspective offers a new lens for understanding why certain models are biased, moving beyond correlation to structural entanglement. The ORCA adapter offers a practical, lightweight fix that can be applied to existing frozen models. The paper introduces a geometric framework linking subspace entanglement to demographic disparity in speech encoders, validated by a large-scale synthetic audit grid and a lightweight repair adapter. It makes a significant contribution to the theory of fairness in representation learning by providing provable bounds on disparity based on measurable geometric properties, although its reliance on synthetic data for primary validation limits the strength of its causal claims.