Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-improvement system for audio understanding. To our knowledge, it is the first to evolve the model, waveforms, questions, and difficulty in one closed loop. EvoAudio uses the current model's performance to set the focus and difficulty of the next training data. A library of audio tools then constructs questions whose answers follow from how the audio was made, providing verifiable supervision without new human annotation. Reinforcement learning updates the model, and validation decides whether it enters the next evolution round. Across 13 rounds, EvoAudio improves five models with different audio encoders and language backbones on MMSU, MMAU-Pro, and MMAR. It achieves the highest average for every backbone, raising overall performance by up to 6.3 points. The improvement unfolds over successive rounds, with each stronger model starting the next round.
Primary: The Chinese University of Hong Kong, Shenzhen
All Institutions: The Chinese University of Hong Kong, Shenzhen, Tsinghua University, Tencent Hunyuan, Amphion Technology Co., Ltd.
EvoAudio introduces a recursive self-improvement system that jointly evolves the model, waveforms, questions, and difficulty in a closed loop, achieving significant gains in audio understanding across multiple LALM backbones. The paper presents a rigorous methodology for generating verifiable synthetic data and adapting curricula based on model performance, demonstrating that recursive self-improvement is effective for enhancing perceptual skills in audio language models.
The paper proposes EvoAudio, a recursive self-improvement framework for Large Audio Language Models (LALMs). The core innovation is a closed-loop system where the model's performance on a fixed set of verifiable skills dictates the curriculum for the next training round. Unlike previous methods that use static synthetic data or rely on stronger teacher models for labeling (which introduces bias and error propagation), EvoAudio generates training data using a library of 24 audio tools that construct waveforms with known ground-truth properties (e.g., pitch, tempo, speaker count). This ensures "verifiable supervision" without human annotation. The system uses GRPO (Group Relative Policy Optimization) for reinforcement learning, where rewards are derived from the verifiable answers. A key methodological strength is the adaptive difficulty adjustment: the system monitors the "mixed rate" (variance in model responses) to ensure questions remain in the model's learning zone, avoiding saturation or frustration. The separation of the "verifier" (checking if the answer matches the construction) and the "acoustic check" (ensuring the rendered audio actually contains the intended cue) is a robust design choice that mitigates synthesis errors.
The experiments are extensive and rigorous. The authors evaluate five different LALM backbones (Qwen2.5-Omni, Audio Flamingo 3, Kimi-Audio, MiMo-Audio, MiniCPM-o) across three major benchmarks (MMSU, MMAU-Pro, MMAR). The results show consistent improvements over base models and strong baselines (Static-profile GRPO, Pooled GRPO). The ablation studies are particularly valuable, isolating the impact of adaptive difficulty, skill quotas, and acoustic verification. The finding that removing acoustic verification leads to "reward hacking" (where the model learns to exploit rendering errors) is a significant insight. The improvement of up to 6.3 points on average is substantial in the context of audio understanding, where progress is often incremental. The comparison against "Pooled GRPO" effectively demonstrates the necessity of the recursive, adaptive nature of the curriculum rather than just having a large pool of synthetic data.
The paper provides high-level details on the system architecture, the number of tools (24), question types (47), and training hyperparameters (learning rate, KL coefficient, number of rounds). However, the specific implementation of the "audio tools" library is not fully detailed in the text, and the code repository is not linked in the provided text (only a demo page). The reliance on specific TTS models (Qwen3-TTS) and datasets (LibriSpeech, FSD50K) is clear, but the exact prompts for the LLM proposer and the specific logic for the "acoustic checks" (e.g., how pitch is remeasured) would require the code for full reproduction. The demo page suggests some availability, but a full open-source release of the tool library would be necessary for high reproducibility.
The primary limitation is the scope of the tool library. The system can only improve on skills that can be synthesized and verified by the current 24 tools. Complex semantic or cultural reasoning (tested in MMAR) sees less improvement because these aspects are harder to construct synthetically with verifiable ground truth. The method is also computationally expensive, requiring 13 rounds of training and multiple rollouts per question. Additionally, the "verifiable" nature of the data means the model may not generalize well to open-ended audio questions that do not have a single correct answer derived from construction parameters.
This work has significant implications for the training of multimodal models. The concept of "self-evolving" curricula with verifiable rewards is applicable beyond audio to other modalities where ground truth can be constructed (e.g., code, math, visual geometry). It addresses the bottleneck of high-quality, diverse training data for perceptual tasks. By demonstrating that models can improve their own perceptual abilities without human labels, it offers a scalable path for developing more robust audio understanding systems. The insights into reward variance and curriculum adaptation are also valuable for the broader reinforcement learning community. EvoAudio introduces a recursive self-improvement system that jointly evolves the model, waveforms, questions, and difficulty in a closed loop, achieving significant gains in audio understanding across multiple LALM backbones. The paper presents a rigorous methodology for generating verifiable synthetic data and adapting curricula based on model performance, demonstrating that recursive self-improvement is effective for enhancing perceptual skills in audio language models.
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Distillation (M$^{2}$-OPD) to transfer language capabilities and develop native audio skills. Act uses self-evolving executable environments and multi-granularity rollouts for Group Relative Policy Optimization (GRPO), teaching the model to use tools, interpret feedback, and complete tasks. Speak and Coordinate aligns how, when, and whether the assistant speaks or acts. We evaluate audio reasoning, multilingual understanding, tool use, conversational behavior, full-duplex interaction, and safety. Compared with Qwen-Audio-3.0-Realtime, 3.1 raises overall task success from 78.4% to 82.0% on our half-duplex speech-to-text adaptation of $ฯ$-Voice. On speech-to-speech Full-Duplex-Bench v1.5, the response rate to background speech falls from 73.0% to 13.0%. We also present a separate Voice Harness prototype, using Qwen-Audio-3.0-Realtime as its foreground, that extends spoken interaction to persistent tasks through foreground--background coordination and memory.
Primary: Alibaba Group
All Institutions: Alibaba Token Foundry, Alibaba Group
The paper presents a comprehensive framework for reliable agentic voice interaction, combining novel on-policy distillation techniques with self-evolving executable environments for tool use. It demonstrates significant improvements in task success, multilingual understanding, and conversational etiquette, particularly in handling background speech and full-duplex interactions, marking a substantial step forward in the development of practical, safety-conscious voice agents.
The paper proposes a comprehensive framework for real-time voice agents, structured around three layers: Think (foundation post-training), Act (agentic tool use), and Speak/Coordinate (conversational policy). The "Think" layer introduces a sophisticated post-training pipeline combining Core-Cocktail SFT with a novel "M$^2$-OPD" (Multimodality and Multi-Teacher On-Policy Distillation) strategy. This involves using a Text Teacher and a frozen Audio Reference to supervise student-generated trajectories, effectively transferring text-based reasoning capabilities to native audio models while preserving audio-specific nuances. The "Act" layer is particularly strong, introducing self-evolving executable environments for Group Relative Policy Optimization (GRPO). By using code agents to build, validate, and evolve tasks based on model performance (difficulty gating), the system learns robust tool use, state verification, and grounded progress communication. The "Speak/Coordinate" layer formalizes the decision-making process for when to speak, act, or remain silent, addressing critical issues in full-duplex interaction like background speech handling and turn-taking.
The evaluation is extensive, covering intelligence (audio reasoning, multilingual ASR, long-context), action (tool use, retrieval), interaction (persona, empathy, full-duplex behavior), and safety. Key results include a significant improvement in task success on the $\tau$-Voice benchmark (78.4% to 82.0%) and a dramatic reduction in response rate to background speech on Full-Duplex-Bench (73.0% to 13.0%), indicating much better conversational etiquette. The model also shows strong gains in multilingual audio understanding (BBA) and safety metrics. The inclusion of a "Voice Harness" prototype for persistent tasks adds depth to the system-level contribution.
The paper provides detailed descriptions of the training pipeline, data scales (approx. 1M hours), and evaluation protocols. However, as a technical report from a major industry lab, specific hyperparameters, exact dataset compositions, and code are not fully released, limiting independent reproduction. The reliance on in-house benchmarks (Long-AMC, WebSearch1K, VoiceChat) also restricts external verification.
The primary limitation is the lack of open-source code and data, making it difficult for the community to replicate the results. The evaluation relies heavily on in-house benchmarks and automatic judges (e.g., GPT-4o-mini, Qwen-Plus), which may introduce bias. The "Voice Harness" is presented as a prototype and is not fully integrated into the core model evaluation. Additionally, the paper focuses on a specific adaptation of $\tau$-Voice (half-duplex S2T) rather than the official full-duplex S2S protocol, which may limit direct comparability with other full-duplex systems.
This work significantly advances the state-of-the-art in real-time voice assistants by integrating robust agentic capabilities with natural conversational policies. The M$^2$-OPD technique and self-evolving environment for GRPO are valuable contributions to the broader field of multimodal LLM post-training. The focus on safety and reliable interaction (e.g., handling background speech, refusing unsupported requests) is crucial for the deployment of such systems in real-world scenarios. The paper presents a comprehensive framework for reliable agentic voice interaction, combining novel on-policy distillation techniques with self-evolving executable environments for tool use. It demonstrates significant improvements in task success, multilingual understanding, and conversational etiquette, particularly in handling background speech and full-duplex interactions, marking a substantial step forward in the development of practical, safety-conscious voice agents.
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
Primary: KAIST
All Institutions: KAIST, VGG, University of Oxford
The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
The paper employs a rigorous mechanistic interpretability framework, combining Representational Similarity Analysis (RSA) and Causal Mediation Analysis (CMA) to dissect the internal workings of Audio-Visual LLMs. The authors successfully identify a three-stage symbolic binding mechanism (Anchor ID Retrieval, Target ID Selection, Feature Retrieval) that relies on modality-specific symbolic variables (Temporal IDs for audio, Position IDs for vision). This is a sophisticated approach that moves beyond black-box evaluation to understand *how* models process cross-modal information. The proposed intervention, using an off-the-shelf Active Speaker Detection (ASD) model to overlay bounding boxes, is a clever, training-free (and lightly fine-tuned) solution that directly targets the identified bottleneck in audio-visual alignment.
The experiments are extensive and well-structured. The authors validate their mechanistic findings across four different AVLLM architectures (video-SALMONN2+, Qwen2.5-Omni, MiniCPM-o-4.5) using both synthetic toy datasets and real-world benchmarks (SocialOmni, AVSpeaker, DailyOmni, etc.). The ablation studies effectively isolate the contribution of the ASD prompting, showing that it outperforms other training-free decoding methods and standard fine-tuning. The generalization to broader audio-visual benchmarks (DAVE, OmniBench, WorldSense) further strengthens the claim of the method's utility.
The paper provides detailed implementation specifics, including LoRA ranks, training steps, and dataset construction methods. The use of standard off-the-shelf models and clear descriptions of the prompting strategy enhances reproducibility. However, the specific synthetic video generation pipeline and the exact ASD model used (referenced as [CITATION]) would need to be clearly specified in the final publication for full reproducibility.
The primary limitation is the reliance on a synthetic toy dataset for the core mechanistic analysis, which may not fully capture the complexity of real-world multi-speaker scenarios, although the authors do validate on real-world data. Additionally, the method depends on the accuracy of the external ASD model; if the ASD model fails, the prompting strategy may degrade performance. The fine-tuning, while lightweight, still requires access to the model's weights and a GPU, which may not be feasible for all users.
This work has significant implications for the development of more robust multimodal AI systems. By identifying specific failure modes in cross-modal binding, it provides actionable insights for improving model architectures and training strategies. The proposed ASD prompting method is a practical, low-cost solution that can be easily integrated into existing AVLLM pipelines, potentially improving performance in applications like video conferencing, accessibility tools, and video understanding. The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.
Primary: Sapienza University of Rome
All Institutions: Sapienza University of Rome, Moises Systems, Inc., Paradigma
SAGE introduces a compact, spectrogram-native VAE with semantic distillation that achieves state-of-the-art reconstruction and latent semantics at low inference cost. The paper demonstrates that adapting vision transformers (SwinV2) to audio spectrograms, combined with careful curriculum learning for semantic alignment, can outperform larger, waveform-based models in both fidelity and utility for downstream generative tasks.
The paper proposes SAGE, a 105M parameter Variational Autoencoder (VAE) for audio that operates on the complex Short-Time Fourier Transform (STFT) rather than raw waveforms or magnitude spectrograms. The core architectural innovation is the adaptation of the SwinV2 vision transformer backbone to audio, utilizing rectangular patches and attention windows to handle the time-frequency structure of spectrograms. A key methodological contribution is the "semantic distillation" strategy, where the latent space is aligned with embeddings from a frozen CLAP (Contrastive Language-Audio Pretraining) model. This alignment is introduced via a delayed schedule (detached warm-up) to ensure reconstruction quality is established before semantic constraints are applied, preventing the semantic loss from degrading fidelity. The training process is divided into two phases: pretraining with a full adversarial objective (using a WavTokenizer discriminator) and a fine-tuning phase where the encoder is frozen and the decoder is refined. The use of a sum-and-difference loss for stereo imaging is also a notable technical detail, addressing the common issue of stereo collapse in autoencoders.
The experimental evaluation is rigorous and comprehensive. The authors compare SAGE against five strong baselines (Stable Audio Open, SAME-L, SAME-S, CoDiCodec, Music2Latent) across five different held-out datasets, including in-domain (FMA) and out-of-domain (MoisesDB, MusicCaps, Song Describer) sets. Metrics cover perceptual (CLAP similarity), distributional (Frรฉchet Audio Distance in MERT, PANN, and CLAP spaces), and sample-exact (SI-SDR, STFT distance) categories. Crucially, the paper includes a MUSHRA listening test with 21 raters, showing SAGE is statistically indistinguishable from the much larger SAME-L model. The semantic evaluation is extensive, using 19 probing tasks (genre, artist, instrument, etc.) in the MAEB format, where SAGE achieves state-of-the-art results on all tasks. The ablation studies on the semantic distillation schedule and adversarial components provide strong evidence for the design choices.
The paper offers high reproducibility. The authors provide a link to the GitHub repository containing code, weights, and the evaluation harness. The training data consists of publicly available corpora (FMA, MTG-Jamendo, M4Singer), and the paper details the specific splits and preprocessing steps. The hyperparameters for both training phases are listed in the appendix. The use of standard libraries and the release of the evaluation harness allows other researchers to easily verify the results and compare their own models.
The primary limitation is the reliance on the CLAP model for semantic alignment; if the CLAP embeddings are biased or limited in their musical understanding, SAGE's latent space will inherit these limitations. The model is trained on music, so its performance on non-music audio (speech, sound effects) is not evaluated and likely poor. The inference cost, while competitive, is still higher than simple convolutional codecs, which may be a barrier for real-time mobile applications. The paper focuses on music, so generalization to other audio domains is an open question.
SAGE provides a lightweight, high-fidelity, and semantically rich audio encoder that can serve as a drop-in replacement for existing autoencoders in latent diffusion pipelines for music generation. Its semantic structure could enable better control and conditioning in generative models. The efficient inference cost makes it suitable for interactive applications. The comprehensive benchmarking harness contributes to the standardization of audio autoencoder evaluation. SAGE introduces a compact, spectrogram-native VAE with semantic distillation that achieves state-of-the-art reconstruction and latent semantics at low inference cost. The paper demonstrates that adapting vision transformers (SwinV2) to audio spectrograms, combined with careful curriculum learning for semantic alignment, can outperform larger, waveform-based models in both fidelity and utility for downstream generative tasks.
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps short accent exemplars into a conditioning space using an exemplar encoder supervised through learned accent prototypes, requiring neither accent labels at inference time nor post-synthesis waveform conversion. On seen accents, increasing accent guidance improves accent-probe accuracy from 6.5% to 19.6%. More importantly, a single DEFINE model generalizes accent control beyond its training accent set: on seen and out-of-domain accents, though not on held-out accents, it matches the accent transfer performance of a two-model TTS-voice-conversion cascade while achieving higher speaker similarity and comparable predicted speech quality. These results demonstrate that speaker identity and accent can be independently controlled from audio exemplars within a single zero-shot TTS model, including for accents unseen during training.
Primary: Ca' Foscari University of Venice
All Institutions: Ca' Foscari University of Venice, Kandinsky Lab, Sleeping AI, Singapore University of Technology and Design
The paper introduces DEFINE, a framework that decouples speaker identity and accent in zero-shot TTS by using separate audio exemplars and a prototype-anchored encoder, achieving comparable accent transfer to a two-model cascade while better preserving speaker identity. The technical contribution is significant, with a novel supervision mechanism for accent embeddings and a continuous inference-time control for accent strength. The experiments are thorough, including objective metrics and a small-scale listening test, demonstrating the effectiveness of the approach. The work is well-positioned to influence future research on disentangled speech synthesis and fine-grained control in TTS systems.
The paper proposes DEFINE, a framework for disentangling speaker identity and accent in zero-shot TTS. The core technical contribution is the "prototype anchoring" mechanism, which uses a learned table of accent prototypes to supervise an exemplar encoder (based on frozen XLS-R features). This addresses the weak supervision signal for accent representations in flow-matching objectives. The method injects the accent embedding into the F5-TTS backbone via additive shifts to timestep and text embeddings, scaled by RMS to ensure proper magnitude. A key feature is the inference-time guidance weight $w$ that allows continuous control over accent strength without retraining. The use of LoRA for parameter-efficient adaptation is standard but effective. The methodology is sound, though the reliance on a specific backbone (F5-TTS) and the specific injection point (adaLN modulation via embedding shifts) limits generalizability to other architectures.
The experiments are comprehensive, evaluating seen, held-out, and out-of-domain accents. The use of a 15-way logistic regression probe on XLS-R features for accent accuracy is a reasonable objective metric, though it is a proxy for perceptual accent. The comparison against a Seed-VC cascade is strong, showing that DEFINE matches accent transfer performance while improving speaker similarity. The listening test with 11 listeners is a positive addition, confirming that the accent changes are perceptually real and that the guidance weight controls the perceived accent strength. However, the sample size for the listening test (11 listeners, 10 trials) is relatively small, and the statistical significance, while present, is based on a limited number of votes. The WER and UTMOS metrics are standard and reported, showing that intelligibility and quality are maintained.
The paper provides a GitHub link, which is a strong indicator of reproducibility. The training data sources (Common Voice, in-house studio recordings) are described, but the in-house data is not publicly available, which limits full reproducibility. The hyperparameters for LoRA, the prototype table size, and the guidance weight are specified. The use of frozen XLS-R and specific layer selection (layer 15) is detailed. Overall, reproducibility is good, assuming access to the code and the ability to recreate the in-house speaker prompts or use a similar public dataset.
The main limitation is the reliance on a specific TTS backbone (F5-TTS), which may not generalize to other architectures without significant modification. The accent control is limited to the accents seen during training or those that can be represented by the exemplar encoder; truly novel accents not in the training distribution may not be well-captured. The listening test is small-scale, and the objective accent probe may not fully capture perceptual accent nuances. The method requires separate exemplar clips for accent, which adds a data requirement at inference time.
The work has significant implications for personalized TTS systems, allowing users to control not just the voice but also the accent of the synthesized speech. This could be useful for language learning, accessibility, and creative applications. The decoupling of identity and accent is a step towards more fine-grained control over speech synthesis, which is a key goal in the field. The method's ability to generalize to out-of-domain accents is particularly promising for real-world applications where the target accent may not be in the training set. The paper introduces DEFINE, a framework that decouples speaker identity and accent in zero-shot TTS by using separate audio exemplars and a prototype-anchored encoder, achieving comparable accent transfer to a two-model cascade while better preserving speaker identity. The technical contribution is significant, with a novel supervision mechanism for accent embeddings and a continuous inference-time control for accent strength. The experiments are thorough, including objective metrics and a small-scale listening test, demonstrating the effectiveness of the approach. The work is well-positioned to influence future research on disentangled speech synthesis and fine-grained control in TTS systems.
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.
Primary: Unknown
All Institutions: Unknown
The paper introduces VoxMem, a comprehensive benchmark for evaluating multi-session spoken conversational memory across acoustic evidence types and memory operations. It provides a principled taxonomy and rigorous construction pipeline, revealing significant gaps in current LALMs' ability to retain non-lexical acoustic information over long contexts.
The paper proposes a principled two-dimensional taxonomy for spoken conversational memory, crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal evolution tracking, and answer refusal). This is a significant methodological advance over prior benchmarks that treated memory as a single-session, lexical-only problem. The construction pipeline is rigorous, utilizing a three-stage process (planning, dialogue writing, speech synthesis) with specific controls to ensure "acoustic necessity" (i.e., the answer cannot be derived from the transcript alone). The use of Higgs-TTS-3 and VCTK voices, along with mixed environmental sounds from ESC-50, provides a controlled yet realistic audio environment. The inclusion of "haystack" and "filler" sessions to create multi-session histories of varying lengths (8K-64K tokens) is a strong design choice that isolates the effect of context length from question difficulty.
The evaluation covers 15 Large Audio Language Models (LALMs), including both open-weight and proprietary models. The results reveal a clear hierarchy of difficulty: models perform significantly better on speech semantics than on audio-native cues (speaker, paralinguistic, environmental). The finding that no model exceeds 40% accuracy at the 32K token budget highlights a critical gap in current LALM capabilities. The error analysis is particularly valuable, distinguishing between binding failures (wrong speaker) and retention failures (lost cue), which provides actionable insights for future model development. The controlled scaling analysis shows that performance degrades as history length increases, with non-lexical information degrading faster than lexical information.
The paper provides detailed descriptions of the construction pipeline, quality control checks, and evaluation protocol. The use of specific TTS systems, voice datasets, and sound event datasets enhances reproducibility. The release of the benchmark instances, question templates, and judge prompts (mentioned in the text) further supports reproducibility. However, the reliance on proprietary LLMs (Gemini-3.7-Flash, GPT-5.6-Luna) for generation and evaluation introduces a dependency on external services, which may limit full reproducibility for all researchers.
The benchmark relies on synthetic speech generated by TTS systems, which may not fully capture the variability and noise of real-world human speech. The "acoustic necessity" check, while rigorous, uses a specific LLM (Gemini-3.7-Flash) for validation, which could introduce bias if that model has specific strengths or weaknesses in audio understanding. The paper does not extensively discuss the computational cost of evaluating models on 64K token audio histories, which may be prohibitive for some researchers. Additionally, the focus on English speech (implied by the datasets used) limits the generalizability of the findings to other languages.
VoxMem addresses a fundamental challenge in the development of long-term conversational AI systems. By providing a standardized benchmark for multi-session, audio-native memory, it enables fair comparison of LALMs and guides future research toward improving non-lexical acoustic memory. The taxonomy and benchmark are likely to become a standard reference in the field, driving progress in speaker identification, paralinguistic understanding, and environmental sound recognition within conversational contexts. The findings on the distinct failure modes for different evidence types will inform the design of specialized memory modules in LALMs. The paper introduces VoxMem, a comprehensive benchmark for evaluating multi-session spoken conversational memory across acoustic evidence types and memory operations. It provides a principled taxonomy and rigorous construction pipeline, revealing significant gaps in current LALMs' ability to retain non-lexical acoustic information over long contexts.
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
Primary: KAIST
All Institutions: KAIST, VGG, University of Oxford
The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
The paper employs a rigorous mechanistic interpretability framework, combining Representational Similarity Analysis (RSA) and Causal Mediation Analysis (CMA) to dissect the internal workings of Audio-Visual LLMs. The authors successfully identify a three-stage symbolic binding mechanism (Anchor ID Retrieval, Target ID Selection, Feature Retrieval) that relies on modality-specific symbolic variables (Temporal IDs for audio, Position IDs for vision). This is a sophisticated approach that moves beyond black-box evaluation to understand *how* models process cross-modal information. The proposed intervention, using an off-the-shelf Active Speaker Detection (ASD) model to overlay bounding boxes, is a clever, training-free (and lightly fine-tuned) solution that directly targets the identified bottleneck in audio-visual alignment.
The experiments are extensive and well-structured. The authors validate their mechanistic findings across four different AVLLM architectures (video-SALMONN2+, Qwen2.5-Omni, MiniCPM-o-4.5) using both synthetic toy datasets and real-world benchmarks (SocialOmni, AVSpeaker, DailyOmni, etc.). The ablation studies effectively isolate the contribution of the ASD prompting, showing that it outperforms other training-free decoding methods and standard fine-tuning. The generalization to broader audio-visual benchmarks (DAVE, OmniBench, WorldSense) further strengthens the claim of the method's utility.
The paper provides detailed implementation specifics, including LoRA ranks, training steps, and dataset construction methods. The use of standard off-the-shelf models and clear descriptions of the prompting strategy enhances reproducibility. However, the specific synthetic video generation pipeline and the exact ASD model used (referenced as [CITATION]) would need to be clearly specified in the final publication for full reproducibility.
The primary limitation is the reliance on a synthetic toy dataset for the core mechanistic analysis, which may not fully capture the complexity of real-world multi-speaker scenarios, although the authors do validate on real-world data. Additionally, the method depends on the accuracy of the external ASD model; if the ASD model fails, the prompting strategy may degrade performance. The fine-tuning, while lightweight, still requires access to the model's weights and a GPU, which may not be feasible for all users.
This work has significant implications for the development of more robust multimodal AI systems. By identifying specific failure modes in cross-modal binding, it provides actionable insights for improving model architectures and training strategies. The proposed ASD prompting method is a practical, low-cost solution that can be easily integrated into existing AVLLM pipelines, potentially improving performance in applications like video conferencing, accessibility tools, and video understanding. The paper identifies emergent symbolic trimodal binding mechanisms in Audio-Visual LLMs and proposes a simple, effective audio-visual prompting method using Active Speaker Detection to mitigate identified binding failures, demonstrating significant performance gains across multiple benchmarks.
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.
Primary: Sakana AI
All Institutions: Sakana AI
The paper introduces a symbiotic architecture that decouples audio prefilling from the LLM backbone by directly generating the KV cache via a lightweight injector, thereby reducing computational cost and preventing catastrophic forgetting. The technical contribution is significant in addressing two major bottlenecks in audio-LLM development, though the experimental validation is limited to small-scale models, leaving the full scalability benefits to be verified in future work.
The paper proposes a "symbiotic" architecture where an audio injector module generates the Key-Value (KV) cache for a frozen Large Language Model (LLM), bypassing the need to pass audio embeddings through the LLM's self-attention layers during prefilling. The injector is a CNN-based module (using LConv blocks) that maps audio encoder outputs directly to the LLM's KV space. Key technical contributions include a "KV Scale Matching" strategy to align the injector's output distribution with the LLM's internal KV distribution, and "Noisy RoPE" training to improve robustness to sequence lengths unseen during training. The approach is theoretically sound, leveraging the fact that the KV cache is the interface between context and generation, allowing modality-specific processing to be decoupled from the backbone.
Experiments are conducted on a compact model (Qwen3-0.6B) with WavLM as the audio encoder. The paper evaluates ASR (LibriSpeech), Audio QA (Clotho), and Acoustic Scene Classification (CochlScene). The proposed method outperforms the "Encoder-only" (SLM-style) baseline significantly on non-ASR tasks and approaches the performance of a fine-tuned "Monolithic" model while using fewer active parameters during audio prefilling. Crucially, it preserves text-only performance (WikiText-2, HellaSwag, GSM8K) by construction, avoiding catastrophic forgetting. However, the evaluation is limited to a single small backbone, and the speedup is modest (157s vs 199s) in the tested regime.
The paper provides detailed architectural descriptions, including specific hyperparameters for the injector (kernel size, subsampling stride), training configurations (learning rates, batch size, steps), and specific techniques for stabilization (RMSNorm initialization, noisy RoPE parameters). The use of standard open-source components (WavLM, Qwen3) enhances reproducibility. However, code availability is not explicitly confirmed in the text, and the specific implementation of the KV injection into the inference engine is not detailed.
The primary limitation is the scale of the experiments; using a 0.6B LLM does not fully validate the scalability claims for larger models where the prefilling bottleneck is more severe. The speedup demonstrated is relatively small in the current setup, and the authors acknowledge that the CNN-centric injector may limit performance on complex acoustic tasks compared to attention-based injectors. The method relies on specific internal details of the backbone (e.g., Qwen3's RMSNorm), which may require adaptation for other LLM families.
This work offers a practical solution to the high computational cost of multimodal prefilling and the risk of catastrophic forgetting in LLM fine-tuning. By decoupling audio processing from the backbone, it enables the use of larger, more capable LLMs for audio tasks without proportional increases in inference latency for the audio portion. This could facilitate the deployment of ALMs in edge devices or real-time applications where prefilling latency is critical. The paper introduces a symbiotic architecture that decouples audio prefilling from the LLM backbone by directly generating the KV cache via a lightweight injector, thereby reducing computational cost and preventing catastrophic forgetting. The technical contribution is significant in addressing two major bottlenecks in audio-LLM development, though the experimental validation is limited to small-scale models, leaving the full scalability benefits to be verified in future work.
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.
Primary: Sungkyunkwan University
All Institutions: Sungkyunkwan University
The paper introduces a practical and effective method for reducing KV cache memory in full-duplex speech models by converting acoustic history to text during listening slack, demonstrating significant memory savings with preserved conversational quality.
The paper proposes a novel "acoustic-to-text KV compression" strategy for full-duplex speech models. The core idea is to exploit "listening-time slack"โthe computational gap between audio unit arrivalsโto run a lightweight transcription side channel (implemented via LoRA) that converts incoming speech into text tokens. These text tokens are retained in the KV cache while older acoustic KV states are evicted. The methodology is technically sound, leveraging the existing language model backbone rather than a separate ASR model. A significant methodological strength is the use of knowledge distillation from the frozen original model to preserve native listening/speaking behaviors (turn-taking, interruption) which would otherwise be disrupted by the new transcription objective. The training setup uses word-level forced alignments and a composite loss function including cross-entropy for ASR and KL divergence for behavior preservation.
The experiments are conducted on the MiniCPM-o 4.5 model. The evaluation covers three key areas: 1) Long-form speech understanding (LongSpeech benchmark), showing improved WER and QA accuracy compared to native streaming, and competitive performance with external ASR cascades. 2) Full-duplex interaction (Full-Duplex-Bench), demonstrating that the distillation component is critical for maintaining natural conversational dynamics (pause handling, turn-taking). 3) Efficiency, showing a 64.6% reduction in peak KV cache size with no real-time deadline misses on H100 GPUs. The ablation study effectively isolates the impact of the distillation loss, showing that without it, the model fails at turn-taking. The comparison with external ASR cascades is particularly relevant, showing that the proposed method achieves similar quality with fewer added parameters and integrated memory management.
The paper provides sufficient detail for reproduction, including the specific model (MiniCPM-o 4.5), training data (LibriSpeech train-clean), hyperparameters (LoRA rank 16, learning rate, batch size), and the specific eviction strategy (5-unit retention window). The use of standard benchmarks (LongSpeech, Full-Duplex-Bench) facilitates comparison. However, the code is not explicitly linked in the text provided, and the specific implementation details of the "listening-time slack" scheduling might require access to the source code for precise replication.
The method is evaluated primarily on a single model architecture (MiniCPM-o 4.5), so generalizability to other full-duplex models is not established. The transcription side channel relies on the model's ability to predict the start of the ASR segment; errors in this gating mechanism could lead to missed transcriptions. The method assumes a consistent "listening-time slack" of ~900ms, which may not hold for all hardware configurations or model variants. Additionally, the WER (12.9%) is higher than dedicated streaming ASR models (e.g., FastConformer at 10.3%), suggesting a trade-off between integration and pure transcription accuracy.
This work addresses a critical bottleneck in deploying full-duplex speech agents: memory consumption during long conversations. By converting acoustic history to text, it enables longer context windows without proportional memory increases. This is highly relevant for real-time voice assistants and interactive agents. The approach of using "slack" time for auxiliary tasks (transcription) is a generalizable principle for efficient inference in streaming multimodal models. The paper introduces a practical and effective method for reducing KV cache memory in full-duplex speech models by converting acoustic history to text during listening slack, demonstrating significant memory savings with preserved conversational quality.
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.
Primary: Microsoft
All Institutions: Microsoft, University of Michigan
The paper presents a robust adaptation of a personalized speech enhancement model for online audio-visual target-speaker extraction, demonstrating significant improvements in target recovery and listening quality over existing extractors on both synthetic and real-world meeting data. By leveraging the prior knowledge of a deployed enhancer and integrating visual cues through a gated conditioning mechanism, the method effectively addresses the selection challenge while maintaining high speech quality, supported by extensive objective and subjective evaluations that highlight its practical deployability and the trade-offs involved in such adaptation.
The paper proposes a pragmatic and effective adaptation of a deployed personalized speech enhancement model (PVQE) into an online audio-visual target-speaker extraction system (V+E). The core methodological contribution is the integration of visual cues (mouth motion via AV-HuBERT) into the existing speaker-conditioning path of the enhancer, rather than training a separator from scratch. The use of a gated residual connection to fuse enrollment embeddings with visual features is a sound architectural choice that allows the model to retain its enhancement capabilities while learning selection. The fine-tuning strategy, which involves removing lookahead operations and rearranging decoder filters for causal processing, is well-described and technically rigorous. The approach effectively leverages the prior knowledge of the enhancement model to solve the extraction problem, addressing the "selection" gap in personalized enhancement.
The experimental evaluation is extensive and high-quality. The authors test on both synthetic mixtures (LRS3, VoxCeleb2) and, crucially, recorded meeting corpora (AMI, MCoRec, MTM, UniTalk), which is a significant strength as it moves beyond the standard synthetic benchmark limitations. The inclusion of a large-scale human listening test (170 listeners, P.835 protocol) provides strong evidence for the perceptual quality claims. The results show consistent improvements over the baseline extractor (AVASE) in both objective metrics (SI-SNRi, WAcc) and subjective quality (MOS). The preservation and rejection tests further validate the system's deployability, showing it does not distort the target when no interference is present. The analysis of the UniTalk results, where background noise suppression degrades, is honest and provides valuable insight into the trade-offs of this adaptation strategy.
The paper provides sufficient detail for reproducibility, including model architecture dimensions, training data sources, and fine-tuning procedures. The use of standard public datasets (LRS3, VoxCeleb2, AMI, MCoRec) aids reproducibility, although the internal MTM dataset is not publicly available. The specific hyperparameters and initialization strategies (e.g., zero-initialization of the enrollment projection) are clearly stated. However, the code and model weights are not explicitly linked in the provided text, which is a minor drawback for immediate reproducibility.
The primary limitation is the degradation in background noise suppression on in-the-wild data (UniTalk) compared to the original enhancement model, indicating that the adaptation to extraction comes at a cost to general noise robustness. The system relies on the presence of synchronized video, which may not be available in all deployment scenarios. Additionally, the comparison with PVQE (the starting model) is somewhat confounded because PVQE is not designed for extraction, so the "quality retention" claim is relative to a model that fails at the primary task (selection) in many cases.
This work has significant practical impact for real-time communication systems (e.g., video conferencing). By demonstrating that a deployed enhancement model can be adapted for extraction with minimal architectural changes and high perceptual quality, it offers a viable path for improving user experience in noisy or multi-speaker environments. The focus on low-latency, online processing makes it highly relevant for industrial applications. The rigorous evaluation on recorded meetings sets a new standard for how such systems should be tested, encouraging the field to move beyond synthetic benchmarks. The paper presents a robust adaptation of a personalized speech enhancement model for online audio-visual target-speaker extraction, demonstrating significant improvements in target recovery and listening quality over existing extractors on both synthetic and real-world meeting data. By leveraging the prior knowledge of a deployed enhancer and integrating visual cues through a gated conditioning mechanism, the method effectively addresses the selection challenge while maintaining high speech quality, supported by extensive objective and subjective evaluations that highlight its practical deployability and the trade-offs involved in such adaptation.
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39~dB, compared with 2.49~dB for our strongest baseline. See the demo page at https://benschou.com/notesep.
Primary: Purdue University
All Institutions: Purdue University, Loyola University Chicago, University of Michigan
The paper introduces the first deep learning framework for score-informed note separation, achieving state-of-the-art performance by combining a dual-stream extraction network with a novel joint energy reallocation mechanism. By addressing the specific challenges of harmonic overlap and mixture consistency, NoteSep provides a robust tool for note-level audio manipulation, significantly advancing the field of music source separation beyond instrument-level stems.
The paper proposes NoteSep, a two-stage framework for score-informed note separation. The first stage, NoteGrab, utilizes a dual-stream TFC-TDF U-Net architecture to process harmonic and percussive components (via HPSS) separately, linked by bidirectional cross-attention. A key methodological contribution is the "selective harmonic gating," which conditionally suppresses lower-octave interference based on score alignment, addressing a specific physical challenge in polyphonic music. The second stage, Adaptive Set Ownership (ASO), is a novel joint separation module that reallocates mixture energy among concurrent note estimates to enforce consistency, using a learned gate and log-gain prediction. This approach effectively tackles the "double counting" problem inherent in independent note extraction.
The authors curate a substantial training set (SCNS-Train) and a disjoint evaluation set (SCNS-Eval) covering 16 instruments. The evaluation is rigorous, comparing against both traditional methods (Score-Informed NMF) and commercial software (Melodyne). The results show a significant improvement in SI-SDR (7.39 dB vs 2.49 dB for the strongest baseline). Furthermore, the paper evaluates the system on real-world ensemble recordings (PHENICX-Anechoic, Bach10) by aggregating note estimates into instrument stems, demonstrating competitive performance against specialized instrument separation models. The inclusion of an editing evaluation (SCNS-Edit) further validates the utility of the separated notes for downstream tasks.
The paper commits to releasing code, model weights, and datasets. The architectural details are described with sufficient specificity (e.g., TFC-TDF v3, FiLM conditioning, specific loss functions), and the demo page provides audio examples. The use of standard libraries (librosa) for preprocessing aids reproducibility.
The method relies heavily on accurate score alignment (pitch, onset, offset); errors in the input score will propagate to the separation quality. The computational cost is non-trivial, requiring K passes for K notes plus an ASO pass, which may limit real-time application on consumer hardware. The evaluation on real recordings is indirect (via instrument stem aggregation), so the fidelity of individual notes in complex, noisy real-world scenarios is not directly quantified with the same rigor as the synthetic benchmark.
This work enables precise, note-level audio editing, which has significant applications in music production, education (isolating specific notes for learning), and restoration. It bridges the gap between symbolic music information and audio signal processing, potentially facilitating more granular music information retrieval tasks. The paper introduces the first deep learning framework for score-informed note separation, achieving state-of-the-art performance by combining a dual-stream extraction network with a novel joint energy reallocation mechanism. By addressing the specific challenges of harmonic overlap and mixture consistency, NoteSep provides a robust tool for note-level audio manipulation, significantly advancing the field of music source separation beyond instrument-level stems.
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic information is encoded by individual latents or by structured groups of features. We find that PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings. This recoverability is not reducible to lexical memorisation, and Open and Closed PoS classes differ substantially. Categories are supported by compact groups of sparse latents, with substantial variation across tags. These groups remain stable on held-out data, while also showing overlap between related categories. Our results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
Primary: University of Pisa
All Institutions: University of Pisa
The paper demonstrates that Part-of-Speech categories in LLMs are encoded as distributed but compact groups of SAE latents, with significant structural differences between open and closed classes. By combining probing, feature salience, and coverage analysis, the authors provide a detailed map of how morpho-syntactic information is organized in sparse latent spaces, offering valuable insights into the interpretability of modern language models.
The paper employs a rigorous three-step interpretability pipeline to analyze Part-of-Speech (PoS) encoding in Sparse Autoencoder (SAE) latents. First, it establishes recoverability using L1-regularized logistic regression probes on SAE activations from LLaMA-3-8B. Second, it moves beyond binary accuracy to localize information by analyzing feature salience (classifier coefficients) and coverage (the minimal number of latents required to cover 95% of instances for a given PoS tag). Third, it validates these localized groups on held-out data and a controlled synthetic dataset to test stability and additivity. The methodology is sound, moving logically from "is the information there?" to "where is it?" and "is it stable?". The use of a controlled dataset with minimal templates is a strong methodological choice to isolate lexical effects from syntactic ones.
The experiments are comprehensive, utilizing the GUM Treebank for natural data and a custom controlled dataset for validation. The results clearly demonstrate that PoS information is distributed across compact groups of latents rather than single monosemantic units. The distinction between Open-class (nouns, verbs) and Closed-class (determiners, conjunctions) PoS is well-supported, with closed classes showing more compact and stable latent groups. The ablation studies, including random label controls and layer-wise analysis, effectively rule out trivial explanations like lexical memorization or layer-specific artifacts. The finding that a small union of these latents (498 out of 131,072) supports strong multi-class classification is a significant quantitative result.
The authors provide high reproducibility by releasing code, data, and the specific SAE checkpoint used. The GitHub repository contains the pipeline for extracting activations and running the probes. The use of standard libraries (Scikit-learn, Sparsify) and public models (LLaMA-3-8B, EleutherAI SAE) ensures that other researchers can easily replicate the findings. The detailed description of the subword-to-token alignment strategy (leftmost subword anchoring) is crucial for reproducibility in this domain.
The study is limited to a single model (LLaMA-3-8B) and a single SAE variant, which may limit the generalizability of the findings to other architectures or sparsity regimes. The analysis is restricted to English, and it is unclear how these patterns would manifest in morphologically rich languages. The controlled dataset, while useful, is small and covers only a subset of syntactic structures. Additionally, the reliance on linear probes means that non-linear interactions between latents are not captured, potentially underestimating the complexity of the representation.
This work contributes to the broader field of mechanistic interpretability by providing a structured framework for analyzing how linguistic categories are represented in sparse latent spaces. It challenges the assumption of strict monosemanticity for linguistic features, suggesting instead a distributed but localized organization. This has implications for how interpretability tools are designed and how we understand the internal representations of LLMs. The findings could inform the development of more targeted interpretability methods that focus on groups of features rather than individual units. The paper demonstrates that Part-of-Speech categories in LLMs are encoded as distributed but compact groups of SAE latents, with significant structural differences between open and closed classes. By combining probing, feature salience, and coverage analysis, the authors provide a detailed map of how morpho-syntactic information is organized in sparse latent spaces, offering valuable insights into the interpretability of modern language models.
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
Primary: Xiamen University
All Institutions: Xiamen University
[One sentence main contribution]. The paper introduces EditVoice, a variable-length non-autoregressive zero-shot TTS model using Edit Flows and Complementary Prompt Sampling, achieving competitive performance in TTS and speech editing with high inference efficiency.
The paper proposes EditVoice, a non-autoregressive (NAR) zero-shot TTS model based on Edit Flows. The core methodological contribution is the application of Edit Flows to speech token sequences, enabling variable-length generation through insertions, deletions, and substitutions without pre-specifying sequence length. The authors introduce a speech-infilling training objective that unifies TTS and speech editing. A key heuristic contribution is Complementary Prompt Sampling (CPS), which leverages the complementary predictions of prefix and suffix prompt placements to improve generation quality, along with a local consistency rule to prevent artifacts. The method also exploits the model's ability to generalize to editing plausible (non-noise) token sequences for post-generation refinement.
Experiments are conducted on Seed-TTS Eval EN and LibriSpeech-PC for TTS, and RealEdit for speech editing. The model is trained on 10K hours of GigaSpeech. Results show competitive WER and speaker similarity compared to autoregressive baselines like CosyVoice2 and other NAR models. The 16-NFE variant achieves a low RTF of 0.0989. Ablations confirm the effectiveness of CPS and post-generation refinement. Subjective evaluations (NMOS, SMOS) are included.
The paper provides detailed architectural specifications (14-layer LLaMA-style Transformer, Conformer text encoder) and training hyperparameters. It specifies the use of frozen S3Tokenizer2 and CosyVoice2 acoustic decoder. However, the code repository is not explicitly linked in the text provided (only a demo page), which may limit immediate reproducibility compared to papers with open-source code.
The model relies on a specific tokenizer (S3Tokenizer2) and acoustic decoder (CosyVoice2), which may limit generalizability to other speech tokenization schemes. The CPS heuristic, while effective, is a manual rule-based approach that may not scale optimally to all conditions. The training data size (10K hours) is significantly smaller than some state-of-the-art systems (1M+ hours), though the paper argues for efficiency.
This work contributes to the efficiency of NAR TTS by removing the need for length prediction, a common bottleneck. The unification of TTS and editing in a single model is a step towards more flexible speech manipulation tools. The use of Edit Flows for speech is a novel application of this generative framework. [One sentence main contribution]. The paper introduces EditVoice, a variable-length non-autoregressive zero-shot TTS model using Edit Flows and Complementary Prompt Sampling, achieving competitive performance in TTS and speech editing with high inference efficiency.
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-improvement system for audio understanding. To our knowledge, it is the first to evolve the model, waveforms, questions, and difficulty in one closed loop. EvoAudio uses the current model's performance to set the focus and difficulty of the next training data. A library of audio tools then constructs questions whose answers follow from how the audio was made, providing verifiable supervision without new human annotation. Reinforcement learning updates the model, and validation decides whether it enters the next evolution round. Across 13 rounds, EvoAudio improves five models with different audio encoders and language backbones on MMSU, MMAU-Pro, and MMAR. It achieves the highest average for every backbone, raising overall performance by up to 6.3 points. The improvement unfolds over successive rounds, with each stronger model starting the next round.
Primary: The Chinese University of Hong Kong, Shenzhen
All Institutions: The Chinese University of Hong Kong, Shenzhen, Tsinghua University, Tencent Hunyuan, Amphion Technology Co., Ltd.
EvoAudio introduces a recursive self-improvement system that jointly evolves the model, waveforms, questions, and difficulty in a closed loop, achieving significant gains in audio understanding across multiple LALM backbones. The paper presents a rigorous methodology for generating verifiable synthetic data and adapting curricula based on model performance, demonstrating that recursive self-improvement is effective for enhancing perceptual skills in audio language models.
The paper proposes EvoAudio, a recursive self-improvement framework for Large Audio Language Models (LALMs). The core innovation is a closed-loop system where the model's performance on a fixed set of verifiable skills dictates the curriculum for the next training round. Unlike previous methods that use static synthetic data or rely on stronger teacher models for labeling (which introduces bias and error propagation), EvoAudio generates training data using a library of 24 audio tools that construct waveforms with known ground-truth properties (e.g., pitch, tempo, speaker count). This ensures "verifiable supervision" without human annotation. The system uses GRPO (Group Relative Policy Optimization) for reinforcement learning, where rewards are derived from the verifiable answers. A key methodological strength is the adaptive difficulty adjustment: the system monitors the "mixed rate" (variance in model responses) to ensure questions remain in the model's learning zone, avoiding saturation or frustration. The separation of the "verifier" (checking if the answer matches the construction) and the "acoustic check" (ensuring the rendered audio actually contains the intended cue) is a robust design choice that mitigates synthesis errors.
The experiments are extensive and rigorous. The authors evaluate five different LALM backbones (Qwen2.5-Omni, Audio Flamingo 3, Kimi-Audio, MiMo-Audio, MiniCPM-o) across three major benchmarks (MMSU, MMAU-Pro, MMAR). The results show consistent improvements over base models and strong baselines (Static-profile GRPO, Pooled GRPO). The ablation studies are particularly valuable, isolating the impact of adaptive difficulty, skill quotas, and acoustic verification. The finding that removing acoustic verification leads to "reward hacking" (where the model learns to exploit rendering errors) is a significant insight. The improvement of up to 6.3 points on average is substantial in the context of audio understanding, where progress is often incremental. The comparison against "Pooled GRPO" effectively demonstrates the necessity of the recursive, adaptive nature of the curriculum rather than just having a large pool of synthetic data.
The paper provides high-level details on the system architecture, the number of tools (24), question types (47), and training hyperparameters (learning rate, KL coefficient, number of rounds). However, the specific implementation of the "audio tools" library is not fully detailed in the text, and the code repository is not linked in the provided text (only a demo page). The reliance on specific TTS models (Qwen3-TTS) and datasets (LibriSpeech, FSD50K) is clear, but the exact prompts for the LLM proposer and the specific logic for the "acoustic checks" (e.g., how pitch is remeasured) would require the code for full reproduction. The demo page suggests some availability, but a full open-source release of the tool library would be necessary for high reproducibility.
The primary limitation is the scope of the tool library. The system can only improve on skills that can be synthesized and verified by the current 24 tools. Complex semantic or cultural reasoning (tested in MMAR) sees less improvement because these aspects are harder to construct synthetically with verifiable ground truth. The method is also computationally expensive, requiring 13 rounds of training and multiple rollouts per question. Additionally, the "verifiable" nature of the data means the model may not generalize well to open-ended audio questions that do not have a single correct answer derived from construction parameters.
This work has significant implications for the training of multimodal models. The concept of "self-evolving" curricula with verifiable rewards is applicable beyond audio to other modalities where ground truth can be constructed (e.g., code, math, visual geometry). It addresses the bottleneck of high-quality, diverse training data for perceptual tasks. By demonstrating that models can improve their own perceptual abilities without human labels, it offers a scalable path for developing more robust audio understanding systems. The insights into reward variance and curriculum adaptation are also valuable for the broader reinforcement learning community. EvoAudio introduces a recursive self-improvement system that jointly evolves the model, waveforms, questions, and difficulty in a closed loop, achieving significant gains in audio understanding across multiple LALM backbones. The paper presents a rigorous methodology for generating verifiable synthetic data and adapting curricula based on model performance, demonstrating that recursive self-improvement is effective for enhancing perceptual skills in audio language models.
Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited labels for their intent. We present a pipeline for synthesizing intent-labeled, two-channel conversational speech from relational event lists. An LLM authors each event's speaker, text, conversational act, and attachment to an earlier event without predicting absolute timestamps. Events are synthesized independently, aligned with their source text, and placed on a shared clock, so turn-taking landmarks are measured from the rendered signal while silence durations are specified or sampled from turn-taking distributions. The pipeline covers 42 phenomena across eight families in English and Mandarin, derives frame-level system actions from authored intent, and promotes diversity using small, diverse sets of prior examples and batch prompts that request alternatives with self-reported probabilities. Ablations show gains in each targeted diversity dimension. On a four-action label space for taking, holding, releasing, and not holding the conversational floor, a semantic voice-activity detector using only current and past audio reaches start-speaking and start-listening F1 scores of 0.819 and 0.802. When generating its own responses, the full-duplex speech model Moshi takes 0.85 of the reference turns after fine-tuning on the generated corpus, compared with 0.44 before fine-tuning. Its frame-level precision for predicting system-floor occupancy rises from 0.46 to 0.88. With reference context at each step, its frame-level floor F1 rises from 0.893 to 0.962. These results show that controlled synthesis can provide learnable and transferable supervision for full-duplex turn management.
Primary: Tencent Americas
All Institutions: Tencent Americas, School of Electronic Information Wuhan University, Columbia University
The paper presents a novel pipeline for synthesizing intent-labeled, full-duplex conversational speech by decoupling LLM-based semantic authoring from acoustic realization via forced alignment. This approach effectively addresses the supervision gap in existing corpora by providing precise, intent-aware labels for turn-taking events, leading to significant improvements in the performance of full-duplex speech models like Moshi in managing conversational floor occupancy.
The paper proposes a robust pipeline for synthesizing full-duplex conversational data by decoupling semantic authoring from acoustic realization. The core innovation is the "relational event representation," where an LLM defines conversational acts (turn, barge-in, backchannel) and their temporal relationships to previous events without predicting absolute timestamps. This is a significant methodological improvement over previous approaches that attempted to have LLMs predict precise timing, which is unreliable. The pipeline uses forced alignment to measure actual speech boundaries and inserts silence based on turn-taking distributions, ensuring that the acoustic timing is natural while the intent labels remain precise. The use of "diversity hubs" to prevent mode collapse in LLM generation is a clever application of negative memory. The separation of intent (semantic) from activity (acoustic) allows for the creation of frame-level labels that distinguish between a barge-in (which should stop the system) and a backchannel (which should not), a critical distinction for full-duplex models that voice activity detection (VAD) alone cannot provide.
The evaluation is strong and directly addresses the paper's claims. The authors train a semantic voice-activity detector and fine-tune the Moshi full-duplex speech model on the generated corpus. The results show significant improvements: Moshi's ability to take the correct turn increases from 0.44 to 0.85, and frame-level floor occupancy precision rises from 0.46 to 0.88. These metrics are highly relevant to the task of turn-taking management. The ablation studies on diversity mechanisms further validate the pipeline's components. The use of both English and Mandarin data demonstrates the pipeline's cross-lingual applicability.
The paper provides detailed descriptions of the pipeline stages, including the specific models used (DeepSeek-V4-Pro for authoring, Qwen3-TTS for synthesis, Qwen3-ForcedAligner for alignment). The mathematical formulation for timeline assembly is clear. However, the specific hyperparameters for the turn-taking distributions and the exact prompt templates are referenced in appendices that are not fully visible in the provided text, which slightly limits immediate reproducibility. The reliance on specific commercial/proprietary LLMs (DeepSeek, Qwen) may also pose accessibility challenges for some researchers, though these models are increasingly open-source.
The pipeline relies on the quality of the underlying TTS and LLM models; if the TTS produces unnatural prosody, the "naturalistic" claim is weakened. The evaluation is primarily on the Moshi model; generalization to other full-duplex architectures is not tested. The "intent" labels are derived from the LLM's authoring, so any errors in the LLM's understanding of conversational acts will propagate to the labels. The paper does not extensively discuss the computational cost of the LLM-based authoring and validation loop for large-scale corpus generation.
This work addresses a critical bottleneck in training full-duplex dialogue systems: the lack of high-quality, intent-labeled data. By providing a scalable pipeline to generate such data, the paper enables the development of more natural and responsive conversational AI. The focus on distinguishing between different types of overlap (barge-in vs. backchannel) is particularly important for user experience, as incorrect handling of these events leads to frustrating interactions. The methodology could be extended to other modalities or languages with appropriate configuration overlays. The paper presents a novel pipeline for synthesizing intent-labeled, full-duplex conversational speech by decoupling LLM-based semantic authoring from acoustic realization via forced alignment. This approach effectively addresses the supervision gap in existing corpora by providing precise, intent-aware labels for turn-taking events, leading to significant improvements in the performance of full-duplex speech models like Moshi in managing conversational floor occupancy.
Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.
Primary: People Make Things
All Institutions: People Make Things
The paper introduces TACT, a novel intent-conditioned benchmark and scoring framework for full-duplex spoken dialogue models that replaces binary turn-taking metrics with a strictly proper, continuous ranked probability score. By conditioning evaluation on the speaker's latent intent and behavioral memory, the authors demonstrate that current state-of-the-art models exhibit significant rigidity and over-eagerness, failing to adapt to conversational context as humans do, thereby providing a more rigorous and human-aligned standard for assessing dialogue timing.
The paper proposes a theoretically grounded shift from binary, fixed-window turn-taking metrics to a continuous, intent-conditioned scoring framework (TACT). The core methodological contribution is the use of a threshold-weighted Continuous Ranked Probability Score (twCRPS) where the weights are derived from intent-specific human floor-transfer-offset (FTO) distributions. This is a sophisticated approach that correctly identifies that "silence" and "overlap" are not inherently good or bad, but depend on the speaker's latent intent (e.g., rhetorical vs. answer-seeking). The authors provide rigorous proofs for boundedness, strict propriety, and the reduction of their metric to existing binary metrics in degenerate cases, which is a strong theoretical foundation. The integration of a "memory profile" for each speaker to condition the intent posterior is a novel and practical addition that addresses the variability in human conversational styles.
The experimental setup is robust, utilizing a large-scale benchmark (9,728 episodes, 73.2 hours) constructed from five diverse public corpora. The evaluation of eleven frontier systems (including open-source and proprietary APIs) provides a comprehensive landscape of current full-duplex model capabilities. The finding that models exhibit an "over-eagerness pathology" and fail to adapt to speaker-specific memory profiles is a significant empirical contribution. The validity study, showing a Spearman correlation of 0.81 with human judgments compared to 0.46 for binary metrics, strongly supports the utility of the proposed benchmark.
The paper claims that data, kernels, and prompts are available in the Supplementary Material, which is a positive sign for reproducibility. However, the reliance on proprietary APIs (GPT-4o, Gemini) for some evaluations limits full reproducibility for those specific data points. The detailed description of the annotation protocol and the mathematical formulation of the scoring rules enhances the potential for independent verification.
The primary limitation is the discretization of intent into six classes, which may oversimplify the continuous nature of human intent. The reliance on an LLM judge for fusing the intent posterior introduces potential biases or failure modes shared with the underlying LLM. Additionally, the kernels are fitted on English corpora, limiting the immediate generalizability to other languages without refitting. The "memory profile" mechanism, while effective in ablation, adds complexity to the inference pipeline for evaluated systems.
This work has high potential impact on the development of natural spoken dialogue systems. By providing a more nuanced and human-aligned evaluation metric, it can guide the training of models that are not just fast, but contextually appropriate. The identification of the "over-eagerness" problem highlights a critical gap in current model architectures that future research must address. The framework could be extended to other interactive AI domains where timing and intent are crucial. The paper introduces TACT, a novel intent-conditioned benchmark and scoring framework for full-duplex spoken dialogue models that replaces binary turn-taking metrics with a strictly proper, continuous ranked probability score. By conditioning evaluation on the speaker's latent intent and behavioral memory, the authors demonstrate that current state-of-the-art models exhibit significant rigidity and over-eagerness, failing to adapt to conversational context as humans do, thereby providing a more rigorous and human-aligned standard for assessing dialogue timing.
End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipulation. We formalize imperceptible attacks on full-duplex agents as optimization over additive perturbations confined beneath the psychoacoustic masking threshold of the carrier speech, under three goals: targeted semantic hijacking, response suppression, and policy jailbreaking. Against an undefended Moshi-style agent, white-box attacks succeed in up to 91.7% of trials. We then introduce psychoacoustically aligned latent smoothing (PALS), which injects anisotropic Gaussian noise shaped by local codebook covariance at the residual-vector-quantized latent interface, with input noise shaped by the masking threshold constraining the attacker and trained by a Kullback--Leibler consistency objective. Deployed with no inference-time cost, PALS reduces hijack to 8.3%, mute to 11.2%, and jailbreak to 9.1% at clean quality within 2.3%. A Monte Carlo-smoothed variant certifies an ellipsoidal latent radius up to 0.616, a guaranteed floor that the empirical robustness far exceeds.
Primary: People Make Things
All Institutions: People Make Things
The paper introduces a principled, psychoacoustically aligned latent smoothing defense for full-duplex speech agents, significantly reducing adversarial attack success rates while maintaining utility, though its certified guarantees are conservative and its institutional provenance is obscure.
The paper proposes Psychoacoustically Aligned Latent Smoothing (PALS), a defense mechanism for full-duplex speech-to-speech (S2S) models. The core innovation lies in applying randomized smoothing at the latent interface of the Residual Vector Quantization (RVQ) codec, using anisotropic Gaussian noise shaped by local codebook covariance. This is coupled with an input-space augmentation that matches the psychoacoustic masking threshold, derived via a maximum-entropy argument. The training objective incorporates a KL consistency term (similar to TRADES) to ensure robustness without requiring adversarial examples during training. The theoretical contribution includes a certified robustness radius for the smoothed operator, though the authors correctly note that the certificate is a lower bound and the empirical robustness of the deployed model (PALS-Train) is the primary metric of interest. The methodology is sound, leveraging the geometric properties of the quantizer to define the noise distribution, which is a clever and non-trivial design choice.
The experiments are comprehensive, evaluating three distinct attack types (semantic hijacking, response suppression/muting, and jailbreaking) against a Moshi-style agent. The paper reports significant reductions in attack success rates (e.g., hijacking drops from 91.7% to 8.3%) while maintaining clean quality metrics (UTMOS, WER, Turn-Taking F1). Ablations clearly demonstrate the contribution of each component (anisotropic covariance, input-space dual, KL consistency). The inclusion of over-the-air (OTA) tests with reverberation and babble noise adds practical relevance. However, the evaluation relies on a specific "Moshi-style" agent architecture; generalizability to other S2S architectures (e.g., those without RVQ or with different codec designs) is not tested.
The paper provides detailed hyperparameters, training objectives, and attack configurations. It mentions a released `certify.py` harness. However, the code and model checkpoints are not explicitly linked in the provided text (no GitHub URL found), and the "Moshi-style" agent is described rather than provided as a specific open-source artifact. The reliance on specific proprietary or large-scale datasets (CANDOR, Seamless Interaction) may limit immediate reproducibility for smaller groups, though the methodology is clearly specified.
The certified robustness radius is relatively small (0.616 latent units, corresponding to -26 dB input budget), which is far below the attack budgets tested empirically. The authors acknowledge this, positioning the certificate as a "floor" rather than a guarantee against strong attacks. The defense is specific to RVQ-based codecs; its applicability to continuous latent spaces or other quantization schemes is unclear. The "People Make Things" affiliation is obscure, raising questions about the scale of resources behind the work, though the technical depth suggests significant effort.
As full-duplex voice agents become prevalent, securing them against imperceptible audio attacks is critical. This work provides a strong baseline for latent-space defenses in speech, potentially influencing future designs of secure speech interfaces. The psychoacoustic alignment of the defense with the attack constraint is a principled approach that could be adapted to other sensory modalities. The paper introduces a principled, psychoacoustically aligned latent smoothing defense for full-duplex speech agents, significantly reducing adversarial attack success rates while maintaining utility, though its certified guarantees are conservative and its institutional provenance is obscure.
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage $ฮ$ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks $ฮ$ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
Primary: People Make Things
All Institutions: People Make Things
The paper introduces a geometric framework linking subspace entanglement to demographic disparity in speech encoders, validated by a large-scale synthetic audit grid and a lightweight repair adapter. It makes a significant contribution to the theory of fairness in representation learning by providing provable bounds on disparity based on measurable geometric properties, although its reliance on synthetic data for primary validation limits the strength of its causal claims.
The paper proposes a rigorous geometric framework for auditing demographic fairness in speech encoders. The core contribution is the definition of "axis-fidelity" and "subspace leakage" (principal angles between semantic and voice subspaces). The theoretical contribution is a proposition linking aggregate leakage to average probe disparity and a corollary linking peak leakage to worst-case disparity. This is a strong methodological advance in moving fairness from black-box output metrics to white-box representation geometry. The proposed repair mechanism, ORCA, uses a residual adapter with orthogonal contrastive heads, which is a sound architectural choice for disentangling factors without destroying utility.
The experimental setup is extensive, utilizing a custom "TRIAD" grid of 28,800 synthetic utterances to control for confounding variables (content, affect, demographics). The evaluation covers 10 open-weights encoders and 2 closed-source models. The correlation between measured leakage and disparity (r=0.93) is a compelling empirical validation of the theory. The ablation studies effectively isolate the contribution of the orthogonality penalty. However, the reliance on a single TTS system for the grid introduces a significant validity threat, which the authors acknowledge but only partially mitigate with real-speech validation checks.
The paper claims that the grid, estimators, and code are available in the Supplementary Material. However, no specific URLs or repository links are provided in the text. The detailed hyperparameters (AdamW, lr=1e-3, 20k steps) and dataset construction details (Gemini 3.1 Flash TTS) are provided, which aids reproducibility if the supplementary materials are accessible. The use of a specific, potentially proprietary or rapidly changing TTS model (Gemini 3.1) may limit long-term reproducibility.
The primary limitation is the reliance on synthetic data for the main audit grid. While the authors perform validation on real speech (CANDOR, SSSD, etc.), the causal attribution of disparity to demographic attributes is weaker in synthetic data due to potential TTS artifacts. The theory is limited to linear probes; non-linear disparities are not bounded by the proposed leakage metrics. The institution "People Make Things" is obscure, raising questions about the resources and peer-review rigor compared to major academic or industrial labs.
This work provides a principled tool for diagnosing bias in speech models, which is critical as these models are deployed in high-stakes applications (healthcare, legal, customer service). The geometric perspective offers a new lens for understanding why certain models are biased, moving beyond correlation to structural entanglement. The ORCA adapter offers a practical, lightweight fix that can be applied to existing frozen models. The paper introduces a geometric framework linking subspace entanglement to demographic disparity in speech encoders, validated by a large-scale synthetic audit grid and a lightweight repair adapter. It makes a significant contribution to the theory of fairness in representation learning by providing provable bounds on disparity based on measurable geometric properties, although its reliance on synthetic data for primary validation limits the strength of its causal claims.
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Distillation (M$^{2}$-OPD) to transfer language capabilities and develop native audio skills. Act uses self-evolving executable environments and multi-granularity rollouts for Group Relative Policy Optimization (GRPO), teaching the model to use tools, interpret feedback, and complete tasks. Speak and Coordinate aligns how, when, and whether the assistant speaks or acts. We evaluate audio reasoning, multilingual understanding, tool use, conversational behavior, full-duplex interaction, and safety. Compared with Qwen-Audio-3.0-Realtime, 3.1 raises overall task success from 78.4% to 82.0% on our half-duplex speech-to-text adaptation of $ฯ$-Voice. On speech-to-speech Full-Duplex-Bench v1.5, the response rate to background speech falls from 73.0% to 13.0%. We also present a separate Voice Harness prototype, using Qwen-Audio-3.0-Realtime as its foreground, that extends spoken interaction to persistent tasks through foreground--background coordination and memory.
Primary: Alibaba Group
All Institutions: Alibaba Token Foundry, Alibaba Group
The paper presents a comprehensive framework for reliable agentic voice interaction, combining novel on-policy distillation techniques with self-evolving executable environments for tool use. It demonstrates significant improvements in task success, multilingual understanding, and conversational etiquette, particularly in handling background speech and full-duplex interactions, marking a substantial step forward in the development of practical, safety-conscious voice agents.
The paper proposes a comprehensive framework for real-time voice agents, structured around three layers: Think (foundation post-training), Act (agentic tool use), and Speak/Coordinate (conversational policy). The "Think" layer introduces a sophisticated post-training pipeline combining Core-Cocktail SFT with a novel "M$^2$-OPD" (Multimodality and Multi-Teacher On-Policy Distillation) strategy. This involves using a Text Teacher and a frozen Audio Reference to supervise student-generated trajectories, effectively transferring text-based reasoning capabilities to native audio models while preserving audio-specific nuances. The "Act" layer is particularly strong, introducing self-evolving executable environments for Group Relative Policy Optimization (GRPO). By using code agents to build, validate, and evolve tasks based on model performance (difficulty gating), the system learns robust tool use, state verification, and grounded progress communication. The "Speak/Coordinate" layer formalizes the decision-making process for when to speak, act, or remain silent, addressing critical issues in full-duplex interaction like background speech handling and turn-taking.
The evaluation is extensive, covering intelligence (audio reasoning, multilingual ASR, long-context), action (tool use, retrieval), interaction (persona, empathy, full-duplex behavior), and safety. Key results include a significant improvement in task success on the $\tau$-Voice benchmark (78.4% to 82.0%) and a dramatic reduction in response rate to background speech on Full-Duplex-Bench (73.0% to 13.0%), indicating much better conversational etiquette. The model also shows strong gains in multilingual audio understanding (BBA) and safety metrics. The inclusion of a "Voice Harness" prototype for persistent tasks adds depth to the system-level contribution.
The paper provides detailed descriptions of the training pipeline, data scales (approx. 1M hours), and evaluation protocols. However, as a technical report from a major industry lab, specific hyperparameters, exact dataset compositions, and code are not fully released, limiting independent reproduction. The reliance on in-house benchmarks (Long-AMC, WebSearch1K, VoiceChat) also restricts external verification.
The primary limitation is the lack of open-source code and data, making it difficult for the community to replicate the results. The evaluation relies heavily on in-house benchmarks and automatic judges (e.g., GPT-4o-mini, Qwen-Plus), which may introduce bias. The "Voice Harness" is presented as a prototype and is not fully integrated into the core model evaluation. Additionally, the paper focuses on a specific adaptation of $\tau$-Voice (half-duplex S2T) rather than the official full-duplex S2S protocol, which may limit direct comparability with other full-duplex systems.
This work significantly advances the state-of-the-art in real-time voice assistants by integrating robust agentic capabilities with natural conversational policies. The M$^2$-OPD technique and self-evolving environment for GRPO are valuable contributions to the broader field of multimodal LLM post-training. The focus on safety and reliable interaction (e.g., handling background speech, refusing unsupported requests) is crucial for the deployment of such systems in real-world scenarios. The paper presents a comprehensive framework for reliable agentic voice interaction, combining novel on-policy distillation techniques with self-evolving executable environments for tool use. It demonstrates significant improvements in task success, multilingual understanding, and conversational etiquette, particularly in handling background speech and full-duplex interactions, marking a substantial step forward in the development of practical, safety-conscious voice agents.
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can faithfully realize. While recent reasoning models have shown that intermediate "thinking" tokens improve output quality, this paradigm has been confined to the text modality. In this work, we extend reasoning to the audio token space by training a LALM with reinforcement learning to reason over its own speech output. The model first generates a draft speech as a form of audio-token reasoning, critiques its own generation by reflecting on the acoustic realization in text, and then produces a refined version conditioned on both the first-pass speech and the critique, all within a single model. After RL training, the refined two-hop outputs achieve a relative improvement of 7.15\% on the InstructTTSEval benchmark, demonstrating the model's reflective ability.
Primary: National Taiwan University
All Institutions: National Taiwan University, NVIDIA Research
The paper introduces a novel RL-based self-refinement framework for speech synthesis that enables a LALM to reason over its own audio output. By combining a generate-critique-refine pipeline with a specifically designed GRPO reward function that accounts for iterative improvement, the method achieves significant gains in instruction-following capability, demonstrating the viability of audio-token reasoning for enhancing generative quality.
The paper proposes a "Listen, Critique, and Refine" framework for instruction-following speech synthesis using Large Audio Language Models (LALMs). The core methodological contribution is extending the chain-of-thought/reasoning paradigm from text to audio tokens. The model generates a draft speech ($v_1$), generates a textual critique ($c$) by "listening" to $v_1$, and then generates a refined speech ($v_2$) conditioned on both. The training utilizes Group Relative Policy Optimization (GRPO) with a novel reward function. A key technical insight is the use of a non-linear transformation for the improvement term in the reward to prevent cancellation of the constant first-pass baseline within the GRPO group-relative advantage calculation. This is a sound and well-motivated technical detail that addresses a specific failure mode in applying GRPO to iterative refinement tasks.
The experiments are conducted on the InstructTTSEval benchmark, covering in-domain (DSD) and out-of-domain (APS, RP) tasks. The evaluation uses both objective metrics (CLSP score, WER) and subjective metrics (LALM-as-a-judge, human evaluation). The results show consistent improvements over zero-shot and RL one-hop baselines. The ablation studies effectively isolate the contributions of the non-linear reward shaping, speaker retrieval, and the critique mechanism. The human evaluation with 5 senior researchers adds credibility to the subjective gains. However, the training dataset is small (1000 examples from ParaSpeechCaps), which raises questions about the generalizability and robustness of the learned refinement behavior.
The authors provide a GitHub repository link, which is a positive step. The paper details the model architecture (Step-Audio-2-mini), hyperparameters, and implementation details (LoRA, vLLM, NCCL). However, the reliance on specific proprietary or less common models (Step-Audio-2-mini, CLSP) might limit reproducibility for the broader community compared to using more standard open-source components. The small training set size is a specific detail that is reproducible but may not be representative of large-scale applications.
The primary limitation is the small training dataset (1000 samples), which may lead to overfitting or limited generalization. The method relies on the base model's ability to generate meaningful critiques, which may not always be accurate or useful. The two-hop inference process doubles the computational cost compared to single-pass generation. The evaluation is limited to English speech. The reliance on CLSP for style adherence may not capture all nuances of complex instructions.
This work demonstrates the potential of self-refinement and reasoning in audio generation, a promising direction for improving the controllability and quality of LALMs. The technique of using non-linear reward shaping to handle iterative refinement under GRPO could be applicable to other multi-step generation tasks in other modalities. The framework provides a blueprint for integrating understanding and generation in a unified reasoning loop. The paper introduces a novel RL-based self-refinement framework for speech synthesis that enables a LALM to reason over its own audio output. By combining a generate-critique-refine pipeline with a specifically designed GRPO reward function that accounts for iterative improvement, the method achieves significant gains in instruction-following capability, demonstrating the viability of audio-token reasoning for enhancing generative quality.
Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio representations into a large language model (LLM) backbone to enable multimodal reasoning. Recent test-time reinforcement learning (TTRL) methods further improve LLM reasoning capability by leveraging unlabelled test data after pre-training. However, the importance of the perceptual capability of LALMs remains underexplored, particularly how much acoustic evidence is integrated and relied upon during reasoning, and how this contributes to final task performance. This gap limits the development of effective post-training methods like TTRL for audio reasoning. In this work, we first analyse how audio information is integrated and utilised during reasoning process. We quantify layer-wise perceptual reliance and show that stronger acoustic reliance is associated with higher accuracy and a larger performance gain attributable to the audio input. Building on this, we propose Perception-Grounded TTRL (PG-TTRL), which aligns label-free test-time optimisation with perceptually grounded reasoning, encouraging the model to structure its reasoning more strongly on the audio input. Experiments across LALMs and benchmarks show that PG-TTRL consistently improves reasoning performance over both the base models and standard TTRL, showing the value of perceptual-grounding optimisation for test-time audio reasoning.
Primary: University of Maryland, College Park
All Institutions: University of Maryland, College Park, University of Illinois Urbana-Champaign, University of Washington
The paper presents a rigorous analysis of perceptual reliance in LALMs and proposes PG-TTRL, a novel test-time reinforcement learning method that aligns policy optimization with acoustic grounding, resulting in consistent performance gains over standard TTRL and base models.
The paper introduces a novel diagnostic framework for Large Audio-Language Models (LALMs) that quantifies "perceptual reliance" by comparing hidden representations under full audio access versus masked audio attention. This analysis reveals that acoustic evidence integration peaks in intermediate layers and correlates with task accuracy. Building on this, the authors propose PG-TTRL (Perception-Grounded Test-Time Reinforcement Learning), which modifies the Group Relative Policy Optimization (GRPO) advantage function. Instead of relying solely on majority vote correctness, PG-TTRL computes a trajectory-level grounding score based on the log-probability gap between original and masked audio conditions. This score is used to derive a reliability weight that scales the advantage, effectively penalizing updates that rely on language priors rather than acoustic evidence. The method is theoretically sound, addressing a critical flaw in standard TTRL where models may reinforce linguistic shortcuts. The integration of causal intervention (attention masking) into the reward signal is a creative and rigorous approach to grounding multimodal reasoning.
Experiments are conducted on two state-of-the-art LALMs (Qwen2.5-Omni-3B and 7B) across two major benchmarks (MMAR and MMAU). The results show consistent improvements over both base models and standard TTRL, with gains up to 4.9 points on MMAU. The analysis of Pass@k metrics is particularly compelling, demonstrating that PG-TTRL preserves diverse correct trajectories better than standard TTRL, which tends to collapse to incorrect majority votes on hard instances. The ablation study on layer-wise changes confirms that the method specifically enhances late-stage reasoning utilization of audio, validating the hypothesis. However, the evaluation is limited to multiple-choice tasks, and the sample size for correlation analysis is modest.
The paper provides detailed hyperparameters, including LoRA ranks, learning rates, and specific bounds for the reliability weight. The algorithm is clearly defined, and the reproducibility checklist is filled out, indicating that code will be released. The use of standard benchmarks and open-source models (Qwen) enhances reproducibility. However, the specific implementation of the attention masking for the grounding score requires careful handling of the model's internal attention mechanisms, which may be non-trivial to replicate without the provided code.
The primary limitation is the restriction to multiple-choice audio reasoning tasks; the method's applicability to open-ended generation or other modalities is not explored. The grounding score relies on teacher-forcing passes, which may not perfectly reflect the autoregressive generation process. Additionally, the method adds computational overhead due to the extra forward passes required for grounding score calculation. The correlation analysis excludes small-sample tasks, which may limit the generalizability of the findings.
This work addresses a fundamental issue in multimodal AI: the tendency of models to ignore perceptual inputs in favor of language priors. By providing a mechanism to explicitly reward perceptual grounding, the paper offers a pathway to more robust and trustworthy multimodal systems. The insights into layer-wise perceptual reliance are valuable for the broader community developing multimodal architectures. The method could be extended to other modalities (vision, video) and tasks, potentially improving the reliability of AI systems in safety-critical applications where accurate perception is crucial. The paper presents a rigorous analysis of perceptual reliance in LALMs and proposes PG-TTRL, a novel test-time reinforcement learning method that aligns policy optimization with acoustic grounding, resulting in consistent performance gains over standard TTRL and base models.
Constraint-following music generation asks a score to satisfy several user-specified properties at once, each checkable programmatically (key, meter, length, range, final note, rhythm, motion and form), yet no existing benchmark isolates this capability. We construct MusicConstraintBench, 2,180 items over eight constraint families, on which current models fail once a few constraints are combined. The natural remedy is reinforcement learning with these verifiers as reward, yet we observe that a reward paid only when every property holds leaves most training groups without a learning signal: over the first 50 updates, 0.550 of rollout groups score identically and receive no gradient, even though a failing score typically misses only one requested property. Under the joint criterion, rollouts for a prompt tend to fail together, so a binary reward cannot separate a nearly correct score from a malformed one. We therefore introduce MusicRLVR, which pays graded per-property credit behind a hard validation gate that rejects malformed outputs, plus a joint-satisfaction bonus, requiring no human annotation, learned reward model, or music-domain fine-tuning. On MusicConstraintBench, MusicRLVR lifts Qwen3-4B-Instruct from 0.160 to 0.807 on mixed constraints and leads every zero-shot baseline including Llama-3.1-70B at 0.380. It also generalises to property combinations unseen in training and to out-of-range parameter values, showing that verifiable rewards need not presuppose a target output.
Primary: Anonymous (Double-Blind Review)
All Institutions: Anonymous
The paper introduces a verifier-driven reinforcement learning framework and benchmark for constraint-following symbolic music generation, demonstrating that graded, deterministic rewards can significantly improve multi-constraint compliance in large language models without requiring music-domain supervised fine-tuning.
The paper proposes MusicRLVR, a reinforcement learning framework that uses deterministic, programmatic verifiers as rewards for symbolic music generation. The core methodological contribution is the design of a graded reward function that combines a hard validity gate (rejecting malformed ABC notation) with per-family partial credit and a joint-satisfaction bonus. This addresses the "sparse reward" problem in GRPO where binary success/failure signals lead to zero-variance groups and no gradient updates. The approach is technically sound, leveraging the fact that musical constraints (key, meter, form, etc.) are programmatically checkable. The use of GRPO (Group Relative Policy Optimization) is appropriate for this setting, and the ablation studies effectively demonstrate the necessity of the graded credit component over binary rewards.
The authors introduce MusicConstraintBench, a new benchmark of 2,180 items covering eight constraint families. The experiments are rigorous, comparing the proposed method against strong zero-shot baselines (including Llama-3.1-70B) and supervised fine-tuning (SFT) baselines. The results show significant improvements in constraint compliance, particularly on multi-constraint tasks where baselines fail. The evaluation includes tests for compositional generalization (unseen constraint combinations) and parameter extrapolation, which are critical for assessing the robustness of the learned policy. The statistical significance is supported by McNemar tests and Holm correction.
The paper provides extensive details on the training setup, hyperparameters, and data construction. It explicitly states that code, benchmark suites, and model outputs will be released under an MIT license. The reproducibility statement is detailed, mentioning specific scripts for asset generation and evaluation. The use of standard tools like vLLM and ms-swift enhances reproducibility.
The primary limitation is the domain specificity; the method is tailored to symbolic music in ABC notation and may not directly transfer to other domains without significant adaptation. Additionally, the paper acknowledges that compliance scores do not measure perceptual musical quality, which is a significant gap for practical applications. The reliance on a specific corpus (IrishMAN) for SFT data might introduce stylistic biases, although the RL phase is designed to be independent of this.
This work demonstrates that RLVR (Reinforcement Learning with Verifiable Rewards) can be applied to open-ended generation tasks where the output space is large but the properties of interest are verifiable. This has implications for other structured generation tasks (e.g., code, math, structured data) where deterministic checks are possible. It provides a template for building benchmarks and reward functions for property-constrained generation. The paper introduces a verifier-driven reinforcement learning framework and benchmark for constraint-following symbolic music generation, demonstrating that graded, deterministic rewards can significantly improve multi-constraint compliance in large language models without requiring music-domain supervised fine-tuning.