Audio ML Papers

Last 7 Days (September 20 - September 27, 2026)

Subcategories: All (19) | Speech Synthesis (2) | Music Synthesis (0) | Ambient Synthesis (0) | Quality Evaluation (0) | Enhancement (2) | Asr (0) | Llm Audio (8) | Midi Generation (1) | Generative Conditioning (0) | Other (6)
← Previous Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 81)
Yuxiang Wang, Shengbo Cai, Yingda Shen ... ยท The Chinese University of Hong Kong, Shenzhen +6 ยท arXiv
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefo...
#2 TOP PAPER (Score: 80)
Lujia Bao, Qian Chen, Luyao Cheng ... ยท Alibaba Group +1 ยท arXiv
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multi...
#3 TOP PAPER (Score: 79)
Jihoo Jung, Youngjoon Jang, โ˜… Joon Son Chung ยท KAIST +2 ยท NeurIPS 2026
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate ho...
Saturday, September 26, 2026
Francesco Brigante, Luca Cerovaz, Davide Marincione ... ยท Sapienza University of Rome +3 ยท arXiv
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference spee...
Ambuj Mehrish, Abhinaba Roy, Alex Ivanov ... ยท Ca' Foscari University of Venice +3 ยท arXiv
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identi...
Yang Xiao, Vidhyasaharan Sethu, Eun-Jung Holden ... ยท arXiv
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be r...
Friday, September 25, 2026
Jihoo Jung, Youngjoon Jang, โ˜… Joon Son Chung ยท KAIST +2 ยท NeurIPS 2026
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate ho...
Yotaro Kubo, Qi Sun, Yujin Tang ยท Sakana AI ยท arXiv
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LL...
Yejin Lee, Seungbeom Kim, Yongha Lee ... ยท Sungkyunkwan University ยท arXiv
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time sla...
Thursday, September 24, 2026
Rayhan Rashed, Senja Filipi, Ross Cutler ยท Microsoft +1 ยท arXiv
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested....
Benjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos ... ยท Purdue University +2 ยท arXiv
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-...
Alessandro Bondielli, Lucia Passaro, Serena Auriemma ... ยท University of Pisa ยท EMNLP 2026
Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test case to study whether morpho-syntactic informati...
Hongyao Deng, Wenhao Guan, Xuetao Lin ... ยท Xiamen University ยท arXiv
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Ed...
Wednesday, September 23, 2026
Yuxiang Wang, Shengbo Cai, Yingda Shen ... ยท The Chinese University of Hong Kong, Shenzhen +6 ยท arXiv
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefo...
Matthew Sun, Vinay Kothapally, Meng Yu ... ยท Tencent Americas +2 ยท arXiv
Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited c...
Kian Shamsaie, Iman Modarressi ยท People Make Things ยท IEEE SLT 2026
Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap,...
Kian Shamsaie, Iman Modarressi ยท People Make Things ยท IEEE SLT 2026
End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipulation. We formalize imperceptible attacks on full-duplex agents as optimization over additive perturbations confined beneath the p...
Kian Shamsaie, Iman Modarressi ยท People Make Things ยท IEEE SLT 2026
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band...
Monday, September 21, 2026
Lujia Bao, Qian Chen, Luyao Cheng ... ยท Alibaba Group +1 ยท arXiv
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multi...
Chee-En Yu, Yi-Cheng Lin, Sung-Feng Huang, โ˜… Hung-yi Lee ... ยท National Taiwan University +1 ยท arXiv
Large Audio Language Models (LALMs) can follow diverse instructions to synthesize speech in specified styles. However, complex instructions that require simultaneous control over pitch dynamics, speaking rate, and emotional tone often exceed what a single-pass generation can fait...
Sunday, September 20, 2026
Jiaheng Dong, Xiaofeng Yu, Jean Honorio ... ยท University of Maryland, College Park +4 ยท AAAI 2027 (Submission)
Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio representations into a large language model (LLM) backbone to enable multimodal reasoning. Recent test-time reinforcement learning (TTR...
Haoyue Liu, Ye Chen, Zhichao Wang ... ยท Anonymous (Double-Blind Review) +1 ยท ICLR 2027 (inferred from template `iclr2027_conference`)
Constraint-following music generation asks a score to satisfy several user-specified properties at once, each checkable programmatically (key, meter, length, range, final note, rhythm, motion and form), yet no existing benchmark isolates this capability. We construct MusicConstra...