Audio ML Papers

Last 7 Days (September 13 - September 20, 2026)

Subcategories: All (16) | Speech Synthesis (1) | Music Synthesis (2) | Ambient Synthesis (0) | Quality Evaluation (1) | Enhancement (0) | Asr (3) | Llm Audio (7) | Midi Generation (0) | Generative Conditioning (0) | Other (2)
← Previous Week

๐Ÿ† Top Papers This Week

#1 TOP PAPER (Score: 82)
โ˜… Haohe Liu, Varun Nagaraja, Gael Le Lan, โ˜… Abdelrahman Mohamed ... ยท Meta AI ยท arXiv
Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in w...
#2 TOP PAPER (Score: 78)
Yanfeng Shi, Yan Song, Junhui Li ... ยท University of Science and Technology of China +1 ยท arXiv
Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as time...
#3 TOP PAPER (Score: 77)
Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein ... ยท University of Wisconsin-Madison +1 ยท arXiv
What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some o...
Saturday, September 19, 2026
Nishit Anand, Jiaqi Su, Ke Chen, โ˜… Rithesh Kumar ... ยท University of Maryland +2 ยท Interspeech 2026
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralingu...
Jing Peng, Zichao Nie, Zhisheng Zhang ... ยท Tsinghua University ยท INTERSPEECH 2026
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated. Existing benchmarks often focus on narrow domains or evaluate encoders outside LALM contexts. To add...
Friday, September 18, 2026
Haolin He, Yunfei Chu, Qi Chen ... ยท Alibaba Group +4 ยท arXiv
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, with...
Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang ... ยท Meta Reality Labs ยท Interspeech 2026
Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearab...
Piotr Masztalski, Michaล‚ K. Grzeszczyk, Olaf Sikorski ยท Samsung R&D Institute Poland +1 ยท Interspeech 2026
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-dev...
Thursday, September 17, 2026
Thomas J Stoll, Ross K Maddox ยท University of Michigan +2 ยท arXiv
Computational models of auditory physiology commonly target specific responses or stages of the auditory pathway, limiting their ability to integrate findings across experimental paradigms and neural timescales. We present a foundation model of human auditory electrophysiology: a...
Zifan Guan, Longyu Lu, Junan Zhang ... ยท The Chinese University of Hong Kong +2 ยท arXiv
Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspe...
Wednesday, September 16, 2026
Xiuwen Zheng ยท University of Illinois Urbana-Champaign ยท arXiv
Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $ฯ„$ that boun...
Mohan Shi, Zilai Wang, Natarajan Balaji Shankar ... ยท University of California Los Angeles ยท IEEE SLT 2026
Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting t...
Tuesday, September 15, 2026
Rongxiang Wang, Berkin Durmus, Aysegul Orhon ... ยท Argmax, Inc. +4 ยท arXiv
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and spea...
Monday, September 14, 2026
โ˜… Haohe Liu, Varun Nagaraja, Gael Le Lan, โ˜… Abdelrahman Mohamed ... ยท Meta AI ยท arXiv
Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in w...
Yanfeng Shi, Yan Song, Junhui Li ... ยท University of Science and Technology of China +1 ยท arXiv
Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as time...
Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein ... ยท University of Wisconsin-Madison +1 ยท arXiv
What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some o...
Daxin Tan, Dehua Tao, Chengxi Deng ... ยท Huawei (AI Lab, Leibniz Research Center) +5 ยท arXiv
Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables streaming generation with explicit textual guidance, generated acoustic tokens become part of the con...
Sunday, September 13, 2026
Tingzhen Xiong, Rilin Chen, Weiwei Li ... ยท Tencent ยท IEEE Spoken Language Technology Workshop (SLT) 2026
When reinforcement learning (RL) is used for post-training automatic speech recognition (ASR), the reward almost always lives in the text space: it compares a hypothesis with the reference and never checks whether the hypothesis is supported by the audio. On highly regular speech...
Quoc-Huy Trinh, Minh-Van Nguyen, Debesh Jha ยท Aalto University +3 ยท arXiv
Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large...