Audio ML Papers

Last 7 Days (September 09 - September 16, 2026)

Subcategories: All (16) | Speech Synthesis (2) | Music Synthesis (2) | Ambient Synthesis (2) | Quality Evaluation (0) | Enhancement (1) | Asr (1) | Llm Audio (4) | Midi Generation (0) | Generative Conditioning (0) | Other (4)
← Previous Week | Current Week

🏆 Top Papers This Week

#1 TOP PAPER (Score: 82)
★ Haohe Liu, Varun Nagaraja, Gael Le Lan, ★ Abdelrahman Mohamed ... · Meta AI · arXiv
Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in w...
#2 TOP PAPER (Score: 81)
Bin Lin, Bo Zhao, Boyang Wang ... · StepFun · arXiv
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is...
#3 TOP PAPER (Score: 80)
Zhongjie Duan, Shengchuan Gao, Hong Zhang ... · Alibaba Group +1 · arXiv
Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music (https://modelscope.cn/models/DiffSynth-Studio/DiffSynth-Music), a framework that adds composable ...
Tuesday, September 15, 2026
Rongxiang Wang, Berkin Durmus, Aysegul Orhon ... · Argmax, Inc. +4 · arXiv
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and spea...
Monday, September 14, 2026
★ Haohe Liu, Varun Nagaraja, Gael Le Lan, ★ Abdelrahman Mohamed ... · Meta AI · arXiv
Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in w...
Yanfeng Shi, Yan Song, Junhui Li ... · University of Science and Technology of China +1 · arXiv
Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as time...
Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein ... · University of Wisconsin-Madison +1 · arXiv
What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some o...
Daxin Tan, Dehua Tao, Chengxi Deng ... · Huawei (AI Lab, Leibniz Research Center) +5 · arXiv
Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables streaming generation with explicit textual guidance, generated acoustic tokens become part of the con...
Sunday, September 13, 2026
Tingzhen Xiong, Rilin Chen, Weiwei Li ... · Tencent · IEEE Spoken Language Technology Workshop (SLT) 2026
When reinforcement learning (RL) is used for post-training automatic speech recognition (ASR), the reward almost always lives in the text space: it compares a hypothesis with the reference and never checks whether the hypothesis is supported by the audio. On highly regular speech...
Quoc-Huy Trinh, Minh-Van Nguyen, Debesh Jha · Aalto University +3 · arXiv
Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large...
Saturday, September 12, 2026
Bin Lin, Bo Zhao, Boyang Zhang ... · StepFun · arXiv
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Perception captures rich acoustic cues to interpret ...
Ruixiang Zhao, Hualei Wang, Renhe Sun ... · Ant Group +1 · arXiv
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-V...
Bojro Das · Cornell University · arXiv
Asked to describe what one of six speakers in a recording talks about, audio language models describe the right one on 6 to 16% of trials, below the 16.7% a guess would give. Adding a fixed bias to the attention logits of a hundred heads, under a tenth of the model's and with no ...
Friday, September 11, 2026
Bin Lin, Bo Zhao, Boyang Wang ... · StepFun · arXiv
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is...
Zhongjie Duan, Shengchuan Gao, Hong Zhang ... · Alibaba Group +1 · arXiv
Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music (https://modelscope.cn/models/DiffSynth-Studio/DiffSynth-Music), a framework that adds composable ...
Chengli Feng, Zhiyue Wu, Jiahao Song ... · StepFun +3 · arXiv
We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantica...
Thursday, September 10, 2026
Luca Della Libera, Cem Subakan, ★ Mirco Ravanelli · Mila-Quebec AI Institute +2 · arXiv
Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present Z...
Wednesday, September 09, 2026
Yupei Li, Qiyang Sun, Mohamed Mady ... · Imperial College London +4 · AACL 2026
Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than g...
Mingyu Zhao, Zhiyong Wu · Tsinghua University +1 · NCMMSC 2026
We present UniStream, a fully causal 48 kHz neural audio codec for streaming speech, music, and environmental sounds. At its core is Multi-Expert Residual Vector Quantization (ME-RVQ), which replaces the single shared codebook in each residual quantization layer with four expert ...