arXiv 每日论文精读

📡 eess.AS / cs.SD
Audio and Speech Processing, Sound
2026年07月15日
LLM: deepseek-v4-pro
29
论文总数
16
跨领域
29
成功解读
0
待处理
#1
eess.AScs.SD

The Sound of Absence: Audio-Language Embedding Models Struggle with Negation 解读失败跨领域

Chun-Yi Kuan, Hung-yi Lee
Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
Comments: Manuscript in progress
查看摘要
Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated captions to nearly identical representations. To expose this blind spot, we introduce NegEval-Audio, a framework that converts existing datasets into two negation-aware tasks, Retrieval-Neg and Multiple-Choice Negation (MCQ-Neg), to probe whether models distinguish present from absent events. On AudioCaps and Clotho, performance degrades sharply under negation, with negation-type MCQ accuracy falling far below chance, and the failure persists even for a recent multimodal LLM-based embedding model. While a training-free steering method improves MCQ-Neg, it yields marginal gains for Retrieval-Neg. This indicates that affirmation bias is a fundamental flaw in the representation geometry, necessitating explicit negation-aware training objectives.

📖 深度解读

[PDF 下载失败,无法解读]

#2
eess.AS

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching 解读失败

Jihwan Kim, Nam Soo Kim
Audio and Speech Processing (eess.AS)
Comments: Accepted to Interspeech 2026. 5 pages, 1 figure, 3 tables
查看摘要
Zero-shot dialog TTS benefits from flow-matching, but minute-scale generation on dense mel-spectrograms causes severe memory bottlenecks, often forcing unnatural chunked synthesis. We propose ZipL-Dialog, which shifts conditional flow-matching into a 4x time-compressed (25 Hz) latent space. To preserve acoustic fidelity under compression, we employ a deterministic mel autoencoder with auxiliary mel-domain supervision and optimize the ZipFormer's hierarchical downsampling schedule. Experiments show that ZipL-Dialog reduces maximum peak GPU memory by 11.22x and accelerates inference by 2.23x over the baseline, substantially lowering the memory footprint of single-pass multi-minute dialog synthesis while maintaining perceptual naturalness.

📖 深度解读

[PDF 下载失败,无法解读]

#3
eess.AS

Listen first: Output-based multi-microphone speech enhancement 解读失败

Panos Apostolidis, Svend Feldt, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen
Audio and Speech Processing (eess.AS)
Comments: Accepted at the International Workshop on Acoustic Signal Enhancement (IWAENC) 2026
查看摘要
Traditionally, hearing-aid speech enhancement (SE) algorithms rely on input-based feature estimation, often derived by a voice activity detector (VAD), to configure beamformers. Yet features extracted from noisy microphone signals can become unreliable in challenging acoustic scenes where users most need help. We introduce a novel paradigm in which the settings of a sound processing system are determined by evaluating characteristics of its output. To demonstrate this idea, we employ an output-based system that selects among a set of minimum power distortionless response (MPDR) beamformers. Although MPDR beamformers are typically avoided due to their sensitivity to steering errors, we show that they become effective within an output-based framework. We compare the proposed system to a conventional input-based minimum variance distortionless response (MVDR) baseline. Experimental results show that the proposed system consistently outperforms the MVDR baseline, particularly at low SNRs, in terms of SNR, ESTOI and PESQ.

📖 深度解读

[PDF 下载失败,无法解读]

#4
eess.AS

Investigating the Integration of Spatial Information in Foundation-Model-Based Speaker Diarization 解读失败

Marc Deegen, Adrian Meise, Reinhold Haeb-Umbach
Audio and Speech Processing (eess.AS)
Comments: Accepted at IWAENC 2026
查看摘要
Spatial information gleaned from multi-channel input has been shown to lead to improvements in meeting processing tasks like diarization and source separation. At the same time, diarization based on features extracted by large pretrained single-channel foundation models, such as WavLM, achieved state-of-the-art performance. This work compares three approaches to integrate spatial features into foundation model-based diarization systems: the cascade of a beamformer and a single-channel foundation model, a multi-channel foundation model, and the conditioning of the downstream network on explicitly extracted spatial features. Results show that the beamformer front-end is even detrimental to diarization performance in regions of overlapped speech, while best performance is achieved with the conditioning, demonstrating that the incorporation of explicit spatial features is a competitive approach to foundation-model-supported diarization. This approach is further subjected to a detailed error analysis showing that the conditioning system removes errors to a good extent that would occur when either only spectral or only spatial features were used.

📖 深度解读

[PDF 下载失败,无法解读]

#5
eess.AS

Audio Diarization: A New Paradigm for Exploring Audio Recordings with Unknown Event Classes 解读失败

Alexander Werning, Reinhold Haeb-Umbach
Audio and Speech Processing (eess.AS)
Comments: accepted at IWAENC 2026
查看摘要
We propose a new task, audio diarization. The motivation is that there are applications, such as audio monitoring in an unknown environment, where initially the sound event classes to be recognized are unknown. For such a scenario, we propose to first localize in time relevant sound events and to classify them, e.g., by comparing with known event classes, in a second step. This contribution is dedicated to the first step, which we call audio diarization, as it is reminiscent of the speaker diarization stage that precedes and simplifies the second stage, speech recognition, in multi-talker conversational speech processing. In this contribution, we define audio diarization as detecting onset and offset times of sound events with overlap for an open set of classes and without user prompts. We show how a speaker diarization system can be adjusted for audio diarization and propose an evaluation setup. Compared to a closed-set sound event detection system, the proposed system achieves similar performance with the additional ability to detect novel sounds.

📖 深度解读

[PDF 下载失败,无法解读]

#6
eess.AS

Spatial-Frequency Cued Generative Fixed-Filter Active Noise Control Based on Deep Learning in Reverberant Environments 解读失败

Boxiang Wang, Haowen Li, Dongyuan Shi, Junwei Ji, Ziyi Yang 等 (7 人)
Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
查看摘要
Generative fixed-filter active noise control (GFANC) effectively attenuates noise with diverse frequency characteristics through the combination of sub control filters. However, it does not incorporate the spatial information of the noise source, which limits its performance, particularly in reverberant environments. To address this limitation, this paper proposes a novel spatial-frequency cued GFANC (SF-GFANC) method that exploits both three-dimensional (3D) spatial and frequency information of the noise source. Specifically, a multi-task convolutional recurrent neural network (CRNN) is designed to estimate the source distance, elevation angle, and azimuth angle as spatial cues, while predicting the combination weights of sub control filters as frequency cues. These spatial-frequency cues jointly guide the generation of the appropriate control filter. In addition, a theoretical analysis of the optimal control filter in reverberant environments is presented, highlighting the importance of 3D spatially conditioned control filter design. Evaluations using both simulated and measured acoustic paths demonstrate that the CRNN is robust to unseen acoustic environments and noise types. Furthermore, the results confirm that SF-GFANC outperforms representative ANC algorithms when handling noise sources across diverse 3D locations and frequency characteristics in reverberant environments.

📖 深度解读

[PDF 下载失败,无法解读]

#7
eess.AS

Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification 解读失败跨领域

Pravina Mylvaganam, Ting Dang, Eliathamby Ambikairajah, Vidhyasaharan Sethu, Jingyao Wu
Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
Comments: Accepted by Interspeech 2026
查看摘要
Language identification is an important step toward integrating endangered Australian Aboriginal languages (AALs) into speech technologies supporting language revitalisation and digital inclusion. However, extreme data scarcity limits model performance. Transfer learning from high-resource languages shows promise but often suffers from catastrophic forgetting when adapting to new languages. Continual learning (CL) can mitigate this issue, though it remains challenging with very limited data. To address this, we propose two hybrid continual learning methods: Replay Augmented Elastic Weight Consolidation and Constraint Guided Knowledge Distillation to adapt pretrained speech models for AAL identification while preserving previously learned knowledge. Experiments on Warlpiri, Dalabon and Dharawal show that the proposed methods outperform fine-tuning and existing CL baselines, improving adaptation to multiple AALs while maintaining performance on previously learnt high-resource languages.

📖 深度解读

[PDF 下载失败,无法解读]

#8
eess.AScs.SD

PolarBM: Complex-valued Boltzmann Machine for Modeling Audio Signals in Polar and Log-polar Coordinates 解读失败跨领域

Toru Nakashika, Kohei Yatabe
Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
Comments: Submitted to IEEE Trans. ASLP
查看摘要
Although vast amounts of data, such as audio signal spectra, are naturally represented using complex numbers, conventional machine learning methods often simplify complex-domain problems by employing frameworks designed for real-valued variables. While this simplification offers computational benefits, it discards structural information regarding the inherent relationship between amplitude and phase. In this paper, we propose a novel Boltzmann machine (BM), named PolarBM, capable of naturally handling complex-valued variables in the polar coordinate (i.e., an amplitude-phase representation). PolarBM defines a probability density function for complex variables in which the phase explicitly depends on the amplitude, thereby capturing the physically important relationships of complex-valued signals. Furthermore, to process audio signals in accordance with human auditory perception, we propose LogPolarBM, which models amplitude on a logarithmic scale. This extension yields a flexible conditional probability density function, a power-weighted noncentral complex Gaussian (PW-NCCG) distribution, whose marginal amplitude distribution encompasses the Rice, Nakagami, and noncentral chi distributions as special cases. For practical applications, we also introduce the restricted variants of these proposed models: PolarRBM and LogPolarRBM. Experimental results demonstrate that by explicitly modeling the dependency between amplitude and phase, the proposed RBMs achieve superior modeling accuracy compared to conventional models, including deep neural networks. Although our experiments focus on audio signals, the utility of the proposed BMs is not limited to audio applications; their potential extends widely across various fields of science and engineering that involve complex-valued data, such as wireless communications and quantum mechanics.

📖 深度解读

[PDF 下载失败,无法解读]

#9
eess.AScs.SD

DiffAU: Diffusion-Based Ambisonics Upscaling 跨领域

Amit Milstein, Nir Shlezinger, Boaz Rafaely
Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)
查看摘要
Spatial audio enhances immersion by reproducing 3D sound fields, with Ambisonics offering a scalable format for this purpose. While first-order Ambisonics (FOA) notably facilitates hardware-efficient acquisition and storage of sound fields as compared to high-order Ambisonics (HOA), its low spatial resolution limits realism, highlighting the need for Ambisonics upscaling (AU) as an approach for increasing the order of Ambisonics signals. In this work we propose DiffAU, a cascaded AU method that leverages recent developments in diffusion models combined with novel adaptation to spatial audio to generate 3rd order Ambisonics from FOA. By learning data distributions, DiffAU provides a principled approach that rapidly and reliably reproduces HOA in various settings. Experiments in anechoic conditions with multiple speakers, show strong objective and perceptual performance.

📖 深度解读

好的,我将以资深学术论文解读专家的身份,为您解读这篇论文。

1. 一句话总结

这篇论文提出了一种名为 DiffAU 的新方法,利用扩散模型将低分辨率的一阶 Ambisonics 音频信号“超分辨率”生成为高分辨率的三阶信号,从而用更便宜的设备获得更逼真的 3D 空间音频体验。

2. 研究背景与动机

  • 核心问题:如何将低阶 Ambisonics(如 4 通道的 FOA)信号提升为高阶 Ambisonics(如 16 通道的 HOA)信号,即 Ambisonics 上变换(AU)问题。
  • 问题的重要性:高阶 Ambisonics 能提供更精细的空间分辨率和沉浸感,但其录制需要昂贵庞大的麦克风阵列。如果能从硬件简单、成本低廉的低阶信号中恢复出高阶信号,就能大幅降低高质量空间音频的制作门槛,对 VR/AR、游戏、影视等行业意义重大。
  • 现有方法的不足
    • 基于模型的方法(如压缩感知):依赖于“声源在空间中稀疏分布”的物理假设。在理想自由场下有效,但在真实混响、多声源等复杂环境中,这个假设不成立,性能会急剧下降。
    • 数据驱动的方法(如卷积神经网络):虽然不依赖稀疏假设,但现有网络(如 Conv-TasNet)在感知听感测试中,与真实高阶信号仍有明显差距,缺乏生成式模型那种对数据分布进行建模的能力。

3. 核心方法

  • 方法/模型名称:DiffAU,一个基于扩散模型的级联式 Ambisonics 上变换框架。
  • 关键创新点
    1. 将 AU 问题重新定义为条件生成任务:不再将其视为一个确定性映射,而是学习在给定低阶信号条件下,高阶信号的概率分布。这为解决 AU 问题的“病态性”(一个低阶信号可能对应多种高阶信号)提供了更合理的概率框架。
    2. 为空间音频定制的扩散模型:将扩散模型适配到 Ambisonics 信号格式,包括使用特定的时频变换(STFT)和非线性幅度变换来归一化信号,使其适合神经网络处理。
    3. 级联式上变换架构:不是一步到位从 1 阶生成 3 阶,而是通过两个串联的扩散模块,逐阶生成(1→2→3)。这种模块化设计更灵活,可扩展到任意阶数,且每个模块任务更简单。
  • 核心思路直觉解释
    你可以把一阶 Ambisonics 想象成一张分辨率很低的模糊照片,而三阶 Ambisonics 是一张细节丰富的高清照片。传统方法试图通过猜测(如假设照片里只有几个物体)来“算”出高清图,但常常猜错。DiffAU 的方法则像一个顶尖的“AI 图像修复师”,它先学习了无数张“高清照片”长什么样。当给它一张“模糊照片”时,它不是直接计算,而是从一个充满噪点的“纯噪声图”开始,一步步地、有引导地去噪,每一步都参考那张“模糊照片”作为条件,最终“生成”出一张既清晰、又与模糊照片内容一致的“高清照片”。级联的方式则是先修复到中等分辨率,再基于中等分辨率图修复到高分辨率,分步完成。

4. 实验与结果

  • 数据集:基于 WSJ0 语音语料库,在自由场(无混响)条件下,随机生成 1 到 4 个说话人、不同方位的 3 阶 Ambisonics 信号。
  • 对比基线
    • PWD CS:基于压缩感知的平面波分解方法(模型驱动)。
    • Conv-TasNet AU:基于 Conv-TasNet 架构的波形域编解码网络(数据驱动)。
  • 主要实验结果
    • 客观指标(STFT-SDR,越高越好):DiffAU 在所有声源数量下都显著优于基线。总体平均得分 24.7 dB,远超 Conv-TasNet 的 14.0 dB 和 PWD CS 的 12.6 dB。即使在稀疏假设成立的自由场,DiffAU 也表现最佳。
    • 主观听感测试(MUSHRA 评分):DiffAU 生成的三阶信号(79.0 分)与真实的三阶信号(78.4 分)在统计上没有显著差异,听感几乎一致,远高于一阶锚点信号(21.8 分)。
  • 消融实验:论文未提供传统意义上的消融实验(如移除某个模块),但通过对比不同方法,间接证明了扩散模型和级联架构的有效性。

5. 优势与局限

  • 本文方法的主要优势
    1. 性能卓越:在客观和主观指标上均大幅超越现有方法,达到了与真实高阶信号难以区分的感知质量。
    2. 无显式假设:不依赖声源稀疏性等物理假设,通过学习数据分布来解决问题,泛化能力更强。
    3. 架构灵活:级联式设计模块化,便于扩展到更高阶数的上变换,也支持逐阶训练和解释。
  • 局限性
    1. 场景受限:所有实验仅在自由场(无混响)、无噪声的条件下进行。论文明确指出,在更真实的混响和噪声环境下的性能有待验证。
    2. 计算成本:扩散模型通常需要多步迭代去噪,推理速度可能慢于前馈式神经网络(如 Conv-TasNet),论文未讨论实时性。
    3. 训练数据依赖:作为数据驱动方法,其性能依赖于训练数据的覆盖度,对于未见过的复杂声学环境可能表现不佳。

6. 关键结论与启发

  • 最重要的 Takeaway:扩散模型这种生成式方法,为解决 Ambisonics 上变换这类病态逆问题提供了强大且有效的新范式,其性能在理想条件下已能“以假乱真”。
  • 对后续研究的启发或延伸方向
    1. 向真实世界扩展:最直接的延伸是将 DiffAU 应用于带混响和噪声的真实录音环境,这需要构建相应的数据集并可能调整模型结构。
    2. 加速推理:研究如何利用扩散模型的加速采样技术(如 DDIM)或模型蒸馏,使其满足实时音频处理的需求。
    3. 与其他任务结合:可以将此框架与声源分离、去混响等任务结合,实现从复杂声学场景的廉价录音中直接重建高保真空间音频。
👀 推荐值得看
🏷️ 领域空间音频 > 空间上混/重制
💡 创新基于扩散模型的条件生成框架——将Ambisonics上变换问题重新定义为条件生成任务,并采用级联式架构逐阶提升空间分辨率
⭐ 评分7.5
#10
eess.AS

Sky-Ear: An Unmanned Aerial Vehicle-Enabled Victim Sound Detection and Localization System 解读失败跨领域

Yi Hong, Mingyang Wang, Yalin Liu, Yaru Fu, Kevin Hung 等 (6 人)
Audio and Speech Processing (eess.AS)
查看摘要
Unmanned Aerial Vehicles (UAVs) are increasingly deployed in search-and-rescue (SAR) missions, yet continuous and reliable victim detection and localization remain challenging due to on-board hardware constraints. This paper designs an UAV-Enabled Victim Sound Detection and Localization System (called ``Sky-Ear'' for brevity) to achieve energy-efficient acoustic sensing and sound detection for SAR. Sky-Ear enables the ``ear'' of the UAV with a circular-shaped microphone array, and the array conducts continuous audio recordings during the UAV's flight. In Sky-Ear, a two-stage (Sentinel and Responder) audio processing method is developed for energy-consuming and highly reliable sound detection. In the Sentinel stage, a Masking autoencoder (MAE)-based sound detection mechanism is designed to analyze frequency-time acoustic features. For improved precision, a continuous localization method is designed by optimizing detected directions from multiple observations. Extensive simulation experiments are conducted to validate the system's performance in terms of victim detection accuracy and localization error.

📖 深度解读

[PDF 下载失败,无法解读]

#11
eess.AScs.SD

The evolution of inharmonicity and noisiness in contemporary popular music 解读失败跨领域

Emmanuel Deruty, David Meredith, Stefan Lattner
Sound (cs.SD); Audio and Speech Processing (eess.AS)
Comments: 44 pages, 23 figures
查看摘要
Much of Western classical music relies on instruments based on acoustic resonance, which produce harmonic or quasi-harmonic sounds. In contrast, since the mid-twentieth century, popular music has increasingly been produced in recording studios, where it is not bound by the constraints of harmonic sounds. In this study, we use modified MPEG-7 features to explore and characterise the evolution of noise and inharmonicity in popular music since 1961. We place this evolution in the context of other broad categories of music, including Western classical piano music, orchestral music, and musique concrète. We introduce new features that distinguish between inharmonicity caused by noise and that resulting from interactions between discrete partials. Our analysis reveals that the history of popular music since 1961 can be divided into three phases. From 1961 to 1972, inharmonicity in popular music, initially only slightly higher than in orchestral music, increased significantly. Between 1972 and 1986, this rise in inharmonicity was accompanied by an increase in noise, but since 1986, both inharmonicity and noise have moderately decreased. In recent years (up to 2020), popular music has remained much more inharmonic than popular music from the 1960s or orchestral music involving acoustic resonance instruments. However, it has become less noisy, with noise levels comparable to those of orchestral music. We relate these trends to the evolution of music production techniques. In particular, the use of multi-tracking may explain the higher inharmonicity in popular music compared to orchestral music. We illustrate these trends with analyses of key artists and tracks.

📖 深度解读

[PDF 下载失败,无法解读]

#12
eess.AScs.SD

Insights on Harmonic Tones from a Generative Music Experiment 解读失败跨领域

Emmanuel Deruty, Maarten Grachten
Sound (cs.SD); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
Comments: 15th International Workshop on Machine Learning and Music, September 9, 2024, Vilnius, Lithuania
查看摘要
The ultimate purpose of generative music AI is music production. The studio-lab, a social form within the art-science branch of cross-disciplinarity, is a way to advance music production with AI music models. During a studio-lab experiment involving researchers, music producers, and an AI model for music generating bass-like audio, it was observed that the producers used the model's output to convey two or more pitches with a single harmonic complex tone, which in turn revealed that the model had learned to generate structured and coherent simultaneous melodic lines using monophonic sequences of harmonic complex tones. These findings prompt a reconsideration of the long-standing debate on whether humans can perceive harmonics as distinct pitches and highlight how generative AI can not only enhance musical creativity but also contribute to a deeper understanding of music.

📖 深度解读

[PDF 下载失败,无法解读]

#13
eess.AScs.SD

An introduction to pitch strength in contemporary popular music analysis and production 解读失败跨领域

Emmanuel Deruty
Sound (cs.SD); Audio and Speech Processing (eess.AS)
Comments: In Music 2024, Innovation in Music Conference, 14-16 June, 2024, Kristiania University College, Oslo, Norway
查看摘要
Music information retrieval distinguishes between low- and high-level descriptions of music. Current generative AI models rely on text descriptions that are higher level than the controls familiar to studio musicians. Pitch strength, a low-level perceptual parameter of contemporary popular music, may be one feature that could make such AI models more suited to music production. Signal and perceptual analyses suggest that pitch strength (1) varies significantly across and inside songs; (2) contributes to both small- and large-scale structure; (3) contributes to the handling of polyphonic dissonance; and (4) may be a feature of upper harmonics made audible in a perspective of perceptual richness.

📖 深度解读

[PDF 下载失败,无法解读]

#14
cs.SD

An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge 解读失败

Shuming Fang, Shuifei Zeng
Sound (cs.SD); Artificial Intelligence (cs.AI)
Comments: Accepted to INTERSPEECH 2026. 4 pages + references. Technical description of our 2nd MLC-SLM Challenge Task 1 submission
查看摘要
We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time. On the official Development set (150 conversations, 21 language/accent categories) the system attains a macro tcpMER of 29.27%, versus 79.15% for the official baseline; on the Evaluation set it scores 50.23%. We also analyze two engineering choices that substantially affect tcpMER. First, embedding-based speaker clustering outperforms an end-to-end-style alternative that assigns speakers from ASR <sc> turn markers alone. Second, overlap-aware segmentation, although intended to raise diarization recall, increases tcpMER because overlapped speech is transcribed twice.

📖 深度解读

[PDF 下载失败,无法解读]

#15
cs.SD

UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection 解读失败

Pengxiang Gao, Yu Qiu, Yanzhi Song
Sound (cs.SD)
Comments: 5 pages, 3 figures, Interspeech 2026
查看摘要
Anomalous Sound Detection (ASD) aims to determine whether faults have occurred by monitoring sounds. Existing methods detect a limited range of anomalies, exhibit poor generalization, or train a separate model for each machine. Diffusion models possess strong generalization and can generate specific data with condition guidance. We propose a unified diffusion model only with a small module. The audio is first transformed into log-Mel spectrograms. The lightweight module embeds machine IDs into condition embeddings, guiding the model to reconstruct data for specific machines. Then diffusion model reconstructs data with condition, using Gaussian Mixture Models to fit the distributions of reconstruction errors. Our unified model could monitor multiple machine types and learn more fundamental feature spaces with cross-domain learning. Experiments on DCASE2022 Challenge Task 2 show that our model achieves 3.44% AUC and 2.52% pAUC improvements over baseline, validating its effectiveness.

📖 深度解读

[PDF 下载失败,无法解读]

#16
cs.SD

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction 解读失败

Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani
Sound (cs.SD); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Multimedia (cs.MM)
Comments: Accepted at ACM IH&MMSec 2026
查看摘要
The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transitional regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering.

📖 深度解读

[PDF 下载失败,无法解读]

#17
cs.SD

What is a Musical Scale? Regularity and Convention in the Organization of Pitch 解读失败

John M McBride
Sound (cs.SD)
Comments: 13 pages, 3 figures, includes a 3-page statistical reporting checklist
查看摘要
Musical scales are near-universal in human music, and most readers will feel they already know what a scale is. On closer inspection, however, the literature lacks a consensus definition: which conditions are necessary and sufficient shifts across disciplines and traditions, and the term turns out to cover several distinct objects. I argue this is less a failure of rigour than a sign that ``scale'' names several related objects: prescriptive abstractions, instrument tunings, statistical regularities in performed pitch, perceptual categories, social conventions. I adopt an empirical definition -- a scale as a statistical regularity in pitch organisation relative to a tonic -- that is portable across traditions and computable from recordings, and situate it alongside the other senses of the term. Even this empirical core is not purely observational, as convention enters in deciding which pitches belong to a scale. And a further step of grouping scales into named categories is a separate convention, which I approach through prototype theory and illustrate with examples from Irish music. Separating these layers provides a basis from which scales can be re-examined empirically and cross-culturally.

📖 深度解读

[PDF 下载失败,无法解读]

#18
cs.SD

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling 解读失败

Haowei Lou, Junda Wu, Chengkai Huang, Tong Yu, Hye-young Paik 等 (7 人)
Sound (cs.SD)
查看摘要
State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In professional scenarios such as film dubbing, game voice acting, and video content generation, users often need to modify a specific style category, such as emotion, age, or gender, while preserving all others. Existing style-controllable TTS methods typically rely on either text-described styles or speech-reference style transfer, making it difficult to jointly control explicit semantic attributes and preserve subtle, text-undescribed prosodic details. We propose AutoSIFT, a controllable speech generation framework for category-level style editing. AutoSIFT decomposes speaking style into known text-describable categories and unknown residual styles that capture non-verbal prosody and speaker-specific nuances. It consists of a generalized Style Disentangler, which extracts category-aware style prototypes from reference speech, and an Arbitrary Style Infiller, which selectively infills unspecified style categories from the reference. By replacing only text-specified style categories while preserving residual speech-derived styles, AutoSIFT enables natural, expressive, and highly customizable speech generation.

📖 深度解读

[PDF 下载失败,无法解读]

#19
cs.SD

Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs 解读失败

Emmanouil Karystinaios
Sound (cs.SD)
Comments: In proceedings of the 29th International Conference on Digital Audio Effects (DAFx) 2026
查看摘要
Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and realtime behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.

📖 深度解读

[PDF 下载失败,无法解读]

#20
cs.SD

ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation 解读失败

Jhen-Ke Lin
Sound (cs.SD); Artificial Intelligence (cs.AI)
查看摘要
A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Reference-note agreement therefore measures reconstruction, not the full design problem. We introduce ChartGenEval, a six-question evaluation framework with an automatic, corruption-tested core. It leaves note choice open while anchoring timing to the song: the matched official chart supplies only its authored timing map, never target notes. We test each core output with dose-controlled failures rather than assume that a familiar statistic measures chart quality. Across 80 held-out song groups, seven output axes satisfy prespecified sensitivity and invariance criteria in nine nonredundant tests. Complementary stress tests on the 40-song development panel expose two broader lessons. A chart-wide phase estimate recovers injected shifts of 15, 30, and 60 ms while chart-only outputs remain essentially unchanged. Common-pattern rewriting lowers mean language-model perplexity by 37%, and loop collapse raises mean self-similarity by 62%. ChartGenEval therefore reports separate, role-specific signals instead of one proxy or total score. This profile provides automatic feedback for comparing and iterating generators; selected outputs are candidate optimization targets or constraints after task-specific stress testing.

📖 深度解读

[PDF 下载失败,无法解读]

#21
cs.SD

Low-Latency Neural Models for Real-Time Music Enhancement 解读失败

Emmanouil Karystinaios, Jonathan Greif, David Nadrchal, Paul Primus, Gerhard Widmer
Sound (cs.SD)
查看摘要
Music recordings and live streams are often affected by noise, reverberation, spectral imbalances, or artifacts that degrade listening quality. While speech enhancement has matured into a well-defined research area, music enhancement is less established because musical signals combine overlapping sources, wide bandwidths, strong dynamics, and intentional production effects. We study real-time music enhancement under strict causal and low-latency constraints. We formulate the task around recovery of the intended produced mix from acoustic and production-oriented degradations, adapt compact causal networks to music, and compare speech-derived real-time baselines, an external music-denoising model, an offline restoration reference, and a music-specific MusicFilterNet-MS variant. On the tested hardware, all causal models run faster than real time, but improvements depend strongly on the dataset, degradation type, and metric family; under several objective criteria, indiscriminate enhancement can worsen the degraded input. The main contribution is therefore a benchmark and an analysis rather than a universal best model: real-time music enhancement is feasible, but robust improvement requires degradation-aware modeling, stereo-aware processing, identity-preserving correction, and evaluation beyond a single objective score.

📖 深度解读

[PDF 下载失败,无法解读]

#22
cs.SD

Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems 解读失败跨领域

Kazushi Kato, Koji Inoue, Taiga Mori, Divesh Lala, Tatsuya Kawahara
Human-Computer Interaction (cs.HC); Sound (cs.SD)
Comments: Accepted by 28th ACM International Conference on Multimodal Interaction (ICMI '26), Long paper
查看摘要
In human dialogue, we achieve smooth communication by expressing nonverbal cues such as eye contact, nodding, and facial expressions with precise timing. It is expected for conversational avatars to express these cues appropriately to realize natural and human-like interactions. This study focuses on nodding, which is crucial for demonstrating active listening and encouraging further user utterances. We propose a model that predicts both timing and kinematic parameters representing the motion features of listener nodding in real time. The proposed model consists of a timing prediction module and a kinematic parameter prediction module. Each implements a dyadic attention network over the speaker and listener channels based on the technique of Voice Activity Projection (VAP). Unlike conventional models, this approach enables real-time prediction of kinematic parameters based on the specific context of the dialogue rather than just predicting the timing. Furthermore, we demonstrate the effectiveness of fine-tuning the kinematic parameter prediction module initialized from the trained timing prediction module. The proposed model is lightweight and capable of real-time operation, and it has been integrated into an avatar dialogue system. Subjective evaluation experiments shows that our proposed method significantly outperforms both a baseline with stochastic timing and another with fixed-motion nodding. The code and trained models are available at this https URL .

📖 深度解读

[PDF 下载失败,无法解读]

#23
cs.SD

Open-Source Intelligence and Music Information Retrieval for Geographic Attribution of Musical Affect and the Ecological Limits of Population Inference 解读失败跨领域

Mohammadreza Rashidi
Cryptography and Security (cs.CR); Sound (cs.SD)
Comments: 16 pages, 12 figures
查看摘要
A common intuition holds that a region's music mirrors the temperament of its people, so that melancholic melodies mark melancholic populations. We test the measurable half of that intuition and reject the inferential half. Using the Essen Folksong Collection, a corpus of thousands of notated folk melodies, we extract real melodic and affect-related features from 2393 deduplicated melodies spanning 16 countries and 7 geographic regions, with the analysis performed on symbolic scores rather than audio. The mode of each melody is computed with a key-finding algorithm rather than read from the file, because the collection's own documentation warns its major and minor labels are unreliable. Cross-country differences in melodic structure are large and highly significant. All 8 tested features differ across countries at p<0.001, with the leap-related features reaching p<10^-90, and China carries a distinctive wide-leap, high-activity signature (arousal composite +1.24 standard deviations, mean absolute interval 2.77 semitones against Germany's 2.17). We then test the inferential half. We correlate the regional musical-affect measures with two published, validated national indices, the World Happiness Report ladder score and the Hofstede individualism index. None of the 6 correlations is significant (0 of 6). The geography of musical affect is real and measurable, but it does not predict how happy or how individualist a population is, and any claim that it does is an ecological fallacy. We release the full extraction and analysis pipeline, and a fail-closed checker re-derives every number in this paper from the data.

📖 深度解读

[PDF 下载失败,无法解读]

#24
cs.SD

Traceback Translators Against Forgetting in Continual Fake Speech Detection 解读失败跨领域

Enrico Gottardis, Mattia Tamiazzo, Simone Milani
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD)
Comments: Accepted at EUSIPCO 2026
查看摘要
Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to decreased performance on previously seen samples (catastrophic forgetting). In this work, we propose a forgetting-resilient solution based on the adoption of domain translators within a frozen detector, which remaps the new feature spaces into the original ones by means of a traceback translator network. Experimental results show that this strategy enables the achievement of high detection rates with respect to traditional retraining, while minimizing the computational effort and preserving the detection accuracy on previous data.

📖 深度解读

[PDF 下载失败,无法解读]

#25
cs.SD

Contrasting statistical patterns in melodic and molecular evolution reveal distinctive constraints in a culturally evolving system 解读失败跨领域

John M McBride, W Tecumseh Fitch
Populations and Evolution (q-bio.PE); Sound (cs.SD); Physics and Society (physics.soc-ph)
Comments: 13 pages, 3 figures, 12 extra pages of supplementary information
查看摘要
Evolved sequences can be used to infer the rules of evolution. Orally transmitted folk melodies are evolved sequences whose similarity to protein sequences (one-dimensional, drawn from a limited alphabet) invites application of bioinformatics methods to study cultural evolution. A major obstacle is that melodies encode rhythm, which breaks some assumptions of standard sequence-alignment algorithms. We develop a rhythm-aware alignment method and apply it to \num{40000} Irish dance tune variants, enabling the first large-scale automated melodic alignment. Four canonical bioinformatics analyses -- mutability, substitution matrices, positional conservation, and covariance -- reveal patterns distinct from those of molecular evolution, revealing the forces that shape each domain: biochemical and biophysical constraints for proteins; memory, motor, and social biases for melodies. Together the results show that bioinformatics provides a powerful framework -- conceptual as much as algorithmic -- for studying cultural evolution. Although the cultural transmission of music has been discussed for centuries, here we show how to analyze it at large scale.

📖 深度解读

[PDF 下载失败,无法解读]

#26
cs.SD

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model 解读失败跨领域

Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal
Artificial Intelligence (cs.AI); Sound (cs.SD)
Comments: 10 pages, 2 figures, 6 tables
查看摘要
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discrete diffusion rather than the absorbing-mask scheme common to recent diffusion language models. A frozen Whisper encoder supplies acoustic features, a lightweight projector maps them into the model embedding space, and low-rank adapters let the frozen backbone attend to the new modality. About 42M parameters are trained, which is 0.16 percent of the backbone. We find that the natural training objectives fail to ground the audio because their gradient reaches the projector only through attention that has already dismissed it. A connectionist temporal classification loss applied through the frozen output head breaks this deadlock. The resulting model reaches 6.6 percent word error rate on LibriSpeech test-clean, transcribes in roughly eight parallel steps regardless of utterance length, and uses a single adapter trained on six languages, which we evaluate here on English, Hindi, and Mandarin.

📖 深度解读

[PDF 下载失败,无法解读]

#27
cs.SD

Methods for pitch analysis in contemporary popular music: multiple pitches from harmonic tones in Vitalic's music 解读失败跨领域

Emmanuel Deruty, David Meredith, Maarten Grachten, Pascal Arbez-Nicolas, Andreas Hasselholt Jørgensen 等 (8 人)
Sound (cs.SD)
Comments: 26 pages
查看摘要
Aims. This study suggests that the use of multiple perceived pitches arising from a single harmonic complex tone is an active and intentional feature of contemporary popular music. The phenomenon is illustrated through examples drawn from the work of electronic artist Vitalic and others. Methods. Two listening tests were conducted: (1) evaluation of the number of simultaneous pitches perceived from single harmonic tones, and (2) manual pitch transcription of sequences of harmonic tones. Relationships between signal characteristics and pitch perception were then analyzed. Results. The synthetic harmonic tones found in the musical sequences under study were observed to transmit more perceived pitches than their acoustic counterparts, with significant variation across listeners. Multiple ambiguous pitches were associated with tone properties such as prominent upper partials and particular autocorrelation profiles. Conclusions. Harmonic tones in a context of contemporary popular music can, in general, convey several ambiguous pitches. The set of perceived pitches depends on both the listener and the listening conditions.

📖 深度解读

[PDF 下载失败,无法解读]

#28
cs.SD

MelT: A Portable, Single-GEMM Mel Audio Frontend via Non-Uniform DFT with Measured Latency and Energy Gains on GPUs 跨领域

Augusto Camargo, Marcelo Finger
Sound (cs.SD)
Comments: 17 pages, 9 figures, 10 tables. v3: corrected author affiliations
查看摘要
Modern neural audio models run on accelerators whose peak throughput comes from dense matrix multiplication, increasingly at the edge and in datacenters. The conventional acoustic frontend, however -- a Short-Time Fourier Transform (STFT) followed by sparse Mel aggregation -- remains a multi-stage pipeline centered on the Fast Fourier Transform (FFT), with execution overheads unlike the dense linear algebra dominating the inference stack. This work introduces MelT, a portable single-stage Mel frontend that precomputes Mel-spaced Non-Uniform Discrete Fourier Transform (NDFT) bases and applies them to time-domain frames through General Matrix Multiplication (GEMM). The contribution is a computational design principle: decoupling Mel feature extraction from vendor-specific FFT primitives and lowering it onto the matrix-multiplication substrate accelerators already optimize. It is not a new spectral operator. MelT's direct projection performs more arithmetic than the FFT pipeline. Yet in the compact-resolution regime of neural audio frontends, it achieves a 1.64-times to 3.29-times latency reduction and up to a 3.03-times reduction in measured active energy, from the Apple A18 Pro to the NVIDIA H100. All gains are within-platform comparisons, accompanied by task-level validation. Word error rate stays statistically equivalent to the native frontend's on frozen Whisper models of medium size and larger; speaker-attribute classification on VoxCeleb1 is non-inferior. The cepstral extension MFCCT preserves utility on a clinical respiratory-insufficiency classification task (SPIRA) while improving on the MFCC baseline. These results indicate that, in practical regimes on the accelerators evaluated here, hardware alignment rather than arithmetic count can govern the realized cost of feature extraction.

📖 深度解读

好的,我将按照您提供的框架,为您解读这篇名为《MelT: GEMM-Native NDFT for Efficient Single-Stage Audio Frontends on Modern Accelerators》的论文。

1. 一句话总结

这篇论文提出了一种名为 MelT 的新型音频前端,它将传统的“傅里叶变换+梅尔滤波器组”两步走流程,替换为一步到位的密集矩阵乘法,从而在现代AI加速器上实现了最高3.75倍的推理加速和3.52倍的能耗降低。

2. 研究背景与动机

  • 核心问题:现代AI加速器(如GPU、NPU)的算力巅峰在于密集矩阵乘法(GEMM),但经典的音频前端(STFT + 梅尔滤波器组)却是一个多阶段、结构异构的流程,与硬件特性不匹配。
  • 问题的重要性:音频前端是几乎所有语音识别、音频分类等系统的第一步。这种软硬件不匹配会导致额外的内存带宽消耗、内核启动延迟和中间张量分配开销,成为整个系统的性能瓶颈,尤其是在追求低延迟和低功耗的边缘设备上。
  • 现有方法的不足
    • 传统方法(STFT+Mel):依赖FFT算法,虽然数学上高效,但在现代加速器上执行时,其稀疏的滤波器聚合和分步操作无法充分利用矩阵计算单元。
    • 可学习前端(如LEAF):虽然尝试用神经网络替代固定前端,但计算成本高昂,且论文指出它们“未能持续超越固定的梅尔滤波器组”,也未报告硬件能效。

3. 核心方法

  • 方法/模型/框架MelT,一个基于GEMM原生的非均匀离散傅里叶变换(NDFT)的单阶段音频前端框架。其核心思想是直接计算梅尔尺度上的频谱能量,而不是先算线性频谱再转换。
  • 关键创新点
    1. 计算流程统一化:将“分帧、加窗、FFT、取能量、梅尔滤波”这一系列操作,统一为一次密集矩阵乘法 X * W^T,其中 W 是预先计算好的、直接对应梅尔频率的投影矩阵。
    2. 硬件原生设计:整个前端被表达为纯粹的GEMM和逐元素操作,这是现代加速器最擅长、最成熟的执行模式,能最大化利用Tensor Cores等矩阵计算单元。
    3. 代数差异的明确:明确指出这种“直接投影”与传统的“先FFT再聚合”在代数上并不完全等价,它是一种“相干投影”,而非“非相干聚合”,并以此作为新前端来评估,而非旧方法的简单重排。
    4. 扩展到倒谱域:提出了MFCCT,在MelT的基础上增加对数压缩和DCT-II变换,作为传统MFCC的硬件原生替代品,证明了该框架的可扩展性。
  • 直觉解释:想象你要统计一篇文章中特定几个关键词(比如“人工智能”、“深度学习”、“神经网络”)的出现频率。传统方法像是先把所有词(所有频率)都数一遍,再挑出你关心的那几个。而MelT的方法是,你只关心这几个词,所以直接拿着这几个词的列表去文章里找,一步到位。在现代硬件上,拿着列表直接查找(矩阵乘法)比先做全量统计再筛选(FFT+滤波)要快得多。

4. 实验与结果

  • 数据集/基准
    • 性能/能耗基准:使用LibriSpeech语音数据集,时长从1秒到160秒。
    • 下游任务验证:VoxCeleb1(性别分类)和SPIRA(新冠肺炎临床呼吸音分类)。
  • 对比基线:传统的STFT+Mel前端(使用librosa实现)。
  • 主要实验结果
    • 速度与能耗:在Apple A18 Pro(边缘设备)上,处理160秒音频,MelT实现了3.75倍的推理加速和3.52倍的能耗降低。在NVIDIA H100(数据中心)上,也实现了1.92倍加速和2.74倍能耗降低。
    • 下游任务表现:在VoxCeleb1性别分类任务上,MFCCT准确率(97.84%)与传统MFCC(97.95%)几乎持平。在SPIRA呼吸音分类任务上,MFCCT的F1分数(0.9860)甚至略高于传统MFCC(0.9737)。
  • 消融实验揭示了什么
    • 梅尔滤波器数量(M)的影响:随着M从40增加到512,MelT的加速优势会单调递减。在M=128(常用配置)时仍有1.75倍加速,但在M=512时优势消失(1.01倍)。这验证了其O(NM)的复杂度,并表明该方法最适用于M较小的常见场景。
    • 方法泛化性:MFCCT(倒谱扩展)在所有硬件平台上都保持了与MelT相似的加速和节能趋势,证明性能提升来自核心的“直接梅尔投影”矩阵化操作,而非特定输出格式。

5. 优势与局限

  • 本文方法的主要优势
    1. 显著的硬件效率:通过将计算模式与硬件对齐,在延迟和能耗上实现了数倍的提升,尤其在边缘设备上效果突出。
    2. 保持下游性能:在追求极致效率的同时,没有牺牲特征的表征能力,下游分类任务准确率与传统方法相当甚至更优。
    3. 实现简单,即插即用:该方法本质上是将固定的投影矩阵预计算好,然后替换掉原有的前端模块,无需训练,易于集成到现有系统中。
  • 局限性
    1. 可扩展性受限:其计算复杂度为O(NM),当梅尔滤波器数量M非常大时,计算成本会超过O(NlogN)的FFT,失去优势。论文仅在M≤512的范围内验证。
    2. 代数不等价性:MelT的输出与传统Mel谱在数学上不完全等价,虽然实验证明下游任务表现良好,但在某些对传统特征有严格依赖的场景下可能无法直接替换。
    3. 评估任务有限:下游任务验证仅在小规模分类任务上进行,未在更复杂的大规模任务(如语音识别ASR)上验证其表征的通用性。

6. 关键结论与启发

  • 最重要的Takeaway在AI加速器时代,算法的“计算复杂度”不再是衡量效率的唯一标准,甚至不是最重要的标准。将算法表达为硬件最擅长的计算模式(如GEMM),即使理论计算量更大,也可能在墙钟时间和能耗上获得巨大收益。 这是一个“软件-硬件协同设计”的胜利。
  • 对后续研究的启发或延伸方向
    1. 前端设计的范式转变:鼓励研究者重新审视更多经典的信号处理流程(不仅仅是音频),将它们重构为矩阵原生的形式,以适应现代硬件。
    2. 可学习前端的硬件感知设计:可以将MelT的矩阵思想与可学习前端结合,让网络学习一个硬件高效的投影矩阵,而不是复杂的卷积核,从而兼顾性能和效率。
    3. 更广泛的验证:在大型语音识别(ASR)、音频生成等更复杂的下游任务中验证MFCCT的普适性,并探索其在更多硬件平台(如FPGA、专用NPU)上的表现。
👀 推荐值得看
🏷️ 领域音频理解 > 声学场景分类
💡 创新基于GEMM原生的非均匀离散傅里叶变换(NDFT)的单阶段音频前端——将传统多阶段STFT+梅尔滤波器组流程统一为一次密集矩阵乘法
⭐ 评分7.5
#29
cs.SD

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers 跨领域

Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang, Shubin Zhang 等 (8 人)
Sound (cs.SD); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
查看摘要
Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of diffusion transformer blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a novel instruction-guided audio editing framework based on rectified flow matching (RFM), named RFM-Editing 2, built on a hybrid two-stage diffusion transformer. The proposed model performs joint attention over audio and text tokens to establish coarse semantic alignment at the low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at the high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency.

📖 深度解读

好的,我将按照要求为您解读这篇论文。

1. 一句话总结

这篇论文提出了一种混合扩散Transformer架构,通过“先粗后精”的两阶段设计,在保证编辑质量的同时,大幅提升了基于文本指令的音频编辑效率。

2. 研究背景与动机

  • 核心问题:如何根据一句自然语言指令(如“添加鸟叫声”),精确修改音频中的特定内容,同时完整保留其他未提及的背景声音。
  • 问题的重要性:这能让用户像用“美图秀秀”P图一样,直观地“P音频”,无需专业软件,在音效设计、后期制作等场景有巨大应用潜力。
  • 现有方法的不足
    • 基于U-Net的方法:依赖局部卷积操作,难以理解跨越长时间的整体语义和精确定位编辑区域。
    • 基于普通DiT的方法:虽然全局建模能力强,但通常在所有层都对拼接后的音频和文本token做联合注意力计算,导致计算复杂度随token长度呈平方级增长,效率低下,难以平衡精度与速度。

3. 核心方法

  • 方法/模型:基于整流流匹配混合两阶段扩散Transformer音频编辑框架。
  • 关键创新点
    1. 粗-细两阶段架构:将编辑过程分为低分辨率“粗对齐”和高分辨率“精修”两个阶段。
    2. 混合注意力模块:在不同阶段使用不同的Transformer模块,低分辨率阶段用双流联合注意力MMDiT,高分辨率阶段交替使用联合注意力MMDiT交叉注意力DiT
    3. 双重条件注入:同时使用全局条件(通过AdaLN-Zero调制)和token级条件(通过交叉注意力),增强编辑可控性和非编辑内容的保留。
  • 核心思路直觉解释
    想象你要在一幅复杂的画作上修改一个小细节。你不会一上来就拿着放大镜在画布上找位置,而是会先退后几步,看清整体构图(低分辨率粗对齐),确定要修改的大致区域;然后再走近,用精细的笔触交替进行观察和修改(高分辨率精修)。这篇论文的方法正是如此:
    • 低分辨率阶段:将音频和文本指令“压缩”到低分辨率,让模型快速、低成本地建立“声音”和“文字”的整体对应关系。
    • 高分辨率阶段:恢复到原始分辨率,交替使用“联合注意力”(同时看音频和文字,做深度融合)和“交叉注意力”(以文字为指导,专门修改音频),像用精细的画笔一样完成局部编辑,同时避免在所有层都进行高成本的全联合计算,从而提升效率。

4. 实验与结果

  • 数据集:作者基于AudioCaps和AudioSet等公开数据集,合成了用于“添加”、“删除”、“替换”三种任务的音频编辑数据集,共约10.6万对样本。
  • 基线方法:对比了4种代表性方法,包括2种免训练的(Zero-Shot, AudioEditor)和2种需训练的(AUDIT, RFM-Editing)。
  • 主要实验结果
    • 质量与一致性:在多项客观指标(如LSD, FD, FAD, KL)上,本方法在多数任务中取得了最佳或次佳结果,表明其编辑后的音频在频谱保真度和整体分布一致性上表现优异。
    • 效率:本方法的平均编辑时间仅为5.07秒,远快于所有对比方法(例如,比AudioEditor的101.87秒快了约20倍,比RFM-Editing的11.23秒快了约2倍)。
    • 模型紧凑性:本方法仅需7861万可训练参数,远少于AUDIT的8.59亿,与RFM-Editing的7009万相当。
  • 消融实验
    • 混合架构是关键:去掉混合设计(如仅用单流DiT)会导致语义对齐和分布一致性指标明显下降。
    • 交替策略优于顺序堆叠:在高分辨率阶段,交替使用不同模块比先堆叠一种再堆叠另一种效果更好。
    • 多阶段设计有益:完全移除低分辨率粗对齐阶段,直接在高分辨率操作,性能会下降,证实了“先粗后精”策略的有效性。

5. 优势与局限

  • 主要优势
    1. 效率显著提升:通过“先粗后精”和混合注意力设计,在保证质量的同时,将推理速度提升到了实用级别。
    2. 编辑质量均衡:在语义对齐、频谱保真度和分布一致性等多个维度上取得了很好的平衡,避免了顾此失彼。
    3. 模型紧凑:以相对较小的参数量实现了有竞争力的性能。
  • 局限性
    1. 指令复杂度有限:论文主要验证了“添加”、“删除”、“替换”三种原子操作,对更复杂、开放式的现实世界指令(如“让音乐听起来更悲伤”)的泛化能力有待考证。
    2. 数据集为合成数据:训练数据是通过规则合成的,与真实世界中千变万化的音频编辑需求可能存在差距。
    3. 依赖外部模型:整个流程依赖VAE、文本编码器、声码器等多个预训练模型,系统的整体性能受限于这些组件。

6. 关键结论与启发

  • 最重要的Takeaway:在音频编辑的扩散Transformer中,“一刀切”地在所有层使用昂贵的联合注意力是低效的。通过“先粗后精”的架构设计和不同注意力机制的混合使用,可以在不牺牲甚至提升编辑质量的前提下,大幅提升计算效率。
  • 对后续研究的启发
    1. 架构设计思路:这种“粗-精”两阶段、混合注意力的设计范式可以推广到其他条件生成任务,如视频编辑、3D内容生成等。
    2. 效率优化方向:为在资源受限设备上部署高性能生成模型提供了新思路,即通过算法架构的改进而非单纯的模型压缩来实现提速。
    3. 可能的延伸方向:论文提到未来会探索更复杂的声学场景和更开放的指令遵循,这可以结合大语言模型来解析复杂指令,并规划出一系列原子编辑操作,从而完成更高级的音频创作。
👀 推荐值得看
🏷️ 领域通用音频生成 > 音频编辑/修复
💡 创新混合两阶段扩散Transformer架构——基于整流流匹配,引入粗-细两阶段设计和混合注意力机制,提升文本引导音频编辑的效率与质量
⭐ 评分7.5