| # | Title | Categories | Authors | Abstract |
|---|---|---|---|---|
| cs.MM 11 papers | ||||
| 120 |
ASSEMBLE: Atomic Skills for Evidence-Grounded Video Reasoning
2609.32466
|
cs.MM
|
Xiyang Wu, Zongxia Li, Shengxin Zhang, Zhichao Liu, Dinesh Manocha |
Complex video reasoning often depends on evidence scattered across distant moments, entities, and events, yet a correct answer alone does not reveal whether a model relied on the right parts of the video. We introduce ASSEMBLE, a framework that makes supportin...Complex video reasoning often depends on evidence scattered across distant moments, entities, and events, yet a correct answer alone does not reveal whether a model relied on the right parts of the video. We introduce ASSEMBLE, a framework that makes supporting evidence explicit throughout long-video reasoning. ASSEMBLE organizes local observations and cross-clip narratives into timestamped evidence catalogs traceable to the source video. A grounding-aware reader then composes question-specific ...
|
| 121 |
ReVR: Dual-Path Concept Reasoning for Multimodal Fake News Detection
2609.33195
|
cs.MM
|
Zhikai Tan, Yuzhou Yang, Qichao Ying, Pinjie Xu, Sheng Li |
Vision-language models (VLMs) support multimodal fake news detection (FND) by producing explicit analyses. Recent methods further improve interpretability by organizing verification knowledge into explicit concepts. However, two questions remain: how to improv...Vision-language models (VLMs) support multimodal fake news detection (FND) by producing explicit analyses. Recent methods further improve interpretability by organizing verification knowledge into explicit concepts. However, two questions remain: how to improve the reliability and applicability of verification concepts, and how to effectively apply reusable concepts to verify unseen news. We propose \textbf{ReVR}, a dual-path reasoning framework that constructs and applies reusable verification ...
|
| 122 |
Learning Through Game: Skewed Transfer of Tabular Knowledge to Strengthen Image Model
2609.32272
|
cs.MM
|
Longfei Huang, Shangdong Yang, Yang Yang |
Multimodal tabular-image learning is gaining growing attention, yet it faces challenges due to tabular data unavailable at test time. A practical solution involves transferring tabular knowledge to images during training to enhance the performance of image mod...Multimodal tabular-image learning is gaining growing attention, yet it faces challenges due to tabular data unavailable at test time. A practical solution involves transferring tabular knowledge to images during training to enhance the performance of image models at inference. However, the overlooked yet important challenges lie in the modality imbalance between images and tables, as well as their asymmetric modality relationship in cross-modal transfer, which limits the auxiliary role of tabula...
|
| 123 |
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion
2609.32540
|
cs.MM
|
Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos |
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconst...Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less-noisy key--value (KV) cache by additional forwards to build the cache without advancing an output latent. However, every denoising forward itself already computes the in-flight KV of the current chunk. We introduce Flas...
|
| 124 |
Re:Cognize -- Open-Set Comic Character Re-Identification
2609.34032
|
cs.MM
|
Aaditya Baranwal, Madhav Kataria, Yogesh S Rawat, Shruti Vyas |
A manga reader meets a character on one page and knows them on sight a hundred pages later, without ever being handed a cast list. Re-identifying comic characters demands the same, open-set and sequential: pages arrive as a stream in reading order, new faces a...A manga reader meets a character on one page and knows them on sight a hundred pages later, without ever being handed a cast list. Re-identifying comic characters demands the same, open-set and sequential: pages arrive as a stream in reading order, new faces appear before anyone names them, and the cast is assembled as the story is read. $\textbf{Re:Cognize}$ evaluates recognition as the story is read, not against a cast handed over in advance: four protocols on one query stream, from closed-set...
|
| 125 |
SyncRA: Learning Temporal Correspondence in Omni-Modal Models
2609.34363
|
cs.MM
|
Zelong Xu, Yan Li, Wenhe Hu, Xiyang Hu |
Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the w...Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To addre...
|
| 126 |
Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
2609.35143
|
cs.MM
|
Gunin Gupta, Nirmit Arora, Pavan Kalyan Tankala |
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw produc...AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a contai...
|
| 127 |
SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale
2609.35225
|
cs.MM
|
Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai, Rei Kawakami |
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling bet...Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise...
|
| 128 |
RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning
2607.11044
|
cs.MM
|
Ruoxuan Zhang, Qiyun Zheng, Siyu Wu, Ling Zou, Hongxia Xie |
Vision-Language Models (VLMs) are widely used for visual understanding, yet current evaluation protocols fail to assess whether these capabilities are grounded in physical reasoning. To address this gap, we introduce Retrospective Physical Process Reasoning, a...Vision-Language Models (VLMs) are widely used for visual understanding, yet current evaluation protocols fail to assess whether these capabilities are grounded in physical reasoning. To address this gap, we introduce Retrospective Physical Process Reasoning, a new evaluation paradigm to reason backward from outcomes under explicit physical constraints. Building on the paradigm, we present RetroHolmes, the first real-world benchmark for Retrospective Physical Process Reasoning, comprising object-...
|
| 129 |
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
2605.06643
|
cs.MM
|
Hao Dong, Hongzhao Li, Shupan Li, Muhammad Haris Khan, Eleni Chatzi |
Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation protocols. Current ...Despite the growing popularity of Multimodal Domain Generalization (MMDG) for enhancing model robustness, it remains unclear whether reported performance gains reflect genuine algorithmic progress or are artifacts of inconsistent evaluation protocols. Current research is fragmented, with studies varying significantly across datasets, modality configurations, and experimental settings. Furthermore, existing benchmarks focus predominantly on action recognition, often neglecting critical real-world...
|
| 130 |
Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models
2608.05126
|
cs.MM
|
Yuezhang Peng, Yuxin Liu, Changfeng Gao, Zhifu Gao, Qian Chen |
Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for closed-set tasks after in-domain supe...Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for closed-set tasks after in-domain supervised fine-tuning, it faces significant challenges in leveraging in-context learning for open-domain tasks due to its ambiguous rule definitions. This work proposes Spoken Function Calling (SFC), a novel semantic understanding perspective ...
|
| cs.SD 95 papers | ||||
| 1 |
Open-Qwen-Music: An Auditable Framework for LLM-Based Music Composition and Diffusion Rendering
2609.31652
|
cs.SDeess.AS
|
Yangbin Yu, Mingyu Yang |
We present Open-Qwen-Music, an open reconstruction of Qwen-Music and a fully specified research system for text-to-music generation that couples LLM-based semantic composition with diffusion-based acoustic rendering. The system comprises a 25 Hz single-codeboo...We present Open-Qwen-Music, an open reconstruction of Qwen-Music and a fully specified research system for text-to-music generation that couples LLM-based semantic composition with diffusion-based acoustic rendering. The system comprises a 25 Hz single-codebook music tokenizer, a 3B-parameter autoregressive Music LLM, and a diffusion renderer producing 48 kHz stereo audio, following the cross-module interfaces reported by Qwen-Music. The strongest systems of this design remain closed, and promin...
|
| 2 |
Normalise or condition? Noise-floor front-ends for on-board keyword spotting under UAV rotor ego-noise
2609.31699
|
cs.SDeess.AS
|
Yida Lin, Bing Xue, Mengjie Zhang, Sam Schofield, Richard Green |
A microphone on the airframe of a small multi-rotor UAV is dominated by rotor ego-noise, so spoken flight commands arrive at negative signal-to-noise ratio (SNR). We study small-footprint keyword spotting (KWS) for a ten-word command vocabulary under real ego-...A microphone on the airframe of a small multi-rotor UAV is dominated by rotor ego-noise, so spoken flight commands arrive at negative signal-to-noise ratio (SNR). We study small-footprint keyword spotting (KWS) for a ten-word command vocabulary under real ego-noise, training on one quadrotor and testing on another. Besides per-clip accuracy we measure the streaming false-alarm rate on 4.4 h of continuous rotor noise. We compare classical noise-robust front-ends (CMN, PCEN, spectral subtraction),...
|
| 3 |
Video-to-Music Generation for Gameplay Videos
2609.31810
|
cs.SDcs.MM
|
Felipe Marra, Lucas N. Ferreira |
Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are r...Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are rendered graphics, music is mostly synthetic audio, and soundtracks loop across entire levels rather than following on-screen events. We introduce a new dataset of 217.6 hours of Super Nintendo (SNES) gameplay video paired with 485 hours of ...
|
| 4 |
CORD-KWS: Calibrated, Order-Aware Detection for Open-Vocabulary Keyword Spotting
2609.31869
|
cs.SDeess.AS
|
Ramesh Gundluru, Adarsh Arigala, Sri Rama Murty Kodukula |
Open-vocabulary keyword spotting (KWS) must detect arbitrary keywords without retraining. Cross-attention-based models achieve state-of-the-art performance but require pairwise interaction between the audio and each keyword, recomputing the fused representatio...Open-vocabulary keyword spotting (KWS) must detect arbitrary keywords without retraining. Cross-attention-based models achieve state-of-the-art performance but require pairwise interaction between the audio and each keyword, recomputing the fused representation for every audio--keyword pair. Embedding-based models avoid this computational cost through independent encoding and similarity scoring, yet remain inferior on challenging benchmarks such as LibriPhrase-hard. We show that this gap is prim...
|
| 5 |
NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
2609.31892
|
cs.SDeess.AS
|
Qiaolin Wang, Pedro Sandoval-Segura, Anunaya Joshi, Edvardas Jurkonis, Jake Downie |
While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal co...While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an ...
|
| 6 |
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
2609.31948
|
cs.SDeess.AS
|
Chengqian Ma, Wenhao Feng, Weixuan Jin, Gaole Dai, Tianyu Xie |
Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather ...Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or fo...
|
| 7 |
VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale
2609.32016
|
cs.SDeess.AS
|
Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby, Felix Friedrich, Maurice Kraus |
Expressive speech synthesis has outpaced expressive speech perception: systems now render fine-grained vocal performances that no public benchmark can score. Most benchmarks for this inverse problem stop at six to nine basic emotion categories, largely on acte...Expressive speech synthesis has outpaced expressive speech perception: systems now render fine-grained vocal performances that no public benchmark can score. Most benchmarks for this inverse problem stop at six to nine basic emotion categories, largely on acted speech. This paper introduces VoiceNet, a human-annotated representation-level benchmark for voice performance understanding on permissively-licensed in-the-wild speech. VoiceNet has two subsets: VoiceNet-Emo applies a 40-emotion taxonomy...
|
| 8 |
Tracing Decoder Artifacts for Compact Synthetic Speech Screening
2609.32050
|
cs.SDeess.AS
|
Yi Chen Liu, Jian Liu |
Recent advances in speech synthesis and voice cloning have increased the need for reliable synthetic-speech detection, yet high-accuracy detectors increasingly rely on large pretrained models that are costly to invoke on every recording. Rather than replacing ...Recent advances in speech synthesis and voice cloning have increased the need for reliable synthetic-speech detection, yet high-accuracy detectors increasingly rely on large pretrained models that are costly to invoke on every recording. Rather than replacing such detectors, we investigate a compact front-end screen that processes all inputs cheaply and forwards only suspicious recordings for more expensive analysis. To enable lightweight screening without a large learned encoder, we exploit spe...
|
| 9 |
Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
2609.32536
|
cs.SDeess.AScs.MM
|
Yanjie Zhang, Nanchen Hu, Yushi Sun |
Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across ...Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each item uses a shared action space comprising silence, a tool call, and a natural-language answer. Speaker-switch pairs hold the specified words fixed while source, distance rendering, a...
|
| 10 |
SAGE: Semantic Audio Generative Encoder
2609.32755
|
cs.SDeess.AS
|
Francesco Brigante, Luca Cerovaz, Davide Marincione, Giorgio Strano, Luca Zhou |
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space...Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that s...
|
| 11 |
DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS
2609.32777
|
cs.SDeess.AS
|
Ambuj Mehrish, Abhinaba Roy, Alex Ivanov, Tawsif Ahmed, Dorien Herremans |
Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditi...Zero-shot text-to-speech (TTS) can reproduce an unseen speaker from a short reference recording, but typically entangles speaker identity and accent within the same reference. We introduce DEFINE, an end-to-end framework that decouples these factors by conditioning speaker identity and target accent on separate audio exemplars. A single inference-time guidance weight continuously controls accent strength without retraining. Built on F5-TTS with parameter-efficient LoRA adaptation, DEFINE maps sh...
|
| 12 |
Finding Emotions Where They Belong: Rethinking Audio Emotion Recognition through Masked Temporal Affective Grounding
2609.32804
|
cs.SDeess.AS
|
Abdelrahman Mohamed, Lars Kai Hansen, Zheng-Hua Tan |
Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal A...Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal Affective Grounding (TAG) task that associates emotions with temporally bounded speech spans and vocal tone descriptions. To support this formulation, we curate temporally annotated versions of existing emotion recognition datasets and const...
|
| 13 |
Whisper-Flash: Acoustically Conditioned Parallel Drafting for Faster Whisper Decoding
2609.32869
|
cs.SDeess.AS
|
Huapeng Zhou, Huayu Wang, Junkai Wu, Kangqi Wang, Xinyu Wang |
Whisper is a widely used encoder-decoder model for speech recognition. Its encoder reads an utterance in one parallel pass, but its decoder writes the transcript one token at a time, which dominates inference time. Speculative decoding shortens such loops with...Whisper is a widely used encoder-decoder model for speech recognition. Its encoder reads an utterance in one parallel pass, but its decoder writes the transcript one token at a time, which dominates inference time. Speculative decoding shortens such loops without changing their output: a small drafter guesses several upcoming tokens, and the original model verifies them all in one forward pass. We present Whisper-Flash, a two-layer drafter built on a property of speech recognition: the words sti...
|
| 14 |
SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings
2609.33265
|
cs.SDeess.AS
|
Yiheng Lu, Hao-Wen Dong |
Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rol...Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rolls to audio features before mask prediction. We introduce SCISSOR (Score-Conditioned Instrument Source Separation for Orchestral Recordings), which uses the score to form a frame-wise query for each source. Each query matches a shared audio...
|
| 15 |
Identity-Assisted Association of Unordered DOA Estimates for Neural Speech Source Tracking
2609.33373
|
cs.SDeess.AS
|
Bing Yang, Di Liang, Xiaofei Li |
Tracking speech sources remains a challenge due to ambiguous data association arising from intermittent speech, close spatial proximity, and complex acoustic conditions. To address these issues, we propose an identity-assisted association that maps unordered d...Tracking speech sources remains a challenge due to ambiguous data association arising from intermittent speech, close spatial proximity, and complex acoustic conditions. To address these issues, we propose an identity-assisted association that maps unordered direction-of-arrival (DOA) estimates to speaker-consistent source trajectories for reliable speech source tracking. Specifically, speaker identity embeddings are directly integrated into the model input as a complementary cue to spatial feat...
|
| 16 |
What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection
2609.33375
|
cs.SDeess.AS
|
Jiajun Xu, Menglu Li, Xiao-Ping Zhang |
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discri...The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchica...
|
| 17 |
CORA: A Protocol for Diagnosing Boundary Robustness in Text-to-Audio Retrieval under Query Reformulations
2609.33433
|
cs.SDeess.AS
|
Jae Min Woo, Kyongmin Kong, Bogyung Jeong, Minjeong Kim, HaeJun Yoo |
Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source...Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source caption into five intent preserving forms (Command, Question, Indirect, Key phrase, and Statement) while fixing the target audio. By tracking the same target across query forms, CORA defines RankDrop, a metric revealing failures hidden by ...
|
| 18 |
Mask-Induced Displacement in Audio XAI via Logit Trajectory Decomposition
2609.33486
|
cs.SDeess.AS
|
Nico Garc\'ia-Peguinho (School of Electronic Engineering and Computer Science, Queen Mary University of London), David Kelly (Department of Informatics, King's College London), Fabrizio Smeraldi (School of Electronic Engineering and Computer Science |
Perturbation-based XAI methods for audio classifiers often estimate feature importance by masking spectrogram regions and crediting output changes to the retained signal. Yet they typically assume the fill (the mask replacement) is negligible. We propose logit...Perturbation-based XAI methods for audio classifiers often estimate feature importance by masking spectrogram regions and crediting output changes to the retained signal. Yet they typically assume the fill (the mask replacement) is negligible. We propose logit-space trajectory decomposition to examine this assumption. An on-axis component captures output along a line connecting the fully filled (occluded) spectrogram to the fully retained original; an off-axis component captures perpendicular di...
|
| 19 |
DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation
2609.33742
|
cs.SDeess.AS
|
Yayue Deng, Dingdong Wang, Yuxuan Hu, Jinyu Li, Yanqing Liu |
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems large...Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an ex...
|
| 20 |
Transformer-based Neural Beamforming for Real-Time Speech Enhancement on Smart Low-Power Hearable Devices
2609.33755
|
cs.SDeess.AS
|
Luca Bompani, Marco Fariselli, Giovanni Oltrecolli, Francesco Conti |
Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses signifi...Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resource-constrained microcontroller units (MCUs) used in hearables. We present an optimized methodology for the real-time execution of a neural-network-based minimum variance distortionless response (MVDR...
|
| 21 |
Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech
2609.33774
|
cs.SDeess.AS
|
Kanghwi Lee, Kyeongseok Jeong, Jeongmin Liu |
Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates direct...Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation. Its main weakness is retokenization: decoding generated speech to a waveform and encoding it again can change token identities and erode the watermark. To make the watermark robust to these changes, we propose Redwing,...
|
| 22 |
Controlling Speaking Rate in Autoregressive TTS via Activation Steering
2609.33810
|
cs.SDeess.AS
|
Francesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet |
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation a...Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direc...
|
| 23 |
Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations
2609.34030
|
cs.SD
|
Cecilia Bola\~nos, Luciana Ferrer, Magdalena Fuentes |
Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts...Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly in practice. We propose a pipeline to uncover shortcut learning in audio classifiers by discovering recurring concepts in their temporal explanations. Specifically, we isolate audio segments that expl...
|
| 24 |
Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: Exploration toward Age, Gender, and Accent Steering
2609.34052
|
cs.SD
|
Shih-Heng Wang, Tiantian Feng, Aditya Kommineni, Huang-Cheng Chou, Bowen Yi |
Neural audio codecs (NACs) are widely used in speech generation and audio-language modeling, yet how they encode speaker-trait information remains poorly understood. Prior work applied sparse autoencoders (SAEs) to investigate accent information in NACs throug...Neural audio codecs (NACs) are widely used in speech generation and audio-language modeling, yet how they encode speaker-trait information remains poorly understood. Prior work applied sparse autoencoders (SAEs) to investigate accent information in NACs through task-level analysis. Here, we extend this analysis to the waveform level and to age, gender, and accent, using SAE steering to probe trait-related information in sparse activations. We identify trait-associated dimensions, modify their ac...
|
| 25 |
SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former
2609.34347
|
cs.SDeess.AS
|
Zhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou, Woon-Seng Gan |
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fu...Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scene...
|
| 26 |
Prox-Friendly Log-Magnitude Prior on Complex-Valued Signal
2609.34445
|
cs.SD
|
Kazuki Matsumoto, Keidai Arai, Kohei Yatabe |
The logarithmic transform is essential in audio signal processing since human auditory perception is approximately logarithmic with respect to magnitude. However, directly incorporating prior knowledge about signals (e.g., harmonic structure) in the log-magnit...The logarithmic transform is essential in audio signal processing since human auditory perception is approximately logarithmic with respect to magnitude. However, directly incorporating prior knowledge about signals (e.g., harmonic structure) in the log-magnitude domain into optimization problems solved by standard proximal splitting algorithms remains challenging. To address this issue, this paper proposes a novel regularizer termed EPILOG (Exponential Penalty for Imposing priors on LOG-magnitu...
|
| 27 |
SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows
2609.34648
|
cs.SD
|
Tianxin Xie, Pengfei Zhang, Kai Jiang, Zelin Zhao, Li Liu |
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly ...Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the...
|
| 28 |
Unsupervised Speech Enhancement via Drifting
2609.34662
|
cs.SDeess.AS
|
Diego Caviedes-Nozal, Liang Xu, Rasmus Kongsgaard Olsson, W. Bastiaan Kleijn |
This paper addresses unsupervised speech enhancement in the unpaired setting using drifting methods, where training relies on separate collections of degraded and clean audio without corresponding pairs. While recent drifting approaches enable unpaired trainin...This paper addresses unsupervised speech enhancement in the unpaired setting using drifting methods, where training relies on separate collections of degraded and clean audio without corresponding pairs. While recent drifting approaches enable unpaired training, they do so at a heavy cost: because the objective optimizes only a marginal prior over clean speech, the enhancer gradually loses the input's linguistic content and speaker identity. To fix this, we introduce input-conditioned drifting. ...
|
| 29 |
On Temporal Binding in Large Audio Language Models
2609.34806
|
cs.SD
|
Paul Primus, Gerhard Widmer |
Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model...Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model components that may need improvement. Using mechanistic interpretability, we investigate how temporal information is represented and bound to sound events in three open-source LALMs. We find that across all three, event-specific location b...
|
| 30 |
SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living
2609.34907
|
cs.SD
|
Debolina Chowdhury, Suman Samui, Sujoy Saha |
Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environm...Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environment. This work introduces \dataset{}, a seven-class bathroom acoustic-event dataset containing 21{,}387 annotated clips recorded across five environments, and proposes SincDPNet, a compact raw-waveform classifier with a learnable sinc filte...
|
| 31 |
JazzSAMBA: A Synchronous and Asynchronous Multi-take Band Audio Dataset of Jazz Standards for Live Music Models
2609.34931
|
cs.SDeess.AScs.MM
|
Phillip Long, Jacob Nguyen, Jace Hosto, Gage Hosto, Jett Takazawa |
Machine learning has made strong progress on music tasks, both as assistive tools and as creative partners. However, most systems train on multitrack corpora that emphasize pop and rock. Jazz, with improvisation at the core of its practice, still lacks a well-...Machine learning has made strong progress on music tasks, both as assistive tools and as creative partners. However, most systems train on multitrack corpora that emphasize pop and rock. Jazz, with improvisation at the core of its practice, still lacks a well-annotated corpus of clean per-stem combo recordings on standards. We introduce JazzSAMBA (Jazz Synchronous and Asynchronous Multi-take Band Audio) to fill this gap: the first originally recorded jazz-combo multitrack dataset of standards wi...
|
| 32 |
Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device
2609.35005
|
cs.SD
|
Pawe{\l} Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefa\'nski, Szymon Klimaszewski |
Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables r...Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term...
|
| 33 |
RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing
2609.35118
|
cs.SD
|
Yu Wang, Haixin Guan, Shuang Wei, Yanhua Long |
Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable....Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverage...
|
| 34 |
Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions
2609.35345
|
cs.SD
|
Michel Olvera (S2A, LTCI, IDS), Paraskevas Stamatiadis (S2A, LTCI |
Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, s...Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation ...
|
| 35 |
GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection
2609.35411
|
cs.SD
|
Zelin Zhao, Guanjie Huang, Danny Hin Kwok Tsang, Li Liu |
Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, th...Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate ...
|
| 36 |
Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities
2609.35613
|
cs.SDcs.MM
|
Xinyuan Qian, Yanghao Zhou, Ziyang Jiang, Yu Chen, Xinjia Zhu |
Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems cond...Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially, enrollment speech as a cue has inherent limitations. Its reliability degrades when the target and interfering speakers have similar voice characteristics, when intra-speaker variab...
|
| 37 |
CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
2609.35645
|
cs.SDeess.AS
|
Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols |
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational ...Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to e...
|
| 38 |
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
2609.35672
|
cs.SD
|
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein |
Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. ...Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot emb...
|
| 39 |
OneVoice: An Intermediate Representation for Agentic Speech Pipelines
2609.31673
|
cs.SDeess.AScs.MM
|
Vipul Charugundla, Dancheng Liu, Jinjun Xiong |
Agentic speech systems must exchange more than text, including speaker identity, timing, phonetic information, and behavioral annotations. Yet these signals are often produced in incompatible tool-specific formats, making agent-to-agent handoff fragile. We pre...Agentic speech systems must exchange more than text, including speaker identity, timing, phonetic information, and behavioral annotations. Yet these signals are often produced in incompatible tool-specific formats, making agent-to-agent handoff fragile. We present OneVoice, a lightweight, JSON-native intermediate representation that provides a shared semantic structure for speech pipelines. OneVoice organizes heterogeneous speech evidence into validated session records with stable identifiers, l...
|
| 40 |
DiffVQE2: An Efficient Low-delay Diffusion Model for Acoustic Echo and Noise Control
2609.31703
|
cs.SDeess.AS
|
Haljan Lugo, Ernst Seidel, Pejman Mowlaee, Ziyue Zhao, Tim Fingscheidt |
Hands-free communication devices and speakerphones are inherently affected by acoustic echo and background noise. To mitigate these impairments, end-to-end discriminatively trained neural networks have emerged as the best-performing approach in research and de...Hands-free communication devices and speakerphones are inherently affected by acoustic echo and background noise. To mitigate these impairments, end-to-end discriminatively trained neural networks have emerged as the best-performing approach in research and deployment. While recent advancements in generative methods have provided remarkable results for various speech enhancement tasks, diffusion-based acoustic echo control (AEC) research is still restricted to non-causal, utterance-level process...
|
| 41 |
RadarVox: Radar-Audio Multimodal Cocktail-Party Speech Separation with Speaker-Aware Cross-Modal Matching
2609.31708
|
cs.SDeess.AScs.MM
|
Yanlin Xu, Yiwei Ru, Mupei Li, Yongji Liu, Jie Wang |
In embodied voice interaction, cocktail-party speech perception requires both speech separation and speaker attribution across machine-generated and human speech sources. However, conventional audio-only blind source separation remains permutation ambiguous, m...In embodied voice interaction, cocktail-party speech perception requires both speech separation and speaker attribution across machine-generated and human speech sources. However, conventional audio-only blind source separation remains permutation ambiguous, making the correspondence between separated streams and physical speakers unclear. This paper presents RadarVox, a radar-audio multimodal benchmark for identity-aware cocktail-party speech separation. RadarVox provides acoustic mixtures and ...
|
| 42 |
Distributional Metrics for Evaluating Spoken Conversational Systems
2609.31719
|
cs.SDeess.AS
|
Shree Harsha Bokkahalli Satish, Erica Cooper, Patr\'icia Schmidtov\'a, Maike Z\"ufle, \'Eva Sz\'ekely |
Evaluating conversational systems is a difficult and unresolved problem. We introduce the Conversational Distribution Score (CDS), which compares distributions of conversational behaviour using human conversations as a reference. CDS describes speech rate, syl...Evaluating conversational systems is a difficult and unresolved problem. We introduce the Conversational Distribution Score (CDS), which compares distributions of conversational behaviour using human conversations as a reference. CDS describes speech rate, syllabic rhythm, and turn interaction through eight interpretable features plus a separate two-feature semantic baseline. We compare conversations with two reference scales: one based on conversational success within human dialogue and another...
|
| 43 |
Acoustic domain shift in spoken language identification from systematic domain generalization evaluation to real-world application
2609.31759
|
cs.SDeess.AS
|
Francois Derrida (X), Rapha\"el Duroselle (X), Thomas Courtat (X), Jean-Fran\c{c}ois Bonastre (X) |
Domain Generalization (DG) aims to develop models that remain robust to conditions unseen during training. While DG has been systematically studied in computer vision through controlled benchmarks and diverse distribution shifts, its evaluation in spoken langu...Domain Generalization (DG) aims to develop models that remain robust to conditions unseen during training. While DG has been systematically studied in computer vision through controlled benchmarks and diverse distribution shifts, its evaluation in spoken language recognition remains less structured. Existing speech datasets provide valuable benchmarks for robustness to real-world acoustic conditions, but are primarily designed around specific scenarios and large scale rather than as general-purp...
|
| 44 |
Oracle Complementarity Is Not Realizable Complementarity in Frozen-Encoder Audio-Visual Emotion Recognition
2609.31764
|
cs.SDeess.AS
|
Benjamin Hurt |
Complementarity analyses of multimodal systems commonly report an oracle ceiling (the fraction of examples on which at least one unimodal branch is correct) and interpret the gap between it and realized fusion accuracy as recoverable headroom. We show that thi...Complementarity analyses of multimodal systems commonly report an oracle ceiling (the fraction of examples on which at least one unimodal branch is correct) and interpret the gap between it and realized fusion accuracy as recoverable headroom. We show that this ceiling is not a fusion target. Using frozen self-supervised audio and visual encoders with a trained head, we measure the ceiling and realized gain on CREMA-D under two audio encoders (one whose fine-tuning lineage includes CREMA-D, one ...
|
| 45 |
PRIME-ANC: Path-Ratio-Informed Modeling for Efficient Neural Filter Synthesis in Active Noise Control
2609.31772
|
cs.SDeess.AS
|
Yaokun Huang, Chunyang Xu, Haowen Hua, Sen Lin, Shichao Hu |
Changes in listener acoustics require active noise control (ANC) filters to be redesigned for new acoustic paths. We introduce PRIME-ANC, a shared neural synthesizer that learns a bounded, path-dependent logmagnitude correction to a regularized path-ratio base...Changes in listener acoustics require active noise control (ANC) filters to be redesigned for new acoustic paths. We introduce PRIME-ANC, a shared neural synthesizer that learns a bounded, path-dependent logmagnitude correction to a regularized path-ratio base. Minimum-phase reconstruction and truncation produce finite-impulse-response (FIR) filters; training optimizes their noise-control performance. Across ten random training/test splits per dataset, PRIME-ANC achieves average held-out one-thi...
|
| 46 |
Cross-Modal Knowledge Distillation for Acoustic Pedestrian Detection
2609.31785
|
cs.SDeess.AS
|
Yonghyun Kim, Chaeyeon Han, Sancho Gatungay, Subhrajit Guhathakurta, Alexander Lerch |
Audio-only pedestrian detection is attractive for urban sensing but limited by weak acoustic cues. An appealing strategy is cross-modal knowledge distillation, in which a video teacher supervises the audio student during training so that the deployed model run...Audio-only pedestrian detection is attractive for urban sensing but limited by weak acoustic cues. An appealing strategy is cross-modal knowledge distillation, in which a video teacher supervises the audio student during training so that the deployed model runs on audio alone. Under this task's severe class imbalance and wide video-audio modality gap, however, what such distillation contributes is unclear. We introduce Trust-Filtered Distillation (TFD), which selectively suppresses teacher super...
|
| 47 |
MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus
2609.31898
|
cs.SDeess.AS
|
K M Naimul Hassan, Ali Alavi, Donald S. Williamson |
Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Spe...Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise a...
|
| 48 |
Improving Audiovisual Speech Recognition through Synthetic Visual Data Augmentation
2609.31961
|
cs.SDeess.AS
|
Pol Buitrago, Pol G\`alvez, Javier Hernando |
Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability o...Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability of labeled audiovisual (AV) datasets. This work explores the use of synthetic visual data as a solution, using an audio-driven talking-head pipeline to generate lip-synchronized visual content from existing audio data. We evaluate the effect...
|
| 49 |
Perceptually Motivated Alignment and Interpolation of Pitch-Aligned Time-Frequency Representations
2609.31971
|
cs.SDeess.AS
|
Shahan Nercessian, Jeff Sontag, Alejandro Koretzky |
We propose a framework for the alignment and interpolation of pitch-aligned time-frequency representations. Building on the tonal interval vector, we introduce a series of extensions that reformulate it as an invertible operator, culminating in a new feature e...We propose a framework for the alignment and interpolation of pitch-aligned time-frequency representations. Building on the tonal interval vector, we introduce a series of extensions that reformulate it as an invertible operator, culminating in a new feature extractor that embeds perceptual consonance priors within a pitch-aligned representation. Accordingly, we develop methods for aligning and interpolating between musical structures. For alignment, we cast the problem as a permutation search u...
|
| 50 |
Binaural Audio-Visual Instance Segmentation
2609.32180
|
cs.SDeess.AS
|
Saijun Wang, Guanfeng Tang, Hongbo Zhao, Zhicheng Lei, Yutong Zhang |
Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic corresponde...Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually similar instances of the same semantic class. In contrast, humans naturally exploit binaural hearing, where interaural differences and direction-dependent acoustic filtering introduced...
|
| 51 |
Audio Preprocessing Effects on Stuttering Detection: A Class-Specific Analysis
2609.32285
|
cs.SDeess.AS
|
Anisha Pattanayak, Hanie Kang, Huang-Cheng Chou, Sudarsana Reddy Kadiri |
Audio preprocessing can affect how well a system detects stuttering. We study a simulated chain of denoising,loudness normalisation, Opus coding, and voice activity detection on SEP-28k. We use frozen WavLM Base+ features and report pointwise confidence interv...Audio preprocessing can affect how well a system detects stuttering. We study a simulated chain of denoising,loudness normalisation, Opus coding, and voice activity detection on SEP-28k. We use frozen WavLM Base+ features and report pointwise confidence intervals from episode-level bootstrap resampling. At a fixed threshold of 0.5, the chain reduces block F1 from 0.638 to 0.465, with smaller decreases for the other four classes. ROC-AUC decreases for all five classes. Blocks show the largest F1 ...
|
| 52 |
Toward Human-Aligned Judgement of Speech Emotion Similarity
2609.32504
|
cs.SDeess.AS
|
Yun-Shao Tsai, Yi-Cheng Lin, Chih-Kai Yang, Ho-Jung Cheng, Tsun-Yi Chang |
Evaluating emotion preservation in expressive speech generation involves assessing how closely generated speech matches a reference in emotion. Human listening tests assess this similarity, but their cost motivates automatic measures aligned with human judgmen...Evaluating emotion preservation in expressive speech generation involves assessing how closely generated speech matches a reference in emotion. Human listening tests assess this similarity, but their cost motivates automatic measures aligned with human judgments. To support the development and evaluation of such measures, we introduce SES-Bench, a speech emotion similarity benchmark built from human comparisons of two candidate utterances against a shared reference. These comparisons record whic...
|
| 53 |
Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations
2609.32522
|
cs.SDeess.AS
|
Wenxu Jia, Xize Cheng, Zihan Zhang, Dongjie Fu, Linjun Li |
Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting ...Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom. To this end, we propose VoxPolyMem, an interaction-aware multimodal memory framework combining incremental speaker ident...
|
| 54 |
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
2609.32607
|
cs.SDeess.AS
|
Yang Xiao, Vidhyasaharan Sethu, Eun-Jung Holden, Ting Dang |
Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio si...Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold acros...
|
| 55 |
Getting Motif-ated: Controllable AI Compositions from Injected Motif Prompts
2609.32788
|
cs.SDeess.AScs.MM
|
Chao Peter Yang, Cynthia Rudin, Yue Jiang, Simon Mak, Stephen Ni-Hahn |
Deep learning has transformed symbolic music generation by borrowing the training paradigms of large language models, with systems such as NotaGen now producing complete, stylistically convincing classical scores from a short prompt. These systems could become...Deep learning has transformed symbolic music generation by borrowing the training paradigms of large language models, with systems such as NotaGen now producing complete, stylistically convincing classical scores from a short prompt. These systems could become powerful creative partners, helping musicians generate endless possibilities. However, current systems expose almost no control handles on the music itself. In principle, control handles could be built into a foundation model trained from ...
|
| 56 |
WhisperVC-AV: Audio-Visual Content Restoration for Noise-Robust Whisper-to-Normal Voice Conversion
2609.32843
|
cs.SDeess.AS
|
Ziyue Yin, Dong Liu, Ming Li |
Environmental noise obscures the content cues needed for whisper-to-normal conversion. We propose WhisperVC-AV, which restores acoustic content features while leaving the original WhisperVC conversion module unchanged. Its context-guided restoration module com...Environmental noise obscures the content cues needed for whisper-to-normal conversion. We propose WhisperVC-AV, which restores acoustic content features while leaving the original WhisperVC conversion module unchanged. Its context-guided restoration module combines temporal acoustic context with synchronized lip features through attention and gated residual correction. Experiments on AISHELL6-Whisper show lower character error rates (CERs) than WhisperVC across three ASR systems, on clean speech...
|
| 57 |
Overview and Analysis of the RecSys Challenge 2026: Conversational Music Recommendation
2609.33045
|
cs.SDcs.MM
|
Seungheon Doh, Sergio Oramas, Bruno Sguerra, Abhinav Bohra, Claudio Pomo |
The RecSys Challenge 2026 studies conversational music recommendation as a joint item recommendation and response generation problem: given a multi-turn dialogue, systems must retrieve relevant tracks from a large catalog and produce a grounded natural-languag...The RecSys Challenge 2026 studies conversational music recommendation as a joint item recommendation and response generation problem: given a multi-turn dialogue, systems must retrieve relevant tracks from a large catalog and produce a grounded natural-language response. This paper presents the challenge task, dataset, evaluation protocol, and official results. Beyond the leaderboard, we analyze the 16 accepted systems through a common retrieve--rerank--generate framework and examine how recomme...
|
| 58 |
CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification
2609.33212
|
cs.SDeess.AS
|
Massa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan, Rita Singh |
Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decis...Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-SV}, a speaker language model that combines strong speaker discrimination with structured, acoustically grounded comparison reports. By connecting a pretrained speaker encoder to a language model and s...
|
| 59 |
Language Discrimination Improves Linguistic Learning in Multilingual Speech Models
2609.33345
|
cs.SDeess.AS
|
Maureen de Seyssel, Jie Chi, Zakaria Aldeneh |
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate lang...Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT ...
|
| 60 |
From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS
2609.33362
|
cs.SDeess.AS
|
Kangxiang Xia, Xinfa Zhu, HangRui Hu, Kexin Huang, Wenjie Tian |
Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to sat...Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS...
|
| 61 |
Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends
2609.33443
|
cs.SDeess.AS
|
Seonghyeon Go, Yongwoo Kim, Hyeonjin Cha, Jaeho Shin |
Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trap...Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving them unable to access real-time information and tool execution. Furthermore, even when Large Language Models (LLM) retrieve information, many duplex speech models process it within a compressed latent sp...
|
| 62 |
An Efficient Parametric Codec for Low-Bitrate First-Order Ambisonics
2609.33554
|
cs.SDeess.AS
|
Wei-Ting Lai, Amy Bastine, Lachlan Birnie, Thushara D. Abhayapala, Prasanga N. Samarasinghe |
Driven by the rapid growth of immersive teleconferencing and generative spatial audio, efficient low-bitrate coding of first-order Ambisonics (FOA) has become increasingly important. In this work, we develop a lightweight parametric FOA codec that retains the ...Driven by the rapid growth of immersive teleconferencing and generative spatial audio, efficient low-bitrate coding of first-order Ambisonics (FOA) has become increasingly important. In this work, we develop a lightweight parametric FOA codec that retains the standard Directional Audio Coding analysis and synthesis while redesigning the spatial metadata quantization scheme. Rather than quantizing direction-of-arrival (DOA) and diffuseness independently, we combine them into a 3-D directivity vec...
|
| 63 |
DGS-MLDG: Domain Gradient Surgery Guided Meta-Learning for Domain Generalization in Speech Deepfake Detection
2609.33706
|
cs.SDeess.AS
|
Siqing Qin, Kong Aik Lee, Youzhi Tu, Eng Siong Chng, Man-Wai Mak |
Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often...Speech deepfake detection faces significant challenges due to domain shifts. Domain generalization (DG), particularly meta-learning for domain generalization (MLDG), offers a promising solution by simulating and mitigating domain shifts. However, MLDG is often hindered by conflicting gradients between its meta-train and meta-test objectives, leading to suboptimal performance. To address this problem, we propose domain gradient surgery (DGS), a meta-learning method that resolves conflicts through...
|
| 64 |
Domain-Adaptive Dual-Gating Mixture of Experts for Generalizable Speech Deepfake Detection
2609.33709
|
cs.SDeess.AS
|
Siqing Qin, Zhe Li, Kong Aik Lee, Man-Wai Mak |
Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel do...Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD under unseen attack types and acoustic conditions. Our innovative dual-gating mechanism leverages Sinc-layer-based filters to process both low-level acoustic signals (raw waveforms) ...
|
| 65 |
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
2609.33757
|
cs.SDeess.AS
|
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou |
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at fron...Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons usi...
|
| 66 |
Unified Target-Speaker ASR with Text and Enrollment Speech Cues
2609.33853
|
cs.SDeess.AS
|
Yuxiang Mei, Yuchen Yan, Dongxing Xu, Jiaen Liang, Yanhua Long |
Target-speaker automatic speech recognition (TS-ASR) aims to recognize a designated speaker while suppressing interfering speech in multi-talker environments. Conventional TS-ASR typically relies on an enrollment utterance, whereas text-guided methods use know...Target-speaker automatic speech recognition (TS-ASR) aims to recognize a designated speaker while suppressing interfering speech in multi-talker environments. Conventional TS-ASR typically relies on an enrollment utterance, whereas text-guided methods use known lexical content, such as a wake word, to identify the target speaker from the observed mixture. These two cues provide complementary information but are usually studied separately. We propose a Unified Dual-Cue TS-ASR framework that suppo...
|
| 67 |
In-Context Adaptation of Encoder-Decoder Models in Speech Recognition
2609.33865
|
cs.SDeess.AS
|
Yen Meng, Sharon Goldwater, Hao Tang |
In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models ar...In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, co...
|
| 68 |
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
2609.33999
|
cs.SDeess.AS
|
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee |
Speaker verification (SV) models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However...Speaker verification (SV) models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However, by establishing a human perceptual alignment metric and conducting systematic analysis, we demonstrate that perceptual alignment is governed far more by how a model is trained (its learning objective) than by how well it performs (EER). N...
|
| 69 |
SPEAR-Gen: Generation-Aware Pre-training for Unified Speech Representations
2609.34147
|
cs.SDeess.AS
|
Xiaoyu Yang, Arthur Hinsvark, Antonios Alexos, Osama Hanna, Philip C. Woodland |
Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a sing...Speech understanding and generation place different demands on speech representations, and existing models are typically optimised towards one capability or the other. To reduce this gap, we introduce SPEAR-Gen, a speech representation model that learns a single representation for both capabilities. Task-aligned feature aggregation consolidates complementary linguistic and paralinguistic information across a frozen encoder into discrete targets for masked prediction, while a coarse-to-fine objec...
|
| 70 |
Uncovering Ordinal-Matching Bias in Audio-Visual LLMs
2609.34223
|
cs.SD
|
Jihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo, Joon Son Chung |
This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end...This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by sim...
|
| 71 |
Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
2609.34381
|
cs.SDeess.AScs.MM
|
Abhinav Sharma, Sai Karthik Navuluru, Wang Wei, Daksh Dangi, Xiangbo Gao |
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the ...Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy...
|
| 72 |
Harmonizing Spectral Evolution in Conditional Flow Matching for TTS
2609.34431
|
cs.SD
|
Isha Pandey Varad Deshpande Abhijat Bharadwaj Ganesh Ramakrishnan |
Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to th...Conditional Flow Matching (CFM) models for text-to-speech (TTS) suffer from incoherent frequency evolution during inference. While similar spectral imbalances are addressed in diffusion models for other domains, those generic solutions fail to generalize to the inherently uncoordinated acoustic dynamics of CFM. We demonstrate that this issue can be effectively mitigated by introducing a novel training-free frequency-selective boosting strategy. Using the Discrete Wavelet Transform (DWT), our met...
|
| 73 |
SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences
2609.34582
|
cs.SD
|
Mingyue Huo, Shivam Mehta, Bhavin Jawade, Yinghong Lan, Haoqi Li |
Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which ...Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-l...
|
| 74 |
From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram
2609.34735
|
cs.SD
|
Josef Pavl\'i\v{c}ek, Petra Pavl\'i\v{c}kov\'a, Irena \v{S}trausov\'a |
Contemporary AI-based music generation can produce compositions that satisfy formal requirements of tonality and musical coherence. However, whether musical expression can be described by mathematical properties alone remains a fundamental question. Human comp...Contemporary AI-based music generation can produce compositions that satisfy formal requirements of tonality and musical coherence. However, whether musical expression can be described by mathematical properties alone remains a fundamental question. Human composers operate within personal and cultural contexts that influence harmonic decisions and deliberate departures from established patterns. This study investigates six narrative-driven popular songs by Bob Dylan, Johnny Cash, and Ritchie Val...
|
| 75 |
Domain-Incremental Learning for Generative Speech Enhancement
2609.34901
|
cs.SDeess.AS
|
Manjunath Mulimani, Annamaria Mesaros, Minje Kim, Jesper Rindom Jensen |
We propose a domain-incremental learning framework for generative speech enhancement (SE) that learns from a sequence of datasets or domains recorded under diverse acoustic conditions. Fine-tuning a pretrained model on continuously evolving domains leads to ca...We propose a domain-incremental learning framework for generative speech enhancement (SE) that learns from a sequence of datasets or domains recorded under diverse acoustic conditions. Fine-tuning a pretrained model on continuously evolving domains leads to catastrophic forgetting of previously acquired knowledge, while zero-shot generalization often fails to adequately adapt to unseen domains. To address these challenges, we first develop a novel language model-based generative SE model that we...
|
| 76 |
Localized time-frequency representation learning for bioacoustic classification in complex soundscapes
2502.13440
|
cs.SDeess.AS
|
Simen Hexeberg, Mandar Chitre, Matthias Hoffmann-Kuhnt, Bing Wen Low |
Prevailing bioacoustic classifiers assign species labels to fixed time-frequency windows rather than to individual vocalizations. When multiple vocalizations occur within the same window, predictions cannot be unambiguously linked to specific calls, which limi...Prevailing bioacoustic classifiers assign species labels to fixed time-frequency windows rather than to individual vocalizations. When multiple vocalizations occur within the same window, predictions cannot be unambiguously linked to specific calls, which limits analyses at the level of individual vocalizations. This work introduces a framework for time-frequency localized bird classification. A Local-Context Classifier (LCC) identifies species from localized time-frequency events (TFEs), while ...
|
| 77 |
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
2507.02915
|
cs.SDeess.AS
|
Ludovic Tuncay (IRIT-SAMoVA), Etienne Labb\'e (IRIT-SAMoVA), Emmanouil Benetos (QMUL), Thomas Pellegrini (IRIT-SAMoVA) |
Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio Joint-Embedding Predictive Ar...Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio Joint-Embedding Predictive Architecture), tailored specifically for audio data. Audio-JEPA uses a simple Vision Transformer backbone to predict latent representations of masked spectrogram patches rather than reconstructing raw audio. We pre-train on unlabeled AudioSet...
|
| 78 |
AlignBeat: A Latent Variable Model for Multi-Class Beat Tracking from Partially Labeled Data
2510.14391
|
cs.SD
|
Jaehoon Ahn, Tae Gum Hwang, Moon-Ryul Jung |
Recent neural beat trackers predict beats and downbeats with two independent frame-wise binary classification heads, a convention adopted because a single multi-class head cannot fully exploit datasets that annotate only beats. The heads can disagree, so Beat ...Recent neural beat trackers predict beats and downbeats with two independent frame-wise binary classification heads, a convention adopted because a single multi-class head cannot fully exploit datasets that annotate only beats. The heads can disagree, so Beat This moves each predicted downbeat to the nearest predicted beat. We instead predict a sparse set of grid points, each with an event time and one distribution over downbeat, beat and no event, so a downbeat is a beat by construction and no ...
|
| 79 |
EnvSSLAM-FFN: Lightweight Layer-Fused System for ESDD 2026 Challenge
2512.20369
|
cs.SDeess.AS
|
Xiaoxuan Guo, Hengyan Huang, Jiayi Zhou, Renhe Sun, Jian Liu |
Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under unseen generators (Track ...Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under unseen generators (Track 1) and black-box low-resource detection (Track 2) conditions. We propose EnvSSLAM-FFN, which integrates a frozen SSLAM self-supervised encoder with a lightweight FFN back-end. To effectively capture spoofing artifacts under severe data imba...
|
| 80 |
TADA! Tuning Audio Diffusion Models through Activation Steering
2602.11910
|
cs.SD
|
{\L}ukasz Staniszewski, Katarzyna Zaleska, Mateusz Modrzejewski, Kamil Deja |
Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remains challenging, as their internal mechanisms for representing high-level concepts are poorly understood. In this work,...Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remains challenging, as their internal mechanisms for representing high-level concepts are poorly understood. In this work, we use activation patching to demonstrate that recent audio diffusion architectures exhibit a semantic bottleneck, where a small, shared subset of consecutive attention layers controls distinct musical concepts, such as the presence of spe...
|
| 81 |
Patient-level validation of a foundation-model pipeline for pediatric lung sounds: physician-annotated adventitious events are recognized, disease groups are not reliably predicted
2603.15688
|
cs.SD
|
Izzet Turkalp Akbasli, Oguzhan Serin |
Background: Lung-sound classifiers are usually evaluated on recording- or event-level splits, although each child contributes many recordings. We evaluated a foundation-model pipeline (PulmoVec) with the patient as the unit of partitioning and asked whether di...Background: Lung-sound classifiers are usually evaluated on recording- or event-level splits, although each child contributes many recordings. We evaluated a foundation-model pipeline (PulmoVec) with the patient as the unit of partitioning and asked whether disease group can be predicted beyond age and sex. Methods: We analyzed 19693 physician-annotated respiratory events from 736 children in the public SPRSound database. A frozen Health Acoustic Representations (HeAR) encoder with low-rank adap...
|
| 82 |
When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal
2603.20433
|
cs.SDeess.AS
|
Yen-Ting Piao, Jay Chiehen Liao, Wei-Tang Chien, Toshiki Ogimoto, Shang-Tse Chen |
While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evalua...While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples with audio remains understudied. To address this gap, we design a three-stage evaluation pipeline that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability in the audio modality. Evaluating six LALMs across four audio understanding tasks under two output constraint categories...
|
| 83 |
Channel-Preserving Representation Alignment for EEG-to-Music Reconstruction
2606.04040
|
cs.SDeess.AS
|
Jiaxin Qing, Junwei Lu, Lexin Li |
Reconstructing naturalistic music from EEG requires extracting music-specific information from weak signals distributed across electrodes. Existing spectrogram regression and diffusion approaches provide audio synthesis mechanisms, but learning the EEG-to-musi...Reconstructing naturalistic music from EEG requires extracting music-specific information from weak signals distributed across electrodes. Existing spectrogram regression and diffusion approaches provide audio synthesis mechanisms, but learning the EEG-to-music mapping faces an information--estimation tradeoff. Combining electrode measurements into fewer features can ease estimation from limited recordings while obscuring distinctions between music segments. We propose channel-preserving represe...
|
| 84 |
A Second-Order Cepstral Signature of Contact-Vibration Sounds Reproduced by Laptop Loudspeakers: A Synthetic Case Study
2606.04475
|
cs.SDcs.MM
|
Jim Salsman |
A mobile phone vibrating on a hard surface often sounds qualitatively unlike ordinary audiovisual recordings when reproduced through laptop loudspeakers. We propose that part of this perceptual distinctiveness can be described as a nested periodicity: a first-...A mobile phone vibrating on a hard surface often sounds qualitatively unlike ordinary audiovisual recordings when reproduced through laptop loudspeakers. We propose that part of this perceptual distinctiveness can be described as a nested periodicity: a first-order cepstral structure reflecting the vibration period and its multiples, and a second-order cepstral structure reflecting repeated spacing within the first-order cepstrum. Treating the perceptual effect as real and using a deliberately t...
|
| 85 |
Time-frequency localization of bird calls in dense soundscapes
2606.10407
|
cs.SD
|
Simen Hexeberg, Fanghui Tong, Hari Vishnu, Mandar Chitre |
Passive acoustic monitoring enables large-scale wildlife observation. Most bioacoustic classifiers predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses. We formulate time-freque...Passive acoustic monitoring enables large-scale wildlife observation. Most bioacoustic classifiers predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses. We formulate time-frequency localization of bird calls as object detection on spectrograms and compare three computer vision model families (YOLO11, SAM 3, RF-DETR) against a non-learnable baseline. We introduce Intersection over Minimum (IoMin), an evaluation met...
|
| 86 |
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
2608.02673
|
cs.SDeess.AS
|
Hankun Wang, Bohan Li, Shi Lian, Xiaoyu Gu, Jing Peng |
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, ...Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, parameters, or target region underspecified. We study a precise and explicit interface for speech editing: a transcript-grounded structural edit instruction with XML-style tags explicitly specifies typed operations and localizes them to tra...
|
| 87 |
ALAS: An Automatic Latent Alignment Score for Audio Language Models
2505.19937
|
cs.SDeess.AS
|
Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli, Cem Subakan |
Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) behavior. Yet despite a growth of fusion strategies, there is no standard way to mea...Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) behavior. Yet despite a growth of fusion strategies, there is no standard way to measure how well a Speech-LLM internally binds audio frames to text tokens. We introduce ALAS (Automatic Latent Alignment Score), a model- and task-agnostic metric that probes the LLM's per-layer hidden states, scoring the cross-modal cosine s...
|
| 88 |
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
2604.02102
|
cs.SDeess.AS
|
Haitong Sun, Stephen McIntosh, Kwanghee Choi, Eunjung Yeo, Daisuke Saito |
Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast...Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast in S3M representations via minimal pairs. We introduce prosodic ABX, an extension of this framework to evaluate prosodic contrast with only a handful of examples and no explicit labels. Also, we build and release a dataset of English and J...
|
| 89 |
PoDAR: Power-Decoupled Audio Representation for Generative Modeling
2605.10084
|
cs.SDeess.AS
|
Alejandro Luebs, Mithilesh Vaidya, Ishaan Kumar, Sumukh Badam, Stephen W. Bailey |
The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of...The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of audio codecs, we demonstrate that latent modelability can be significantly improved through explicit factor decoupling. We present PoDAR (Power-Decoupled Audio Representation), a framework that utilizes a randomized power augmentation and ...
|
| 90 |
$C^3$ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection
2607.03018
|
cs.SD
|
Jin Hong, Jisoo Park, Junseok Kwon |
Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrade under real-world corruptions such as background noise, occlusion, or simultane...Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrade under real-world corruptions such as background noise, occlusion, or simultaneous modality degradation. We attribute this limitation to the absence of explicit consistency constraints that promote robust, semantically aligned representations across modalities. Without such guidance, models tend to learn fragile modal...
|
| 91 |
Why Do You Say It Like That? A Phoneme-level Framework for Explainable Speech Deepfake Detection
2607.08586
|
cs.SDeess.AS
|
Anna Taylor, Michele Panariello, Massimiliano Todisco, Chiara Galdi, Nicholas Evans |
As the accuracy of speech deepfake detection improves with the use of self-supervised representations such as wav2vec 2.0 and HuBERT, understanding why the speech is classified as bona fide or deepfake remains an open challenge. In pursuit of more trustworthy ...As the accuracy of speech deepfake detection improves with the use of self-supervised representations such as wav2vec 2.0 and HuBERT, understanding why the speech is classified as bona fide or deepfake remains an open challenge. In pursuit of more trustworthy and interpretable artificial intelligence, we introduce a phoneme-level analysis framework that connects model predictions to phonetic units. Our post-hoc explainability method is generally applicable to a variety of speech deepfake detecti...
|
| 92 |
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge: A Benchmark with Natural Mixtures and Degraded Video
2608.23759
|
cs.SDeess.AS
|
Kai Li, Wenze Ren, Junjie Li, Cheng Yu, Peijun Yang |
Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge a...Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-talker mixtures, which lack a clean reference, with reference-available remixes of the same speakers; Track~2 additionally degrades the target video in five ways and adds 3-m far-f...
|
| 93 |
DuplexJail: Spoken Interruption Attacks on Full-Duplex Speech Models
2609.09420
|
cs.SD
|
Jaechul Roh, Deepak Chandran, Amir Houmansadr, Andrea Fanelli |
Full-duplex speech models accept user speech while generating responses, making input timing a potential safety concern. We introduce DuplexJail, which delivers fixed, request-independent spoken jailbreak prompts through the user audio channel. Across four ope...Full-duplex speech models accept user speech while generating responses, making input timing a potential safety concern. We introduce DuplexJail, which delivers fixed, request-independent spoken jailbreak prompts through the user audio channel. Across four open-source models and 720 harmful requests from AdvBench and HarmBench, we compare fixed-delay and refusal-triggered interruption with request-end and post-response controls. On AdvBench, Guided Completion at a 1.0 s delay raises whole-respon...
|
| 94 |
Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling
2609.14990
|
cs.SDeess.AS
|
Homayoon Beigi, Grace Conneely |
A finite-difference bowed-string model with implicitly resolved Stribeck friction is presented, with a regime diagnostic. Without implicit resolution no stick phase forms at any bow force. With friction, impedance and quality factor from published measurement ...A finite-difference bowed-string model with implicitly resolved Stribeck friction is presented, with a regime diagnostic. Without implicit resolution no stick phase forms at any bow force. With friction, impedance and quality factor from published measurement rather than fitted, all four strings return a stick fraction of 89.1% against an ideal 90%. Schelleng's maximum bow force is recovered on every string. His minimum, $F_{\min} \propto Z^2 v_b \beta^{-2}$, is replaced by a law, $F_{\min} = C ...
|
| 95 |
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
2609.29816
|
cs.SD
|
Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu |
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training o...Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computa...
|
| eess.AS 24 papers | ||||
| 96 |
Optimal transport meets speech: a tutorial review
2609.31787
|
eess.AS
|
Xugang Lu, Yu Tsao |
Optimal Transport (OT) provides a principled framework for comparing and transforming probability distributions while preserving geometric structure. Recently, OT has gained significant attention in machine learning due to its ability to measure discrepancies ...Optimal Transport (OT) provides a principled framework for comparing and transforming probability distributions while preserving geometric structure. Recently, OT has gained significant attention in machine learning due to its ability to measure discrepancies between distributions, even when their supports do not overlap, making it effective for tasks such as generative modeling, domain adaptation, and transfer learning. Despite its success in fields such as computer vision and natural language ...
|
| 97 |
Acoustic Progress Propagation for Long-Horizon Speculative Decoding in ASR
2609.33245
|
eess.AS
|
Yuanyuan Jia, Qianqian Yang |
Speculative decoding accelerates autoregressive automatic speech recognition (ASR), but the acceptance length of alignment-aware drafters can saturate as the draft horizon increases. We propose a progress-aware speculative drafter that recurrently propagates a...Speculative decoding accelerates autoregressive automatic speech recognition (ASR), but the acceptance length of alignment-aware drafters can saturate as the draft horizon increases. We propose a progress-aware speculative drafter that recurrently propagates an acoustic progress state across draft steps and feeds it back into audio cross-attention to guide token generation. We jointly train the drafter and progress predictor over variable draft horizons. On five ASR test sets, our method achieve...
|
| 98 |
Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
2609.33645
|
eess.AS
|
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang |
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LL...Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned ...
|
| 99 |
GAMF: Learned and Analytical Array Transfer Function Matching for Array-Generic Direction-of-Arrival Estimation
2609.34216
|
eess.AS
|
Zhiheng Jin, Shichao Hu, Chunyang Xu, Mengyao Zhu |
Microphone positional encoding supports cross-array direction-of-arrival (DOA) estimation, but coordinates alone cannot fully describe device shadowing or microphone directivity. We propose a Generalizable ATF Matching Framework (GAMF) for DOA estimation acros...Microphone positional encoding supports cross-array direction-of-arrival (DOA) estimation, but coordinates alone cannot fully describe device shadowing or microphone directivity. We propose a Generalizable ATF Matching Framework (GAMF) for DOA estimation across array geometries and microphone counts, using array transfer functions (ATFs) as acoustic descriptors. The learned branch incorporates ATF embeddings into geometry-conditioned neural estimation to match acoustic observations with candidat...
|
| 100 |
Explainable and Generalisable LLM-based Cognitive Decline Detection with Spontaneous Speech
2609.34217
|
eess.AS
|
Ziyun Cui, Wen Wu, Chuan Shi, Shuguang Yang, Xueying Gui |
Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To ad...Alzheimer's disease (AD) and mild cognitive impairment (MCI), which may precede AD, manifest early through subtle linguistic and acoustic alterations. Traditional diagnostics, however, are often resource-intensive and lack scalability for mass screening. To address these challenges, we introduce a novel bilingual speech large language model framework for automated, explainable cognitive screening. Unlike conventional pipelines that rely on error-prone automatic speech recognition, our system dir...
|
| 101 |
Audio Tokens as a Budgeted Resource: Marginal-Utility Allocation for Scalable Audio Representations
2609.34337
|
eess.AS
|
Mingyu Zhao, Jinchao Zhang, Zhiyong Wu |
Discrete audio tokens are widely used as a representation interface, yet fixed-depth RVQ tokenizers allocate equal capacity to every frame despite varying refinement value. We introduce UniAdapt, which learns marginal utility of RVQ refinements on a frozen cod...Discrete audio tokens are widely used as a representation interface, yet fixed-depth RVQ tokenizers allocate equal capacity to every frame despite varying refinement value. We introduce UniAdapt, which learns marginal utility of RVQ refinements on a frozen codec and allocates them under exact serialized-bit budgets. A rate-independent causal controller predicts acoustic utility, while an optional semantic head supports speech-only utterance-level allocation; measured acoustic and semantic margin...
|
| 102 |
Beyond Textual Chain-of-Thought: JEPA-Conditioned Latent Reasoning for Large Audio Language Models
2609.34407
|
eess.AS
|
Donghang Wu, Haoyang Zhang, Yizhou Peng, Shreyas Gopal, Yi-Wen Chao |
Explicit textual Chain-of-Thought (CoT) has improved the reasoning ability of large audio language models (LALMs). However, textual CoTs are often constructed from text captions of audio and provide limited access to the acoustic evidence, which can introduce ...Explicit textual Chain-of-Thought (CoT) has improved the reasoning ability of large audio language models (LALMs). However, textual CoTs are often constructed from text captions of audio and provide limited access to the acoustic evidence, which can introduce problems like hallucination. To address this modality-gap issue, we introduce JELAR, a Joint-Embedding Predictive Architecture (JEPA)-based latent reasoning framework that conditions latent reasoning supervision on acoustic representations ...
|
| 103 |
CharDuplex: Building Character-Consistent Full-Duplex Spoken Dialogue Models
2609.34461
|
eess.AS
|
Donghang Wu, Yisi Liu, Chen Chen, Hexin Liu, Eng Siong Chng |
Full-duplex speech models are moving voice interaction beyond conventional turn-taking, yet natural conversation is shaped not only by when an agent speaks, but also by how it behaves as a conversational character. We present CharDuplex, a character-driven ful...Full-duplex speech models are moving voice interaction beyond conventional turn-taking, yet natural conversation is shaped not only by when an agent speaks, but also by how it behaves as a conversational character. We present CharDuplex, a character-driven full-duplex speech model that combines real-time spoken interaction with persona-conditioned behavior. We first adapt GLM-4-Voice to an always-on dual-stream architecture and train the model for full-duplex conversation. Then a fully automated...
|
| 104 |
Measurement-Based Bitrate-Energy-Quality Analysis of Neural Audio Codec Decoders on Laptop and Phone Platforms
2609.34524
|
eess.AS
|
Seunghyeon Shin, Seokjin Lee |
Neural audio codecs can achieve similar objective quality at lower bitrates than conventional codecs, but their decoder-side computational cost may offset this bitrate advantage on battery-powered client devices. This paper presents a measurement-based rate--e...Neural audio codecs can achieve similar objective quality at lower bitrates than conventional codecs, but their decoder-side computational cost may offset this bitrate advantage on battery-powered client devices. This paper presents a measurement-based rate--energy--quality analysis of four neural audio codecs--EnCodec, DAC, HILCodec, and SNAC--and two conventional baselines, AAC-LC and Opus, on laptop and phone platforms. Speech and music are evaluated separately using the original-reference Vi...
|
| 105 |
Perceptual Quality Loss or Loss of Perceptual Quality?
2609.35054
|
eess.AS
|
Danilo de Oliveira, Tal Peer, Maur\'icio do V. M. da Costa, Timo Gerkmann |
Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily...Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily correlate with an improved listening experience. Through objective and subjective experiments, we assess the performance of SE models trained with two different types of auxiliary PESQ loss terms. The numerical evaluation on a suite of sta...
|
| 106 |
Simulation-Based Inference for Plate Reverb System Identification
2609.35295
|
eess.AS
|
Dylan Sechet, Marc Evrard, Matthieu Kowalski |
We address Task A of the 1st DAFx Parameter Estimation Challenge, which aims to retrieve the physical parameters of a plate model from an impulse response. To do so, we use the Simulation-Based Inference (SBI) framework, in which we train a neural network to e...We address Task A of the 1st DAFx Parameter Estimation Challenge, which aims to retrieve the physical parameters of a plate model from an impulse response. To do so, we use the Simulation-Based Inference (SBI) framework, in which we train a neural network to estimate a density over plate parameters given an impulse response, using a dataset generated by the simulator. Inference for a new impulse response then requires only a forward pass through the network, without involving the simulator. For ...
|
| 107 |
Jev Matches 7B Language Models for Speech-Neuroprosthesis Rescoring
2609.33538
|
eess.AS
|
Gabriele Cin\`a |
A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is h...A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is hard: general language models asked to pick one sentence from a list answer from where a label sits in the list rather than from the sentence itself. We pose rescoring as a single typed decision, one call that returns a probability for every...
|
| 108 |
SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents
2609.34247
|
eess.AS
|
Wenyi Yu, Siyin Wang, Terumi Chiba, Xianzhao Chen, Xiaohai Tian |
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringen...Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from delibe...
|
| 109 |
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
2412.01488
|
eess.AS
|
Hugo Malard, Michel Olvera, Stephane Lathuiliere, Slim Essid |
Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corre...Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corresponding to objects heard in an audio signal. Most existing approaches tackle this problem by fine-tuning pre-trained models or by training additional modules specifically for the task. We adopt a different strategy: we introduce a training...
|
| 110 |
SSIMuse: Assessing Data Replication in Symbolic Music via an Adapted Structural Similarity Index Measure
2509.13658
|
eess.AS
|
Shulei Ji, Zihao Wang, Le Ma, Jiaxing Yu |
AI-generated music may inadvertently replicate training samples and raise plagiarism concerns. Similarity measures can quantify such replication, thereby offering supervision and guidance for music generation models. Existing similarity measures for symbolic m...AI-generated music may inadvertently replicate training samples and raise plagiarism concerns. Similarity measures can quantify such replication, thereby offering supervision and guidance for music generation models. Existing similarity measures for symbolic music mainly target monophonic melodies and provide limited support for complex musical textures, which restricts their ability to assess replication in polyphonic music. To address this limitation, we explore a simple cross-domain idea and ...
|
| 111 |
ImmersiveFlow: Stereo-to-7.1.4 spatial audio generation with flow matching
2601.12950
|
eess.AS
|
Zining Liang, Runbang Wang, Yujia Xiao, Xuzhou Ye, Qiuqiang Kong |
Immersive spatial audio has become increasingly critical for applications ranging from AR/VR to home entertainment and automotive sound systems. However, existing generative methods remain constrained to low-dimensional formats such as binaural audio and First...Immersive spatial audio has become increasingly critical for applications ranging from AR/VR to home entertainment and automotive sound systems. However, existing generative methods remain constrained to low-dimensional formats such as binaural audio and First-Order Ambisonics (FOA). Binaural rendering is inherently limited to headphone playback, while FOA suffers from spatial aliasing and insufficient resolution for high-frequency. To overcome these limitations, we introduce ImmersiveFlow, the ...
|
| 112 |
HRIR-Former: Grid-Free Time-Domain Reconstruction of Head-Related Impulse Responses with a Spatially Encoded Transformer
2603.27998
|
eess.AS
|
Shaoheng Xu, Chunyi Sun, Jihui Zhang, Amy Bastine, Prasanga N. Samarasinghe |
Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a listener, predict HRIRs...Individualized head-related impulse responses (HRIRs) enable binaural rendering, but dense per-listener measurements are costly. We address HRIR spatial up-sampling from sparse per-listener measurements: given a few measured HRIRs for a listener, predict HRIRs at unmeasured target directions. Prior learning methods often work in the frequency domain, rely on minimum-phase assumptions or separate timing models, and use a fixed direction grid, which can degrade temporal fidelity and spatial contin...
|
| 113 |
Task-Aware Answer Preservation under Audio Compression for Large Audio Language Models
2605.06631
|
eess.AS
|
Amir Ivry |
Large audio language models (LALMs) increasingly reason over long audio clips, motivating audio compression to reduce memory use and inference latency. However, audio compression can leave the model's overall answer accuracy acceptable while severely degrading...Large audio language models (LALMs) increasingly reason over long audio clips, motivating audio compression to reduce memory use and inference latency. However, audio compression can leave the model's overall answer accuracy acceptable while severely degrading accuracy for particular query families. We introduce a framework for auditing a given audio compression method by measuring the increase in an LALM's answer error relative to uncompressed audio. We formalize an acceptance criterion that li...
|
| 114 |
Data Augmentation for L2 English Speaking Assessment using TTS
2607.10790
|
eess.AS
|
Stefano Bann\`o, Penny Karanasou, Mengjie Qian, Kate M. Knill, Mark J. F. Gales |
Automated assessment of second language (L2) speaking proficiency requires substantial annotated speech data, which are scarce compared to written learner corpora. We investigate whether written L2 responses can be transformed into useful synthetic speech for ...Automated assessment of second language (L2) speaking proficiency requires substantial annotated speech data, which are scarce compared to written learner corpora. We investigate whether written L2 responses can be transformed into useful synthetic speech for proficiency assessment using text-to-speech (TTS) and voice cloning. Using COREFL, a corpus of paired spoken and written responses from L2 learners of English, we systematically study two factors: how written responses should be transformed...
|
| 115 |
Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis
2607.11157
|
eess.AS
|
Mingyue Huo, Yuheng Zhang, Hao Zhang |
Speech enhancement (SE) can substantially improve perceptual quality, yet enhanced speech does not necessarily improve automatic speech recognition (ASR). Common explanations such as enhancement artifacts and over-suppression remain qualitative and do not loca...Speech enhancement (SE) can substantially improve perceptual quality, yet enhanced speech does not necessarily improve automatic speech recognition (ASR). Common explanations such as enhancement artifacts and over-suppression remain qualitative and do not localize which enhancement component affects recognition. We study inference-time polar projection, which transforms an STFT mask $M=Ae^{j\phi}$ into $M_{\alpha,\gamma}=A^\alpha e^{j\gamma\phi}$ and independently varies magnitude strength and e...
|
| 116 |
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
2609.21465
|
eess.AS
|
Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng |
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the a...We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two cons...
|
| 117 |
Performance Analysis and Design of an Optomechanical Microphone Using an Integrated Photonic Waveguide Interferometer
2609.24073
|
eess.AS
|
Xiaoyu Niu, Yuqi Meng, Zihuan Liu, Ehsan Vatankhah, Neal Hall |
We present an optomechanical microphone based on a diaphragm-integrated photonic waveguide Mach-Zehnder interferometer. Acoustic pressure deforms the MEMS diaphragm, inducing strain in the sensing waveguide and changing its optical path length. We analytically...We present an optomechanical microphone based on a diaphragm-integrated photonic waveguide Mach-Zehnder interferometer. Acoustic pressure deforms the MEMS diaphragm, inducing strain in the sensing waveguide and changing its optical path length. We analytically evaluate the optical and mechanical transduction mechanisms and key figures of merit, including signal-to-noise ratio, dynamic range, acoustic overload pressure, and minimum detectable pressure. Two design cases are considered: a MEMS micr...
|
| 118 |
Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time
2605.01766
|
eess.AS
|
Itai Allouche, Joseph Keshet |
Multimodal large language models (MLLMs) achieve strong performance on vision- and audio-language tasks, yet can generate responses that conflict with the given visual or auditory inputs, a problem known as multimodal hallucinations. Prior work suggests that t...Multimodal large language models (MLLMs) achieve strong performance on vision- and audio-language tasks, yet can generate responses that conflict with the given visual or auditory inputs, a problem known as multimodal hallucinations. Prior work suggests that this occurs when models rely more on textual cues and learned language patterns than on evidence from the perceptual input. To obtain a more direct account of this imbalance, we apply Layer-wise Relevance Propagation (LRP), which attributes ...
|
| 119 |
Sometin Beta Pass Notin: Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
2605.17710
|
eess.AS
|
Sewade Ogun |
Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and French. Nigerian languages present unique modelling hurdles, incl...Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind resource-rich languages such as English and French. Nigerian languages present unique modelling hurdles, including acute data scarcity, inconsistent orthography, tonal diacritics, diverse accents, frequent code-switching, and localised named entities. To address these challenges, we developed a multilingual ASR framework using a two-stage distilla...
|