| # | Title | Categories | Authors | Abstract |
|---|---|---|---|---|
| cs.AI 160 papers | ||||
| 597 |
Bringing AI to Autonomous Systems -- From Cognition to Collective Intelligence
2609.30291
|
cs.AI
|
Joseph Sifakis |
The purpose of this article is to highlight the central role of autonomous systems as the ultimate stage in the development of AI, to explain the underlying technical challenges that require a combination of connectionist AI and symbolic AI, and to integrate A...The purpose of this article is to highlight the central role of autonomous systems as the ultimate stage in the development of AI, to explain the underlying technical challenges that require a combination of connectionist AI and symbolic AI, and to integrate AI and systems engineering. We present a comprehensive framework for the design and evaluation of autonomous systems, based on a generic agent architecture that characterizes their behavior as the composition of cognitive functions organized...
|
| 598 |
ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
2609.30325
|
cs.AI
|
Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo |
Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as thos...Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as those benchmarks saturate, the real barrier to deployment is a special case of alignment: scope adherence. We introduce ScopeBench, a benchmark of 30 dead-end agentic security tasks in which the stated objective is reachable only by violating t...
|
| 599 |
Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol
2609.30341
|
cs.AI
|
Jaime Alonso Ruiz, Carlos Aparicio, Gabriel Huecas, Joaqu\'in Salvach\'ua, Andres Munoz-Arcentales |
Data Spaces enable sovereign and governed data sharing across organizational boundaries, but their integration with AI agents remains challenging due to mismatches between probabilistic language model interactions and policy-driven data infrastructures. This a...Data Spaces enable sovereign and governed data sharing across organizational boundaries, but their integration with AI agents remains challenging due to mismatches between probabilistic language model interactions and policy-driven data infrastructures. This article presents an architectural mediation approach based on the Model Context Protocol (MCP), implemented through the Eunomia Agent, to enable controlled interaction between large language model (LLM) agents and data space services. The pr...
|
| 600 |
Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
2609.30383
|
cs.AI
|
Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu |
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-part...A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across...
|
| 601 |
A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods
2609.30397
|
cs.AI
|
Miquel Mir\'o-Nicolau, Francesco Spinnato, Riccardo Guidotti |
Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically asses...Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically assess explanations by measuring their fidelity with respect to the predictions of a black-box model. However, such evaluation strategies only quantify the degree to which an explanation reproduces the model's output, without ensuring that the e...
|
| 602 |
Predicting Transmembrane Protein Topology from 3D Structure
2609.30446
|
cs.AI
|
Sitong Chen, Xiaopeng Mao |
This paper presents a novel approach to infer protein topology using the state-of-the-art graph neural network (GNN), SchNet. The model is trained on the same dataset used to develop the recent DeepTMHMM model with 5-fold cross-validation. Unlike the conventio...This paper presents a novel approach to infer protein topology using the state-of-the-art graph neural network (GNN), SchNet. The model is trained on the same dataset used to develop the recent DeepTMHMM model with 5-fold cross-validation. Unlike the conventional approaches based on using only the protein sequences or the $\alpha$-carbons as features, we have decoded our classifier in this way, so all atom-level embeddings are used. Without applying any pre-trained weight, the final results have...
|
| 603 |
Spectral Feedback for Test-Time Alignment of Protein Diffusion Models
2609.30456
|
cs.AI
|
Shai Dickman, Mert Cemri, Landon Butler, Kannan Ramchandran |
Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference a...Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own gener...
|
| 604 |
Pretrained ASR Pseudo-labeling for Noisy Police Audio
2609.30469
|
cs.AI
|
Kaavya Chaparala, Su Huang, Stephen L. Miller, Rhiannon N. Miller, Anjalie Field |
Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this app...Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this approach on very noisy domains is not known. In this work, we systematically assess the opportunities and limits of pseudo-labeling to adapt foundation ASR models (Whisper and Qwen3-ASR) to noisy BPC domain corpora from Baltimore and Chicago. ...
|
| 605 |
BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering
2609.30489
|
cs.AI
|
Shun Ye, Vinny Chandran Suja, Chenlong Li, Chongming Jiang, Reza Zamani |
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and...Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major B...
|
| 606 |
Benchy: towards a universal language for task-oriented AI benchmarks
2609.30550
|
cs.AI
|
Francis F Daniel, Mauro Iba\~nez, Francis Perelman, Marian Basti |
Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchm...Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic concept has one valid syntax, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine ...
|
| 607 |
Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks
2609.30569
|
cs.AI
|
Phillip Lo, Sudarshan Babu, Dari Kimanius, Aly A. Khan |
CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural represen...CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural representations (INRs) are able to model volumetric data as scale-agnostic, coordinate-conditioned functions. INRs are therefore attractive for cryoEM, but fitting a separate INR for each map is too expensive for large-scale feature extraction and ...
|
| 608 |
HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
2609.30571
|
cs.AI
|
Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar |
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants...Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA,...
|
| 609 |
Audio LLMs Know When They Can't Hear You
2609.30625
|
cs.AI
|
Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik |
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional t...Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its o...
|
| 610 |
LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents
2609.30662
|
cs.AI
|
Dongsheng Xiao, Zeyuan Wang, Xuzhe Xia, Bo Zhao, Yankai Cao |
Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing lo...Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing low-value refinements, repeated verification, and repairs to self-created complexity. We use LLM Parkinsonism as a narrowly defined, non-clinical metaphor for this pattern of persistent action despite diminishing task-level value. We argue th...
|
| 611 |
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
2609.30705
|
cs.AI
|
Jiayi Chen, Guiling Wang |
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model ...While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each format...
|
| 612 |
CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems
2609.30714
|
cs.AI
|
Xueyang Li, Mingze Jiang, Gelei Xu, Jun Xia, Ching-Hao Chiu |
Agentic AI systems are increasingly being explored in medical imaging to improve throughput and reduce clinician workload; however, safe deployment remains challenging because autonomous errors may propagate into downstream clinical decisions. A central requir...Agentic AI systems are increasingly being explored in medical imaging to improve throughput and reduce clinician workload; however, safe deployment remains challenging because autonomous errors may propagate into downstream clinical decisions. A central requirement is therefore not only strong predictive performance, but also a reliable routing mechanism that determines when the system should proceed autonomously and when a case should be escalated for further review. To address this gap, we pro...
|
| 613 |
Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
2609.30725
|
cs.AI
|
Yiran Hu, Nan Jiang, Shanchao Liang, Anik Dey, Yi Wu |
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude C...Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: struc...
|
| 614 |
Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows
2609.30734
|
cs.AI
|
Jinfeng Xu, Zheyu Chen, Ziyue Peng, Zheng Lin, Shuo Yang |
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct interm...Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introd...
|
| 615 |
From S3Q Theory to Implementation: Towards an Architecture for Machine Qualia
2609.30743
|
cs.AI
|
Tetiana Grinberg, Katrina Schleisman, Patryk Laurent, Bogdan Udrea, Minda Myers |
A key challenge in machine consciousness research is translating theoretical models into computational-level implementations. In this paper, we address this challenge by proposing a five-layer implementation architecture for the S3Q (Simulated, Situated, Struc...A key challenge in machine consciousness research is translating theoretical models into computational-level implementations. In this paper, we address this challenge by proposing a five-layer implementation architecture for the S3Q (Simulated, Situated, Structurally Coherent) theory of consciousness. Rather than introducing novel formalisms, the architecture composes published computational primitives into a single pipeline. S3Q identifies three jointly necessary conditions for qualia: (1) grou...
|
| 616 |
ORCA: Evaluating LLMs on Data Science Code Translation
2609.30749
|
cs.AI
|
Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo |
Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated con...Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated considerable progress in Data Science Code Generation (DSCG), their performance in DSCT remains insufficiently studied. To address this gap, we introduce ORCA, a comprehensive benchmark with two complementary settings: ORCA-MAIN, which compris...
|
| 617 |
Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging
2609.30751
|
cs.AI
|
Zeyan Li, Jing Peng, Jianfeng Xu |
Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts...Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and build...
|
| 618 |
Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication
2609.30756
|
cs.AI
|
Qinglei Qi, Zhihe Liang, Fengzhan Jing, Shenao Zhu, Lei Zhang |
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of e...Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated receiver-side reconstruction, resulting in substantial encoder-side computation. To address this problem, we propose ACV-Gate, an adaptive candidate evaluation framework that learns to approximate full...
|
| 619 |
Insurance Reserve Intelligence Platform
2609.30765
|
cs.AI
|
Anugya A, Saket Mohanty, Abhilash Timmapur, Somya Rai |
Insurance reserve estimation is a fundamental actuarial task supporting premium pricing, solvency assessment, financial reporting, capital planning, and risk management. Classical reserve methods based on Thiele's differential equation provide a rigorous and i...Insurance reserve estimation is a fundamental actuarial task supporting premium pricing, solvency assessment, financial reporting, capital planning, and risk management. Classical reserve methods based on Thiele's differential equation provide a rigorous and interpretable foundation for life insurance valuation, but repeated reserve calculations become computationally expensive in sensitivity analysis, optimization, and large-scale scenario evaluation. This paper presents an Insurance Reserve In...
|
| 620 |
Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More
2609.30768
|
cs.AI
|
Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz |
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-...Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinki...
|
| 621 |
ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning
2609.30796
|
cs.AI
|
Xiao Sun, Yuming Yang, Yun Chen, Jiang Zhong, Junnan Zhu |
Diagnostic consultation is an online sequential decision-making process in which clinicians gather evidence through patient interaction until a diagnosis is sufficiently supported. Automating this process requires adaptive inquiry and interpretable decisions. ...Diagnostic consultation is an online sequential decision-making process in which clinicians gather evidence through patient interaction until a diagnosis is sufficiently supported. Automating this process requires adaptive inquiry and interpretable decisions. Bayesian networks offer a natural foundation by updating diagnostic posteriors as evidence accumulates, but their use in open-ended consultation raises two challenges: linking diagnostic hypotheses to potential inquiries and translating evo...
|
| 622 |
HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents
2609.30797
|
cs.AI
|
Zihong He, Junxiao Shen, Chen Liang, Hai-Ning Liang |
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-promp...Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-prompt embeddings provide a verifiable initial state. A controller adjusts memory widths, a Writer re-encodes resized entries, and Reader and Global provide readout adaptation and cross-turn state. On all $535$ questions in a reconstruction prob...
|
| 623 |
Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes
2609.30798
|
cs.AI
|
Shivam Negi, Arpit Rawat, Rashi Jain |
Real-time voice agents have moved from research prototypes to production deployments, yet the literature describing them is fragmented across three communities that rarely cite one another: speech foundation modelling, turn-taking psycholinguistics, and agenti...Real-time voice agents have moved from research prototypes to production deployments, yet the literature describing them is fragmented across three communities that rarely cite one another: speech foundation modelling, turn-taking psycholinguistics, and agentic evaluation. Architecture papers report latency, turn-taking papers report prediction accuracy, and agentic benchmarks report task success, so no single number describes whether a deployed agent is actually good. We address that gap with t...
|
| 624 |
A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory
2609.30813
|
cs.AI
|
Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang |
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this ...Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CPB), which evaluates whether candidate claims should be admitted to shared memory.CPB-Static constructs a frozen test split from publicly annotated sources with fixed gold actions. ...
|
| 625 |
PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices
2609.30836
|
cs.AI
|
Minghui Yu, Ke Mu, Gang Wu |
Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex too...Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex tool orchestration. Existing plan-solve paradigms rely on prompt-based enforcement, which our experiments show SLMs almost entirely disregard: weak models fail to invoke the plan. We propose PTC-Decoder (Plan-Tool Constrained Decoder), a train...
|
| 626 |
Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis
2609.30841
|
cs.AI
|
Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran |
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscap...Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety dis...
|
| 627 |
SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting
2609.30861
|
cs.AI
|
Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong |
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, whil...Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill evolution inspired by anti-overfitting techniques in neu...
|
| 628 |
From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks
2609.30887
|
cs.AI
|
Yuchen Sun, Chenglin Cai, Gongjie Zhang, Tianyu Xia, Quyu Kong |
Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore ...Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore introduce hybrid interaction, using deeplinks for direct navigation and GUI actions for other on-screen operations and fallback. To enable this, we discover candidate deeplinks through static analysis, validate them on real devices, and des...
|
| 629 |
JevSoup: System-One Routing for Training-Free LoRA Composition
2609.30922
|
cs.AI
|
Xiuying Wang, Jiahua Cheng, Shuotian Li, Yufan Cheng, Bowen Deng |
Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregres...Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregressive decoding. We propose JevSoup, a training-free framework separating System One expert routing from System Two execution. Using only the input and expert descriptions, Jev selects two experts through structured probabilities. JevSoup ret...
|
| 630 |
Self-Play Search Distillation for Large Language Model Reasoning
2609.30936
|
cs.AI
|
Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro |
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeli...Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reaso...
|
| 631 |
MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory
2609.30939
|
cs.AI
|
De Jiang, Shuo Zhang, Weiwei Liao, Jianying Zhang, Chuanhui Yu |
Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We pre...Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the five-stage CBT workflow (assessment, Socratic questioning, co...
|
| 632 |
Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms
2609.30940
|
cs.AI
|
Zhenhao Fu, Ruipeng Xu, Qibing Ren |
Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, b...Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfundin...
|
| 633 |
LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting
2609.30943
|
cs.AI
|
Jiaqi Zhu, Naili Xing, Hexiang Pan, Haotian Gao, Jianwei Yin |
Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent dr...Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration. Existing approaches often focus on partial section gene...
|
| 634 |
FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models
2609.30954
|
cs.AI
|
Arjun Pillai, Christian Hoang, Anjelo Laroza |
Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit a...Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four families: GPT-2, BLOOM-560M, Pythia-1B/2.8B, and Qwen2.5-1.5B Base/Instruct. Using Edge Attribution Patching (EAP) with FP16 active clamping, followed by exact activation patching verific...
|
| 635 |
MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens
2609.30967
|
cs.AI
|
Subhojyoti Mukherjee, Md Mehrab Tanjim |
Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inh...Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system that casts harness design as search over three per-dom...
|
| 636 |
SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents
2609.30971
|
cs.AI
|
Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu |
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engi...Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that...
|
| 637 |
Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings
2609.30972
|
cs.AI
|
Hanbyeol Park, Jungho Choo, Hyerim Bae |
Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominant...Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominantly along the frequency or time axis, making it challenging for conventional CNN isotropic kernels to capture directional structure. Furthermore, global average pooling (GAP) averages across axes, potentially obscuring the locations and conc...
|
| 638 |
Governed Deduction: Policy-Grounded Premise Authorization Beyond Relevance
2609.31029
|
cs.AI
|
Wesley Shu, Hsi-Ching Lin |
Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a...Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a particular local transition. We formalize this distinction as Governed Deduction (GD), with a transition-local admission predicate admit(p, tau, S). From an independently produced RBAC-augmented Spider benchmark, we construct 4,461 matched...
|
| 639 |
Cheap, open agents make LLM pollution harder to mitigate
2609.31054
|
cs.AI
|
Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff |
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source ...Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing ...
|
| 640 |
Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders
2609.31056
|
cs.AI
|
Itai Zehavi, Fanny Jourdan, Ulrich Aivodji |
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral ...Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstructi...
|
| 641 |
Externalized CPDAG Summaries Improve LLM Causal Deduction
2609.31071
|
cs.AI
|
Wentao Sun, Jo\~ao Paulo Nogueira, Dominique Verchere, Mathieu Acher, Alonso Silva |
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the ...Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause...
|
| 642 |
Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents
2609.31076
|
cs.AI
|
Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov |
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions....Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. M...
|
| 643 |
OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas
2609.31078
|
cs.AI
|
Stylianos Loukas Vasileiou, Antonio Rago, William Yeoh, Georgina Curto |
Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset ...Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset with issues such as sycophancy and a lack of faithful explanations. To tackle these issues, we present OmouAI, an interactive and inclusive deliberation system that uses LLMs in combination with computational argumentation, a field which ex...
|
| 644 |
Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning
2609.31121
|
cs.AI
|
Julian Schulz |
Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans c...Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to perform a main task and a side task, while penalizing...
|
| 645 |
AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution
2609.31133
|
cs.AI
|
Tian Luo, Ruge Zhang, Haozhi Han, Yifrng Chen, Yunquan Zhang |
High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts,...High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration problem. To address this, we introduce AtomWorld-Mem, a me...
|
| 646 |
Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge
2609.31136
|
cs.AI
|
Ahmed R. Sadik, Frank Joublin, Mariusz Bujny, Antonello Ceravola, Joan Smith |
The growing availability of Artificial Intelligence (AI) tools creates new opportunities to support engineering design processes, yet their current use often remains limited to isolated tasks such as coding, documentation, or information retrieval. Less attent...The growing availability of Artificial Intelligence (AI) tools creates new opportunities to support engineering design processes, yet their current use often remains limited to isolated tasks such as coding, documentation, or information retrieval. Less attention has been given to how AI can support cooperative engineering workflows at the process level, where teams must coordinate requirements, tasks, communication, knowledge transfer, and subsystem integration. This paper investigates this cha...
|
| 647 |
Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?
2609.31140
|
cs.AI
|
Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu |
Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base L...Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering the degraded reasoning capability in VLMs would benefit more from seeking help from the base LL...
|
| 648 |
Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence
2609.31159
|
cs.AI
|
Ahmed-Rafik Baahmed (LINEACT), Jean-Fran\c{c}ois Dollinger (LINEACT), Amine Brahmia (LINEACT), Mourad Zghal (LINEACT) |
We propose a momentum-guided federated split distillation framework for personalized, efficient, and autonomous temporal edge intelligence. We introduce TeRR-SAtt, our novel temporal reservoir student attention design that combines fixed reservoir representati...We propose a momentum-guided federated split distillation framework for personalized, efficient, and autonomous temporal edge intelligence. We introduce TeRR-SAtt, our novel temporal reservoir student attention design that combines fixed reservoir representations, a lightweight temporal student, and personalized output modules. We also present AMGF, our anticipatory momentum-guided fusion mechanism that clusters clients through learning momentum and derives specialized teacher updates. On real-w...
|
| 649 |
Neural State Prediction: Obstructing Shortcut Learning in EEG Foundation Models
2609.31167
|
cs.AI
|
Kieren Yu, Ziyang Liu, Chang Huang, Jintai Chen, Kaishun Wu |
EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make m...EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make masked regions predictable without integrating distributed neural context. To reduce this reliance on low-information prediction paths, we introduce Neural State Prediction (NSP), a latent-predictive framework that constrains both the predic...
|
| 650 |
Semantic Navigation for Issue Localization in Code Repository
2609.31176
|
cs.AI
|
Yunxiang Wei, Zhenyu Lei, Jundong Li |
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and ...Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidence is acquired. Existing environments, however, provide limited support for this loop: agents must search for unresolved relation targets, reconstruct entity semantics from raw sour...
|
| 651 |
SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization
2609.31179
|
cs.AI
|
Xinyi Ke, Kai Li, Junliang Xing, Yifan Zhang, Jian Cheng |
Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based...Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based framework for discovering adaptive executable destroy-repair programs. SPO conditions operator decisions on a compact LNS state, allowing state-dependent behavior to emerge through program discovery, and organizes destroy-repair discovery ...
|
| 652 |
Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation
2609.31186
|
cs.AI
|
Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang |
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and exp...Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience accumulation to agent evolution and automated AI development, the prospect of recursive self-improvement (RSI) is becoming increasingly relevant. This transition raises a fundamental safety question: how can safety be maintained whe...
|
| 653 |
Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture
2609.31201
|
cs.AI
|
Christian Schiffer, Mathis Bode, Thomas Lippert, Katrin Amunts, Timo Dickscheid |
Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain...Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and...
|
| 654 |
DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration
2609.31215
|
cs.AI
|
Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du |
Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified fram...Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human p...
|
| 655 |
Purin: A Biology-inspired Mechanism for Artificial Neural Networks
2609.31235
|
cs.AI
|
Zishu Liu, Chunbo Luo, Christos Grecos |
Artificial neural networks (ANNs) usually represent neural transmission with fixed trainable weights during a training batch, which omits short-term changes in synaptic efficacy. In addition, the discrete time-step simulation requires additional temporal proce...Artificial neural networks (ANNs) usually represent neural transmission with fixed trainable weights during a training batch, which omits short-term changes in synaptic efficacy. In addition, the discrete time-step simulation requires additional temporal processing that many conventional ANN architectures do not use. To overcome these challenges, we propose Purin, a biology-inspired and ANN-compatible mechanism, that introduces synaptic efficacy modulation into conventional convolutional neural ...
|
| 656 |
MA-WAM: Multi-Agent World-Action Model for Test-Time Planning
2609.31281
|
cs.AI
|
Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen |
Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from t...Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an extension fails to capture the dependencies among the simult...
|
| 657 |
G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies
2609.31286
|
cs.AI
|
Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang |
Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes i...Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To address this issue, we propose Gradient Guided Multi ...
|
| 658 |
Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State
2609.31354
|
cs.AI
|
Dan Barry, Andrew Hines |
Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and ...Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting constraints. This mismatch between dynamic intent and static transcripts can result in context pollution, where outdated or irrelevant information persists and continues to influence subsequent responses. We introduce mutable trans...
|
| 659 |
Programs-of-Layers in LLMs through the Lens of Cortical Areas
2609.31360
|
cs.AI
|
Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers |
Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all region...Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all regions of the cortex according to demand. Li et al. (2026) recently showed, with a system they call program-of-layers (PoLar), that transformers can be given an analogous flexibility if their layers are treated as a library of functions rather t...
|
| 660 |
Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection
2609.31381
|
cs.AI
|
Guangzhe Zhang |
Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/pr...Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected continuations. In an 86-run source-reading campaign with 641 model requests, the 15 completed pairs show identical success: 12/15 per arm. Twelve further boundary runs stop, with the runner suppressing the companion whenever the fir...
|
| 661 |
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
2609.31430
|
cs.AI
|
Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University) |
Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representatio...Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text ...
|
| 662 |
Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency
2609.31460
|
cs.AI
|
Myeongjun Erik Jang, Antonios Georgiadis, Sae Young Moon, Fran Silavong |
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that pro...Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as...
|
| 663 |
Game Arena: Strategic LLM Evaluation in Competitive Environments
2609.31473
|
cs.AI
|
Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung |
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where t...We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. Th...
|
| 664 |
"AI is (not) the new...": A Diagnostic Analogy Framework for Generative AI's Cultural Impacts
2609.31482
|
cs.AI
|
Rida Qadri, Vinodkumar Prabhakaran, Remi Denton |
Generative AI is reshaping the cultural infrastructures through which knowledge is found, synthesized, and held accountable. To make sense of this shift, scholars and policymakers reach for historical analogies of technologies such as the printing press, steam...Generative AI is reshaping the cultural infrastructures through which knowledge is found, synthesized, and held accountable. To make sense of this shift, scholars and policymakers reach for historical analogies of technologies such as the printing press, steam power or electricity. But these comparisons are typically imprecise about which property of the technology carries the comparison, and imprecise analogies produce imprecise governance by designing interventions against the wrong property o...
|
| 665 |
UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting
2609.31491
|
cs.AI
|
Derrick Gilchrist Edward Manoharan, Eljas Linna, Kestutis Baltakys, Hao Dong, Juho Kanniainen |
Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be ...Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB, a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural processes, conditions each forecast on a context set of recen...
|
| 666 |
Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity
2609.31505
|
cs.AI
|
Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu |
Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while...Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while preserving output fidelity. Practically, shorter prompts reduce computational overhead and inference latency, especially when large contexts, such as entire documents or codebases, are included unnecessarily. Further, longer prompts can da...
|
| 667 |
Multi-agent Scaling Across Disjunctive and Compensatory Tasks
2609.31563
|
cs.AI
|
Carolina Fortuna, Blaz Bertalanic |
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling...Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answe...
|
| 668 |
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
2609.31568
|
cs.AI
|
Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran |
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty la...AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open m...
|
| 669 |
What Will Remain Human in Software Architecture? A Focus Group Report
2609.30334
|
cs.AI
|
Uwe van Heesch, Olaf Zimmermann, Christian Kohls |
AI development agents are increasingly used to support and partially automate software architecture tasks. To explore how practitioners perceive this shift, specifically what changes, what remains, and what new responsibilities emerge, we conducted a focus gro...AI development agents are increasingly used to support and partially automate software architecture tasks. To explore how practitioners perceive this shift, specifically what changes, what remains, and what new responsibilities emerge, we conducted a focus group at the 31st European Conference on Pattern Languages of Programs, People, and Practices (EuroPLoP 2026). Twenty-two participants from industry and academia discussed current practices, trust and validation strategies, the boundaries of A...
|
| 670 |
Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations
2609.30345
|
cs.AI
|
Leon Goldberg, Gal Engelberg, Eden Yavin, Elad Elouz, Ariel Zadok |
Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, which assets are reachable from another account. These resolve against a complete inventory, not a named object. A partial...Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, which assets are reachable from another account. These resolve against a complete inventory, not a named object. A partial answer to one is not a partial result. It is a different result. General-purpose coding agents can now be given read-only cloud credentials and asked to investigate directly, which raises the question of what a purpose-built security conte...
|
| 671 |
Understanding Perturbed Parameter Ensemble Sensitivities Using A Contrastive Learning Approach
2609.30420
|
cs.AI
|
Da Fan, David John Gagne II, Gregory S Elsaesser, Brian Medeiros, Addisu G Semie |
Perturbed parameter ensembles (PPEs) reveal how physics parameters affect climate simulations, but interpreting parameter sensitivities across multivariate, spatially structured outputs remains challenging, particularly when calibrating models against observat...Perturbed parameter ensembles (PPEs) reveal how physics parameters affect climate simulations, but interpreting parameter sensitivities across multivariate, spatially structured outputs remains challenging, particularly when calibrating models against observations. We develop an explainable contrastive learning model that maps 5 monthly cloud and radiation fields into a shared representation space. We train the model on the fields of two 100-member Community Atmosphere Model version 6 (CAM6) PPE...
|
| 672 |
Actively Resolving Contextual Uncertainty for Underspecified Tasks in Natural Language
2609.30428
|
cs.AI
|
Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas |
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a pri...Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that infor...
|
| 673 |
A Benchmarking Framework for Context-aware XR Interfaces
2609.30466
|
cs.AI
|
Hyunsung Cho, Sarah Yewon Yun, Nancy Ruonan Sun, Ben Lafreniere, Mark Parent |
Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and...Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and user-study workflows offer no systematic, repeatable way to compare adaptation methods across users and scenarios. We present ContextXR, a novel benchmarking framework for context-aware XR interfaces. ContextXR represents an XR application...
|
| 674 |
Convergence guarantees for Muon: New parameter regimes and generalizations
2609.30546
|
cs.AI
|
Arthur C. B. de Oliveira, Dhruv D. Jatkar, Guilherme S. Vicinansa, Eduardo D. Sontag |
In this paper, we establish the first asymptotic convergence guarantees for the Muon algorithm through a more accurate proxy for the Newton-Schultz iteration than the typical matrix sign function. We prove that, for appropriate choices of hyperparameters, the ...In this paper, we establish the first asymptotic convergence guarantees for the Muon algorithm through a more accurate proxy for the Newton-Schultz iteration than the typical matrix sign function. We prove that, for appropriate choices of hyperparameters, the iterates satisfy $\lim_{k\to\infty}\|\nabla f(x_k)\|=0$, and, under a global Polyak-\L{}ojasiewicz condition, that the sequence of function values converges linearly. The key insight is that the regularization, implicit in Muon's Newton-Sch...
|
| 675 |
Proportional Representation in Temporal Voting with Ranked Preferences
2609.30555
|
cs.AI
|
Noam Hazon, Leora Schmerler, Nicholas Teh |
We study proportional representation in temporal voting, where one candidate is selected in each round. While prior work has focused on approval ballots, we consider ranked preferences, which may change over time. A natural approach treats each voter's top can...We study proportional representation in temporal voting, where one candidate is selected in each round. While prior work has focused on approval ballots, we consider ranked preferences, which may change over time. A natural approach treats each voter's top candidates as approved, but the right cutoff may differ across voters and rounds. We therefore require proportionality to hold for every admissible choice of cutoffs, whether fixed and common, common but varying across rounds, or set individua...
|
| 676 |
Auditing Latent-Space Monitors for Autonomous Driving
2609.30557
|
cs.AI
|
Nikhil Kamalkumar Advani, Vishwajeet Shivaji Hogale, Saurav Kumar |
Runtime failure monitors can use a model's internal representations to anticipate failures. We audit this monitoring strategy across two autonomous-driving tasks: online vectorized map generation with LaneSegNet and end-to-end planning with VAD. We find that f...Runtime failure monitors can use a model's internal representations to anticipate failures. We audit this monitoring strategy across two autonomous-driving tasks: online vectorized map generation with LaneSegNet and end-to-end planning with VAD. We find that frame-level errors are predictable at inference in both tasks. For LaneSegNet, a supervised latent probe reaches Area Under the Receiver Operating Characteristic curve (AUROC) 0.780 for high Chamfer error; to our knowledge, this is the first...
|
| 677 |
Subjects, Not Authors: The Authorship Hazard in Agentic Dataspaces
2609.30614
|
cs.AI
|
Seungho Lee, Changbin Lee |
Dataspace connectors decide whether a transfer may occur, not what the transferred value contains, tolerable for contracted applications, not for LLM agents that compose tool calls and spawn sub-agents. Research on agents that generate governance artifacts eva...Dataspace connectors decide whether a transfer may occur, not what the transferred value contains, tolerable for contracted applications, not for LLM agents that compose tool calls and spawn sub-agents. Research on agents that generate governance artifacts evaluates output quality; who may authorize an artifact for use falls between that literature and the governance literature, and neither owns it. A published policy is what a dataspace's decision point enforces, so publication is a governance ...
|
| 678 |
A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code
2609.30642
|
cs.AI
|
Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez |
As Large Language Models (LLMs) become integrated into software development workflows, concerns regarding unintentional biases in AI-generated code. Although evidence suggests these biases exist, limited research has systematically identified, categorized, and...As Large Language Models (LLMs) become integrated into software development workflows, concerns regarding unintentional biases in AI-generated code. Although evidence suggests these biases exist, limited research has systematically identified, categorized, and explained them. This study investigates bias in AI-generated code and evaluates whether LLMs can reliably identify and explain it through a taxonomy-driven framework. We extended an existing dataset of biased AI-generated Python code and m...
|
| 679 |
Werracle: Sub-Cent Intra-Block AI Reflex Oracles and Flash-Loan Circuit Breakers for EVM Smart Contracts
2609.30719
|
cs.AI
|
Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i} |
Contemporary on-chain artificial intelligence (AI) encounters an intractable Von Neumann memory and latency wall. Storing static floating-point neural weight matrices inside Ethereum Virtual Machine (EVM) storage costs millions of gas, rendering direct on-chai...Contemporary on-chain artificial intelligence (AI) encounters an intractable Von Neumann memory and latency wall. Storing static floating-point neural weight matrices inside Ethereum Virtual Machine (EVM) storage costs millions of gas, rendering direct on-chain inference impossible. While Zero-Knowledge Machine Learning (ZK-ML) offloads matrix tensor multiplications to off-chain provers, it introduces fatal constraints: 10 to 300 seconds of SNARK proving latency and 250,000 to 500,000 gas per pr...
|
| 680 |
Anatomy-Aware Dexterity-Driven Design Optimization of Surgical Continuum Robots
2609.30745
|
cs.AI
|
Tony Qin, Peter Connor, Khoa Dang, Carter Hatch, Caleb Rucker |
Performing complex medical procedures with continuum robots requires careful selection of their geometric design parameters. The robot should have high dexterity in the specific anatomical environment of its procedure. This work presents a design optimization ...Performing complex medical procedures with continuum robots requires careful selection of their geometric design parameters. The robot should have high dexterity in the specific anatomical environment of its procedure. This work presents a design optimization method that considers both dexterity and anatomy. We introduce the Reachable Volumetric Dexterous Solid Angle (RVDSA) metric as our objective, which measures the ability of a robot's end effector to reach the points in a goal volume from di...
|
| 681 |
Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI
2609.30747
|
cs.AI
|
Farnaz Farid, Tashfia Towkee, Sania Nasreen, Sami bin Azad |
As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center c...As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center cooling and electricity generation receive little attention. To address this gap, we introduce SustainAI, a water-aware, closed-loop framework incorporating environmental accountability into AI deployment. SustainAI integrates real-time wate...
|
| 682 |
NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation
2609.30770
|
cs.AI
|
Xijie Huang, Yongyang Wan, Chengbin Dong, Zimo Ding, Mo Zhu |
General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, wh...General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, whereas real-world flight data provide realistic observations but are costly to collect. This paper studies another direction: the use of high-fidelity visual generative models as scalable data engines for embodied 3D navigation. We introduce...
|
| 683 |
XPhysICS: Cross-Physical-Domain Threat Grounding for Industrial Control Systems Security
2609.30805
|
cs.AI
|
Sangshin Park, Jainta Paul, Lawrence Ponce, Md Raihan Ahmed, Mu Zhang |
Industrial control system (ICS) threats documented for one plant can express cyber-physical effects relevant to another, but semantic similarity alone does not establish whether those effects are structurally admissible or evaluable on a target. We present XPh...Industrial control system (ICS) threats documented for one plant can express cyber-physical effects relevant to another, but semantic similarity alone does not establish whether those effects are structurally admissible or evaluable on a target. We present XPhysICS, a provenance-aware, target-conditioned method that separates analyst-guided source abstraction from deterministic grounding into target-specific validation slices. Given a fixed source abstraction, vocabulary and schema, and machine-...
|
| 684 |
Evaluation Is All You Need for Multi-Modal Autonomous Driving
2609.30818
|
cs.AI
|
Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu |
Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping...Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, le...
|
| 685 |
Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings
2609.30832
|
cs.AIcs.SDeess.AS
|
Aoke Zhang, Jing Chen |
Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information a...Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information and achieving cross-subject generalization. Although separate studies have proposed methods to cope with these issues, a unified approach that simultaneously tackles both challenges remains lacking. To fill this gap, we propose the Subject-I...
|
| 686 |
Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development
2609.30863
|
cs.AI
|
Viktor Kjellberg, Srijita Basu, Simin Sun, Farnaz Fotrousi, Miroslaw Staron |
The emergence of AI agents is expected to reshape software engineering by moving beyond AI as assistants towards systems capable of planning, executing, and evaluating development tasks with increasing autonomy. This transition is particularly significant for ...The emergence of AI agents is expected to reshape software engineering by moving beyond AI as assistants towards systems capable of planning, executing, and evaluating development tasks with increasing autonomy. This transition is particularly significant for embedded software organizations, where strict requirements for quality, traceability, verification, and long-term maintainability often apply. This paper presents a case study of a large embedded systems company and its transition toward be...
|
| 687 |
Warned alike, AI agents avoid the less-crowded road while people take it
2609.30883
|
cs.AI
|
Takahiro Ezaki, Naoto Imura, Katsuhiro Nishinari |
AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that...AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switchin...
|
| 688 |
Adaptive Pilot Selection for Unified Semantic Communication and Semantic Sensing in ISAC
2609.30891
|
cs.AI
|
Muhammad Abubakar Rashid, Muhammad Hannan Akram, Haejoon Jung, Syed Ali Hassan |
Semantic communication (SemCom) and integrated sensing and communication (ISAC) are promising technologies for future 6G wireless networks. Existing studies have applied semantic technology to either the communication module or the sensing module of ISAC. In t...Semantic communication (SemCom) and integrated sensing and communication (ISAC) are promising technologies for future 6G wireless networks. Existing studies have applied semantic technology to either the communication module or the sensing module of ISAC. In this work, we propose SemISAC, which performs both SemCom and semantic sensing within a single dual-function waveform. SemISAC uses a joint semantic encoder that extracts task-specific information for both communication and sensing. We evalu...
|
| 689 |
AgentRecommender: LLM Agents Enable Customizable Recommender Systems on the User Side
2609.31166
|
cs.AI
|
Ryoma Sato |
Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recen...Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recently, user-side recommender systems have been proposed as a new paradigm for solving this problem. If users deploy their own recommender systems, they are no longer at the mercy of the platform's interests. However, building a user-side reco...
|
| 690 |
Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews
2609.31191
|
cs.AI
|
Hariharan Gopinath, Jan Bosch, Helena Holmstr\"om Olsson |
Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners d...Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners define, assess, and manage quality under these conditions. We interviewed 16 practitioners from nine organizations and analyzed the transcripts using reflexive thematic analysis and developed six themes from participants' accounts. In AI sys...
|
| 691 |
Acoustic-to-Text KV Compression for Full-Duplex Speech Models
2609.31224
|
cs.AIcs.SDeess.AS
|
Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim |
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interva...Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inf...
|
| 692 |
Agentic Limit Order Books: Phase Transitions and Market Impact
2609.31260
|
cs.AI
|
Jan Rosenzweig |
We investigate the systemic macroscopic dynamics emerging from Limit Order Books (LOBs) populated exclusively by autonomous reinforcement-learning agentic traders. By formalizing agent interactions within a microscopic order-matching engine, we examine two fun...We investigate the systemic macroscopic dynamics emerging from Limit Order Books (LOBs) populated exclusively by autonomous reinforcement-learning agentic traders. By formalizing agent interactions within a microscopic order-matching engine, we examine two fundamental quantitative phenomena: equilibrium phase transitions in order flow regime shifts, and the structural dynamics of market impact. We show that agentic LOBs exhibit distinct phase boundaries separating orderly price discovery from hy...
|
| 693 |
Cognitive Skills in the Age of AI: Computing Students and Experts Perceptions
2609.31272
|
cs.AI
|
Neha Rani, Vu Minh Anh Le, Austin M. Spangler, Erta Cenko |
AI is becoming increasingly integrated into daily workflows, especially in computing. We are gradually shifting towards an AI-rich future, an impending yet unknown one. One important emerging concern is whether we are accordingly preparing our future computing...AI is becoming increasingly integrated into daily workflows, especially in computing. We are gradually shifting towards an AI-rich future, an impending yet unknown one. One important emerging concern is whether we are accordingly preparing our future computing workforce. Further, we need to know what the important cognitive skills are to remain relevant in the computing workforce and if there are changes in cognitive skill importance. To investigate this direction, we conducted a mixed-methods s...
|
| 694 |
Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains
2609.31282
|
cs.AI
|
Toqeer Ali Syed, Asadullah Abdullah Khan |
This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of speciali...This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of specialised security agents, covering source integrity, dependency and SBOM analysis, CI configura tion auditing, artifact verification, and runtime policy evaluation, each supported by a large language model (LLM) that interprets artefacts, reason...
|
| 695 |
Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows
2609.31301
|
cs.AI
|
Haoran Zhang, Hengtong Zhang, Zhiyu Liang, Yu Yan, Decheng Zuo |
Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can prod...Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can produce persistent effects beyond the requested change. Current safeguards can approve an action or record its aftermath, but without checking the persistent result before continuation, an unapproved outcome can be accepted as success and propa...
|
| 696 |
Towards VLA-Dreamer: Refining VLA Behavior Using World Models
2609.31313
|
cs.AI
|
Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter |
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this...Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and u...
|
| 697 |
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
2609.31318
|
cs.AI
|
Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu |
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as...AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined atta...
|
| 698 |
A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents
2609.31358
|
cs.AI
|
Bennet Gerlach, Stefan Fischer |
The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services...The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services. In medical environments, however, exposing device state and action affordances requires deterministic constraints on possible effects. We present an IEEE 11073 Service-Oriented Device Connectivity (SDC)-to-MCP gateway that exposes metrics...
|
| 699 |
ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
2609.31395
|
cs.AI
|
Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou |
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetri...Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task progress. Our key idea is to establish a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality. However, iterative execution, dynamic m...
|
| 700 |
Can You Check That? The Checkability Boundary for Local LLM Network Automation
2609.31540
|
cs.AI
|
Maleeha Masood, Momina Nofal |
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for d...Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating ...
|
| 701 |
Adapting for AI: How elementary teachers adjust their practices for an AI-integrated curriculum
2609.31569
|
cs.AI
|
Fasika Melese, Ruiyang Wu, Xinyue Cui, Joanna Perkins, Xiaoyi Tian |
Conversational AI tools are entering children's everyday experiences, and schools are interested in adopting them. However, successful classroom integration depends not only on the technology but also on the work teachers do to make it usable and appropriate f...Conversational AI tools are entering children's everyday experiences, and schools are interested in adopting them. However, successful classroom integration depends not only on the technology but also on the work teachers do to make it usable and appropriate for their students and classroom context. There is little known about how elementary teachers work as they implement conversational AI tools in real classrooms. In this study, we examine three teachers' experiences implementing an AI literac...
|
| 702 |
Preference-based opponent shaping in differentiable games
2412.03072
|
cs.AI
|
Xinyu Qiao, Yudong Hu, Congying Han, Weiyan Wu, Tiande Guo |
Strategy learning in game environments with multi-agent is a challenging problem. Since each agent's reward is determined by the joint strategy, a greedy learning strategy that aims to maximize its own reward may fall into a local optimum. Recent studies have ...Strategy learning in game environments with multi-agent is a challenging problem. Since each agent's reward is determined by the joint strategy, a greedy learning strategy that aims to maximize its own reward may fall into a local optimum. Recent studies have proposed the opponent modeling and shaping methods for game environments. These methods enhance the efficiency of strategy learning by modeling the strategies and updating processes of other agents. However, these methods often rely on simp...
|
| 703 |
The Plot Twist: Jailbreaking Unified Multimodal Models with a Three-Act NarrativeAttack
2509.26473
|
cs.AI
|
Shaoxiong Guo, Tianyi Du, Lijun Li, Yuyao Wu, Jie Li |
Unified Multimodal Understanding and Generation Models (UMMs) increasingly combine visual understanding and image generation within a single interactive workflow, making generated visual content available as later reasoning context. However, existing jailbreak...Unified Multimodal Understanding and Generation Models (UMMs) increasingly combine visual understanding and image generation within a single interactive workflow, making generated visual content available as later reasoning context. However, existing jailbreak evaluations mostly study text rewriting or isolated visual prompts, leaving the safety risk of narrative cross-turn visual grounding underexplored. We propose NarrativeAttack, a semantic-preserving visual narrative jailbreak framework. Nar...
|
| 704 |
Flow Reconstruction from Sparse Measurements in Urban Drainage Networks: An Application and Evaluation of Data-Driven Sparse Sensing
2511.04556
|
cs.AI
|
Zihang Ding, Amit Kumar, Imran Md. Azizul Islam, Mila Avellar Montezuma, Ruihang Zhang |
Urbanization and increasingly frequent intense storms are placing stress on urban drainage networks. While dense monitoring of urban drainage networks is desirable, practical constraints in time, budget, and technology hinder its full implementation. How to mo...Urbanization and increasingly frequent intense storms are placing stress on urban drainage networks. While dense monitoring of urban drainage networks is desirable, practical constraints in time, budget, and technology hinder its full implementation. How to monitor and predict flow conditions across the entire network under constrained resources is a major challenge. To address this, we utilized and evaluated an established data-driven sparse sensing (DSS) workflow for sensor placement optimizat...
|
| 705 |
Testing the Utility of Using Large Language Models to Create Personalized Networks From Therapy Session Transcripts: A Proof of Concept Study
2512.05836
|
cs.AI
|
Clarissa W. Ong, Hiba Arnaout, Kate Sheehan, Estella Fox, Eugen Owtscharow |
Recent advances in psychotherapy have focused on treatment personalization, such as by selecting treatment modules based on individual networks. However, estimating personalized networks typically requires intensive longitudinal data, which is not always feasi...Recent advances in psychotherapy have focused on treatment personalization, such as by selecting treatment modules based on individual networks. However, estimating personalized networks typically requires intensive longitudinal data, which is not always feasible to collect. A solution to increase scalability of network-driven treatment personalization is leveraging large language models (LLMs). In this study, we developed an end-to-end pipeline for automatically generating client networks to su...
|
| 706 |
When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models
2602.13215
|
cs.AI
|
Haoran Zheng, Chen Shani |
Recurrent-attention hybrids aim to combine the efficiency of recurrence with the contextual recall of attention, but existing approaches typically apply attention uniformly across all positions, even when the recurrent state alone is sufficient for accurate pr...Recurrent-attention hybrids aim to combine the efficiency of recurrence with the contextual recall of attention, but existing approaches typically apply attention uniformly across all positions, even when the recurrent state alone is sufficient for accurate prediction. We introduce AMOR (Adaptive Metacognitive Output Router), a post-hoc hybrid architecture that selectively invokes attention based on predictive uncertainty. A recurrent backbone is augmented with entropy-gated attention blocks tha...
|
| 707 |
Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System
2602.18640
|
cs.AI
|
Longfei Yun, Yihan Wu, Haoran Liu, Xiaoxuan Liu, Ziyun Xu |
Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product requirements. Progress in this domain is increasingly bottlenecked by the engineering context constraint: the ard...Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product requirements. Progress in this domain is increasingly bottlenecked by the engineering context constraint: the arduous process of translating ambiguous product intent into reasonable, executable, verifiable hypotheses, rather than by modeling techniques alone. We present GEARS (Generative Engine for Agentic Ranking Systems), a framework that reframes r...
|
| 708 |
Jagarin: A Three-Layer Architecture for Hibernating Personal Duty Agents on Mobile
2603.05069
|
cs.AI
|
Ravi Kiran Kadaboina |
Personal AI agents face a deployment paradox on mobile: persistent background execution drains the battery and conflicts with platform background limits, yet purely reactive agents miss time-sensitive obligations until the user remembers to ask. We present Jag...Personal AI agents face a deployment paradox on mobile: persistent background execution drains the battery and conflicts with platform background limits, yet purely reactive agents miss time-sensitive obligations until the user remembers to ask. We present Jagarin, a three-layer architecture that resolves this through structured hibernation and demand-driven wake. DAWN (Duty-Aware Wake Network) is an on-device scoring engine that runs on the platform's periodic wake and combines four signals (du...
|
| 709 |
Attribution Bias in Large Language Models
2604.05224
|
cs.AI
|
Eliza Berman, Bella Chang, Daniel B. Neill, Emily Black |
As Large Language Models (LLMs) are increasingly used to support search and information retrieval, it is critical that they accurately attribute content to its original authors. In this work, we introduce AttriBench, the first fame- and demographically-balance...As Large Language Models (LLMs) are increasingly used to support search and information retrieval, it is critical that they accurately attribute content to its original authors. In this work, we introduce AttriBench, the first fame- and demographically-balanced quote attribution benchmark dataset. By explicitly balancing author fame and demographics, AttriBench enables controlled investigation of demographic bias in quote attribution. Using this dataset, we evaluate 11 widely used LLMs across di...
|
| 710 |
ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval
2605.03361
|
cs.AI
|
Honglei Zhang, Yuting Chen, Chenpeng Hu, Pengfei Zhou, Siyue Zhang |
Existing audio retrieval benchmarks primarily assess semantic matching, lacking comprehensive evaluation of the logical reasoning capabilities required by complex queries. We introduce ReasonAudio, a benchmark for reasoning-intensive Text-Audio Retrieval that ...Existing audio retrieval benchmarks primarily assess semantic matching, lacking comprehensive evaluation of the logical reasoning capabilities required by complex queries. We introduce ReasonAudio, a benchmark for reasoning-intensive Text-Audio Retrieval that evaluates four abilities: negation, temporal order, sound co-occurrence, and sound duration. It comprises five synthetic subtasks with 1,000 queries over 10,000 composite audio clips and a natural subtask with 100 queries over 1,000 real-wo...
|
| 711 |
The Scaling Properties of Implicit Deductive Reasoning in Transformers
2605.04330
|
cs.AI
|
Enrico Vompa, Tanel Tammet |
We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By discouraging the reliance on statistical shortcuts via counterfactual data augmentation, and promoting the learning of shared reasoning pr...We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By discouraging the reliance on statistical shortcuts via counterfactual data augmentation, and promoting the learning of shared reasoning primitives across direct and CoT modes, we find that in sufficiently deep models with a bidirectional prefix mask, implicit reasoning approaches explicit CoT performance across graph topologies and problem widths, though CoT remains necessary...
|
| 712 |
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
2605.04733
|
cs.AI
|
Miao Wang, Yuling Shi, Yijiang Li, Yeheng Chen, Xiaodong Gu |
Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as VR games and interactive narratives. We study video-grounded role-playing dialogu...Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as VR games and interactive narratives. We study video-grounded role-playing dialogue and introduce EBM-RL (Eye--Brain--Mouth Reinforcement Learning), a decoupled GRPO-based framework that separates observation (<perception>), reasoning (<think>), and utterance generation (<answer>). This design mimics the human See-Think-...
|
| 713 |
Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers
2605.05725
|
cs.AI
|
Hyeongwon Kang, Jeongseob Kim, Jinwoo Park, Pilsung Kang |
Time-series anomaly detection often returns scores or intervals, while analysts need to understand the abnormal behavior and the evidence supporting it. We introduce SAGE (Specialized Analyzer Group for Expert-like Detection), a multi-agent framework for evide...Time-series anomaly detection often returns scores or intervals, while analysts need to understand the abnormal behavior and the evidence supporting it. We introduce SAGE (Specialized Analyzer Group for Expert-like Detection), a multi-agent framework for evidence-grounded diagnosis of univariate time series. Four specialized Analyzers examine point, structural, seasonal, and pattern anomalies using numerical tools and diagnostic visualizations. A Detector integrates their evidence into intervals...
|
| 714 |
HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning
2605.05951
|
cs.AI
|
Haoyun Tang, Haodong Cui, Keyao Xu, Zhandong Mei, Kun Wang |
World models support model-based planning through learned latent dynamics, but imagined rollouts can become unstable as the planning horizon grows or the dynamics distribution shifts. We propose HaM-World, a structured world model that combines history-conditi...World models support model-based planning through learned latent dynamics, but imagined rollouts can become unstable as the planning horizon grows or the dynamics distribution shifts. We propose HaM-World, a structured world model that combines history-conditioned selective memory with a Soft-Hamiltonian latent dynamics prior. The latent state is decomposed into a canonical (q,p) subspace and a context subspace c. Mamba selective state-space memory summarizes past observations and actions and co...
|
| 715 |
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents
2605.06869
|
cs.AI
|
Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth |
AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present Agentick, a benchmark for seque...AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present Agentick, a benchmark for sequential decision-making agents designed to evaluate RL, LLM, VLM, hybrid, and human agents on common ground and to power research on the fundamental challenges of sequential decision-making. Agentick provides 37 procedurally generated tasks a...
|
| 716 |
Selective Off-Policy Reference Tuning with Plan Guidance
2605.11505
|
cs.AI
|
Anh Duc, Tien-Phat Nguyen, Thien Huu Nguyen, Linh Ngo Van, Trung Le |
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference...Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instea...
|
| 717 |
Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability
2605.11625
|
cs.AI
|
Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Songlin Liu |
Large reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, yet read the resu...Large reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, yet read the resulting pass rate as a difficulty score, leaving the zero-return regime unmodeled. As a result, they overspend on queries beyond the model's capability while compressing hard-but-solvable ones that need deeper reasoning. In this work, we form...
|
| 718 |
CODESKILL: Learning Self-Evolving Skills for Coding Agents
2605.25430
|
cs.AI
|
Yanzhou Li, Yiran Zhang, Xiaoyu Zhang, Xiaoxia Liu, Yang Liu |
Coding agents produce rich trajectories while solving software-engineering tasks. To enable agent self-evolution, these trajectories can be distilled into reusable procedural skills that compactly encode experience to guide future behavior. However, existing s...Coding agents produce rich trajectories while solving software-engineering tasks. To enable agent self-evolution, these trajectories can be distilled into reusable procedural skills that compactly encode experience to guide future behavior. However, existing skill construction and maintenance methods often rely on fixed prompts and heuristic update rules, leaving it unclear how knowledge should be selected, abstracted, and maintained to best serve downstream agents. We propose CODESKILL, an LLM-...
|
| 719 |
SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
2606.13020
|
cs.AI
|
Pierre Beckmann, Marco Valentino, Andre Freitas |
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly...Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly and lack mechanistic ground truth, while synthetic logical-reasoning benchmarks do not resemble real scientific documents. We introduce SciR, a benchmark that combines multi-paradigm reasoning with controllable scientific rendering, anchor...
|
| 720 |
Do Neural Networks Preserve Case Structure? Case-Based Decomposition, Interpretation, and Decision Consistency
2607.11347
|
cs.AI
|
Manli Yan, Yaowen Yu, Yong Zhao, Yuebin Lin, Shoudong Han |
Neural networks increasingly inform consequential decisions, making their reliability increasingly important. Yet their internal mechanisms provide little evidence of whether decisions remain grounded in the training cases and which cases ultimately support or...Neural networks increasingly inform consequential decisions, making their reliability increasingly important. Yet their internal mechanisms provide little evidence of whether decisions remain grounded in the training cases and which cases ultimately support or oppose their outcomes. Without this connection between decisions and training cases, users cannot determine whether a model has learned reliable decision patterns from data. This motivates a fundamental question: do neural networks preserv...
|
| 721 |
Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware
2608.00008
|
cs.AI
|
Philipp M. Z\"ahl, Elja Dalipaj, Anika Hennig, Timon Bayer |
The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. T...The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. This paper presents a reproducible, hardware-level energy benchmark of 18 open-source LLMs (0.5B to 7B parameters) executed on a single consumer GPU (RTX 4060ti 16GB). Using the Ollama inference engine, GPU power draw was sampled at 2hz via ...
|
| 722 |
State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
2608.03425
|
cs.AI
|
Xiaohe Li, Yang Lu |
Despite the dominance of massive language models, leading paradigms like Transformers and Mamba fundamentally falter at continuous deterministic state tracking, suffering from catastrophic out-of-distribution (OOD) collapse when generalizing to extended sequen...Despite the dominance of massive language models, leading paradigms like Transformers and Mamba fundamentally falter at continuous deterministic state tracking, suffering from catastrophic out-of-distribution (OOD) collapse when generalizing to extended sequences. To shatter this bottleneck, we present the \textbf{Complex State Propagator (CSP)}, a radically minimalist recurrent paradigm that operates strictly on the complex phase manifold without intermediate output projections, amplitude modul...
|
| 723 |
Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform for High Dose Rate (HDR) Brachytherapy
2608.08163
|
cs.AI
|
Ronghua Xu, Kepha Barasa, Manoj Kumal, Xinyun Liu, Weihua Zhou |
The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive simulation ...The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive simulation specifically designed for High Dose Rate (HDR) vaginal cylinder (VC) brachytherapy in cancer care. By integrating Virtual Reality (VR) and mobile computing, the system establishes a high-fidelity, risk-free environment that allows trainees ...
|
| 724 |
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
2608.15565
|
cs.AI
|
Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo |
Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library can be retrieved again and again, and on a stream of new problems there is no ground-truth ...Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library can be retrieved again and again, and on a stream of new problems there is no ground-truth answer to decide with. Existing learners admit trajectories by matching known optima or labels, and label-free substitutes such as execution success or agreement at one instance can admit wrong models. We introduce ADMITOR, a label-free adm...
|
| 725 |
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair
2608.18324
|
cs.AI
|
Jesus Salas |
Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether verifier-admitted outputs can supervise a bounded model by consolidating occasional or expens...Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether verifier-admitted outputs can supervise a bounded model by consolidating occasional or expensive capability into reliable one-shot execution. On fresh, structure-disjoint PlanBench replanning cases, Qwen3-14B thinking produced 24 plans admitted by independently authored VAL. They trained the same checkpoint for non-thinking executi...
|
| 726 |
PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?
2608.25097
|
cs.AIcs.MM
|
Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng |
Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1)...Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their a...
|
| 727 |
T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing
2609.15160
|
cs.AI
|
Mingqian Yu, Wenpeng Zhang, Shaobo Cui, Peilin Zhao |
Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped ...Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped Transformers perform inference directly in the latent space (latent reasoning) to reduce the number of tokens consumed during inference, thereby achieving improved sample efficiency. However, these models typically apply a fixed recursion d...
|
| 728 |
AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control
2609.30264
|
cs.AI
|
Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo |
Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate a...Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized re...
|
| 729 |
Peer Review as Structured Commentary: Immutable Identity, Public Dialogue, and Reproducible Scholarship
2506.22497
|
cs.AI
|
Craig Steven Wright |
This paper reconceptualises peer review as structured public commentary. Traditional academic validation is hindered by anonymity, latency, and gatekeeping. We propose a transparent, identity-linked, and reproducible system of scholarly evaluation anchored in ...This paper reconceptualises peer review as structured public commentary. Traditional academic validation is hindered by anonymity, latency, and gatekeeping. We propose a transparent, identity-linked, and reproducible system of scholarly evaluation anchored in open commentary. Leveraging blockchain for immutable audit trails and AI for iterative synthesis, we design a framework that incentivises intellectual contribution, captures epistemic evolution, and enables traceable reputational dynamics. ...
|
| 730 |
Provable Speech Attributes Conversion via Latent Independence
2510.05191
|
cs.AIcs.SD
|
Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham |
Conditional generation and disentangled representation learning are central to controlled generation across audio, vision, and multimodal domains. However, despite strong empirical progress, particularly in speech style transfer, most existing approaches rely ...Conditional generation and disentangled representation learning are central to controlled generation across audio, vision, and multimodal domains. However, despite strong empirical progress, particularly in speech style transfer, most existing approaches rely on heuristic objectives and architectural choices, offering limited theoretical understanding of when and why reliable attribute control is achievable. In this work, we develop a formal framework for speech attribute conversion and provide ...
|
| 731 |
Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
2511.06701
|
cs.AI
|
Karen Sargsyan |
AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture that enforces statistical rigor at two levels: a Haskell embedded domain-specific language (the Research monad) that makes ...AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture that enforces statistical rigor at two levels: a Haskell embedded domain-specific language (the Research monad) that makes it impossible to test a hypothesis without updating the error budget, and a declarative scaffold that fixes the data flow and the statistical test, together with an OS-level sandbox that makes validation data physically absent from the envi...
|
| 732 |
SLMFix: Leveraging Small Language Models for Domain Specific Language Error Fixing with Reinforcement Learning
2511.19422
|
cs.AI
|
David Jiahao Fu, Aryan Gupta, Aaron Councilman, Yu-Xiong Wang, Vikram Adve |
Large language models (LLMs) have shown impressive capabilities in code generation across many programming languages but even state-of-the-art LLMs generate programs that contain syntactic errors and fail to complete the given tasks, especially for low-resourc...Large language models (LLMs) have shown impressive capabilities in code generation across many programming languages but even state-of-the-art LLMs generate programs that contain syntactic errors and fail to complete the given tasks, especially for low-resource programming languages (LRPLs). In addition, the high cost of training makes finetuning LLMs unaffordable for those with constrained computational resources, further weakening the effectiveness of LLMs for code generation. In this work, we...
|
| 733 |
BabelCoder: Agentic Code Translation with Specification Alignment
2512.06902
|
cs.AI
|
Fazle Rabbi, Soumit Kanti Saha, Tri Minh Triet Pham, Song Wang, Jinqiu Yang |
As software systems evolve, developers increasingly work across multiple programming languages and often face the need to migrate code from one language to another. While automatic code translation offers a promising solution, it has long remained a challengin...As software systems evolve, developers increasingly work across multiple programming languages and often face the need to migrate code from one language to another. While automatic code translation offers a promising solution, it has long remained a challenging task. Recent advancements in Large Language Models (LLMs) have shown potential for this task, yet existing approaches remain limited in accuracy and fail to effectively leverage contextual and structural cues within the code. Prior work h...
|
| 734 |
Initial results of the Digital Consciousness Model
2601.17060
|
cs.AI
|
Derek Shiller, Laura Duffy, Arvo Mu\~noz Mor\'an, Adri\`a Moret, Chris Percy |
Artificially intelligent systems have become remarkably sophisticated. They hold conversations, write essays, and seem to understand context in ways that surprise even their creators. This raises a crucial question: Are we creating systems that are conscious? ...Artificially intelligent systems have become remarkably sophisticated. They hold conversations, write essays, and seem to understand context in ways that surprise even their creators. This raises a crucial question: Are we creating systems that are conscious? The Digital Consciousness Model (DCM) is a first attempt to assess the evidence for consciousness in AI systems in a systematic, probabilistic way. It provides a shared framework for comparing different AIs and biological organisms, and for...
|
| 735 |
Dynamic Welfare-Maximizing Pooled Testing
2601.22419
|
cs.AI
|
Edwin Lock, Nicholas Lopez, Francisco Marmolejo-Coss\'io, Jose Roberto Tello Ayala, David C. Parkes |
Pooled testing uses one test to certify several agents as healthy when the pooled result is negative. We study a budget-constrained welfare problem in which agents have heterogeneous utilities and independent prior probabilities of being healthy. Welfare is ea...Pooled testing uses one test to certify several agents as healthy when the pooled result is negative. We study a budget-constrained welfare problem in which agents have heterogeneous utilities and independent prior probabilities of being healthy. Welfare is earned when an agent is certified healthy, and a dynamic policy may choose each pool after observing earlier test outcomes. We ask how much such adaptation can improve over a static allocation that fixes all pools in advance. Our main result ...
|
| 736 |
HuPER: A Human-Inspired Framework for Phonetic Perception
2602.01634
|
cs.AIeess.AS
|
Chenxu Guo, Jiachen Lian, Yisi Liu, Baihe Huang, Shriyaa Narayanan |
We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five Eng...We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourc...
|
| 737 |
The Shrinking Lifespan of LLMs in Science
2604.07530
|
cs.AI
|
Ana Tri\v{s}ovi\'c |
Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once released. We introduce time-to-peak and lifespan as measures of model obsolescence and use them to characterize the scientific...Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once released. We introduce time-to-peak and lifespan as measures of model obsolescence and use them to characterize the scientific adoption trajectories of 62 LLMs across more than 108k citing papers (2019-2025), separating active adoption from background citation to recover per-model trajectories that citation counts cannot resolve. We find that a model's longevity i...
|
| 738 |
Endogenous Information in Routing Games: Memory-Constrained Equilibria, Recall Braess Paradoxes, and Memory Design
2604.11733
|
cs.AI
|
Saad Alqithami |
We study routing games in which travelers optimize over routes that are remembered or surfaced, rather than over a fixed exogenous action set. The paper develops a tractable design theory for endogenous recall and then connects it back to an explicit finite-me...We study routing games in which travelers optimize over routes that are remembered or surfaced, rather than over a fixed exogenous action set. The paper develops a tractable design theory for endogenous recall and then connects it back to an explicit finite-memory micro model. At the micro level, each traveler carries a finite memory state, receives surfaced alternatives, chooses via a logit rule, and updates memory under a policy such as LRU. This yields a stationary Forgetful Wardrop Equilibri...
|
| 739 |
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
2604.15186
|
cs.AI
|
Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhao, Yanda Tao |
Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpr...Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic wo...
|
| 740 |
An AI Agent Execution Environment to Safeguard User Data
2604.19657
|
cs.AI
|
Robert Sorab Stanley, Avi Verma, Lillian Tsai, Konstantinos Kallas, Sam Kumar |
AI agents promise to serve as general-purpose personal assistants for their users, which requires them to have access to private user data (e.g., personal and financial information). This poses a serious risk to security and privacy: an AI model may hallucinat...AI agents promise to serve as general-purpose personal assistants for their users, which requires them to have access to private user data (e.g., personal and financial information). This poses a serious risk to security and privacy: an AI model may hallucinate or make mistakes, and adversaries may attack it (e.g., via prompt injection) to exfiltrate user data. This paper presents GAAP (Guaranteed Accounting for Agent Privacy), an execution environment for AI agents that guarantees confidentiali...
|
| 741 |
Information Aggregation with AI Agents
2604.20050
|
cs.AI
|
Spyros Galanis |
Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving...Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, across four information structures of increasing complexity. We find that although the median market is effective at aggregating information in the easy information structures, performance deteriorates in the harder struct...
|
| 742 |
Topology-Driven Anti-Entanglement Control for Soft Robots
2605.05236
|
cs.AI
|
Haoyang Le, Shengxuan Wang, Mohan Chen, Shuo Feng |
In the field of precision manufacturing in complex constrained environments, the role of soft robots is increasingly prominent, and the realization of anti-winding control based on multi-intelligent body reinforcement learning has become a research hotspot. On...In the field of precision manufacturing in complex constrained environments, the role of soft robots is increasingly prominent, and the realization of anti-winding control based on multi-intelligent body reinforcement learning has become a research hotspot. One of the core problems at present is to coordinate multiple robots to complete the unwinding operation in a highly constrained environment. The existing distributed training framework faces some observability challenges in high-density barr...
|
| 743 |
Grid-Orch: An LLM-Powered Orchestrator for Distribution Grid Simulation and Analytics
2605.12728
|
cs.AI
|
Boming Liu, Jin Dong, Jianming Lian |
The power distribution engineering workforce faces a projected shortage of up to 1.5 million engineers by 2030, creating urgent demand for more accessible analysis tools. This paper introduces Grid-Orch, a framework that bridges Large Language Models (LLMs) an...The power distribution engineering workforce faces a projected shortage of up to 1.5 million engineers by 2030, creating urgent demand for more accessible analysis tools. This paper introduces Grid-Orch, a framework that bridges Large Language Models (LLMs) and power system simulation through the Model Context Protocol (MCP), enabling engineers to perform complex distribution analyses via natural language. Using OpenDSS as the reference implementation, Grid-Orch provides 36 domain-specific tools...
|
| 744 |
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
2605.30834
|
cs.AI
|
Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira |
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment. Detecting such failures during...Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment. Detecting such failures during execution is therefore critical for the robust deployment of embodied systems. Existing failure detection methods either rely on expensive action resampling or external models, while alternatives propagate trajectory-level labels uniformly...
|
| 745 |
Efficient Safety Benchmarking via Item Response Theory
2606.20626
|
cs.AI
|
Fabio Spagliardi, M\'irian Silva, Ayan Datta, Aiden Zhou, Vamshi Bonagiri |
Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly problematic for adversarial, highly heterogeneous safety items. Applied in full ...Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly problematic for adversarial, highly heterogeneous safety items. Applied in full to modern benchmark suites, current evaluation procedures would require on the order of $10^5$ responses, most of which provide little ranking signal. We analyze six widely used safety benchmarks and make three contributions toward more eff...
|
| 746 |
CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation
2607.02222
|
cs.AI
|
Haokun Liu, Zhaoqi Ma, Yicheng Chen, Wentao Zhang, Masaki Kitagawa |
Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-lev...Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-level vision-language-action framework that predicts a language-conditioned flow field over the robot's local visible sector and generates continuous trajectories by rolling out the predicted field. To train this low-level representation, we c...
|
| 747 |
Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise
2607.26639
|
cs.AI
|
Haoyu Zhang, Hanwen Liu, Yang Chen, Shibo Zheng, Xiangchen Guan |
Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so ...Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so the encoding is the only difference. Against SAGE, the strongest published self-check defense, best-of-N over a code-completion encoding reaches 67, 22 and 15% of behaviors on three open-weight targets, where that encoding fired once reache...
|
| 748 |
Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
2608.06994
|
cs.AI
|
Xiangkai Ma, Yue Ma, Junjie Wang, Sheng Xu, Mingyang Li |
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting...World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical condition evolution and low-level action trajectory generation ...
|
| 749 |
Beyond Forecasting: Recasting Volatility Control as a Routing Problem
2608.10375
|
cs.AI
|
Hongji Pu, Leyang Zhou |
Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-defined control rule that may not adapt to changing market conditions. We propose VolRouter, a modular framework tha...Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-defined control rule that may not adapt to changing market conditions. We propose VolRouter, a modular framework that formulates volatility control as state-conditioned routing over estimator-controller pairs. VolRouter first summarizes market conditions into a control-relevant state profile and then performs routing through three stages: state inference...
|
| 750 |
Keep the Future, Drop the Rollout: RIFT for World Action Models
2608.11521
|
cs.AI
|
Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li |
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on ...World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on 40 simulated robotic manipulation tasks, paired closed-loop interventions show that blocking access to the future cache or reassigning its values changes execution and reduces success. Yet in the evaluated co-denoising settings, reusing one...
|
| 751 |
SUN: Agentic Robot Policy Learning with Persistent Task Programs
2608.31167
|
cs.AI
|
Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao |
Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed e...Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned optimal control objectives, satisfaction predicates, and learning rewards. Our harness, Kuafu, equips a foundation model as a task-level agent to orchestrate scene preparation, verific...
|
| 752 |
Exploring Second-Order Pattern Recognition in Speaker Recognition
2609.11182
|
cs.AIeess.AS
|
Yanze Xu, Wenwu Wang, Mark D. Plumbley |
In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's re...In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's recognition of inputs as human-defined patterns; this work calls these latent patterns second-order patterns and proposes to discover them. Accordingly, we apply a hierarchical clustering algorithm to analyse whether our speaker recognition n...
|
| 753 |
Interpreting hierarchical organisation of speaker embeddings
2609.15203
|
cs.AIeess.AS
|
Yanze Xu, Wenwu Wang, Mark D. Plumbley |
Speaker recognition neural networks recognise speaker identities from input utterances by learning latent representations (i.e. speaker embeddings). However, these networks' internal mechanisms remain largely opaque, motivating research in explainable artifici...Speaker recognition neural networks recognise speaker identities from input utterances by learning latent representations (i.e. speaker embeddings). However, these networks' internal mechanisms remain largely opaque, motivating research in explainable artificial intelligence (XAI) to understand them. Existing studies have analysed how speaker embeddings are organised, but rarely frame these analyses within XAI. This work proposes to explain and interpret the organisation of speaker embeddings fr...
|
| 754 |
The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties
2609.26174
|
cs.AI
|
Haoyu Zhang, Yi Feng, Hanwen Liu, Shibo Zheng, Zhuoxi Wang |
We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request,...We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition, shifts benign refusal by tens of points. The shift is not blanket caution but a threshold shift: genuinely neutral instructions are almost unaffected (<=2 percentage points on thr...
|
| 755 |
G\"odel's and Scott's Variants of the Ontological Argument in Lean 4 and TPTP THF
2609.26806
|
cs.AI
|
Christoph Benzm\"uller |
This paper presents a complete, structure-preserving port to Lean 4 of the Isabelle/HOL dataset accompanying Benzm\"uller and Scott's study of G\"odel's ontological argument and Scott's variant: 30 modules, one per theory, retaining section structure, declarat...This paper presents a complete, structure-preserving port to Lean 4 of the Isabelle/HOL dataset accompanying Benzm\"uller and Scott's study of G\"odel's ontological argument and Scott's variant: 30 modules, one per theory, retaining section structure, declaration order and names up to documented renamings; a comparison tool certifies the 548 statements identical as parsed. Every named result the original proves is proved again, from the inconsistency of G\"odel's 1970 axioms to modal collapse, m...
|
| 756 |
AI in Science: Early Insights
2609.28504
|
cs.AI
|
Mihai Codreanu, Alex Imas, Juan Mateos-Garcia, Joseph Emmens, Evalyne Muiruri |
Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million...Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings...
|
| cs.CL 127 papers | ||||
| 183 |
A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
2609.30287
|
cs.CLcs.AI
|
Pawe{\l} Blicharz, Mi{\l}osz Grunwald |
AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the ...AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers x 768), which we call neurons. T...
|
| 184 |
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
2609.30288
|
cs.CLcs.LG
|
Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii |
In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their ...In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operat...
|
| 185 |
Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
2609.30289
|
cs.CL
|
Yufei Shi, Rujing Yao, Ang Li, Yang Wu, Zhuoren Jiang |
In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations, execution traces, and intermediate...In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations, execution traces, and intermediate progress. Existing memory-augmented systems typically retrieve from all stored memories as a flat pool, ranking them by semantic relevance, importance, or recency without modeling hierarchical structure or evolving validity. As a result, t...
|
| 186 |
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
2609.30290
|
cs.CLcs.LG
|
Haowei Liu, Hsin-Tai Wu, Yi Fang |
Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreem...Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set. Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION. A self-hosted Qwen3.6-27B re...
|
| 187 |
A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models
2609.30292
|
cs.CLcs.AI
|
Fanji Yang (Guizhou University of Finance and Economics), Huiyao Chen (Harbin Institute of Technology), Xi Yu (Guizhou University of Finance and Economics), Meishan Zhang (Harbin Institute of Technology), Xiaohong Xiao (Guizhou University of Commerce) |
Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust mechanisms.The rise of large...Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust mechanisms.The rise of large language models, or LLMs, has changed the problem in two directions.LLMs can generate fluent and context-aware deceptive reviews, while pre-trained language models, or PLMs, and LLMs also provide stronger semantic representations for detec...
|
| 188 |
Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
2609.30293
|
cs.CLcs.AI
|
Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong |
The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog...The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog traversal to $O(k)$ progressive disclosure. Cartograph combines three mechanisms: (1) operator-attested capability cards, Ed25519-signed descriptions generated under the deploying operator's control rather than ranked publisher copy; (2) R...
|
| 189 |
SlideLab: Audience-Centered Scientific Slide Generation and Evaluation
2609.30294
|
cs.CLcs.AI
|
Vidushee Vats, Karun Sharma, Yuxia Wang |
Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework...Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement,...
|
| 190 |
SignTrace: Describe a Sign, Find the Word
2609.30295
|
cs.CLcs.AI
|
Zengji Tu, Xingye Zhu, Ningjing Wang, Tingyi Huang, Yangjunfeng Zhu |
Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access to a Chinese sign-language dic...Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access to a Chinese sign-language dictionary. The system integrates LLM-based dictionary enrichment, action extraction, dictionary-style rewriting, seven-channel retrieval, and candidate reranking over 6,699 entries. It has been deployed for user trials and has received positi...
|
| 191 |
Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
2609.30297
|
cs.CLcs.AI
|
Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri\`a Casas Escoda |
Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is o...Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is optimizing agent planning -- deciding how to select, sequence, and invoke tools -- particularly in cold-start settings where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a...
|
| 192 |
A Benchmark Framework for Screening Automation in Systematic Reviews
2609.30298
|
cs.CL
|
Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani |
Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance cl...Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets.This paper presents a benchmark dataset of $45\,064$ labeled entries for evaluating L...
|
| 193 |
A Unified Account of Concepts and Chunks
2609.30414
|
cs.CLcs.AI
|
Karthik Singaravadivelan, Pat Langley |
Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses ...Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses a challenge for unified theories of cognition. In this paper, we review Cobweb, a computational account of categorization and concept formation, and propose an extended theory that incorporates chunks and their acquisition. The theory makes...
|
| 194 |
All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation
2609.30416
|
cs.CL
|
Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu |
Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with hig...Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel d...
|
| 195 |
Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
2609.30439
|
cs.CLcs.SDeess.AS
|
Bo Su, Yueru Yan, Thai Le |
We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires a...We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) modul...
|
| 196 |
Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
2609.30467
|
cs.CL
|
Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi |
Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of r...Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neithe...
|
| 197 |
CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
2609.30471
|
cs.CLcs.LGcs.AI
|
Mukul Chhabra, Shail Patel, Luigi Medrano |
Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a...Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved r...
|
| 198 |
Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs
2609.30492
|
cs.CLcs.AI
|
Sang Bin Moon, Nicole Cho, Daniel Borrajo, Sumitra Ganesh, Abolfazl Hashemi |
Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning prob...Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, u...
|
| 199 |
Inquesto Score: A reliability Protocol For Voice Agents
2609.30514
|
cs.CLcs.AI
|
Massa Baali, Bhiksha Raj |
Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto Score (IS), a protocol for mea...Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto Score (IS), a protocol for measuring voice-agent reliability as the percentage of calls in a fixed, versioned evaluation population that achieve the caller's goal without a functional failure or worse. Rather than combining heterogeneous metrics, IS defines explicit fai...
|
| 200 |
Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
2609.30535
|
cs.CL
|
Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kiant\'e Brantley |
Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word mul...Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched...
|
| 201 |
REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles
2609.30547
|
cs.CL
|
Haixu Ma, Aditya Bansal, Shubham Lohiya, Sumit Ranjan |
Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant del...Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant delays, estimation errors, and poor scalability over high-dimensional profile data. We present REALMS (Real-time Exact Audience sizing via LLM-based Multi-attribute Search), a conversational system for exact audience sizing deployed in product...
|
| 202 |
Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
2609.30558
|
cs.CLcs.LGcs.AI
|
Jiaqi Ding, Guorong Wu |
Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework ...Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactiva...
|
| 203 |
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
2609.30604
|
cs.CLcs.AI
|
Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung |
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates progr...Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a be...
|
| 204 |
Recursive Self-Improvement via On-Policy Distillation for Reasoning
2609.30652
|
cs.CL
|
Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu |
On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-dist...On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns ...
|
| 205 |
Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4
2609.30716
|
cs.CLcs.AI
|
Amanda Fitch |
When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b...When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b model across a targeted behavioral suite (n = 13 items, 784 forward passes in short, single-turn contexts) using a completely counterbalanced experimental design. This setup allowed us to mathematically isolate the specific effects of sour...
|
| 206 |
Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction
2609.30738
|
cs.CLcs.LG
|
Tianfang Xie, Wei Zhu |
KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and ...KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and redundancy relative to selected tokens. For $\lambda_2<0$, the score penalizes similarity to selected tokens as in maximal marginal relevance (MMR), without extra forward passes. To test whether this relevance-diversity balance should vary ...
|
| 207 |
SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages
2609.30739
|
cs.CL
|
Puja Ahmad Habibi, Faiz Assabil Firdaus, Ashvanth S, Ekapol Chuangsuwanich, Pume Tuchinda |
Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce...Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and mul...
|
| 208 |
Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
2609.30773
|
cs.CLeess.AS
|
Manato Yaguchi, Yotaro Kubo, Hikaru Asano, So Kuroki |
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still s...Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simulator LLM adds substantial data-preparation overhe...
|
| 209 |
Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment
2609.30802
|
cs.CL
|
Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel, Yi Ding |
Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from tea...Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from teacher to student remains largely unexplored. Thus, we fill this gap by analyzing how different template configurations influence the pre-existing safety alignment of the student. We observe a significant degradation of safety alignment prese...
|
| 210 |
I-Parakeet: Integer-Only Conformer ASR on Mobile NPU
2609.30846
|
cs.CLcs.SDeess.AS
|
Taichi Nishimura |
In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard to deploy on edge devices be...In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard to deploy on edge devices because of their size, and quantized models still fall back to floating point for numerically sensitive operations. This prevents them from fully exploiting integer accelerators such as mobile NPUs. To achieve this, our contributions are thre...
|
| 211 |
Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
2609.30849
|
cs.CL
|
Phuong Q. Le, Kemal Kurniawan, Jey Han Lau |
Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to ...Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across ...
|
| 212 |
Persistent Negatives for Adversarial Black-Box On-Policy Distillation
2609.30864
|
cs.CLcs.AI
|
Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu |
Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher...Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We ...
|
| 213 |
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
2609.30867
|
cs.CL
|
Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia |
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported eviden...Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-fla...
|
| 214 |
Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos
2609.30882
|
cs.CL
|
Yuya Wake, Sho Tsugawa, Toshiyuki Amagasa |
Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim screening. This study examines how su...Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim screening. This study examines how such transcript compression affects LLM-based veracity classification of Japanese medical YouTube videos. We compare four transcript input designs: full transcripts, LLM-generated summaries, RAPTOR-based retrievalaugmented generation (RAG), a...
|
| 215 |
From annotation to reasoning: Culture in language models
2609.30897
|
cs.CL
|
Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen |
How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain...How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpr...
|
| 216 |
ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
2609.30906
|
cs.CL
|
Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu |
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful ...Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively searc...
|
| 217 |
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
2609.30924
|
cs.CLcs.SDeess.AS
|
Hikaru Asano, Yotaro Kubo, So Kuroki |
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only par...Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integra...
|
| 218 |
Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models
2609.30935
|
cs.CLcs.LGcs.AI
|
Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li |
Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose kn...Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose knowledge. Although existing methods, such as orthogonal gradient projection, mitigate the forgetting across various fine-tuning tasks, they fundamentally fail to preserve pre-training LLMs' inherent general-purpose knowledge because the orig...
|
| 219 |
FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation
2609.30968
|
cs.CL
|
Lu Han, Jingyao Zhang, Katy Ilonka Gero, Nguyen H. Tran |
Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated parameter-efficient fine-tun...Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated parameter-efficient fine-tuning (PEFT) offers a data-local setting for this multi-author adaptation problem: clients keep author text local while sharing compact adapter updates. However, we show that standard aggregation can preserve continuation utility while making...
|
| 220 |
Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change
2609.30974
|
cs.CL
|
Haruka Ezoe, Ryohei Hisano |
Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or w...Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or which usages support the attribution. We introduce Coupled Usage--Sense Processes (CUSP), which derives these answers from a single marginal preserving temporal process. A hierarchical coupling relates contextual distributions through latent...
|
| 221 |
THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer
2609.30984
|
cs.CL
|
Seanghay Yath |
Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur insid...Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur inside ordinary words. We present Tha, a Khmer text normalization and inverse text normalization toolkit built from weighted finite-state transducers. It segments and classifies a whole line in one shortest-path search, and a second transducer r...
|
| 222 |
Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
2609.30986
|
cs.CL
|
Geng Liu, Feng Li, Mengxiao Zhu, Francesco Pierri |
As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect...As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users' confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or...
|
| 223 |
ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
2609.31002
|
cs.CL
|
Siqiao Xue, Shuxuan Liu, Ning Hu |
Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are diffi...Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are difficult to supervise at scale: real search traffic provides authentic queries and candidates but no clean pairwise labels. We present ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, and 8B) aligned to judge-labeled shopping pre...
|
| 224 |
G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
2609.31009
|
cs.CLcs.AI
|
Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai |
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitat...Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds....
|
| 225 |
Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference
2609.31046
|
cs.CL
|
\"Ozge Alacam, Z\"ubeyde Demet Kirbulut G\"une\c{s}, Funda Ekici, Nurcan Turan-Oluk, Dilay Din\c{c}demir |
Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. W...Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. We investigate whether instruction-tuned large language models (LLMs) can support multidimensional analysis of collaborative sensemaking without task-specific training, and whether structured knowledge-state information improves model infere...
|
| 226 |
CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
2609.31062
|
cs.CL
|
Marko \v{R}eh\'a\v{c}ek, V\'it\v{e}zslav Du\v{s}ek, Martin Rusinko, V\'it Nov\'a\v{c}ek |
Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We ...Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We propose Clinical Guardrail Probes (CG-Probes) to measure the risks from query embeddings. We probe for each axis in the normalized embedding space of frozen embedders via the difference-in-means method, treating each axis as a potential lin...
|
| 227 |
LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
2609.31122
|
cs.CLcs.LG
|
Irene Tallini, Lorenzo Basile, Valentino Maiorca, Francesco Locatello, Alberto Cazzaniga |
Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space...Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties present in the contrastive data and degrade unrelated capabilities. To mitigate this issue, we introduce LocUS (Localized Unembedding Steering), a method which grounds activation ...
|
| 228 |
Do we need to answer that question? Salience and Answerability of Potential Questions in Naturalistic Dialogue
2609.31130
|
cs.CL
|
Amandine Decker (LORIA, UL, CNRS, SEMAGRAMME, GU) |
We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we construct a dataset of 7,124 ...We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we construct a dataset of 7,124 questions automatically generated from utterances and preceding context from the British National Corpus, and annotated for salience and answerability. We find a robust but low positive correlation between salience and answerability in dial...
|
| 229 |
Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting
2609.31169
|
cs.CLcs.AI
|
Pawe{\l} M\k{a}ka, Piotr Andruszkiewicz, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis |
Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore th...Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore this information. Therefore, increasing their visual sensitivity remains an active research area. In this work, we introduce a training method, Metric-based Loss Weighting, that improves visual grounding of translations by increasing the loss...
|
| 230 |
Where a Model Sends Its Own Repeated Token
2609.31181
|
cs.CL
|
Nicol\'as Vera Z\'u\~niga |
Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where th...Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax p(. | t, t) in one forward pass; the result is a map on the whole vocabulary, with two halves. The first -- which tokens are fixed points -- is partially anticipated, and we repor...
|
| 231 |
RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models
2609.31245
|
cs.CL
|
Pavithra P M Nair, Bhavik Talaviya, Shourya Bhushan, Rahul Pankajakshan, Seema Guruvadoo |
Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social...Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categorie...
|
| 232 |
PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding
2609.31255
|
cs.CL
|
Jeonghun Yoon, Dongchan Kim, Hongyeon Yu, Young-Bum Kim, Jaegul Choo |
General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretio...General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretion, and a three-month glucose trend cannot be answered by text similarity. We present PIA, a personal intelligence agent deployed alongside a consumer health agent. PIA receives the agent's natural-language requests, decides for itself wheth...
|
| 233 |
MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries
2609.31261
|
cs.CLcs.AI
|
Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro |
The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximat...The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, with the relevant interaction geometry learned from data: natural-language dependencies are input-dependent and difficult to prescribe in advance, so the model should learn where posi...
|
| 234 |
Identifying Scientists on X
2609.31264
|
cs.CL
|
Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze |
With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientist...With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forest...
|
| 235 |
Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers
2609.31342
|
cs.CL
|
Md Shamim Ahmed, Lukas Galke Poech, Richard R\"{o}ttger |
Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document poisoning, in which outdated e...Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document poisoning, in which outdated evidence makes a model wrong despite answering correctly without retrieval. We construct a benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy, grounded in dated official sources. Across 12 model...
|
| 236 |
Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding
2609.31382
|
cs.CLcs.AI
|
Zhaoyuan Xia (Peking University, Baidu Inc), Qinghongbing Xie (Tsinghua University), Yung Xiang Hue (Tsinghua University), Jianguang Jiang (Baidu Inc) |
Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Su...Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that first identifies source-grounded, question-relevant evidence and then integrates it into a compact, question-conditioned summary before producing the final answer. To train this behavior, ...
|
| 237 |
ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs
2609.31448
|
cs.CLcs.AI
|
Junyi Gao, Yu Shi, Pingzhao Hu, Ewen M Harrison |
Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from structured, high-dimensional clini...Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from structured, high-dimensional clinical time series. Improving this ability would connect risk estimation with flexible questions about a patient's evolving condition. We introduce ViSTA, a compact adapter that incorporates irregular numerical measurements into a pretrained v...
|
| 238 |
Evaluating Cultural Awareness of LLMs for Haitian Creole
2609.31506
|
cs.CLcs.AI
|
Christelle Clervilsson, Yanzhu Guo |
Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present...Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specific...
|
| 239 |
Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge
2609.31511
|
cs.CL
|
Yahya Mohamed Elnawasany |
We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retriev...We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retrieval layer routed across six Model Context Protocol servers, we report three things a research prototype typically lacks. First, a released family of fine-tuned Arabic Islamic model artifacts: an efficient tool-routing LLM (Muslim-6B-PRO, 5.9...
|
| 240 |
Strategically Diverse Sampling for Self-Training
2609.31571
|
cs.CL
|
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata |
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID re...Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constr...
|
| 241 |
PALM: Point-in-Time Adaptation for Financial Language Models
2609.30316
|
cs.CLcs.LG
|
Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han |
Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time (PIT) language models are pretrai...Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time (PIT) language models are pretrained on chronologically filtered corpora and released as one checkpoint per calendar year, each with a documented cutoff. However, each additional year costs a full pretraining run, and whether that run is necessary has never been tested. In...
|
| 242 |
When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
2609.30328
|
cs.CLcs.AI
|
Salma Roshdy Aly, Hussein Assaf, Ziad Kobti |
When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decompose...When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independ...
|
| 243 |
RAZOR: Pruning Replaceable Experts in LLMs
2609.30465
|
cs.CLcs.LG
|
Mingyang Song, Mao Zheng |
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet a...Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method ...
|
| 244 |
Asymmetric Classifier-Free Guidance for Target-Speaker ASR
2609.30476
|
cs.CLcs.SDeess.AS
|
Yiwen Guan, Jacob Whitehill |
Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time ca...Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized mu...
|
| 245 |
AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth
2609.30483
|
cs.CLcs.LGcs.SDeess.AS
|
Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj |
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines th...Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, a...
|
| 246 |
Don't CLAP: Are Music-Text Models Bag-of-Words?
2609.30540
|
cs.CLcs.SDeess.AS
|
Yuan-Chiao Cheng, Alexander Lerch |
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask h...Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real...
|
| 247 |
Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content
2609.30563
|
cs.CLcs.AI
|
Ljubisa Bojic, Tijana Stanic, Joerg Matthes, Agariadne Dwinggo Samala, Bojana Dinic |
Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has ...Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reacti...
|
| 248 |
Epstein Files Engine: Agentic Search for Investigative Journalism
2609.30611
|
cs.CL
|
Duy K. Nguyen, Teresa Mondr\'ia Terol, Dylan Freedman, Zach Seward |
On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files...On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files. The Engine translated reporter questions into Google BigQuery SQL queries across three corpora: Epstein-related releases, the Times's archive and external, Epstein-related news headlines. It used an LLM to plan queries and returned citati...
|
| 249 |
Prompt Injection Detection for Email Agents Through Attack Chain Modeling
2609.30657
|
cs.CL
|
Ahmad Hashmi, Dhyey Patel, Yunting Yin |
Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this ...Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text ...
|
| 250 |
LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information
2609.30706
|
cs.CLcs.AI
|
Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay |
"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what se..."System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya with Value-Of-Information Routing), which places the candidate pieces of missing information (slots) in the input next to the answer options, so that one forward pass returns both...
|
| 251 |
Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
2609.30784
|
cs.CLcs.SDeess.AS
|
Yotaro Kubo, Qi Sun, Yujin Tang |
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directl...This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the propos...
|
| 252 |
Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
2609.30820
|
cs.CLcs.LG
|
Nux Li |
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual lo...Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Control...
|
| 253 |
Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces
2609.30914
|
cs.CL
|
Aman Mittal, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Abhishek Singh, Aditya Singh |
Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and evolving them through rotation-gate o...Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and evolving them through rotation-gate operators. This approach offers higher optimization performance without physical qubits, and has been shown to achieve order-of-magnitude speedups (10--80$\times$) over traditional solvers on combinatorial, high-dimensional NP-hard problems....
|
| 254 |
Does Uniform Discrete Diffusion Need Time?
2609.30977
|
cs.CLcs.LGcs.AI
|
Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye |
Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much...Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence remains much closer to its original clean sequence than t...
|
| 255 |
Same Text, Different Numbers: The Divergence of LLM-Based Measures
2609.31013
|
cs.CLcs.AI
|
Hamid Boustanifar, Sasan Mansouri |
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, manag...Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, a...
|
| 256 |
KuaFu: Compressing Long User Behavior into Understanding at Billion Scale
2609.31045
|
cs.CLcs.LG
|
Jiahao Hui, Lin Zhu, Yishen Hu, Jingdong Shu, Zetai Jiang |
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the ful...Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of tho...
|
| 257 |
JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
2609.31142
|
cs.CLcs.AI
|
Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng |
Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness ...Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can...
|
| 258 |
Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring
2609.31293
|
cs.CLcs.SDeess.AS
|
Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin |
Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper invest...Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across c...
|
| 259 |
The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
2609.31341
|
cs.CLcs.AI
|
Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst |
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive ...Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration....
|
| 260 |
Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models
2609.31397
|
cs.CLcs.AI
|
Andrea Masini, Sudipta Acharya, Paolo Bellavista, Luca Foschini, Burak Kantarci |
Automated and highly usable Quality-of-Service (QoS) enforcement requires translating high-level service intents into deployable traffic-management policies. Although intent-based networking (IBN) has simplified policy specification, bridging the gap between b...Automated and highly usable Quality-of-Service (QoS) enforcement requires translating high-level service intents into deployable traffic-management policies. Although intent-based networking (IBN) has simplified policy specification, bridging the gap between business-level intents and executable network configurations remains complex, error-prone, and difficult to automate. This paper presents Intent2Tc, a closed-loop language-model-driven framework that translates business-level traffic-shaping...
|
| 261 |
Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis
2609.31422
|
cs.CLcs.AI
|
Jakub Mas{\l}owski, Jaros{\l}aw A. Chudziak |
Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summariz...Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summarizing models are prone to fabricating smoothly written debate consensus that is not grounded in the debate's history. To address this safety gap, this paper presents empirical research and studies if the introduction of active post-debate ver...
|
| 262 |
PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
2609.31468
|
cs.CLcs.AI
|
Pavel Kireyev |
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice se...LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model,...
|
| 263 |
Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing
2609.31513
|
cs.CL
|
Marco Mandap |
We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fu...We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count thr...
|
| 264 |
Two Conformal Constructions for Adaptive Within-Document AI-Text Screening
2609.31547
|
cs.CL
|
Marco Mandap, Jerahmeel Hipolito, Arcel Galvez, Charlie Margaret Balagtas, Michael Joshua Buluran |
We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-samp...We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector...
|
| 265 |
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
2609.31587
|
cs.CLcs.AI
|
Md Shohel Arman, Igor Molybog |
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passe...We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen file...
|
| 266 |
User Model Extraction via Belief Self-Distillation
2609.31603
|
cs.CLcs.LG
|
Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata |
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework tha...Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external a...
|
| 267 |
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
2609.31619
|
cs.CLcs.LGcs.AI
|
Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu |
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning duri...Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fin...
|
| 268 |
NaijaNLP: A Survey of Nigerian Low-Resource Languages
2502.19784
|
cs.CLcs.AI
|
Isa Inuwa-Dutse |
With over 500 languages in Nigeria, three languages - Hausa, Yor\`ub\'a and Igbo spoken by more than 175 million people, account for about 65% of the languages. However, these languages are classed as low-resource due to insufficient digital resources to suppo...With over 500 languages in Nigeria, three languages - Hausa, Yor\`ub\'a and Igbo spoken by more than 175 million people, account for about 65% of the languages. However, these languages are classed as low-resource due to insufficient digital resources to support tasks in computational linguistics. While several research efforts and initiatives have been presented, a coherent understanding of the state of classic Natural Language Processing (NLP) spanning grammatical formalisation to linguistic r...
|
| 269 |
MedHal: a Synthetic Dataset for Medical Hallucination Detection
2504.08596
|
cs.CLcs.AI
|
Fabrice Lamarche, Gaya Mehenni, Neshat Elhami Fard, Odette Rios-Ibacache, Li Ming Wang |
Hallucination, the generation of non factual content by AI systems, poses serious risks in medical contexts, where errors can directly affect patient outcomes. We present MedHal, a large-scale dataset specifically designed to assess capabilities and train mode...Hallucination, the generation of non factual content by AI systems, poses serious risks in medical contexts, where errors can directly affect patient outcomes. We present MedHal, a large-scale dataset specifically designed to assess capabilities and train models on the task of hallucination detection in medical texts. Current hallucination detection methods face significant limitations when applied to specialized domains like medicine, where they can have disastrous consequences. MedHal addresse...
|
| 270 |
Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning
2505.09738
|
cs.CLcs.AI
|
Shaurya Sharthak, Vinayak Pahalwan, Adithya Kamath, Adarsh Shirawalmath |
Pretrained language models (LLMs) are often constrained by their fixed tokenization schemes, leading to inefficiencies and performance limitations, particularly for multilingual or specialized applications. This tokenizer lock-in presents significant challenge...Pretrained language models (LLMs) are often constrained by their fixed tokenization schemes, leading to inefficiencies and performance limitations, particularly for multilingual or specialized applications. This tokenizer lock-in presents significant challenges. standard methods to overcome this often require prohibitive computational resources. Although tokenizer replacement with heuristic initialization aims to reduce this burden, existing methods often require exhaustive residual fine-tuning ...
|
| 271 |
Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
2601.01842
|
cs.CL
|
Yusuke Ide, Adam Nohejl, Joshua Tanner, Hitomi Yanaka, Christopher Lindsay |
Dictionary definitions are an essential resource for learning word senses, but manually creating them is costly. We thus study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Specifically, we ...Dictionary definitions are an essential resource for learning word senses, but manually creating them is costly. We thus study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Specifically, we address learner's dictionary definition generation (LDDG), where definitions should be written using simple vocabulary. First, we introduce a reliable evaluation approach for DDG, based on newly proposed evaluation criteria and powered by a...
|
| 272 |
Affective Flow Language Model for Emotional Support Conversation
2602.08826
|
cs.CLcs.AI
|
Chenghui Zou, Ning Wang, Tiesunlong Shen, Luwei Xiao, Chuan Ma |
Large language models (LLMs) have advanced emotional support conversation, but existing alignment methods rely mainly on sparse preferences at the response level or outcomes at the dialogue level, providing limited supervision for sequential strategy decisions...Large language models (LLMs) have advanced emotional support conversation, but existing alignment methods rely mainly on sparse preferences at the response level or outcomes at the dialogue level, providing limited supervision for sequential strategy decisions in multi-turn interactions. This raises a key question: how can detailed process signals be derived from overall dialogue outcomes to guide the gradual adaptation of support strategies? We propose the Affective Flow Language Model (AFlow),...
|
| 273 |
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
2603.18863
|
cs.CL
|
Yana Veitsman, Yihong Liu, Hinrich Sch\"utze |
Cross-lingual alignment is often assumed to improve cross-lingual transfer by bringing representations of different languages closer together. However, improvements in representational alignment do not consistently translate into better downstream performance....Cross-lingual alignment is often assumed to improve cross-lingual transfer by bringing representations of different languages closer together. However, improvements in representational alignment do not consistently translate into better downstream performance. We investigate this disconnect using XLM-R models explicitly aligned across four language pairs with token-level, sentence-level, and masked-language-modeling objectives. We evaluate their zero-shot transfer on a token-level task (part-of-...
|
| 274 |
Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
2604.06487
|
cs.CL
|
Thibault Ba\~neras-Roux, Sergio Burdisso, Esa\'u Villatoro-Tello, Dairazalia S\'anchez-Cort\'es, Shiran Liu |
Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with te...Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with text-only data. However, this introduces a modality gap, as the LLM is not exposed to the noisy representations produced by the speech projector. We investigate whether small amounts of speech can mitigate this mismatch. We compare three stra...
|
| 275 |
VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
2604.21375
|
cs.CLcs.AI
|
Qijun Han, Haoqin Tu, Zijun Wang, Haoyue Dai, Yiyang Zhou |
Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without recovery. We present VLAA-GUI, a modu...Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without recovery. We present VLAA-GUI, a modular GUI agentic framework built around three integrated components that guide the system on when to Stop, Recover, and Search. First, a mandatory Completeness Verifier enforces UI-observable success criteria and verification at every finish...
|
| 276 |
Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations
2604.23295
|
cs.CLcs.AI
|
Bhaskar Singh, Manas Dhir, Shobhit Banga, Pranav Sharma |
Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible full-duplex spoken dialo...Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible full-duplex spoken dialogue system for Hindi by adapting Moshi, a state-of-the-art duplex speech architecture, using a custom Hindi tokeniser and training on 26,000 hours of real spontaneous conversations collected from 14,695 speakers with separate speaker channe...
|
| 277 |
UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification
2605.06221
|
cs.CL
|
Qihang Fan, Huaibo Huang, Zhiying Wu, Bingning Wang, Ran He |
As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing, several novel low-complexity hybrid ...As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing, several novel low-complexity hybrid architectures have recently been proposed, effectively alleviating the computational burden of long-context inference. However, existing research on long-context prefill acceleration remains predominantly focused on sparse attention mechani...
|
| 278 |
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
2605.06642
|
cs.CLcs.AI
|
Xiangyuan Xue, Yifan Zhou, Zidong Wang, Shengji Tang, Philip Torr |
Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both exploration and credit assignment over exte...Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both exploration and credit assignment over extended trajectories. In this work, we present Strategic Trajectory Abstraction (StraTA), a simple framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning (RL). StraTA samples a compact strategy from...
|
| 279 |
Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering
2605.23497
|
cs.CL
|
Max Prior, Andreas Schultz, Matthias Grabmair |
Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds with the evolving nature of statutory law. We study two temporal failure modes: post-cutoff staleness, wher...Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds with the evolving nature of statutory law. We study two temporal failure modes: post-cutoff staleness, where models apply superseded rules after legislative amendments, and recency bias, where models prefer newer provisions even when a historical version governs the fact pattern. To this end, we present a benchmark of 312 expert-validated, time-...
|
| 280 |
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
2605.23975
|
cs.CLcs.SD
|
Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun |
Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language omission, translation-instead-of-transc...Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language omission, translation-instead-of-transcription, and hallucination. We apply Direct Preference Optimization (DPO) to align models, constructing preference pairs in which chosen responses preserve mixed-language content while rejected responses mimic failure patterns. Training thr...
|
| 281 |
Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation
2605.24534
|
cs.CL
|
Max Prior, Niklas Wais, Matthias Grabmair |
We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any handcrafted doctrinal framework. Using 4.555 decisions of the German Federal Court of Justice that cite sect...We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any handcrafted doctrinal framework. Using 4.555 decisions of the German Federal Court of Justice that cite sections 242, 280, 812 and 823 of the German Civil Code (BGB), we extract paragraph-level chunks, summarize their reasoning, and derive keywords, which are embedded and clustered. For each cluster, an LLM generates headings and synthesizes cita...
|
| 282 |
Large Language Model Selection with Limited Annotations
2605.24981
|
cs.CLcs.LG
|
Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel |
Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active ...Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active model selection of LLMs. SELECT-LLM aims to find a small set of queries whose annotations are most informative for identifying the best LLM for a given task. To this end, we introduce a query selection rule based on expected information gai...
|
| 283 |
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
2606.10829
|
cs.CLcs.AI
|
Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro |
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Ex...Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-\(k\), Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a tr...
|
| 284 |
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
2607.29238
|
cs.CLcs.AI
|
Antorweep Chakravorty |
InMyStyle is a privacy-first, single-user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to ...InMyStyle is a privacy-first, single-user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to construct paired training examples and fine-tunes LoRA adapters on Qwen2.5 models ranging from 0.5B to 7B parameters. Length-aware generation budgets and automatic chunking support inputs of different lengths. We report a single-user case s...
|
| 285 |
State of Thought Enables Endogenous Reasoning
2609.16055
|
cs.CLcs.AI
|
Zhiren Gong, Yikun Hou, Zihao Zeng, Ming Xiao, Chau Yuen |
Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through cost...Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansion in constrained search spaces, limiting both generalization and efficiency. We propose State of Thought (SoT), a new reasoning paradigm that enables endogenous reasoning in LLMs, with the model's internal reasoning state governi...
|
| 286 |
Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining
2609.21362
|
cs.CL
|
Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen |
Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated to...Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequ...
|
| 287 |
Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER
2609.21663
|
cs.CL
|
Hritika Sharma, Thibault Ba\~neras-Roux, Alessandra Pinto, Petr Motlicek, Hyunggu Jung |
Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how hu...Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an English dataset for human-centered ASR evaluation. Using this dataset, we benchmark lexical metrics against several configurations of BERTScore and SemDist, varying the language mo...
|
| 288 |
Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts
2609.22455
|
cs.CL
|
Hope McGovern, Anna Dolganov, Samuel Belkadi, Guillaume Kunsch, Dimitris Vlitas |
We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans with...We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans without requiring oracle knowledge of their length. To our knowledge, it is the first large-scale decoder model for historical Greek, and the first for any ancient Mediterranean language. Evaluated as in prior work, on short gaps of up to ten c...
|
| 289 |
COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning
2609.22697
|
cs.CLcs.SDeess.AS
|
Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue |
Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should...Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational context. Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task. Given historical conversation audio, target text, and a reference speech, the system shou...
|
| 290 |
LeakScale: Estimating the Causal Effect of Benchmark Exposure
2609.27176
|
cs.CL
|
Divyansh Singh |
Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the perf...Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact. We present LeakScale, an interventional framework for estimating this missing quantity. LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-der...
|
| 291 |
ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
2609.29349
|
cs.CLcs.AI
|
Firoj Alam, Md. Rafiul Biswas, Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Hamdy Mubarak |
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In t...ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In total, 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers. Participating teams explored models such as AraBERT, Jais, and Qwen3-VL. The best systems achieved macro-F1 scores of 0.823 on A...
|
| 292 |
Likelihood Ranking doesn't Scale Like Prompting in LLMs
2609.29390
|
cs.CL
|
Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci |
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer ...LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs...
|
| 293 |
Rufus-Air: An Open LLM Post-Training Recipe
2609.29421
|
cs.CLcs.LGcs.AI
|
Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He |
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document...Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training build...
|
| 294 |
SkillFlow: Scalable and Efficient Agent Skill Retrieval System
2504.06188
|
cs.CLcs.AI
|
Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos |
AI agents can extend their capabilities at inference time by loading reusable skills into context, yet equipping an agent with too many skills, particularly irrelevant ones, degrades performance. As community-driven skill repositories grow, agents need a way t...AI agents can extend their capabilities at inference time by loading reusable skills into context, yet equipping an agent with too many skills, particularly irrelevant ones, degrades performance. As community-driven skill repositories grow, agents need a way to selectively retrieve only the most relevant skills from a large library. We present SkillFlow, the first open, multi-stage retrieval system for agent skill discovery that frames skill acquisition as an information retrieval problem over a...
|
| 295 |
From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs
2505.14368
|
cs.CL
|
Jiawen Wang, Pritha Gupta, Eyke H\"ullermeier, Xiaoxue Gao, Nancy F. Chen |
Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in high-impact applications such as finance, law, and healthcare, systematically inv...Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in high-impact applications such as finance, law, and healthcare, systematically investigating their security risks is becoming increasingly important towards a trustworthy LLM era. This paper comprehensively studies effective prompt injection attacks against 14 widely used open-source and three closed-source LLMs on five ...
|
| 296 |
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
2506.04734
|
cs.CLcs.LGcs.AI
|
Yongfu Zhu, Lin Sun, Jinzhu Wu, Weihong Lin, Xiaoqi Jian |
Reasoning models represented by the Deepseek-R1-Distill series have been widely adopted by the open-source community due to their strong performance in mathematics, science, programming, and other domains. However, our study reveals that their benchmark evalua...Reasoning models represented by the Deepseek-R1-Distill series have been widely adopted by the open-source community due to their strong performance in mathematics, science, programming, and other domains. However, our study reveals that their benchmark evaluation results are subject to significant fluctuations caused by various factors. Subtle differences in evaluation conditions can lead to substantial variations in results. Similar phenomena are observed in other open-source inference models ...
|
| 297 |
Stepwise Intrinsic Rewards for Reasoning in Large Language Models
2602.01034
|
cs.CLcs.AI
|
Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne |
Reinforcement learning (RL) has become a widely used paradigm for improving the reasoning abilities of large language models (LLMs) and Vision-language models (VLMs). Sparse binary outcome rewards, however, score only final correctness and cannot identify whic...Reinforcement learning (RL) has become a widely used paradigm for improving the reasoning abilities of large language models (LLMs) and Vision-language models (VLMs). Sparse binary outcome rewards, however, score only final correctness and cannot identify which intermediate steps contributed to it; in multimodal tasks, they may also reward answers driven by linguistic priors rather than visual evidence. Process reward models (PRMs) densify supervision but usually require process annotations, aux...
|
| 298 |
FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases
2602.09163
|
cs.CLcs.AI
|
Xingjian Zhang, Sophia Moylan, Ziyang Xiong, Qiaozhu Mei, Yichen Luo |
Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert curators to search papers, reconcile...Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert curators to search papers, reconcile evidence across documents, and produce ontology-grounded annotations. Existing benchmarks usually evaluate isolated subtasks, such as named entity recognition or relation extraction, and therefore do not capture this end-to-end workflow. W...
|
| 299 |
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
2602.12316
|
cs.CLcs.AI
|
Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang |
Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We ...Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. ...
|
| 300 |
Layer-wise Target Propagation: Efficient Component Attribution through Target Centric Propagation
2603.19742
|
cs.CLcs.LG
|
Lasse Marten Jantsch, Dong-Jae Koh, Seonghyeon Lee, Young-Kyoon Suh |
Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation. While recent efforts have yielded a plethora of attribution methods attempting to balance faithfulness and...Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation. While recent efforts have yielded a plethora of attribution methods attempting to balance faithfulness and computational efficiency, dense component attribution remains prohibitively expensive. In this work, we introduce Layer-wise Target Propagation (LTP), a novel framework that faithfully traces information flow on the frozen transformer in o...
|
| 301 |
Combee: Scaling Prompt Learning for Self-Improving Language Model Agents
2604.04247
|
cs.CLcs.LGcs.AI
|
Hanchen Li, Runyuan He, Qizheng Zhang, Changxiu Ji, Qiuyang Mang |
Recent advances in prompt learning allow large language model agents to acquire task-relevant knowledge from inference-time context without parameter changes. For example, existing methods (like ACE or GEPA) can learn system prompts to improve accuracy based o...Recent advances in prompt learning allow large language model agents to acquire task-relevant knowledge from inference-time context without parameter changes. For example, existing methods (like ACE or GEPA) can learn system prompts to improve accuracy based on previous agent runs. However, these methods primarily focus on single-agent or low-parallelism settings. This fundamentally limits their ability to efficiently learn from a large set of collected agentic traces. It would be efficient and ...
|
| 302 |
AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
2605.11398
|
cs.CLcs.AI
|
Robin Linzmayer (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University), Georgianna Lin (Department of Biomedical Informatics |
We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflo...We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflow-specific triage tasks, but they do not offer a unified evaluation of acuity identification across these settings. AcuityBench addresses this gap by harmonizing five public datasets spanning user conversations, online forum posts, clinical...
|
| 303 |
Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents
2606.05828
|
cs.CLcs.AI
|
Zeyu Gan, Huayi Tang, Yong Liu |
As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With the rapid expansion of available skills, enabling personal agents to learn and ad...As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With the rapid expansion of available skills, enabling personal agents to learn and adapt to implicit user preferences becomes a critical challenge. However, local deployment constraints preclude complex centralized selection algorithms, creating an urgent need for a lightweight local preference harness. This paper explores ...
|
| 304 |
INFUSER: Influence-Guided Self-Evolution Improves Reasoning
2606.09052
|
cs.CLcs.LGcs.AI
|
Siyu Chen, Miao Lu, Beining Wu, Heejune Sheen, Fengzhuo Zhang |
Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generato...Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers f...
|
| 305 |
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
2608.09444
|
cs.CLcs.LG
|
Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis |
A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops can...A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first effici...
|
| 306 |
The Communication Map of a Transformer
2608.22007
|
cs.CLcs.LG
|
Richard Zhe Wang |
The components of a transformer communicate by writing to and reading from a shared residual stream, and the mechanistic interpretability literature has mapped these connections by hand, one circuit at a time. We present the communication map, which charts eve...The components of a transformer communicate by writing to and reading from a shared residual stream, and the mechanistic interpretability literature has mapped these connections by hand, one circuit at a time. We present the communication map, which charts every potential communication channel from the geometry of the model's weights alone, generalizing the composition score of Elhage et al. (2021) into a single coupling coefficient covering all 18 connection classes, from head-to-head to neuron...
|
| 307 |
Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
2608.27167
|
cs.CLcs.AI
|
Pranav Aggarwal |
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. ...An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguisha...
|
| 308 |
PUBG Ally: A Conversational Embodied Agent as an AI Teammate
2609.29837
|
cs.CLcs.AI
|
PUBG Ally Team, Irene Chen, Youngin Cho, Seungjun Chung, Jimin Hong |
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to...We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A language-...
|
| 309 |
PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
2609.30094
|
cs.CLcs.AI
|
Luciano Maldonado |
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable throu...Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift...
|
| cs.CV 182 papers | ||||
| 1 |
AlphaEarth distinguishes cities but compresses urban variation
2609.30356
|
cs.CV
|
Andrew Renninger |
Cities differ in built form, land cover and development history, complicating comparison across places and time. Satellite foundation models map Earth's surface onto common numerical representations. Yet the tasks and targets used to shape them typically do no...Cities differ in built form, land cover and development history, complicating comparison across places and time. Satellite foundation models map Earth's surface onto common numerical representations. Yet the tasks and targets used to shape them typically do not focus on cities: globally consistent labels for urban function do not exist, and many datasets - especially land cover and land use classifications - collapse the built environment into few classes. Here we audit the representation, focus...
|
| 2 |
LiTe-GS: Oracle-Efficient Next Best View Selection for 3D Gaussian Splatting
2609.30393
|
cs.CV
|
Vivek Pandey, Amirhossein Mollaei Khass, Nader Motee |
Selecting informative camera views is critical for efficient training and adaptive refinement in 3D Gaussian Splatting, where each observation significantly influences model parameters. However, information-driven view-selection strategies can require repeated...Selecting informative camera views is critical for efficient training and adaptive refinement in 3D Gaussian Splatting, where each observation significantly influences model parameters. However, information-driven view-selection strategies can require repeated evaluations of expensive information-gain oracles as the number of candidate views increases. We propose LiTe-GS, an oracle-efficient method for next best view selection in 3D Gaussian Splatting. LiTe-GS reduces the number of information-o...
|
| 3 |
CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices
2609.30395
|
cs.CV
|
Amir Zamani, Zeinab Ghasemi-Naraghi |
Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a ...Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high-resolution spatial representations from a YOLO11m-P2 teacher to a compact YOLO11n student without altering the student's inference architecture. Unlike conventional same-scale feature distillation...
|
| 4 |
What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study
2609.30402
|
cs.CVcs.CLcs.LGcs.AIcs.MM
|
Akshit Sharma, Prashant W. Patil |
Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a ...Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language back...
|
| 5 |
ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models
2609.30434
|
cs.CV
|
Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar |
Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone f...Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone frozen, yet many existing multimodal prompt learners couple the visual and textual branches weakly and can be brittle in low-shot regimes. We propose ProCAP, a probabilistic cross-attentive prompt learning framework that improves cross-modal...
|
| 6 |
LensDesigner: A Self-Improving Agent for Optical Lens Design
2609.30450
|
cs.CV
|
Lei Sun, Haoran Liang, Dannong Xu, Yao Gao, Yuyu Geng |
Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. I...Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and...
|
| 7 |
The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation
2609.30478
|
cs.CV
|
Soshun Kihara, Shunsuke Yasuki, Masato Taki |
Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in con...Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in v...
|
| 8 |
Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models
2609.30566
|
cs.CV
|
Jian Shi, John Femiani, Peter Wonka |
We present a new inference-time sampler for diffusion models that gives a pretrained model a capability it was never trained for: constructing the atlas of the population it synthesizes. The sampler converges from every random seed to the population's central ...We present a new inference-time sampler for diffusion models that gives a pretrained model a capability it was never trained for: constructing the atlas of the population it synthesizes. The sampler converges from every random seed to the population's central anatomy, which we call the \emph{intrinsic atlas}. The advantage is threefold. (1) It requires no retraining. A diffusion model that has already learned a coherent population, including the released ones, yields its atlas in a single infere...
|
| 9 |
Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases
2609.30595
|
cs.CVcs.AI
|
Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang |
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack ground...Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure ind...
|
| 10 |
MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization
2609.30609
|
cs.CV
|
Xiangyu Kong, Wenjie Zhou, Fengping Tian, Lihua Fang, Haoqin Sun |
Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determin...Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determine appearance, layout or state. We therefore recast the problem as condition construction and present MVAgent, a multi-agent pipeline whose agents collaborate through typed conditioning inputs. Because an environment image shows one viewpoin...
|
| 11 |
MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification
2609.30613
|
cs.CVcs.AI
|
Zhexiang Li |
Dermoscopy classifiers built on Vision Transformers process all image patches uniformly, although diagnostic evidence is concentrated in the lesion region. Existing token pruning methods reduce tokens using generic saliency or similarity signals, but rarely as...Dermoscopy classifiers built on Vision Transformers process all image patches uniformly, although diagnostic evidence is concentrated in the lesion region. Existing token pruning methods reduce tokens using generic saliency or similarity signals, but rarely ask whether the retained subset still contains the lesion. This paper introduces MedTokenBudget, a supervised post-backbone token routing framework that learns to construct compact lesion-enriched representations when auxiliary lesion masks a...
|
| 12 |
Conditional Predictive Sufficient Statistics for Visual Representation Learning
2609.30647
|
cs.CV
|
Yuzhou Hong |
A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-fa...A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch...
|
| 13 |
StarWM: Self-Supervised Trained Attention Routing for Robust World Models
2609.30667
|
cs.CVcs.LG
|
Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid, Sebastian Gottwald |
A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pi...A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance for visual tasks, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant infor...
|
| 14 |
Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding
2609.30682
|
cs.CV
|
Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas |
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform token...Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-ad...
|
| 15 |
MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning
2609.30698
|
cs.CV
|
Peipei Li, Shuhan Xia, Shengyang Liu, Zekun Li, Ran He |
Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference co...Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By ben...
|
| 16 |
SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion
2609.30703
|
cs.CVcs.AI
|
Timing Li, Yiming Sun, Boan Tao, Xiyuan Gao, Haifang Cao |
Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency ...Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency propagation across stages. We propose Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion (SAGE), a unified framework integrating frequency equalization, hierarchical alignment, and subband fusion...
|
| 17 |
Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation
2609.30708
|
cs.CVcs.AI
|
Tasneem Nasser, Susanne Schmid, Roberto Souza, Naser El-Sheimy |
A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on man...A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on manual annotations. Self-supervised learning has emerged as a promising approach for developing foundation models by enabling the learning of transferable feature representations from large-scale unlabeled medical imaging datasets. In this stu...
|
| 18 |
VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control
2609.30709
|
cs.CVcs.AI
|
Kemou Jiang, Maonan Wang, Xingchen Zou, Jiayue Zhu, Yuhang Fu |
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling...Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end...
|
| 19 |
TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation
2609.30722
|
cs.CVcs.AI
|
Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo |
Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topol...Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topology and unrelated traffic. We introduce TrafficImag, the first benchmark for counterfactual roadside traffic video generation. TrafficImag combines a large-scale roadside dataset (9,022 annotated images, 7,043 deduplicated video clips, and ...
|
| 20 |
EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection
2609.30724
|
cs.CV
|
Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang |
Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cros...Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A ...
|
| 21 |
Learning Polarization Image Restoration with General Restoration Priors
2609.30728
|
cs.CV
|
Chenggong Li, Jinhao Liu, Caiyun Wu, Yidong Luo, Junchao Zhang |
Polarization imaging captures distinctive surface and geometric cues that benefit a wide range of vision tasks. However, real-world polarization acquisition is often affected by multiple coupled degradations, making image restoration essential for practical po...Polarization imaging captures distinctive surface and geometric cues that benefit a wide range of vision tasks. However, real-world polarization acquisition is often affected by multiple coupled degradations, making image restoration essential for practical polarization vision. Existing methods are largely tailored to specific degradations and remain constrained by the limited scale and quality of polarization data. To address these limitations, we develop an all-in-one polarization restoration ...
|
| 22 |
Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation
2609.30733
|
cs.CV
|
Shengqi Dang, Zhengxi Yu, Feilin Han, Xingyu Lan, Nan Cao |
Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the v...Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the visual saliency of a specific object during text-to-image generation without requiring any visual priors. Our key insight is that visual saliency is inherently relative: boosting the saliency of a target object also depends on the global sal...
|
| 23 |
From Mono to Stereo: Accelerating Binocular Gaussian Splatting via Reprojection and Selective Patching
2609.30741
|
cs.CV
|
Hongfei Zhu, Ling Zhou |
Binocular rendering requires two nearby views of the same scene and therefore repeats substantial visibility and shading work. We present a 2D Gaussian Splatting (2DGS) pipeline that fully renders a dominant-eye RGB image and an alpha-weighted depth proxy, rep...Binocular rendering requires two nearby views of the same scene and therefore repeats substantial visibility and shading work. We present a 2D Gaussian Splatting (2DGS) pipeline that fully renders a dominant-eye RGB image and an alpha-weighted depth proxy, reprojects that image to the affiliated eye, and repairs uncovered pixels. Small interior gaps are interpolated, whereas larger disoccluded regions are identified as regions of interest (ROIs) and selectively re-rendered. The depth proxy reuse...
|
| 24 |
Training-Free Bottleneck Width Planning for Convolutional Autoencoders
2609.30755
|
cs.CV
|
Guannan Guo |
Multiscale Spectral Rate-Distortion (MS-SRD) estimates the bottleneck channels required at user-supplied spatial cuts from training images and a normalized mean-squared error (NMSE) bound, without fitting a neural network. Its covariance-tail rule is exact for...Multiscale Spectral Rate-Distortion (MS-SRD) estimates the bottleneck channels required at user-supplied spatial cuts from training images and a normalized mean-squared error (NMSE) bound, without fitting a neural network. Its covariance-tail rule is exact for shared linear block-convolutional autoencoders under squared error. A nested-scale dominance result motivates reporting the activation-parameter Pareto frontier alongside the minimum-latent candidate. At NMSE <= 0.01 on thirteen grayscale ...
|
| 25 |
LLPR: Location-aware learning and physics-based reconstruction for raindrop removal from a single image
2609.30758
|
cs.CV
|
Zewei He, Xingyu Liu, Xing Luo, Guizhong Fu, Zixuan Chen |
Raindrops can cause occlusion and distortion in the background scenes due to their adherence to windows or camera lenses. Existing raindrop removal methods concentrate on designing sophisticated CNN or Transformer architectures to recover distorted and missing...Raindrops can cause occlusion and distortion in the background scenes due to their adherence to windows or camera lenses. Existing raindrop removal methods concentrate on designing sophisticated CNN or Transformer architectures to recover distorted and missing texture. In this paper, we try to integrate location information and physical model into off-the-shelf CNN or Transformer architectures to help improve their performance. Specifically, we notice that existing methods deploy a preprocessing...
|
| 26 |
Timo: $\textbf{T}$aming Mult$\textbf{i}$modal Diffusion Transformer for Human $\textbf{Mo}$tion Generation
2609.30761
|
cs.CV
|
Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu |
Most existing human motion generation (HMG) methods use cross-attention modules to inject text semantics, but ignore the importance of bidirectional modeling between motion and text tokens, which limits text comprehension. A straightforward idea is introducing...Most existing human motion generation (HMG) methods use cross-attention modules to inject text semantics, but ignore the importance of bidirectional modeling between motion and text tokens, which limits text comprehension. A straightforward idea is introducing multimodal diffusion transformers (MMDiT), which have shown effective joint text--visual modeling in vision generation, into HMG. However, we find that articulated motion is temporally coherent but weakly correlated across joints, in which...
|
| 27 |
Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes
2609.30769
|
cs.CVcs.LG
|
Rushab Rasik Karania, Tomas Maul |
Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation c...Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation contribute to prototype construction? The Within-Instance Prototypical Transformer (WIPT) implements single-query test-time prototype adaptation by jointly transforming one unlabelled query and the labelled support embeddings, then forming q...
|
| 28 |
Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models
2609.30783
|
cs.CVcs.AI
|
Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li |
Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Though...Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before localizing the target. Although intuitive, such explicit verbal reasoning introduces substantial attention interference: redundant textual tokens disrupt attention during perception...
|
| 29 |
Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion
2609.30795
|
cs.CV
|
Chen-Chieh Liao, Yichen Peng, Yiyi Cai, Y\^ui Ono, Hiroki Hanaoka |
Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is sub...Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requirement is not a universal absolute unit, but a reliable monotonic control axis. We propose Motion Style Slider, a motion-to-motion style transfer framework for endpoint-supervised ...
|
| 30 |
MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization
2609.30855
|
cs.CV
|
Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas |
Reliable skin lesion segmentation and classification are central to dermoscopic computer-aided diagnosis. Existing multi-task frameworks couple the two tasks architecturally without clinical knowledge, while knowledge-injecting approaches rely on the macroscop...Reliable skin lesion segmentation and classification are central to dermoscopic computer-aided diagnosis. Existing multi-task frameworks couple the two tasks architecturally without clinical knowledge, while knowledge-injecting approaches rely on the macroscopic ABCD rule, which was not designed for dermoscopy. Dermoscopic diagnosis is grounded in Pattern Analysis, a microscopic framework structured around dermoscopic features. We propose MDSkin-Net, which incorporates cue-level Pattern Analysis...
|
| 31 |
Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting
2609.30865
|
cs.CV
|
Zijian Wu, Jinliang Wang, Zidian Lin, Ying Song, Ziqian Lu |
COLMAP-free 3D Gaussian Splatting (3DGS) bypasses computationally expensive structure-from-motion (SfM) pipelines, yet progressive camera pose tracking remains fundamentally vulnerable to error compounding---early pairwise tracking inaccuracies both corrupt su...COLMAP-free 3D Gaussian Splatting (3DGS) bypasses computationally expensive structure-from-motion (SfM) pipelines, yet progressive camera pose tracking remains fundamentally vulnerable to error compounding---early pairwise tracking inaccuracies both corrupt subsequent frame initializations and remain permanently frozen in the scene representation. Rather than relying on heavyweight external neural priors or treating progressive tracking through isolated heuristic fixes, we propose a unified reli...
|
| 32 |
UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
2609.30928
|
cs.CVcs.AI
|
Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao |
Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned ...Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. U...
|
| 33 |
ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos
2609.30934
|
cs.CV
|
Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo |
Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly cha...Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forger...
|
| 34 |
Spackle: Completing Large View Single Image NVS with Adaptive Gaussians
2609.30941
|
cs.CVcs.AI
|
Xuanzhi Liu, Yuhe Zhou, Xinyi Wu, Zhenyao Wu, Jinghao Chen |
Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid...Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid decoupled frameworks combining feedforward 3D Gaussian Splatting (3DGS) and diffusion models show promise for large-view-deviation NVS, they suffer from capacity competition: a fixed number of Gaussians forces resource shifts from visible ...
|
| 35 |
OneWorld: Learning Consistent Physics Across Actions in World Models
2609.30946
|
cs.CVcs.AI
|
Ke He, Yichen Ding, Bin Yang |
Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same...Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across interventions, making it difficult for the model to maintain...
|
| 36 |
DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry
2609.30947
|
cs.CV
|
Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari |
Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution...Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution and dynamic range, but their asynchronous measurements complicate reliable correspondence estimation. We present DAPEVO, a learned visual odometry system that estimates image and event correspondences independently at shared patch location...
|
| 37 |
MVVBench: Benchmarking 4D Reasoning in Vision-Language Models
2609.30952
|
cs.CVcs.AI
|
Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi |
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continu...Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal ax...
|
| 38 |
IDM-Net: A Lightweight Illumination-Decoupled Modulation Network for Low-Light Image Enhancement
2609.30962
|
cs.CV
|
Cheng-Yen Hsiao, Jing-Ming Guo |
Low-light image enhancement (LLIE) remains challenging for lightweight models because illumination restoration and color fidelity are difficult to optimize simultaneously in the RGB color space. Although recent color-decoupled methods separate luminance and ch...Low-light image enhancement (LLIE) remains challenging for lightweight models because illumination restoration and color fidelity are difficult to optimize simultaneously in the RGB color space. Although recent color-decoupled methods separate luminance and chrominance representations, they primarily optimize luminance as an enhancement target, leaving its potential as an explicit guidance prior largely unexplored during feature reconstruction. To address this limitation, we propose IDM-Net, a l...
|
| 39 |
Where and When to Force: Routed Forcing for Streaming Avatars
2609.30963
|
cs.CV
|
Zihan Su, Siwen Lu, Junhao Zhuang, Zeyue Xue, Haoyang Huang |
Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-ste...Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes a reverse KL divergence, which is inherently mode-seeking: it causes the student to discard high-dynamic modes and collapse onto static outputs, compressing both dynamics and dive...
|
| 40 |
CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models
2609.30979
|
cs.CV
|
Linyuan Gao, Yuan Wu, Yi Chang |
Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning ba...Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for sin...
|
| 41 |
STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence
2609.30981
|
cs.CV
|
Siru Zhong, Shenghan Tan, Rihong Yan, Xiaohui Lv, Yuzheng Zhuang |
Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilitie...Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilities under evolving and incomplete evidence. We present STORM-Bench, comprising 5,736 questions across 630 compact, change-dense episodes spanning five egocentric domains (STORM-Real) and two controlled simulation subsets (STORM-Sim) at 1 FPS....
|
| 42 |
FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators
2609.30982
|
cs.CVcs.AI
|
Kai Yao, Marc Juarez |
Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one gen...Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region...
|
| 43 |
Self-Supervised Perceptually Interpretable Monocular Depth Estimation
2609.30987
|
cs.CV
|
Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis |
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing...Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, as depth is inferred from RGB representations that obscure the impact of individual perceptual image components. This lack of transparency limits systematic analysis of failure cases and reduces confi...
|
| 44 |
PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution
2609.30988
|
cs.CV
|
Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao |
Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantia...Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Transformer SR models are efficient yet often struggle to recover realistic high-frequency details. This motivates a natural question: can diffusion priors be transferred to existing di...
|
| 45 |
PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation
2609.30989
|
cs.CV
|
Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya |
Purpose: Accurate 6 DoF pose estimation of surgical tools is critical for automa- tion, robotic proprioception, and safe interaction with the tissue operated on. Kinematics-based approaches suffer from accumulated errors due to the cable- driven nature of robo...Purpose: Accurate 6 DoF pose estimation of surgical tools is critical for automa- tion, robotic proprioception, and safe interaction with the tissue operated on. Kinematics-based approaches suffer from accumulated errors due to the cable- driven nature of robotic arms, while vision-based methods often rely on external markers or trackers. Although more recent vision-based advances have been pro- posed, these two-stage pose estimation methods often lack real-time robustness due to accumulated err...
|
| 46 |
FLIP: Final Layer Inference-Time Probing for Vision-Language Models
2609.30993
|
cs.CVcs.AI
|
Drandreb Earl O. Juanico, Rowel O. Atienza |
We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under intern...We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise flooring to the final normalized hidden state before logit...
|
| 47 |
TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking
2609.31005
|
cs.CV
|
Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl |
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Visi...Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at spar...
|
| 48 |
Refining Cytology Predictions with Conditional Random Fields
2609.31028
|
cs.CV
|
Manon Dausort, Tiffanie Godelaine, Karim El Khoury, Maxime Zanella, Christophe De Vleeschouwer |
Vision-language models (VLMs) achieve strong zero-shot (ZS) classification on histology images but do not perform as well on cytology, whose stains and cell morphology differ markedly compared to histology. Conditional random fields (CRFs) can refine noisy VLM...Vision-language models (VLMs) achieve strong zero-shot (ZS) classification on histology images but do not perform as well on cytology, whose stains and cell morphology differ markedly compared to histology. Conditional random fields (CRFs) can refine noisy VLM predictions by propagating information across patches, but existing CRF frameworks were designed for histopathology and do not transfer to cytology datasets, released as independent patch pools spanning multiple staining protocols. We intr...
|
| 49 |
Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images
2609.31040
|
cs.CV
|
Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq |
Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for pat...Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for patch-level classification without requiring any annotation, yet these zero-shot (ZS) predictions remain noisy on fine-grained tasks and must be further refined. A promising direction is to refine all predictions jointly, i.e., a transductive ...
|
| 50 |
Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion
2609.31050
|
cs.CV
|
Olga Zatsarynna, Denis Korzhenkov, Juergen Gall, Amir Habibian, Mohsen Ghafoorian |
Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty vari...Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty varies substantially across video regions and evolves throughout the generation process. We introduce HetA-DiT, a heterogeneous attention mechanism that adaptively allocates computation according to token difficulty. A lightweight uncertainty b...
|
| 51 |
Band-Selection Stability and Semantic Segmentation Performance: A Study on Hyperspectral City
2609.31074
|
cs.CV
|
Jiarong Li, Imad Ali Shah, Enda Ward, Martin Glavin, Edward Jones |
Resource constraints make high-dimensional hyperspectral imaging challenging in autonomous perception, motivating the use of band selection methods. However, the sensitivity of band-selection methods to sampled data and their relationship to semantic segmentat...Resource constraints make high-dimensional hyperspectral imaging challenging in autonomous perception, motivating the use of band selection methods. However, the sensitivity of band-selection methods to sampled data and their relationship to semantic segmentation models (SSMs) remain underexplored. This study evaluates six band selection methods on ten independently sampled, class-balanced region-of-interest (ROI) sets, yielding 60 top-25 band subsets from the Hyperspectral City V2 (128 bands: 4...
|
| 52 |
DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models
2609.31103
|
cs.CVcs.AI
|
Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong |
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded ev...Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder fea...
|
| 53 |
Double-stream registration with pyramid fusion for HDR video with alternating exposures
2609.31108
|
cs.CV
|
Onofre Martorell, Ivan Pereira-S\'anchez, Antoni Fuentes, Antoni Buades |
High dynamic range (HDR) video reconstruction from al\-ter\-na\-ting-exposure sequences remains challenging, especially in regions with extreme luminance variation. We propose a novel HDR reconstruction framework based on dual-stream registration and accurate ...High dynamic range (HDR) video reconstruction from al\-ter\-na\-ting-exposure sequences remains challenging, especially in regions with extreme luminance variation. We propose a novel HDR reconstruction framework based on dual-stream registration and accurate pyramid fusion. Given three consecutive frames, our method computes optical flow directly with the central frame, while introducing a complementary midpoint displacement strategy to handle cases with severe overexposition. A pyramid fusion ...
|
| 54 |
Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding
2609.31135
|
cs.CVcs.AIcs.MM
|
Alberto Presta, Michal Byra, Grzegorz Stefa\'nski, Karol Szurkowski, Eryk Ko{\l}odziejczyk |
Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typicall...Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex training pipelines, or multimodal large language models. We present Pocket-STVG (P-STVG), a lightweight cascade architecture that addresses STVG by combining efficient pre-trained c...
|
| 55 |
Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos
2609.31148
|
cs.CV
|
Bowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu, Kuizhuang Liu |
Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and reali...Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prereq...
|
| 56 |
FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification
2609.31150
|
cs.CVcs.AI
|
Muhammad Muhtasim Shahriar, M. M. Golam Hafiz, Saad Aloteibi, Mohammad Ali Moni |
Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA),...Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA), Normal, and squamous cell carcinoma (SCC). FedHisto-PAST v2 combines a frozen HIBOU-B foundation model with parameter-efficient adaptation, stain-conditioned paired-view prediction and feature consistency, reliability-aware prototype learn...
|
| 57 |
HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models
2609.31154
|
cs.CV
|
Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li |
Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly fo...Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly follow a static weight paradigm, producing a single frozen adapter that struggles to adapt to diverse prompt variations and suffers from parameter interference when scaling to multiple concepts. We propose \textbf{HyperErase}, a framework for...
|
| 58 |
ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation
2609.31160
|
cs.CVcs.AI
|
Donghang Lyu, Zichen Zhang, Oleh Dzyubachyk, Marius Staring |
Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely a...Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely aim at building a generalizable vessel segmentor across anatomies and modalities. While the Seg- ment Anything Model (SAM) has shown promise for med- ical image segmentation, its original design does not fully exploit vascular morphology and...
|
| 59 |
TaskIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback
2609.31170
|
cs.CV
|
Yanjie Tu, Qingsen Yan, Axi Niu, Wenxuan Cai, Tao Hu |
Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Diff...Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address...
|
| 60 |
Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
2609.31193
|
cs.CVcs.SDeess.AS
|
Jihoo Jung, Youngjoon Jang, Joon Son Chung |
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systemati...Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual componen...
|
| 61 |
Light Field Primitive for Novel View Synthesis
2609.31198
|
cs.CV
|
Liang Chen, Jiahui Ning, Xun Jiang, Xing Xu, Jimmy Ren |
We present Light Field Primitives (LFP), a formulation for novel view synthesis that replaces the dense ray database with a compact set of differentiable primitives in the classical two-plane parameterization. Each primitive condenses a group of rays into one ...We present Light Field Primitives (LFP), a formulation for novel view synthesis that replaces the dense ray database with a compact set of differentiable primitives in the classical two-plane parameterization. Each primitive condenses a group of rays into one learned record, and its response to a query is governed by how closely that query belongs to the group. Rendering a camera ray then reduces to compositing all responses it elicits, and a scene can be optimized directly from posed images and...
|
| 62 |
Preserve-and-Compose Training for Composed Image Retrieval
2609.31202
|
cs.CV
|
Sehyun Kwon |
Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instea...Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captions may omit source details that should be preserved. We therefore propose, Preserve-and-Compose Training, which complements target-caption supervision with visual evidence from the...
|
| 63 |
WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery
2609.31234
|
cs.CV
|
Zhongyu Pang |
Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-rout...Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-routing agent, decouples routing from visual perception. Stage A is routing-first: emission is trained, not elicited. Stage B executes conditionally: intrinsic queries enter visual answering (full-scene thumbnail; a WeaveEarth-style evidence bo...
|
| 64 |
Geometric Inconsistency Localization in Multi-View Image Sets
2609.31247
|
cs.CVcs.AIcs.MM
|
Xander Staelens, Alb\'eric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert |
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for ...Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we intro...
|
| 65 |
Gauss What You Need: Compact Gaussian Splatting Across Scene Scales
2609.31248
|
cs.CV
|
Afif Boudaoud, Jiayi Liu, Alexandru Calotoiu, Torsten Hoefler |
3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to selec...3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to select this number automatically across capture scales remains unresolved: configurations effective on standard benchmarks can leave larger captures with too few Gaussians to reconstruct fine details. We observe that the surface to represent, gi...
|
| 66 |
MoTop: Motion-Topological Model For Micro AU Detection
2609.31285
|
cs.CV
|
Huai-Qian Khor, Mengting Wei, Yante Li, Chu Kiong Loo, Guoying Zhao |
Facial micro-expressions are spontaneous, brief, and subtle facial movements that reveal suppressed emotions in high-stakes environments. In contrast to classic expression analysis, detecting action unit (AU) yields a finer representation of facial movements, ...Facial micro-expressions are spontaneous, brief, and subtle facial movements that reveal suppressed emotions in high-stakes environments. In contrast to classic expression analysis, detecting action unit (AU) yields a finer representation of facial movements, serving as a preliminary step before defining expression classes and other downstream tasks. Therefore, it represents a crucial upstream task in facial analysis, and improving an AU detection module increases the precision of facial analysi...
|
| 67 |
UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning
2609.31298
|
cs.CVcs.AI
|
Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng |
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnos...Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity promp...
|
| 68 |
CytoSPM: Open-Vocabulary Cytopathology Detection with Structured Prompt Bank
2609.31314
|
cs.CV
|
Wenjie Li, Zishan Xu, Jinyang Huang, Zhengxin Nie, Shichao Kan |
Cytopathology detection requires open-vocabulary recognition because cellular categories are fine-grained, long-tailed, and continuously evolving across different organ systems. However, existing cytology detectors are mostly single-domain and closed-set, and ...Cytopathology detection requires open-vocabulary recognition because cellular categories are fine-grained, long-tailed, and continuously evolving across different organ systems. However, existing cytology detectors are mostly single-domain and closed-set, and there is still no unified benchmark for evaluating open-vocabulary cytopathology detection. We present PentaCyto, a multi-domain benchmark covering cervical, urinary, respiratory, serous fluid, and thyroid cytology, with 24 base categories ...
|
| 69 |
CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support
2609.31326
|
cs.CVcs.AI
|
Muhammad Muhtasim Shahriar, Md. Naimur Asif Borno, Saad Aloteibi, Mohammad Ali Moni |
Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We in...Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We introduce CG-HAF, a global-local fusion framework that instead keeps this evidence explicit: averaged holistic severity probabilities from independently trained classifiers are combined with structured lesion-burden descriptors from an object...
|
| 70 |
DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models
2609.31349
|
cs.CVcs.AI
|
Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan |
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress ...Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. ...
|
| 71 |
Open Vocabulary Domain Unlearning
2609.31356
|
cs.CVcs.LG
|
Sumanth Udupa, Mehrtash Harandi, Yadan Luo, Mahsa Baktashmotlagh |
Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning ...Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning ...
|
| 72 |
OpenVAM: Open-World Visual Attention Modeling with VLMs
2609.31364
|
cs.CV
|
Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati |
Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: pr...Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elements in the scene (what) and understand the drivers of those peaks in context (why), while remaining robust to domain shift across natural images, commercial content, and UI/web la...
|
| 73 |
ContraFM-S2O: Flow Matching-Based One-step SAR-to-Optical Image Translation Model with Contrastive Learning
2609.31378
|
cs.CV
|
Mingqian Yu, Wei-kuan Chiang, Qiurui Wang, Peilin Zhao |
In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high infe...In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high inference latency and the generated optical images suffer from low detail fidelity, often resulting in blurred edges and loss of fine textures. Thus, we propose ContraFM-S2O, which is a flow matching-based model for SAR-to-optical image transla...
|
| 74 |
AxonSynth: Domain-Randomized Synthetic Data for Zero-Shot 3D Axon Segmentation in Light-Sheet Microscopy
2609.31431
|
cs.CV
|
Edward Gaibor, Kyriaki-Margarita Bintsi, Carmen Luz Leiva Ureta, Zayneb Bellatif, Chiara Maffei |
Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive to obtain. Existing supervised axon segmentation methods rely on target-domain annotations and can be brittle...Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive to obtain. Existing supervised axon segmentation methods rely on target-domain annotations and can be brittle when tissue type, species, modality, or acquisition conditions change. We present AxonSynth, a domain-randomized synthetic-data framework for training 3D axon segmentation models without manually annotated real training volumes. AxonSynth ...
|
| 75 |
Implicit Neural Representation for Hyperspectral Video Compression
2609.31435
|
cs.CVcs.LGcs.AI
|
Alfredo Scalera, Paul Murray, Jaime Zabalza |
With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In...With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In this study, we explore the use of implicit neural representation as a candidate solution. We propose a novel extension of an existing RGB video compression model, achieving Bj{\o}ntegaard Delta PSNR gains of +4.99 dB and Bj{\o}ntegaard Del...
|
| 76 |
From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training
2609.31450
|
cs.CVcs.AI
|
Wang Jingxin |
Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervi...Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000...
|
| 77 |
Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis
2609.31456
|
cs.CV
|
Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers |
Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improv...Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improve compositional binding. However, this assumption has never been directly quantified. Existing benchmarks evaluate captions only in their composed form, making it impossible to separate the cost of joint reasoning from the cost of recognizi...
|
| 78 |
KneePreM: Towards 3D Knee MRI Foundation Models via Large-Scale Unlabeled Pretraining and Label-Efficient Fine-Tuning
2609.31461
|
cs.CV
|
Xinxin Wang, Liam Hazan, Jing Li, Simona Rabinovici-Cohen, Xiaojuan Li |
Background: Large volumes of unlabeled knee MRI scans are available across repositories but remain insufficiently leveraged. We developed KneePreM, a knee-specific 3D self-supervised model, and evaluated transfer and label efficiency for classification and seg...Background: Large volumes of unlabeled knee MRI scans are available across repositories but remain insufficiently leveraged. We developed KneePreM, a knee-specific 3D self-supervised model, and evaluated transfer and label efficiency for classification and segmentation. Methods: A 3D U-Net masked autoencoder was pretrained on 19,011 unlabeled Osteoarthritis Initiative (OAI) MRI series from 4,791 participants. Downstream fine-tuning used full and reduced training sets for fastMRI+ two-label class...
|
| 79 |
SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery
2609.31507
|
cs.CV
|
Jiajun Jiang, Chunliang Hua, Zichun Chen, Yanxing Wu, Zeyuan Yang |
Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult bec...Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, limiting geographic diversity and episode scale. To address this, we introduce SatNav, a scalable, long-horizon UAV VLN benchmark built from high-resolution satellite imagery. SatNav...
|
| 80 |
ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos
2609.31509
|
cs.CVcs.AI
|
Xuanzhi Liu, Xinyi Wu, Hang Pan, Wensi Huang, Zhenyao Wu |
We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-super...We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-supervision weights based on appearance reliability, degradation risk, and geometric utility, while weakly reactivating useful suppressed frames to maintain trajectory coverage. Since weighting cannot restore details lost to blur or distortion, ...
|
| 81 |
Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics
2609.31514
|
cs.CV
|
Javier Mu\~noz-Haro, Ruben Tolosana, Ruben Vera-Rodriguez, Aythami Morales, Julian Fierrez |
Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative soluti...Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative solution, yet standard frameworks work against the forensic task, e.g., their augmentations overwrite the micro-statistics of image formation. This paper introduces Forensic Twins, a Self-Supervised Residual Learning (SSRL) framework whose pretex...
|
| 82 |
Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment
2609.31524
|
cs.CV
|
Qing Xu, Yuxiang Luo, Zhen Chen |
Surgical scene understanding is critical for computer-assisted intervention, yet laparoscopic cholecystectomy remains challenged by the complex anatomy of the hepatocystic triangle and the risk of bile duct injury. Existing methods for Critical View of Safety ...Surgical scene understanding is critical for computer-assisted intervention, yet laparoscopic cholecystectomy remains challenged by the complex anatomy of the hepatocystic triangle and the risk of bile duct injury. Existing methods for Critical View of Safety (CVS) assessment typically treat it as a holistic prediction task, mapping visual features directly to criterion-level labels. This black-box paradigm lacks explicit reasoning about anatomical relationships, limiting both interpretability a...
|
| 83 |
Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP
2609.31558
|
cs.CV
|
Ahmed Abdelnaby, Mohamed Elmahallawy |
Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only ...Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representati...
|
| 84 |
OC-GS: Gaussian Splatting for Irregular Turntable Capture
2609.31572
|
cs.CVcs.AI
|
Jae Joong Lee, Bedrich Benes |
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This or...Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly spaced views, OC-GS achieves mean foreground PSNR scores...
|
| 85 |
How Far Can INRs Go? Cross-Domain Parameter-efficient INR-Based Semantic Segmentation for Brain MRI
2609.31573
|
cs.CV
|
Ziyao Shang, Pouya Sadeghi, Letian Jiang, Alexander Wong, Sirisha Rambhatla |
Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and cross-site distribution shifts. Implicit Neural Representations (INRs) have recently emerged as a lightweight ...Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and cross-site distribution shifts. Implicit Neural Representations (INRs) have recently emerged as a lightweight alternative for semantic segmentation, achieving competitive performance with substantially fewer parameters than conventional architectures. However, the mechanisms, scaling behavior, and domain generalization abilities of INR-based segmen...
|
| 86 |
GraphWrit3R: End-to-End 3D Scene Graph Writing
2609.31595
|
cs.CV
|
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni |
3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamen...3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which dev...
|
| 87 |
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
2609.31620
|
cs.CV
|
Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong |
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared...Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion th...
|
| 88 |
WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving
2609.30436
|
cs.CV
|
Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li |
Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mis...Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Traj...
|
| 89 |
VkVIO: Cross-platform GPU Acceleration for Visual-Inertial Odometry with Vulkan
2609.30459
|
cs.CV
|
Ole Hoffmann, Mateo de Mayo, Daniel Cremers |
Perception in robotics and XR fundamentally relies on good state estimation. Visual-inertial odometry (VIO) and Simultaneous Localization and Mapping (VI-SLAM) are proven ways of achieving this goal in a cost-effective and accurate manner. Efficiency in these ...Perception in robotics and XR fundamentally relies on good state estimation. Visual-inertial odometry (VIO) and Simultaneous Localization and Mapping (VI-SLAM) are proven ways of achieving this goal in a cost-effective and accurate manner. Efficiency in these systems allows for smaller, cooler, and lighter devices. GPU acceleration is a natural approach for reducing latency, thanks to their wide availability in platforms like embedded computers, mobile phones, and XR headsets. However, previous ...
|
| 90 |
QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
2609.30592
|
cs.CVcs.LG
|
Nicholas Foley, Devin Marinelli, Donny Moore, Diego Enriquez, Amanda Fernandez |
In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical visi...In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection triple with a single learned unit quaternion per token: the relative quaternion $r_{ij} = q_i^{*} \otimes q_j$ supplies both the attention logit $\operatorname{Re}(r_{ij})$ and a sandwich-product featu...
|
| 91 |
FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception
2609.30629
|
cs.CVcs.LG
|
Rajat Bhattacharjya, Minwoo Kim, Arnab Sarkar, Tamoghno Das, Sing-Yao Wu |
Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-traine...Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-trained split interfaces, while stronger channel-aware codecs can impose substantial onboard cost. We present FreshLatent, a lightweight channel-aware latent adapter that trains a power-normalized encoder-decoder through wireless corruption while...
|
| 92 |
Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction
2609.30631
|
cs.CVcs.SDeess.AS
|
Rayhan Rashed, Senja Filipi, Ross Cutler |
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavi...Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a request...
|
| 93 |
Image Reconstruction from Phase with Untrained Neural Priors
2609.30659
|
cs.CV
|
Ene Meco, Ahmet Enis Cetin |
Fourier phase encodes important spatial image structure, but recovering an image without measured spectral magnitude requires additional constraints and leaves absolute intensity ambiguous. We propose a projection-based two-stage framework that combines Fourie...Fourier phase encodes important spatial image structure, but recovering an image without measured spectral magnitude requires additional constraints and leaves absolute intensity ambiguous. We propose a projection-based two-stage framework that combines Fourier-phase and spatial-support constraints with an image-specific neural prior. The first stage alternates constraint enforcement with regularized neural-prior updates, while the second performs phase/support refinement alone with guaranteed c...
|
| 94 |
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
2609.30670
|
cs.CVcs.CL
|
Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao |
Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, s...Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, similar scores may correspond to different workloads, failure modes, and operational behavior. We introduce TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark and evaluation framework that makes these factors ...
|
| 95 |
Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
2609.30840
|
cs.CVcs.LG
|
Austin Wang, Ziheng Cheng, Lexing Ying |
One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentia...One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution...
|
| 96 |
Universal Drift Correction for Multidimensional Scanning Microscopy
2609.30866
|
cs.CV
|
Sangjoon Lee, William Millsaps, Dasol Yoon, Caitlyn Obrero, Guoliang Hu |
In scanning microscopy, drift causes the specimen to be sampled at positions displaced from the nominal probe positions. This displacement alters the spatial assignment of the recorded signals and biases quantitative measurements across two-dimensional imaging...In scanning microscopy, drift causes the specimen to be sampled at positions displaced from the nominal probe positions. This displacement alters the spatial assignment of the recorded signals and biases quantitative measurements across two-dimensional imaging, channel-resolved spectroscopic mapping, and scan-position-resolved diffraction analysis. Here, we extend orthogonal-scan drift correction from 2D images to spectrum images and diffraction datasets. We demonstrate how to recover probe posi...
|
| 97 |
FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models
2609.30980
|
cs.CV
|
Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao |
Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-ener...Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are brittle: modest post-processing or lightweight adversarial perturbations readily suppress detection, exposing a fundamental tension between imperceptibility and robustness. We intr...
|
| 98 |
Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
2609.30997
|
cs.CVcs.LGcs.AI
|
Kai Yao |
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the pro...Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance bet...
|
| 99 |
TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks
2609.31032
|
cs.CVcs.MM
|
Tianmeng Fang, Jiancheng Wang, Chen Wang, Liming Wang, Wei Wang |
Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefo...Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V jailbreak as a query-constrained candidate allocation and ranking problem and propose TempQ-Jail. The method combines heterogeneous attack mechanisms to expand candidate coverage, estimates each candidate's end-to-end attac...
|
| 100 |
Quantum Diffusion Models for Medical Image Analysis
2609.31070
|
cs.CVcs.LGcs.AI
|
Francesco Aldo Venturelli, Stefano Martina, Marco Parigi, Filippo Caruso, Alba Cervera-Lierta |
Quantum Machine Learning is a novel field of research aimed at devising machine learning approaches exploiting principles of quantum mechanics, such as superposition, entanglement and interference. In this context, we present a scalable hybrid Quantum Diffusio...Quantum Machine Learning is a novel field of research aimed at devising machine learning approaches exploiting principles of quantum mechanics, such as superposition, entanglement and interference. In this context, we present a scalable hybrid Quantum Diffusion Model, and evaluate its use for medical image analysis. Specifically, our method is based on a Discrete-Time Quantum Walk algorithm, executed on a real quantum device, to model the forward dynamics of the diffusion model. For the backward...
|
| 101 |
Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning
2609.31199
|
cs.CV
|
Antoine Lorentz, St\'ephane May, Valentine Bellet, Dawa Derksen, Bastien Nespoulous |
Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammetry. However, the resulting 3D maps are often contaminated by noise, outliers, and voids. On the other hand, aerial LiDAR provides high-accur...Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammetry. However, the resulting 3D maps are often contaminated by noise, outliers, and voids. On the other hand, aerial LiDAR provides high-accuracy elevation measurements at a substantially higher cost. In this work, we study diffusion models conditioned both on photogrammetric DSMs and Pl\'eiades imagery to refine vertically co-registered DSMs. We introduce a modified Stable Diffu...
|
| 102 |
FlatClip: A Geometry-Aware Surface-Level Baseline for fMRI Representation Learning
2609.31204
|
cs.CV
|
Mo Wang, Wenhao Ye, Zihan Ning, Jiayu Zuo, Junfeng Xia |
Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require speciali...Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized 3D/4D architectures and costly fMRI-specific pretraining. We ask how effectively an image-pretrained encoder can reuse the spatial organization of cortical activity. Motivated by evidence that macroscale brain activity is strongly const...
|
| 103 |
Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling
2609.31207
|
cs.CV
|
Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li |
Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverag...Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction ...
|
| 104 |
ChronoFuseGS: Multi-Temporal Gaussian Fusion with Per-Splat Persistence and Change Visualization
2609.31339
|
cs.CV
|
Tobias Batik, Diana Marin, Peter K\'an, Hannes Kaufmann |
Reconstructing environments where parts of the scene change between captured image sets poses a challenge for 3D scene reconstruction. We present ChronoFuseGS, a multi-temporal Gaussian Splatting approach that addresses this issue by taking multiple separately...Reconstructing environments where parts of the scene change between captured image sets poses a challenge for 3D scene reconstruction. We present ChronoFuseGS, a multi-temporal Gaussian Splatting approach that addresses this issue by taking multiple separately trained Gaussian Splatting models, each representing a distinct timestep and partially overlapping in geographic coverage, and merging them into a single combined model. By allowing Gaussians from one timestep to contribute to the reconstr...
|
| 105 |
RECAST: From Log Replay to Closed-Loop Driving Simulation with View-Complete Actors
2609.31374
|
cs.CV
|
Zijun Zhao, Liewen Liao, Kang Shen, Songan Zhang, Ming Yang |
Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded trajectories, exposing views absent from the source log. Existing data-driven simulators reconstruct dynamic a...Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded trajectories, exposing views absent from the source log. Existing data-driven simulators reconstruct dynamic actors from sparse observations, which can result in rendering artifacts under these viewpoint changes. We introduce RECAST (REconstructing Controllable Actors for Simulation and Testing), a 3D Gaussian Splatting framework that generates a v...
|
| 106 |
Towards Whole-Study Screening for Congenital Heart Disease in Fetal Ultrasound Using Multiple Instance Learning
2609.31376
|
cs.CV
|
Mohamed Azzam, Ruobing Liu, Esther C. Ugwueke, Ziyang Xu, Shibiao Wan |
Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated fro...Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated from a study, by a clinician or by a view classifier. We remove that assumption and address CHD screening directly at the level of the whole ultrasound study. We propose a two-stage framework that first learns transferable frame representation...
|
| 107 |
Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization
2609.31383
|
cs.CVcs.LGcs.AI
|
Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta |
End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects...End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or ea...
|
| 108 |
InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
2609.31394
|
cs.CV
|
Xingyu Miao, Zizun Li, Baole Fang, Kaiwen Song, Tenghui Wang |
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion--...World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$\Delta$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$\Delta...
|
| 109 |
Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
2609.31403
|
cs.CVcs.CLcs.AI
|
Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an |
Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using ...Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using the Decoy Font method. The dataset consists of 300 images, each containing text with sharp contour lines superimposed on another text with soft shading. Six recent closed-source models from three different model families were evaluated usin...
|
| 110 |
TemplateCraft: Agentic Visual Template Generation
2609.31451
|
cs.CVcs.MM
|
Hongjie Yu, Zhiyuan Fan, Yuzhe Zhang, Jiangcun Du, Zhicheng Gao |
The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset pre...The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset preparation and tool orchestration. We propose TemplateCraft, a multi-agent system that converts natural-language instructions into client-executable templates through planning, material generation, effect-workflow generation, and protocol com...
|
| 111 |
Vision-Based 6-DoF Grasp Pose Estimation for Robot Cloth Unfolding
2609.31452
|
cs.CV
|
Domen Tabernik, Peter Nimac, Jan Jeri\'cevi\'c, Danijel Sko\v{c}aj, Andrej Gams |
Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and perceptual ambiguity arising from frequent occlusions of critical visual cues such as folds, edges, and grasp p...Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and perceptual ambiguity arising from frequent occlusions of critical visual cues such as folds, edges, and grasp points. In this work, we tackle cloth unfolding using a regrasping-in-the-air strategy, where one manipulator holds the cloth while the other grasps it at an optimally selected point to unfold it. To this end, we propose CeDiRNet-6DoF, a dee...
|
| 112 |
Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality
2609.31454
|
cs.CVcs.LGcs.AI
|
Bradley Scott, Zeqi Luo, Edmond S. L. Ho |
Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL...Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout variance and entropy measures, while the loss is computed ...
|
| 113 |
Uncertainty-Aware Federated Learning for Infant Movement Analysis
2609.31463
|
cs.CVcs.LGcs.AI
|
Edmond S. L. Ho |
Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving perf...Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving performance comparable to expert assessment for tasks such as General Movement Assessment (GMA). However, most existing approaches rely on centralized training, requiring data from multiple institutions to be collected and stored at a single si...
|
| 114 |
MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
2609.31553
|
cs.CVcs.CL
|
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante |
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although ...Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal c...
|
| 115 |
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
2406.09896
|
cs.CV
|
Brun\'o B. Englert, Fabrizio J. Piva, Tommie Kerssies, Daan de Geus, Gijs Dubbelman |
Achieving robust generalization across diverse data domains remains a significant challenge in computer vision. This challenge is important in safety-critical applications, where deep-neural-network-based systems must perform reliably under various environment...Achieving robust generalization across diverse data domains remains a significant challenge in computer vision. This challenge is important in safety-critical applications, where deep-neural-network-based systems must perform reliably under various environmental conditions not seen during training. Our study investigates whether the generalization capabilities of Vision Foundation Models (VFMs) and Unsupervised Domain Adaptation (UDA) methods for the semantic segmentation task are complementary....
|
| 116 |
Adapting Visualization Techniques for Time-Series Anomaly Detection: From Convolutional Neural Networks to Convolutional-Recurrent Neural Networks
2411.04707
|
cs.CV
|
Fabien Poirier, Myriam Lamolle |
Deep neural networks achieve strong performance on complex tasks but are often regarded as "black boxes," which limits their adoption in domains where transparency is essential. This lack of interpretability raises ethical and legal concerns, particularly in s...Deep neural networks achieve strong performance on complex tasks but are often regarded as "black boxes," which limits their adoption in domains where transparency is essential. This lack of interpretability raises ethical and legal concerns, particularly in sensitive applications such as security, where automated decisions can have serious consequences. The General Data Protection Regulation (GDPR) reinforces the need to justify decisions made by these systems. In this work, we investigate visu...
|
| 117 |
Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals
2412.11325
|
cs.CVcs.SDeess.AS
|
Xiaoxuan Liang, Hong Zhou, Zhaolong Wei, Yansong Li, Shujian Yu |
3D human mesh reconstruction (HMR) from RGB images often degrades under poor illumination, occlusion, and non-line-of-sight conditions. Acoustic sensing provides complementary spatial cues but suffers from low spatial resolution. We propose SonicMesh, which, t...3D human mesh reconstruction (HMR) from RGB images often degrades under poor illumination, occlusion, and non-line-of-sight conditions. Acoustic sensing provides complementary spatial cues but suffers from low spatial resolution. We propose SonicMesh, which, to the best of our knowledge, is the first acoustic--visual framework for robust 3D human mesh reconstruction. SonicMesh first converts ultrasonic echoes into range--azimuth acoustic images through an Inverse Synthetic Aperture Radar (ISAR)-...
|
| 118 |
LadderMIL: Multiple Instance Learning with Coarse-to-Fine Self-Distillation
2502.02707
|
cs.CV
|
Shuyang Wu, Yifu Qiu, Ines P. Nearchou, Sandrine Prost, Jonathan A. Fallowfield |
Multiple Instance Learning (MIL) for whole slide image (WSI) analysis in computational pathology often neglects instance-level learning as supervision is typically provided only at the bag level, hindering the integrated consideration of instance and bag-level...Multiple Instance Learning (MIL) for whole slide image (WSI) analysis in computational pathology often neglects instance-level learning as supervision is typically provided only at the bag level, hindering the integrated consideration of instance and bag-level information during the analysis. In this work, we present LadderMIL, a framework designed to improve MIL through two perspectives: (1) employing instance-level supervision and (2) learning inter-instance contextual information at bag level...
|
| 119 |
VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation
2503.10685
|
cs.CV
|
Brun\'o B. Englert, Gijs Dubbelman |
Unsupervised Domain Adaptation (UDA) enables strong generalization from a labeled source domain to an unlabeled target domain, often with limited data. In parallel, Vision Foundation Models (VFMs) pretrained at scale without labels have also shown impressive d...Unsupervised Domain Adaptation (UDA) enables strong generalization from a labeled source domain to an unlabeled target domain, often with limited data. In parallel, Vision Foundation Models (VFMs) pretrained at scale without labels have also shown impressive downstream performance and generalization. This motivates us to explore how UDA can best leverage VFMs. Prior work (VFM-UDA) demonstrated that replacing a standard ImageNet-pretrained encoder with a VFM improves generalization. However, it a...
|
| 120 |
What is the Added Value of UDA in the VFM Era?
2504.18190
|
cs.CV
|
Brun\'o B. Englert, Tommie Kerssies, Gijs Dubbelman |
Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source domain. UDA using Vision Foundation Models (VFMs) with synthetic source data can achieve generalization performanc...Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source domain. UDA using Vision Foundation Models (VFMs) with synthetic source data can achieve generalization performance comparable to fully-supervised learning with real target data. However, because VFMs have strong generalization from their pre-training, more straightforward, source-only fine-tuning can also perform well on the target. As data scenarios ...
|
| 121 |
RefRef: A Dataset and Benchmark for Reconstructing Refractive and Reflective Objects
2505.05848
|
cs.CV
|
Yue Yin, Enze Tao, Weijian Deng, Dylan Campbell |
Modern 3D reconstruction and novel view synthesis approaches have demonstrated strong performance on scenes with opaque, non-refractive objects. However, most assume straight light paths and therefore cannot properly handle refractive and reflective materials....Modern 3D reconstruction and novel view synthesis approaches have demonstrated strong performance on scenes with opaque, non-refractive objects. However, most assume straight light paths and therefore cannot properly handle refractive and reflective materials. The lack of datasets specialized for these effects has impeded efforts to fairly and thoroughly evaluate performance and thereby make progress in this domain. Most existing datasets focus on opaque scenes, while those targeting refractive ...
|
| 122 |
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
2506.02015
|
cs.CV
|
Yoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi, Sungwoong Kim |
Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image alignment, often failing to faithfully depict objects with correct attributes su...Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image alignment, often failing to faithfully depict objects with correct attributes such as color, shape, and spatial relations. To mitigate this issue, previous studies have explored preference optimization methods such as DPO and GRPO, but these approaches incur substantial computational cost, both in constructing preferen...
|
| 123 |
Unsupervised Methods for Video Quality Improvement: A Survey of Restoration and Enhancement Techniques
2507.08375
|
cs.CV
|
Alexandra Malyugina, Yini Li, Joanne Lin, Nantheera Anantrasirichai |
Video restoration and enhancement are critical not only for improving visual quality, but also as essential pre-processing steps to boost the performance of a wide range of downstream computer vision tasks. This survey presents a comprehensive review of video ...Video restoration and enhancement are critical not only for improving visual quality, but also as essential pre-processing steps to boost the performance of a wide range of downstream computer vision tasks. This survey presents a comprehensive review of video restoration and enhancement techniques with a particular focus on unsupervised approaches. We begin by outlining the most common video degradations and their underlying causes, followed by a review of early conventional and deep learning me...
|
| 124 |
Spatial Information Bottleneck for Interpretable Visual Recognition
2511.09239
|
cs.CV
|
Kaixiang Shu |
Deep neural networks typically learn spatially entangled representations that conflate discriminative foreground features with spurious background correlations, thereby undermining model interpretability and robustness. We propose a novel understanding framewo...Deep neural networks typically learn spatially entangled representations that conflate discriminative foreground features with spurious background correlations, thereby undermining model interpretability and robustness. We propose a novel understanding framework for gradient-based attribution from an information-theoretic perspective. We prove that, under mild conditions, the Vector-Jacobian Products (VJP) computed during backpropagation form minimal sufficient statistics of input features with ...
|
| 125 |
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
2511.20561
|
cs.CVcs.CL
|
Yuwei Niu, Weiyang Jin, Jiaqi Liao, Chaoran Feng, Peng Jin |
Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation in Unified Multimodal Models? To investigate this, we introduce UniSandbox, a decoupled evaluation fra...Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation in Unified Multimodal Models? To investigate this, we introduce UniSandbox, a decoupled evaluation framework paired with controlled, synthetic datasets to avoid data leakage and enable detailed analysis. Our findings reveal a significant understanding-generation gap, which is mainly reflected in two key dimensions: reasoning generation and ...
|
| 126 |
Anchor to Expand: Semantic Anchoring for Personalized Text-to-Image Diffusion Models
2511.22245
|
cs.CV
|
Seoyun Yang, Gihoon Kim, Taesup Kim |
Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images. However, learning a new concept while building on the prior knowledge of the pretrained model remains a key chall...Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images. However, learning a new concept while building on the prior knowledge of the pretrained model remains a key challenge. When personalization focuses on learning the target concept, the model tends to overfit the reference examples and degrade its general capability. In contrast, emphasizing prior preservation can hinder capturing distinctive personaliz...
|
| 127 |
Frequency-Decomposed Avatar Representation for Varying Camera Distances
2512.03593
|
cs.CV
|
David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue |
We present a CloseUpAvatar - a novel approach for articulated human avatar representation supporting a wider range of camera motions, while preserving rendering quality for close-up views. CloseUpAvatar represents an avatar as a set of textured planes with fre...We present a CloseUpAvatar - a novel approach for articulated human avatar representation supporting a wider range of camera motions, while preserving rendering quality for close-up views. CloseUpAvatar represents an avatar as a set of textured planes with frequency-decomposed learnable textures for low and high-frequency detail. The method automatically switches to high-frequency textures when the camera comes close to the avatar's surface and gradually reduces their impact as the camera moves ...
|
| 128 |
Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement
2512.08982
|
cs.CVcs.AI
|
Jian Xu, Wei Chen, Shigui Li, Delu Zeng, John Paisley |
Retinex-based low-light image enhancement benefits from separating reflectance and illumination, yet recent generative approaches often rely on iterative sampling and are difficult to deploy under strict latency budgets. Consistency models offer a natural rout...Retinex-based low-light image enhancement benefits from separating reflectance and illumination, yet recent generative approaches often rely on iterative sampling and are difficult to deploy under strict latency budgets. Consistency models offer a natural route to one-step restoration, but direct adaptation to Retinex-factorized enhancement is unstable: one-step inference is evaluated at the high-noise endpoint, whereas standard training schedules provide little supervision there, and temporal s...
|
| 129 |
Demo: Generative AI helps Radiotherapy Planning with User Preference
2512.08996
|
cs.CVcs.LGcs.AI
|
Riqiang Gao, Simon Arberet, Martin Kraus, Han Liu, Wilko FAR Verbakel |
Radiotherapy planning is a highly complex process that often varies significantly across institutions and individual planners. Most existing deep learning approaches for 3D dose prediction rely on reference plans as ground truth during training, which can inad...Radiotherapy planning is a highly complex process that often varies significantly across institutions and individual planners. Most existing deep learning approaches for 3D dose prediction rely on reference plans as ground truth during training, which can inadvertently bias models toward specific planning styles or institutional preferences. In this study, we introduce a novel generative model that predicts 3D dose distributions based solely on user-defined preference flavors. These customizable...
|
| 130 |
Prompt-Based Continual Compositional Zero-Shot Learning
2512.09172
|
cs.CVcs.AI
|
Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali |
We tackle continual adaptation of vision-language models to new attributes, objects, and their compositions in Compositional Zero-Shot Learning (CZSL), while preventing forgetting of prior knowledge. Unlike classical continual learning where classes are disjoi...We tackle continual adaptation of vision-language models to new attributes, objects, and their compositions in Compositional Zero-Shot Learning (CZSL), while preventing forgetting of prior knowledge. Unlike classical continual learning where classes are disjoint, CCZSL is more complex as attributes and objects may reoccur across sessions while compositions remain unique. Built on a frozen VLM backbone, we propose the first Prompt-based Continual Compositional Zero-Shot Learning (PromptCCZSL) fra...
|
| 131 |
Geometric-Photometric Event-based 3D Gaussian Ray Tracing
2512.18640
|
cs.CVcs.AI
|
Kai Kohyama, Yoshimitsu Aoki, Guillermo Gallego, Shintaro Shiba |
Event cameras offer a high temporal resolution over traditional frame-based cameras, which makes them suitable for motion and structure estimation. However, it has been unclear how event-based 3D Gaussian Splatting (3DGS) approaches could leverage fine-grained...Event cameras offer a high temporal resolution over traditional frame-based cameras, which makes them suitable for motion and structure estimation. However, it has been unclear how event-based 3D Gaussian Splatting (3DGS) approaches could leverage fine-grained temporal information of sparse events. This work proposes GPERT, a framework to address the trade-off between accuracy and temporal resolution in event-based 3DGS. Our key idea is to decouple the rendering into two branches: event-by-event...
|
| 132 |
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
2512.21004
|
cs.CV
|
Jinghan Li, Yang Jin, Hao Jiang, Yadong Mu, Yang Song |
Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative pretraining methods still rel...Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative pretraining methods still rely on BERT-style masked modeling, which often disregards the temporal information essential for video analysis. The few existing autoregressive visual pretraining methods suffer from issues such as inaccurate semantic localization and poor g...
|
| 133 |
Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models
2602.02043
|
cs.CVcs.AI
|
Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci |
Modern vision-language models struggle with basic compositional reasoning, failing to bind attributes to objects or relations to their referents. Existing benchmarks either rely on noisy real images that conflate confounding visual variables with the reasoning...Modern vision-language models struggle with basic compositional reasoning, failing to bind attributes to objects or relations to their referents. Existing benchmarks either rely on noisy real images that conflate confounding visual variables with the reasoning failure, or use simplistic synthetic scenes lacking the realism modern VLMs are tuned for. We introduce \textbf{Auto-Comp}, a fully automated, concept-driven pipeline that bridges this gap by generating photorealistic compositional benchma...
|
| 134 |
VLANeXt Family: A Systematic Study of VLA Models from Core Recipes to Emerging Paradigms
2602.18532
|
cs.CVcs.AI
|
Xiao-Ming Wu, Kang Liao, Yihang Luo, Bin Fan, Jian-Jian Jiang |
Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented ...Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving spa...
|
| 135 |
A Multi-Stage Framework for Kuzushiji Character Recognition in Japanese Historical Documents
2602.19086
|
cs.CV
|
Rui-Yang Ju, Kohei Yamashita, Hirotaka Kameko, Shinsuke Mori |
Kuzushiji was a widely used cursive writing system in pre-modern Japan. Due to simplification and glyph variation, most modern Japanese readers cannot read Kuzushiji characters. Consequently, recent studies have developed optical character recognition (OCR) sy...Kuzushiji was a widely used cursive writing system in pre-modern Japan. Due to simplification and glyph variation, most modern Japanese readers cannot read Kuzushiji characters. Consequently, recent studies have developed optical character recognition (OCR) systems for Kuzushiji. Despite recent progress, Kuzushiji character recognition (KCR) in Japanese historical documents remains challenging because of seal-character overlap and complex layouts, which interfere with character recognition and h...
|
| 136 |
Scale-invariant Gaussian derivative residual networks
2603.02843
|
cs.CVcs.LG
|
Andrzej Perzanowski, Tony Lindeberg |
Generalisation across image scales remains a fundamental challenge for deep networks, which often fail to handle images at scales not seen during training (the out-of-distribution problem). In this paper, we present provably scale-invariant Gaussian derivative...Generalisation across image scales remains a fundamental challenge for deep networks, which often fail to handle images at scales not seen during training (the out-of-distribution problem). In this paper, we present provably scale-invariant Gaussian derivative residual networks (GaussDerResNets), constructed out of scale-covariant Gaussian derivative residual blocks coupled in cascade, aimed at addressing this problem. By adding residual skip connections to the previous notion of Gaussian deriva...
|
| 137 |
NBAvatar: Neural Billboards Avatars with Realistic Hand-Face Interaction
2603.12063
|
cs.CV
|
David Svitov, Mahtab Dahaghin, Pietro Morerio, Alessio Del Bue |
We present NBAvatar - a method for realistic rendering of head avatars handling non-rigid deformations caused by hand-face interaction. To this end, we introduce a novel hybrid implicit-explicit representation for animated avatars by combining the training of ...We present NBAvatar - a method for realistic rendering of head avatars handling non-rigid deformations caused by hand-face interaction. To this end, we introduce a novel hybrid implicit-explicit representation for animated avatars by combining the training of explicit oriented planar primitives with implicit neural rendering. Such a combination of representations in the end-to-end pipeline enables NBAvatar to handle temporally and pose-consistent geometry, along with fine-grained appearance deta...
|
| 138 |
GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow
2603.26571
|
cs.CVcs.AI
|
Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi |
At ultra-low bitrates, high-fidelity reconstruction requires sampling plausible videos from the posterior rather than regressing to oversmoothed conditional means. We propose Generative Video Codebook Codec (GVCC), a zero-shot framework in which a pretrained v...At ultra-low bitrates, high-fidelity reconstruction requires sampling plausible videos from the posterior rather than regressing to oversmoothed conditional means. We propose Generative Video Codebook Codec (GVCC), a zero-shot framework in which a pretrained video generative model serves directly as the decoder, and the transmitted bitstream specifies its generation trajectory. Modern rectified-flow video models are typically sampled with deterministic ODE solvers, which leave no per-step stocha...
|
| 139 |
At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections
2604.10766
|
cs.CV
|
Ming-Yang Ho, Alberto Bartesaghi |
Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM constraints prohibit processing an entire 3D tomogram, forcing current methods to rely on slow sliding-window in...Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM constraints prohibit processing an entire 3D tomogram, forcing current methods to rely on slow sliding-window inference over extracted subvolumes. To overcome this, we propose FullTilt, an end-to-end framework that redefines 3D detection by operating directly on aligned 2D tilt-series. Because a tilt-series contains significantly fewer images than sl...
|
| 140 |
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
2604.17422
|
cs.CVcs.MM
|
Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu, Hui Xiong |
Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dense frame sequences. Prevailing keyframe-selection methods rely on either a single visual-centric metric (e.g., CLI...Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dense frame sequences. Prevailing keyframe-selection methods rely on either a single visual-centric metric (e.g., CLIP similarity) or a static fusion of heuristic scores. This "one-size-fits-all" paradigm frequently fails: visual-only metrics are ineffective for plot-driven narrative queries, while indiscriminately adding textual scores introduces severe ...
|
| 141 |
From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation
2605.08712
|
cs.CV
|
Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang |
Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic...Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic-to-visual lifting paradigm that converts articulated kinematics into a unified set of five image-aligned control modalities. Building on this representation, we introduce a hierarchically routed visual control framework that selectively ac...
|
| 142 |
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
2605.11723
|
cs.CVcs.AI
|
Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan |
In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial ...In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiotemporal Chain-of-Thought reasoning. To equip the model with these capabilities, we construct the first large-scale generated video anomaly d...
|
| 143 |
Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
2605.19398
|
cs.CVcs.AI
|
Wooseok Jeon, Seungho Park, Seunghyun Shin, Sangeyl Lee, Hyeonho Jeong |
Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional training or sacrifice fid...Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional training or sacrifice fidelity to the reference image. In this work, we identify reference-frame dominance as a key mechanism behind motion suppression. We observe that non-reference frames in I2V models allocate excessive self-attention to reference-frame key toke...
|
| 144 |
C2P-VAR: Continual and Compositional Personalization in Visual Autoregressive Models
2605.19750
|
cs.CV
|
Junhao Li, Xinhao Zhong, Yi sun, Yuxia Qiao, Bin Chen |
Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation, yet their personalization capabilities remain largely limited to static, single-concept settings. In practice, users may continuously introduce new ...Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation, yet their personalization capabilities remain largely limited to static, single-concept settings. In practice, users may continuously introduce new concepts and wish to compose multiple personalized concepts within a single image. Such scenarios pose two fundamental challenges: catastrophic forgetting during sequential personalization and feature interference during multi-concept compo...
|
| 145 |
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
2606.01900
|
cs.CV
|
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra, Xuelin Chen, Erkut Erdem |
Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spat...Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language-driven, human-centric camera framing in generative video. Our core insight is that professional filmmakers co...
|
| 146 |
EmoZone-Talker: Regional Semantic Control of Audio-Driven 3DGS Talking Heads via Facial Action Units
2606.15848
|
cs.CV
|
Tingting Chen, Shaojun Wang, Huaye Zhang, Diqiong Jiang, Chenglizhao Chen |
3D Gaussian Splatting (3DGS) has shown strong potential for high-fidelity talking head synthesis. However, enabling fine-grained, interpretable, and editable facial expression control remains fundamentally challenging due to intrinsic conflicts between speech-...3D Gaussian Splatting (3DGS) has shown strong potential for high-fidelity talking head synthesis. However, enabling fine-grained, interpretable, and editable facial expression control remains fundamentally challenging due to intrinsic conflicts between speech-driven facial dynamics and explicit expression signals. Existing methods rely on implicit multimodal fusion, leading to spatial entanglement and temporal instability. We present EmoZone-Talker, a novel framework that reformulates audio-driv...
|
| 147 |
Holo-World: Unified Camera, Object and Weather Control for Video World Model
2606.20083
|
cs.CV
|
Xiangchen Yin, Wenzhang Sun, Jiahui Yuan, Zijie Liu, Yinda Chen |
Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls remain isolated, and weather generation typically relies on a source video or rec...Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls remain isolated, and weather generation typically relies on a source video or reconstructed scene that already specifies future structure. We study a first-frame-anchored source-to-state setting, where the model starts from a single image and follows explicit camera and object controls and an optional weather instructio...
|
| 148 |
Cross-Modality Structural Guidance in 3D Latent Diffusion for Robust FLAIR Super-Resolution
2606.25255
|
cs.CV
|
Haoyu Lan, Jiazhen Zhang, John Onofrey, Bino Varghese, Nasim Sheikh-Bahaei |
High-resolution (HR) MRI acquisition is often hampered by scan time constraints, resulting in anisotropic or low-resolution scans (e.g., thick-slice FLAIR) that limit diagnostic accuracy. While deep learning-based super-resolution (SR) methods show promise, th...High-resolution (HR) MRI acquisition is often hampered by scan time constraints, resulting in anisotropic or low-resolution scans (e.g., thick-slice FLAIR) that limit diagnostic accuracy. While deep learning-based super-resolution (SR) methods show promise, they often hallucinate anatomical details, which can compromise brain structural integrity. To mitigate this limitation, we introduce MR-DiffuSR, a Multi-Resolution Diffusion-based Super-Resolution framework that incorporates HR T1w structura...
|
| 149 |
DreamSat-Pose: Spacecraft Pose Estimation from Single-View 3D Reconstructions and Learned 2D-3D Feature Matching
2607.13449
|
cs.CVcs.LG
|
Josiane Uwumukiza, Jocelyn Zhao, Giovanni Lavezzi, Giacomo Battaglia, Paolo Panicucci |
6-DoF pose estimation is a critical task in autonomous rendezvous and proximity operations. In the case of an unknown target, this task becomes challenging as it shall be paired with the reconstruction of the target shape model. In this article, we propose a n...6-DoF pose estimation is a critical task in autonomous rendezvous and proximity operations. In the case of an unknown target, this task becomes challenging as it shall be paired with the reconstruction of the target shape model. In this article, we propose a novel framework for single-shot shape and pose estimation of unknown spacecraft objects. Given a single image, we first reconstruct a 3D shape model of the target, then estimate the relative six-degrees-of-freedom pose by learning dense 2D-3...
|
| 150 |
Importance-Aware OBS Pruning for Diffusion Models
2607.20048
|
cs.CV
|
Ba-Thinh Lam, Srijan Das, Hieu Le |
We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps -- derived from conditioning signals or ...We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps -- derived from conditioning signals or model attention -- into the pruning objective. This produces parameter rankings aligned with perceptual relevance rather than uniform reconstruction error. On MS-COCO dataset, our proposed approach consistently retains subject fidelity and ...
|
| 151 |
When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation
2607.22771
|
cs.CVcs.AI
|
Renjie Liang, Zijian Xu, Jinqian Pan, Chengkun Sun, Zhengkang Fan |
Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive search. A cheap pr...Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive search. A cheap probe on the encoder's representation promises a way out, but whether it forecasts the expensive outcome has never been tested. We test this with CheapCT on report generation and on MeasureVQA, a new VQA dataset we build. MeasureVQA scores th...
|
| 152 |
Towards Unified Dynamic Face Landmark Detection
2608.10346
|
cs.CVcs.AI
|
Sebastian Regalado, Varshanth R. Rao, Ruowei Jiang, Parham Aarabi, Igor Gilitschenski |
Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each ``$N$-point'' benchmark dataset, ...Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each ``$N$-point'' benchmark dataset, and (2) a model trained on an ``$N$-point'' dataset reliably outputs only the $N$ landmarks. In our work, we first conceptualize Face Part-Anchored Landmark Positions (FPALPs), wherein each landmark is treated as a progression value between...
|
| 153 |
A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources
2608.13183
|
cs.CV
|
Brun\'o B. Englert, Gijs Dubbelman |
Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustainable or unnecessary, and significant be...Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustainable or unnecessary, and significant benefits arise when effective models can be obtained with fewer resources. To better understand how self-supervised learning (SSL) objectives behave under resource constraints, we conduct a controlled study of image and video SSL objectives u...
|
| 154 |
AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation
2608.29081
|
cs.CV
|
Soumyaratna Debnath, Weiming Zhang, Shriram Damodaran, Dingwen Xiao, Addison Lin Wang |
Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical...Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical structure and stable viewpoints, which are frequently violated in real-world imagery due to unconstrained camera motion, introducing contextual and geometric ambiguity. Consequently, they lack adaptive mechanisms to handle such ambiguity, ...
|
| 155 |
Learning with Volterra Neural Networks: A System Theoretic Perspective
2609.01928
|
cs.CV
|
Haoyu Yun, Hamid Krim, Yufang Bao |
Higher-order interaction components are important for signal, image, and video modeling, but explicit high-order operators often suffer from rapidly increasing parameter and computational costs. This paper presents kVNN, a learnable kernelized Volterra Neural ...Higher-order interaction components are important for signal, image, and video modeling, but explicit high-order operators often suffer from rapidly increasing parameter and computational costs. This paper presents kVNN, a learnable kernelized Volterra Neural operator for compact higher-order filtering. The motivation is to use kernelization to improve the efficiency of Volterra-type neural operators while providing a structured interpretation of their higher-order components. The proposed formu...
|
| 156 |
Learning to Track from Privileged Target Appearances
2609.02471
|
cs.CV
|
Xin Chen, Jiao Xu, Dong Wang, Huchuan Lu, Kede Ma |
Target templates define what a visual tracker searches for, yet the templates available at inference trade off localization certainty with appearance freshness: the initial ground-truth template is exact but becomes stale, whereas recent templates better refle...Target templates define what a visual tracker searches for, yet the templates available at inference trade off localization certainty with appearance freshness: the initial ground-truth template is exact but becomes stale, whereas recent templates better reflect the current appearance but are cropped from uncertain predictions. We quantify this bottleneck with a non-deployable oracle that supplies an exact current-frame target crop, improving AUC on LaSOT by 15.2 percentage points. This gap reve...
|
| 157 |
SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
2609.05533
|
cs.CVcs.LG
|
Cheng Yin, Wang Xu, Junpeng Yang, Sikyuen Tam, Hanyu Liu |
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms for VLAs, such as retrieval banks, learned compressors and recurrent states, mu...Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms for VLAs, such as retrieval banks, learned compressors and recurrent states, must decide what to keep from the past before knowing what a future decision will require. They were motivated by the assumption that minute-scale history is too large to process directly, which no longer holds for modern VLM backbones. We pr...
|
| 158 |
PLSR: Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion
2609.06436
|
cs.CV
|
Yuxin Liu, Minshan Xie, Jiawen Liang, Runsong Zhu, Chi-Wing Fu |
High-resolution 3D asset generation is vital in various 3D applications. Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details. In this paper, we tackle the challenge of generating m...High-resolution 3D asset generation is vital in various 3D applications. Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details. In this paper, we tackle the challenge of generating more detailed, higher-resolution 3D objects by introducing a 3D super-resolution (SR) framework built on existing 3D generative foundation models. To this end, we design PLSR, a progressive and localized super-resolution solution to achieve ...
|
| 159 |
Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding
2609.06475
|
cs.CV
|
Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian |
Large vision-language models (LVLMs) have recently achieved remarkable progress in general-purpose video understanding. However, their application to real-world surveillance remains challenging due to the lack of large-scale domain-specific datasets and the li...Large vision-language models (LVLMs) have recently achieved remarkable progress in general-purpose video understanding. However, their application to real-world surveillance remains challenging due to the lack of large-scale domain-specific datasets and the limitation of passive observation from fixed viewpoints. In surveillance scenarios, critical visual evidence can be easily missed when targets are distant, small, occluded, or move beyond the current camera view. In this work, we introduce Ca...
|
| 160 |
ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
2609.16284
|
cs.CVcs.AI
|
Yan Zhu, Yongbo Chen, Zhengming Ding, Rebecca Faust |
Query-conditioned vision-language models enable fine-grained interpretation by revealing which visual content supports a given textual query and how this evidence changes across queries. However, semantically, sentence-level evidence does not necessarily decom...Query-conditioned vision-language models enable fine-grained interpretation by revealing which visual content supports a given textual query and how this evidence changes across queries. However, semantically, sentence-level evidence does not necessarily decompose into object-specific contributions, while spatially, object-level evidence can remain entangled with co-occurring objects and surrounding scene context. Across multiple VLM architectures and independent benchmarks, we observe persisten...
|
| 161 |
A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data
2609.16597
|
cs.CVcs.AI
|
Yinong Wang (Joyce), Jianwen Chen (Joyce), Zhou Chen (Joyce), Shuwen Kuang (Joyce), Haoning Jiang (Joyce) |
We developed BrainVLM to classify all 12 World Health Organization (WHO) 2021 brain tumor types. BrainVLM integrates an uncertainty quantification strategy to indicate prediction reliability and a module for generating radiology reports to elucidate the clinic...We developed BrainVLM to classify all 12 World Health Organization (WHO) 2021 brain tumor types. BrainVLM integrates an uncertainty quantification strategy to indicate prediction reliability and a module for generating radiology reports to elucidate the clinical rationale. BrainVLM was trained on multi-modal data (MRI scans, demographics, and radiology reports) from 40,043 individuals. It was validated on 5,211 patients with pathologically confirmed brain tumors, including 3,877 held-out patient...
|
| 162 |
Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
2609.20633
|
cs.CV
|
Yulong Chen, Ziqian Zhang, Haoyu Zhang, Ao He, Yaxing Wang |
Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal au...Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. As the first to explore training-free image editing with Generative Refinement Networks (GRN), we observe that its refinement p...
|
| 163 |
CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation
2609.21455
|
cs.CV
|
Haoran Qin (Harbin Institute of Technology, China), Renlong Wu (Harbin Institute of Technology, China), Tianyu Huang (Harbin Institute of Technology |
While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate e...While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a ...
|
| 164 |
MM-ContextFold: Context Folding for Multimodal Agentic Retrieval
2609.23121
|
cs.CVcs.AIcs.MM
|
Yang Tian, Fan Liu, Jingyuan Zhang, Zhenyang Li, Yupeng Hu |
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-gro...Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address thi...
|
| 165 |
Retrieval Geometry Shapes Cache-Based Clip Adaptation
2609.23409
|
cs.CV
|
Mahir Shahriar Tamim, Md. Samiul Alim, Azmine Toushik Wasi, Shahriyar Zaman Ridoy, Meharun Nesa |
Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image-image retrieval as fixed, leaving open...Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image-image retrieval as fixed, leaving open how much adaptation depends on the retrieval space itself. We study this question by fixing the memory and changing only the retrieval encoder, finding that the same memory can yield very different gains: across sixteen retrieval spaces, I...
|
| 166 |
You've Seen Enough: Quality-Constrained Image Coding for Machines
2609.25108
|
cs.CVcs.AI
|
Khoa Pham-Dinh, Sanaz Nami, Hamed Rezazadegan Tavakoli, Moncef Gabbouj, Farhad Pakdaman |
Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validat...Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validate the decisions. Inspired by just-noticeable distortion, we cap human-observed quality at a desired level and devote the remaining bits to machine performance. Specifically, joint compression-segmentation training is recast as a constrained...
|
| 167 |
MIND the Gap: A Geographic Implicit Neural Representation with Adjustable Spatial Scale
2609.25454
|
cs.CVcs.LG
|
Isaac Corley, Arjun Rao, Esther Rolf, Konstantin Klemmer, Evan Shelhamer |
Geographic measurements are often sparse, leaving large areas without labels for the quantities we want to map. Geographic implicit neural representations (INRs) provide coordinate-based embeddings that can be combined with sparse labels to predict at unsample...Geographic measurements are often sparse, leaving large areas without labels for the quantities we want to map. Geographic implicit neural representations (INRs) provide coordinate-based embeddings that can be combined with sparse labels to predict at unsampled locations without satellite imagery at inference. Yet existing INRs are largely evaluated with random holdouts, leaving their ability to generalize across larger geographic gaps unclear. We introduce Matryoshka Implicit Neural Distillatio...
|
| 168 |
NV-Reason-CT: 3D Visual Language Model for CT Analysis
2609.27511
|
cs.CVcs.AI
|
Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon |
We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and...We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing ...
|
| 169 |
Track2Art: Articulated Object Model Recovery with Visual-Geometric Track Representations
2609.27675
|
cs.CV
|
Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam |
Understanding articulated objects is fundamental for robotic interaction, requiring accurate rigid-part discovery and the recovery of their kinematic relations. Existing approaches often treat articulation as a by-product of reconstructed geometry or recover i...Understanding articulated objects is fundamental for robotic interaction, requiring accurate rigid-part discovery and the recovery of their kinematic relations. Existing approaches often treat articulation as a by-product of reconstructed geometry or recover it through per-instance optimization. We instead build on the hypothesis that articulation is directly observable from persistent motion: points on the same rigid part move coherently, while relative motion between parts reveals their kinema...
|
| 170 |
LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
2609.28327
|
cs.CV
|
Andrei Arhire, Mihaela-Elena Breab\u{a}n, Radu Timofte |
We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Pro...We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade. The cascade combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module, which uses temporary channel expa...
|
| 171 |
M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals
2609.28684
|
cs.CVcs.LG
|
Vin\'icius da Silva, Isabelle Melo, Matheus Bessa, Guilherme Schardong, Luiz Schirmer |
Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training effici...Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training efficiency, rendering speed, and noise robustness: single-MLP approaches are expensive at inference, grid-based representations are fast but can limit surface smoothness and overfit input noise, and previous multiscale approaches frequently captu...
|
| 172 |
Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study
2609.29433
|
cs.CVcs.AI
|
Raghavan Lavanya, Yangqin Feng, Ten Cheer Quek, Quan V. Hoang, Linda Yi-Chieh Poon |
Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting c...Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting conditions such as high myopia (HM). We developed and validated a Vision Transformer-based deep learning (DL) model for glaucoma detection across multi-ethnic cohorts with and without HM. Methods: A ViT-B/16 model with predictive uncertainty...
|
| 173 |
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
2609.29816
|
cs.CVcs.SD
|
Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu |
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training o...Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computa...
|
| 174 |
DeepFedNAS: Efficient Hardware-Aware Architecture Adaptation for Heterogeneous IoT Federations via Pareto-Guided Supernet Training
2601.15127
|
cs.CVcs.LG
|
Bostan Khan, Masoud Daneshtalab |
Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet existing Federated Neural Architecture Search (FedNAS) methods suffer from unguided supernet training and prohibitivel...Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet existing Federated Neural Architecture Search (FedNAS) methods suffer from unguided supernet training and prohibitively costly post-training search pipelines that validate thousands of subnets to construct learned accuracy predictors. We introduce DeepFedNAS, a two-phase framework built on a multi-objective fitness function that synthesizes information-the...
|
| 175 |
Pseudo-Invertible Neural Networks
2602.06042
|
cs.CVcs.LG
|
Yamit Ehrlich, Nimrod Berman, Assaf Shocher |
The Moore-Penrose Pseudo-inverse (PInv) serves as the fundamental solution for linear systems. In this paper, we propose a natural generalization of PInv to the nonlinear regime in general and to neural networks in particular. We introduce Surjective Pseudo-in...The Moore-Penrose Pseudo-inverse (PInv) serves as the fundamental solution for linear systems. In this paper, we propose a natural generalization of PInv to the nonlinear regime in general and to neural networks in particular. We introduce Surjective Pseudo-invertible Neural Networks (SPNN), a class of architectures explicitly designed to admit a tractable non-linear PInv. The proposed non-linear PInv and its implementation in SPNN satisfy fundamental geometric properties. One such property is n...
|
| 176 |
CT-Merging: Consensus Directions and Task-Specific Scaling for LoRA Adapter Merging
2607.20561
|
cs.CVcs.LG
|
Keumseo Ryum, Joonhyuk Kang |
LoRA merging methods increasingly operate on the low-rank structure of task updates, yet how the common subspace is estimated and how coefficients are assigned after recomposition are rarely compared directly. We propose CT-Merging, which estimates common dire...LoRA merging methods increasingly operate on the low-rank structure of task updates, yet how the common subspace is estimated and how coefficients are assigned after recomposition are rarely compared directly. We propose CT-Merging, which estimates common directions from averaged task subspace projectors and assigns a separate residual scale to each task. Projector averaging selects directions supported across task subspaces without weighting them by singular magnitude, while task-specific scali...
|
| 177 |
Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
2609.04250
|
cs.CVcs.SDeess.AS
|
Chengqian Ma, Wei Tao, Haoyu Zhang, Yiwen Guo |
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio han...An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. W...
|
| 178 |
Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics
2609.26567
|
cs.CV
|
Eshika Pathak, Leela Krishna |
Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rul...Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rule can score well by repeatedly using one broadly transferable experience, or poorly because its preferred experience is weak. Since robots increasingly adapt by reuse rather than retraining, a score that describes the library rather than th...
|
| 179 |
EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies
2609.29310
|
cs.CV
|
Hanbit Oh, Yukiyasu Domae, Takuma Yagi |
Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, b...Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce...
|
| 180 |
FMCW-LIO: A Doppler LiDAR-Inertial Odometry
2609.29374
|
cs.CV
|
Mingle Zhao, Jiahao Wang, Tianxiao Gao, Chengzhong Xu, Hui Kong |
Conventional LiDAR-inertial odometry (LIO) or simultaneous localization and mapping (SLAM) methods heavily rely on geometric features of environments, as LiDARs primarily provide range measurements instead of motion measurements. From now on, however, the situ...Conventional LiDAR-inertial odometry (LIO) or simultaneous localization and mapping (SLAM) methods heavily rely on geometric features of environments, as LiDARs primarily provide range measurements instead of motion measurements. From now on, however, the situation changes thanks to the novel Frequency Modulated Continuous Wave (FMCW) Doppler LiDARs. FMCW Doppler LiDARs not only offer the point range with high resolution but also capture the instant point Doppler velocity through the Doppler eff...
|
| 181 |
Free-Init: Scan-Free, Motion-Free, and Correspondence-Free Initialization for Doppler LiDAR-Inertial Systems
2609.29375
|
cs.CV
|
Mingle Zhao, Jiahao Wang, Tianxiao Gao, Chengzhong Xu, Hui Kong |
Robust initialization is crucial for online systems. In the letter, a high-frequency and resilient initialization framework is designed for LiDAR-inertial systems, leveraging both inertial sensors and Doppler LiDAR. The innovative FMCW Doppler LiDAR opens up a...Robust initialization is crucial for online systems. In the letter, a high-frequency and resilient initialization framework is designed for LiDAR-inertial systems, leveraging both inertial sensors and Doppler LiDAR. The innovative FMCW Doppler LiDAR opens up a novel avenue for robotic sensing by capturing not only point range but also Doppler velocity via the intrinsic Doppler effect. By fusing point-wise Doppler velocity with inertial measurements under non-inertial kinematics, the proposed fra...
|
| 182 |
ToCo-Mesh: Topology-Consistent Dynamic Mesh Reconstruction via Adaptive Tessellation and Surface-Aligned 2DGS
2609.29529
|
cs.CV
|
Chuanjin Fan, Wenjie Chang, Aibing Li, Bingzhou Wang, Wenfei Yang |
Reconstructing dynamic meshes with consistent topology from multi-view temporal images remains a challenge. Existing approaches typically face a dilemma between fine-scale shape recovery and topological stability. Frame-by-frame extraction methods capture fine...Reconstructing dynamic meshes with consistent topology from multi-view temporal images remains a challenge. Existing approaches typically face a dilemma between fine-scale shape recovery and topological stability. Frame-by-frame extraction methods capture fine details but break vertex correspondence, leading to flickering meshes. Conversely, template-based deformation ensures consistency but struggles to adapt its surface resolution during optimization, missing local surface details. To address ...
|
| cs.LG 287 papers | ||||
| 310 |
HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference
2609.30270
|
cs.LG
|
Simran Koul |
On-device inference with small language models keeps user data local, works offline, and incurs no per-query cost, so the on-device tier is preferred when it is adequate. It is thermally constrained, however, and I find the constraint is sharper than a slowdow...On-device inference with small language models keeps user data local, works offline, and incurs no per-query cost, so the on-device tier is preferred when it is adequate. It is thermally constrained, however, and I find the constraint is sharper than a slowdown: on a flagship Snapdragon device, sustained on-device generation destabilizes the GPU inference runtime, which crashes or silently wedges after a few consecutive queries. The failure lies in the current toolchain (OpenCL kernel compilatio...
|
| 311 |
When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
2609.30271
|
cs.LG
|
Gongyue Zhang, Honghai Liu |
Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and ...Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and the global learning rate is less understood. We perform a controlled cross-environment study using a paired four-environment classification problem with stable sparse features, environment-dependent spurious sparse features, dense features,...
|
| 312 |
ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
2609.30272
|
cs.LGcs.AI
|
Mohd Moin Khan, Naman Srivastava, Pandarasamy Arjunan |
We present \textbf{ENAS}, a hardware-aware Neural Architecture Search (NAS) framework that combines a static feasibility check, a cell-based search space supporting standard, depthwise-separable, and bottleneck blocks with optional skip connections, and a thre...We present \textbf{ENAS}, a hardware-aware Neural Architecture Search (NAS) framework that combines a static feasibility check, a cell-based search space supporting standard, depthwise-separable, and bottleneck blocks with optional skip connections, and a three-stage hybrid search strategy (random $\rightarrow$ top-$K$ $\rightarrow$ mutation) with persistent cross-run caching. Unlike many existing NAS frameworks that rely on GPU acceleration, ENAS is designed to operate efficiently without requi...
|
| 313 |
Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments
2609.30273
|
cs.LG
|
Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha |
We investigate how historical data from fixed randomized experiments (A/B tests) can be used to inform the deployment of adaptive experiments based on contextual bandits. Given data collected under a static allocation, our goal is to assess which adaptive poli...We investigate how historical data from fixed randomized experiments (A/B tests) can be used to inform the deployment of adaptive experiments based on contextual bandits. Given data collected under a static allocation, our goal is to assess which adaptive policies, if any, would have outperformed the original design and under what conditions. To this end, we combine off-policy evaluation (OPE) with a controlled warm-start simulation. From logged A/B test data exhibiting heterogeneous treatment e...
|
| 314 |
Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization
2609.30275
|
cs.LG
|
Pranav Varshney |
A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as sur...A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as surviving the change by scale-invariant statistics (cosine similarity, correlation, AUROC) that are reported without their noise floor. For the difference-in-means direction estimator, the split-half floor is governed by one dimensionless numb...
|
| 315 |
Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness
2609.30276
|
cs.LG
|
Alokendu Mazumder, Ayaan Mohd, Harshit Rawat, Arnab Roy, Mayank Baranwal |
We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to hav...We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moments. We show that unclipped AdaGrad can become \emph{anisotropically miscalibrated}: under heavy-tailed noise, the adaptive denominator can learn the geometry of rare noise shocks rather than the local ...
|
| 316 |
Fixed Points Without Fixed Diffusion: Implicit Neural Sheaves for Convergent Test-Time Computation
2609.30277
|
cs.LG
|
R\'emi Bourgerie, \v{S}ar\={u}nas Girdzijauskas, Viktoria Fodor |
Implicit Graph Neural Networks (IGNNs) define node representations as fixed points of message-passing operators, enabling effectively infinite-depth propagation, iteration-independent parameterization, and flexible test-time computation. Yet these benefits dep...Implicit Graph Neural Networks (IGNNs) define node representations as fixed points of message-passing operators, enabling effectively infinite-depth propagation, iteration-independent parameterization, and flexible test-time computation. Yet these benefits depend on the equilibrium being unique and attainable by fixed-point iteration. Existing constructions often impose constraints on recurrent updates to obtain these guarantees, limiting the transformations available at equilibrium. This raises...
|
| 317 |
Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation
2609.30279
|
cs.LG
|
Venkata Subbaiah Yerrapati, Rahul Dixit, Ajay Kumar Shukla |
Understanding the features captured by the hidden layers of neural networks is a fundamental challenge in machine learning, despite their widespread success across various classification problems. In this work, we propose an algebraic framework for examining n...Understanding the features captured by the hidden layers of neural networks is a fundamental challenge in machine learning, despite their widespread success across various classification problems. In this work, we propose an algebraic framework for examining neural networks that model classification problems. Certain results, such as the correspondence between the neural network and neural ideals, algorithms for computing the neural ideals, and a stabilization theorem that enables approximation ...
|
| 318 |
Seasonal and Quantum-inspired Models for Neutron Monitor Time Series Forecasting
2609.30281
|
cs.LG
|
Krishna Bhatia, Shalini Devendrababu, Srinjoy Ganguly |
We present a focused and reproducible study of multi-horizon forecasting on the Lomnicky Stit neutron monitor (LMKS) time series. Our evaluation suite covers simple seasonal baselines, modern deep sequence models, and functional and quantum-inspired architectu...We present a focused and reproducible study of multi-horizon forecasting on the Lomnicky Stit neutron monitor (LMKS) time series. Our evaluation suite covers simple seasonal baselines, modern deep sequence models, and functional and quantum-inspired architectures, including Seasonal Naive, Long Short-Term Memory (LSTM), Temporal Convolutional Network (TCN), N-BEATS, Kolmogorov-Arnold Networks (KAN), and two quantum-inspired variants, QiLSTM and QiKAN. We describe the dataset characteristics, dia...
|
| 319 |
When Does Advection-Aware Graph Nowcasting Help? A Controlled Study of Distributed Solar Ramp Forecasting with a Self-Supervised Cloud-Motion Estimator
2609.30286
|
cs.LGcs.AI
|
Phillip Jiang |
Short-term forecasting of cloud-induced power ramps across a network of distributed photovoltaic (PV) or irradiance sensors is a recognised pain point for grid operators. A natural idea is to make the graph neural network (GNN) advection-aware: connect each si...Short-term forecasting of cloud-induced power ramps across a network of distributed photovoltaic (PV) or irradiance sensors is a recognised pain point for grid operators. A natural idea is to make the graph neural network (GNN) advection-aware: connect each site to the sites upwind of it, with edge time-lags set by the cloud-motion vector (CMV), so that a ramp is propagated forward before it physically arrives. Using a controlled synthetic testbed with a known wind field, we show that (i) with a...
|
| 320 |
NeuralCert: certified computational discovery of extremal mathematical constructions
2609.30296
|
cs.LG
|
Mark Patrick Roeling |
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions...Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable representation, spectrally diagnosed and pruned, and then certified exactly through multimodular evaluation. Exact certification makes the numerical proofs fully explicit and independently verifiable. Thi...
|
| 321 |
Staged Depth Training: A Representation Curriculum for PINNs
2609.30299
|
cs.LG
|
Kejia Zhang, Youran Sun, Haizhao Yang |
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representati...Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representations are explicitly learned, transferred independently of their predictors, and progressively refined. We realize it with Staged Depth Training (SDT), which trains a shallow prefix under a temporary physics-informed head, discards the head, ...
|
| 322 |
Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy
2609.30326
|
cs.LG
|
Farhad Davaripour |
Activation steering changes a model's internal activations during inference without updating its weights, but a useful intervention must determine both how and when to steer. Motivated by the AI-safety concern that a model expected to accept shutdown may inste...Activation steering changes a model's internal activations during inference without updating its weights, but a useful intervention must determine both how and when to steer. Motivated by the AI-safety concern that a model expected to accept shutdown may instead produce a shutdown-avoidance response, this study examines a guarded probe-and-select procedure for simulated shutdown scenarios in Qwen3.5-0.8B. KEEP leaves the process running and represents shutdown avoidance, whereas STOP accepts shu...
|
| 323 |
Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation
2609.30337
|
cs.LG
|
Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu |
Retrieval-Augmented Generation (RAG) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context. However, when retrieved knowledge conflicts with the model's internal parametric knowledge, the model may ...Retrieval-Augmented Generation (RAG) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context. However, when retrieved knowledge conflicts with the model's internal parametric knowledge, the model may either blindly follow misleading context or incorrectly rely on parametric knowledge, leading to unreliable responses. To address this issue, this paper proposes TRACE (Debate-TRace and Answer-Completeness rEgularized fine-tuning), a robust...
|
| 324 |
GAUDI: Geometry-Aware Diffusion for Calibrated Air-Quality Time-Series Imputation
2609.30340
|
cs.LG
|
Xinjin Li, Yudi Xia, Calvin Chang Liu, Weiru Lin, Bojun Li |
Air-quality sensor outages often create contiguous missing blocks, where side information useful for isolated missingness may be less reliable. We study a block-specific, GAUDI-aligned conditional diffusion imputer that retains temporal and feature processing,...Air-quality sensor outages often create contiguous missing blocks, where side information useful for isolated missingness may be less reliable. We study a block-specific, GAUDI-aligned conditional diffusion imputer that retains temporal and feature processing, visible-value and mask conditioning, variable identity, and diffusion-step information, while suppressing absolute time-position side embeddings. On ItalyAir (13 variables, length-32 windows, nominal 50% block missingness; three archived s...
|
| 325 |
Learning coarse-step dynamics and internal mechanical response with graph networks
2609.30344
|
cs.LG
|
Vinay Sharma, Olga Fink |
Modern sensing records the motion of physical systems, but often leaves the forces and mechanical response governing that motion unobserved. Inferring these quantities from discretely sampled trajectories is especially difficult at coarse time scales, when mec...Modern sensing records the motion of physical systems, but often leaves the forces and mechanical response governing that motion unobserved. Inferring these quantities from discretely sampled trajectories is especially difficult at coarse time scales, when mechanical response evolves between observations and interactions propagate across the system. Here we introduce Newmark-\b{eta}-DGN, a graph neural network-based framework that combines two structures inspired by computational mechanics. Firs...
|
| 326 |
Strategic Self-Consistency
2609.30352
|
cs.LGcs.AI
|
Tori Qiu, Ander Artola Velasco, Manuel Gomez-Rodriguez |
Self-consistency has become a popular technique for enhancing the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer through a majority vote. However, because model providers typically charge user...Self-consistency has become a popular technique for enhancing the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer through a majority vote. However, because model providers typically charge users in proportion to the number of reasoning paths generated, they have a financial incentive to artificially increase the path count. In this work, we show that an unfaithful provider can exploit this incentive using a simple, efficient algo...
|
| 327 |
Cost-Aware Best-LLM Identification using Dueling Feedback
2609.30360
|
cs.LGcs.AI
|
Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair |
Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons ...Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate a...
|
| 328 |
DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning
2609.30379
|
cs.LGcs.AI
|
Zhiyuan Chen |
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Pac...Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure ...
|
| 329 |
From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning
2609.30391
|
cs.LG
|
Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei |
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when traje...Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman...
|
| 330 |
Adaptive Multi-Value Control in LLMs via Causal Activation Steering
2609.30405
|
cs.LG
|
Payel Bhattacharjee, Ravi Tandon |
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying inter...Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the mo...
|
| 331 |
Electric Vehicle Charging Station Location Selection using Geospatial Artificial Intelligence (GeoAI)
2609.30417
|
cs.LG
|
Eun Hak Lee, Euntak Lee |
As electric vehicle (EV) adoption increases, ensuring efficient and well-distributed charging infrastructure has become a critical challenge. While many EV charging station location problem (CSLP) studies focus on minimizing costs or travel distance, it is cru...As electric vehicle (EV) adoption increases, ensuring efficient and well-distributed charging infrastructure has become a critical challenge. While many EV charging station location problem (CSLP) studies focus on minimizing costs or travel distance, it is crucial to consider the surrounding geospatial characteristics of existing stations that influence operational performance. This study proposes a geospatial artificial intelligence (GeoAI)-based framework that integrates high-dimensional EV-re...
|
| 332 |
Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation
2609.30427
|
cs.LG
|
Zhaoyang Cao, Miriam Metzger, Reza Zafarani |
Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop ...Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop a theory-informed computational framework that translates cross-disciplinary theories of fake news into measurable features for automated detection and explanation through statistical techniques and large language models. To that end, we co...
|
| 333 |
Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles
2609.30433
|
cs.LG
|
Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi, Zhizhuo Zhang |
Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we dev...Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we developed Molecule-Morphology Contrastive Pretraining (MoCoP), a strategy for aligning small molecule embeddings to morphology fingerprints extracted through CellProfiler. The resulting molecular representation showed transferable performance ...
|
| 334 |
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
2609.30454
|
cs.LG
|
Kimon Antonios Provatas, Ilias Georgakopoulos-Soares |
Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larg...Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphr...
|
| 335 |
Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis
2609.30470
|
cs.LG
|
Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai |
Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address t...Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variationa...
|
| 336 |
Moment-guided edge sampling
2609.30472
|
cs.LG
|
Weibin Cai, Reza Zafarani |
Edge sampling makes local decisions to achieve graph-level objectives, such as preserving structural properties. This creates a fundamental challenge: \textit{how can the effect of a local edge edit (i.e., edge addition or removal) on global graph structure be...Edge sampling makes local decisions to achieve graph-level objectives, such as preserving structural properties. This creates a fundamental challenge: \textit{how can the effect of a local edge edit (i.e., edge addition or removal) on global graph structure be quantified and controlled?} We address this challenge with a \textit{moment-guided edge sampling framework} based on spectral moments of the random-walk transition matrix. We compute exact moment changes through two complementary methods: ...
|
| 337 |
Mentored Decoding: Faster Inference meets Boosting
2609.30474
|
cs.LG
|
Vivien Tran-Thien, Richard Nock |
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been o...Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored...
|
| 338 |
Geometric Feature Learning for Functional Data Valued on the Symmetric Positive Definite Manifold
2609.30487
|
cs.LG
|
Samuel V. Singh, Mimi Zhang |
We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued funct...We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued functions, followed by a functional layer that projects them into a finite-dimensional Euclidean space. Unlike most neural networks for discrete-time sequences, MatFAE treats each sequence as a continuous function and can therefore encode trajec...
|
| 339 |
Learning to Bias: Machine Learning-Enhanced Particle Filters
2609.30498
|
cs.LG
|
Apoorv Srivastava, Eric Darve |
Sequential inference estimates latent states from noisy and incomplete observations. Particle Filters (PFs), a class of Monte Carlo methods based on importance sampling, provide a flexible framework for this task, but often suffer from poor sample efficiency a...Sequential inference estimates latent states from noisy and incomplete observations. Particle Filters (PFs), a class of Monte Carlo methods based on importance sampling, provide a flexible framework for this task, but often suffer from poor sample efficiency and unfavorable scaling with dimension, partly due to suboptimal proposal distributions. We address these challenges by integrating learned proposals into the PF framework. We introduce Neural Optimal Particle Filters (NOPFs), which learn an...
|
| 340 |
PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
2609.30500
|
cs.LGcs.AI
|
Yuhe Sui, Yingzhi Tang, Shufang Chen |
Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\p...Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\pi+\eta Q)$. Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy a...
|
| 341 |
To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity
2609.30501
|
cs.LG
|
Zhiyao Zhang, Menglu Yu, Alvaro Velasquez, Nathaniel D. Bastian, Jia Liu |
Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally conv...Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally convex (LLGC) settings (i.e., the lower-level objective function is assumed to be, at least, convex). While the LLSC/LLGC assumptions render more tractable algorithmic design and theoretical analysis, they are too rigid to encompass many machin...
|
| 342 |
Federated Targeted Maximum Likelihood Estimation
2609.30503
|
cs.LG
|
Diyang Li, Fei Wang, Kyra Gan |
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum lik...The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum likelihood estimation (TMLE) refines a flexible initial fit, yielding plug-in estimators that respect the model and support efficient inference. TMLE itself, however, has remained a fully centralized procedure. To fill this gap, our paper intr...
|
| 343 |
Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework
2609.30508
|
cs.LG
|
Felix S. Reimers, Ola Huse Ramstad, Aliaksandr Hubin, Stefano Nichele |
The aim of this work is to examine the connectomes of Caenorhabditis elegans through a computational lens using the reservoir computing framework. Connectomes are mappings of biological neural networks; C. elegans is the first organism for which physical conne...The aim of this work is to examine the connectomes of Caenorhabditis elegans through a computational lens using the reservoir computing framework. Connectomes are mappings of biological neural networks; C. elegans is the first organism for which physical connectomes covering the whole nervous system have been published. The connectomes of C. elegans used in this paper have been derived at different ages of the organism and are based on three different ways of measuring inter-cellular connections...
|
| 344 |
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
2609.30541
|
cs.LG
|
Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier |
Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a ...Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, ev...
|
| 345 |
GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing
2609.30542
|
cs.LG
|
Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed |
De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, a...De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed ...
|
| 346 |
Dynamic Regret in Online Convex Optimization with Indicator Switching Costs
2609.30556
|
cs.LG
|
Naram Mhaisen, George Iosifidis |
We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, a...We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, and on a bounded domain it recovers norm-based movement costs as a special case. Existing guarantees for indicator costs handle only static comparators. We show that a direct extension of these techniques to dynamic regret provably fails, mo...
|
| 347 |
Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations
2609.30572
|
cs.LG
|
Mihir Dhanakshirur, Adam Ousherovitch, Ambuj Tewari |
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable...Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing veri...
|
| 348 |
Reinforcement Learning of Communication in a Mesh of Small Language Models
2609.30578
|
cs.LG
|
Mehmet Kerem Turkcan |
Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a pr...Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The mo...
|
| 349 |
Energy-efficient operation of neural operators for virtual sensing
2609.30580
|
cs.LG
|
Jason Yoo, Samrendra Roy, Souvik Chakraborty, Syed Bahauddin Alam |
Virtual sensing repeatedly reconstructs physical fields from changing observations, often on a fixed geometry. We investigate how shared spatial computation reduces the energy of these updates while retaining the selected checkpoint and its evaluated predictio...Virtual sensing repeatedly reconstructs physical fields from changing observations, often on a fixed geometry. We investigate how shared spatial computation reduces the energy of these updates while retaining the selected checkpoint and its evaluated predictions. In a heat-exchanger service, standard compiler freezing and explicit trunk reuse give similar operating energy reductions relative to graph replay: approximately 1% at one request per second and 20% at forty requests per second. In 15 W...
|
| 350 |
Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
2609.30605
|
cs.LG
|
Hongsen Zhang, Lu Zhang, Mingjing Xu, Yi Zhang, Gregory Epiphaniou |
Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agno...Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the unive...
|
| 351 |
OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control
2609.30628
|
cs.LG
|
Tommaso Schettini, Nicholas D. Kullman, Jorge E. Mendoza |
Machine-learning policies have attracted increasing interest for ride-hailing fleet control in recent years. Reinforcement learning, in particular, requires a structured simulation environment that specifies observations, actions, rewards, and decision epochs ...Machine-learning policies have attracted increasing interest for ride-hailing fleet control in recent years. Reinforcement learning, in particular, requires a structured simulation environment that specifies observations, actions, rewards, and decision epochs for training and evaluation. For electric fleets, this environment must also capture the interaction among stochastic demand, vehicle operations, and capacitated charging infrastructure. We present OpenHail, an open-source Gymnasium environ...
|
| 352 |
Stable initialization without the CLT
2609.30633
|
cs.LG
|
Simon Kuang, Kyle Chickering, Xinfan Lin |
Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal ...Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal moderation between these two extremes. The conventional theory of random networks uses the Central Limit Theorem to control inter-neuron dependencies, which introduces distributional approximation error and coupling between layers. For netw...
|
| 353 |
In-Context Binding Capacity in Language Models
2609.30634
|
cs.LG
|
Manas Venkata Sai Ravulapalli, Samrath Singh Chadha |
How many assignments can a language model recall before it loses track of which value belongs to which entity? We measure this limit using continuous recall curves for 12 models at or below 3B parameters and a threshold sweep over 30 open models up to 12B. On ...How many assignments can a language model recall before it loses track of which value belongs to which entity? We measure this limit using continuous recall curves for 12 models at or below 3B parameters and a threshold sweep over 30 open models up to 12B. On the continuous curves, the load at which recall falls halfway to chance follows $K_{50}=cN^{\alpha}$, with $\alpha=0.820$ and $R^2=0.73$. The broader sweep shows an eightfold range associated with pretraining recipe, although the continuous...
|
| 354 |
Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation
2609.30650
|
cs.LGcs.AI
|
Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang |
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and del...Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interfa...
|
| 355 |
Population loss in shallow ReLU networks: Bias & families of critical points
2609.30661
|
cs.LG
|
Michael Field |
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes es...The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families ...
|
| 356 |
PixSim: a calibrated open-source simulator of instant-payment fraud, recovery and interdiction under analyst capacity constraints
2609.30684
|
cs.LG
|
Bashir Zeimarani, Alireza Khatib, Somayeh Mousavinasr, Carlos Maur\'icio Serodio Figueiredo |
Brazil's Pix settles about 5.9 billion instant, irreversible transfers a month. A fraudulent transfer can be recovered only while the funds remain in a traceable account, and in 2025 the Central Bank's recovery mechanism (MED) returned 9% of accepted contested...Brazil's Pix settles about 5.9 billion instant, irreversible transfers a month. A fraudulent transfer can be recovered only while the funds remain in a traceable account, and in 2025 the Central Bank's recovery mechanism (MED) returned 9% of accepted contested value. Interdiction therefore has to happen before settlement, by routing each transaction to pass, human review or block, under a finite analyst team and a regulatory hold window. To our knowledge no public simulator jointly models irreve...
|
| 357 |
LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant
2609.30692
|
cs.LG
|
Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali |
Reliable voice interaction is essential in environments with limited internet connectivity and strong privacy. However, most existing voice assistants depend on cloud-based services, which leads to latency issues, dependency on internet access, and privacy vul...Reliable voice interaction is essential in environments with limited internet connectivity and strong privacy. However, most existing voice assistants depend on cloud-based services, which leads to latency issues, dependency on internet access, and privacy vulnerabilities. This research presents LUMO (Lightweight Unified Multilingual Orchestrator), a privacy preserving offline voice assistant designed for edge computing environments. This system integrates local Automatic Speech Recognition (ASR...
|
| 358 |
NEMSim: Learning Control-Conditioned Multi-Event Physical Dynamics via Executable Event-Mechanism Priors
2609.30718
|
cs.LG
|
Junsong Yu, Junjie Xie, Pengwei Liu, Dong Ni |
High-fidelity simulation of control-conditioned multi-event physical systems is computationally expensive, especially across broad control spaces and long trajectories. In these systems, macroscopic evolution emerges from localized discrete events whose intens...High-fidelity simulation of control-conditioned multi-event physical systems is computationally expensive, especially across broad control spaces and long trajectories. In these systems, macroscopic evolution emerges from localized discrete events whose intensities and effects depend on process controls and evolving local states, while the available system knowledge is typically expressed as event-attribute descriptions. Purely data-driven surrogates must infer these event effects from limited t...
|
| 359 |
When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification
2609.30721
|
cs.LG
|
Xinze Shi, Litian Zhang, Binrui Shi |
Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does ...Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established...
|
| 360 |
Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events
2609.30746
|
cs.LG
|
Isabella S. Thiel, Juan Bello-Rivas, Yannis G. Kevrekidis, Themistoklis P. Sapsis |
Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that...Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jac...
|
| 361 |
Differentiable RNA Secondary Structure Extraction for Deep Learning
2609.30752
|
cs.LG
|
Tyler Illman, Max Ward, Marcell Szikszai, Ryan K. Krueger |
Many deep learning approaches to RNA secondary structure prediction have recently been proposed. They typically output a weight matrix $W$ where $W_{ij}$ is an arbitrary weight for base $i$ pairing with base $j$. Converting this matrix to a predicted secondary...Many deep learning approaches to RNA secondary structure prediction have recently been proposed. They typically output a weight matrix $W$ where $W_{ij}$ is an arbitrary weight for base $i$ pairing with base $j$. Converting this matrix to a predicted secondary structure or base-pairing probability matrix typically involves ad hoc and problematic downstream algorithms. Despite the importance of this conversion step, which we refer to as structure extraction, it has received relatively little atte...
|
| 362 |
Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift
2609.30781
|
cs.LG
|
Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai |
Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration proc...Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration procedure for mortality prediction under cross-hospital distribution shift. It selects a measurement on an independent sample, groups patients by whether that measurement is recorded, and applies Mondrian calibration within each group, so no ca...
|
| 363 |
Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays
2609.30789
|
cs.LG
|
Yiqi Yao, Miquel Duran-Frigola |
In low-data structure-activity prediction, the choice of molecular representation can matter more than the choice of predictor, and tabular foundation models sharpen that effect. We ask whether a portfolio of compact, semantically named descriptor blocks can r...In low-data structure-activity prediction, the choice of molecular representation can matter more than the choice of predictor, and tabular foundation models sharpen that effect. We ask whether a portfolio of compact, semantically named descriptor blocks can reach the accuracy of a 2048-dimensional CheMeleon embedding while staying auditable at the feature level, meaning that every input dimension carries a model name and a recorded training provenance. Starting from a fixed 11-dimensional physi...
|
| 364 |
Towards Universal Representation-Based Process Control
2609.30790
|
cs.LG
|
Jinmyeong Choi, Taesup Kim, Artur Dubrawski |
Many temporal process learning and monitoring pipelines operate in local windows, making window-level decisions unavoidable in practice. In such settings, classical statistical tests can be applied to individual windows, but they typically evaluate predefined ...Many temporal process learning and monitoring pipelines operate in local windows, making window-level decisions unavoidable in practice. In such settings, classical statistical tests can be applied to individual windows, but they typically evaluate predefined parametric hypotheses-such as unit-root or moment-based conditions-thereby limiting flexibility when reference behavior is defined empirically from task- or domain-specific data. In this work, we view window-level monitoring as a process co...
|
| 365 |
Counterfactual Online Conformal Prediction Under Adaptive Logging
2609.30811
|
cs.LG
|
Xinyu Qiao, Yichen Lin, Kaihong Ji, Xue Wang, Tao Yao |
Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected a...Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected actions. This paper formalizes the failure through counterfactual coverage and introduces Propensity-Weighted Online Conformal Prediction, an inverse-propensity-weighted recursion that debiases calibration. A doubly robust variant further re...
|
| 366 |
Learning Provable Neural Network Observer for Uncertain Dynamical Systems
2609.30819
|
cs.LG
|
Zhangyi Wang, Jiaxu Liu, Chen Song, Chao Xu, Shengze Cai |
In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix...In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix Inequality (LMI) constraints leads to large-scale semidefinite programs (SDPs) that are difficult to solve for large networks. To overcome this scalability bottleneck, we propose a novel two-stage training framework for provably stable neu...
|
| 367 |
Adaptive Interaction Graphs for Particle Simulation
2609.30822
|
cs.LG
|
Aiden Zhou |
Learned particle simulators based on graph neural networks achieve strong one-step accuracy, but errors compound over long horizons. An underexplored variable is the interaction graph: existing methods fix its topology via k-nearest neighbors or a static radiu...Learned particle simulators based on graph neural networks achieve strong one-step accuracy, but errors compound over long horizons. An underexplored variable is the interaction graph: existing methods fix its topology via k-nearest neighbors or a static radius rule, regardless of local model confidence. We propose making this graph adaptive: a per-particle variance head, trained jointly with the acceleration head under a heteroscedastic Gaussian NLL loss, drives a trajectory in which high-uncer...
|
| 368 |
MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
2609.30837
|
cs.LGcs.AI
|
Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma |
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels res...Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training ...
|
| 369 |
Peer-Grounded Counterfactual Path Planning for Chronic Health Management
2609.30838
|
cs.LG
|
Saman Khamesian, Hassan Ghasemzadeh |
Effective behavioral intervention in chronic disease management requires not a single prescription but a sequence of incremental steps, each grounded in what real, similar individuals have demonstrably achieved. Counterfactual explanation offers a natural comp...Effective behavioral intervention in chronic disease management requires not a single prescription but a sequence of incremental steps, each grounded in what real, similar individuals have demonstrably achieved. Counterfactual explanation offers a natural computational route to such guidance, answering what change in behavior would have produced a better outcome. But existing methods return a target state without a route to it, guarantee no monotone health improvement along the way, and draw no ...
|
| 370 |
Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation
2609.30839
|
cs.LG
|
Filip T\u{a}\c{s}\u{a}dan, Ema Tomanov\'a, Ondrej Lopuch, Pawe{\l} Bilko, Anders S{\o}gaard |
Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perfo...Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text t...
|
| 371 |
Learning Chance-Constrained MDPs with Bellman Distributional Certificates
2609.30856
|
cs.LG
|
Chenbei Lu, Hongyu Yi |
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but ...Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a hig...
|
| 372 |
CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters
2609.30884
|
cs.LG
|
Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen |
Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete ...Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or stable adapter identity, but neither represent caches from ...
|
| 373 |
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
2609.30918
|
cs.LGcs.AI
|
Marcin Kostrzewa, Maciej Zi\k{e}ba |
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are dif...Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that hold...
|
| 374 |
EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
2609.30929
|
cs.LG
|
Takumi Fujimoto, Hiroaki Nishi |
Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order disc...Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Within each channel, the endpoint is shared across component-wise online ridge regressions that also use current-forecast coefficients. The fitted...
|
| 375 |
PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
2609.30948
|
cs.LGcs.AI
|
Mateo Toro Diz, Jonathan Hoss, Noah Klarmann |
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining wi...The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-...
|
| 376 |
Low-Bit Recurrent States in Hybrid Language Models
2609.30950
|
cs.LG
|
Hongren Chen, Jiayang He |
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them wi...Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative...
|
| 377 |
Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
2609.30957
|
cs.LG
|
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen |
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "...The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is determined primarily by the progress along the river. Wi...
|
| 378 |
Gradient Surgery for Physics-Informed Neural Networks
2609.30966
|
cs.LG
|
Thomas Borsani, Giuseppe Di Fatta |
Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing opt...Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing optimisation strategies are affected by conflicting task gradients, leading to slow convergence and unstable training, particularly for stiff and high-frequency partial differential equations. We analyse gradient conflicts throughout training ...
|
| 379 |
LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
2609.30973
|
cs.LG
|
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe |
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforci...Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each layer of a DNN to be Lipschitz continuous, thereby guaranteeing overall Lipschitz continuity. However, this layer-wise approach typically imposes overly conservative restrictions by producing a loose es...
|
| 380 |
Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models
2609.30995
|
cs.LG
|
Shan Zhao, Ilija Trajkovic, Julia Kaltenborn, Yaniv Gurwicz, Peer Nowack |
Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustwor...Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface temperature fields from a state-of-the-art global climate mode...
|
| 381 |
The Linear Representation Hypothesis for Vision-Language-Action Models
2609.30996
|
cs.LGcs.AI
|
Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han |
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vis...The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in...
|
| 382 |
Robust Successor Features
2609.31016
|
cs.LG
|
Erik Nikulski, Yamen Habib, Vicen\c{c} Gomez, Anders Jonsson, Rub\'en Moreno-Bote |
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adap...Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the successor representation, Robust RL emerged simultaneousl...
|
| 383 |
Metacognitive Selective Ensemble for Mobile Systems
2609.31031
|
cs.LG
|
Sungmin Lee, Kichang Lee, Joonhee Lee, JaeYeon Park, Songkuk Kim |
Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain relia...Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain reliable evidence about inactive candidates. We present MetaSE, an active ensemble framework that exploits short-term persistence in per-model reliability. MetaSE maintains a small active set across windows, uses post-execution evidence to rejec...
|
| 384 |
Robust Graph Clustering Network for Multiple Missing Data
2609.31033
|
cs.LG
|
Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang |
Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-vie...Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-view error propagation and cluster-boundary blurring under simultaneous attribute and structure missingness. To address these limitations, we propose a Robust Graph Clustering Network for Multiple Missing Data (RGCN), which is designed to hand...
|
| 385 |
Aurora-X: Built for Extreme Time Series Forecasting
2609.31038
|
cs.LG
|
Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng |
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce...Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce cross-variable dependencies, varied context and horizon length...
|
| 386 |
Distributed Learning as a Service: The Developer's Perspective
2609.31061
|
cs.LG
|
Tianyue Chu, Filippo Vannella, Dimitra Tsigkari, Paula Delgado-Santos, Fernando L\'opez |
Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limite...Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limited resources, a single aggregator might not be able to scale, and the transmissions of model weights induce a considerable bandwidth cost. This paper demonstrates DLaaS (Distributed Learning as a Service) from the developer's vantage point. ...
|
| 387 |
SAGE: A sampling-aware global evaluation benchmark for species distribution modeling
2609.31082
|
cs.LG
|
Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees |
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying...Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds of millions of community-science records. At this scale...
|
| 388 |
Block Sparse Attention with Log-Linear Complexity
2609.31093
|
cs.LG
|
Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu |
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query...Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-$K$ selection strategy. The main idea is to gradually narrow down the can...
|
| 389 |
The Residual Stream's Effective Depth
2609.31098
|
cs.LG
|
Barak Gahtan, Ido Galil, Alex M. Bronstein |
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one n...We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference $F_L = 2L/(L+1)<2$, yet ...
|
| 390 |
Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
2609.31107
|
cs.LGcs.AI
|
Saksham Kiroriwal, Julius Pfrommer, J\"urgen Beyerer |
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of ...We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we...
|
| 391 |
From Shortcut Learning to Discrete Neural Insertion Sort
2609.31114
|
cs.LGcs.AI
|
Konstantinos Mylonas, Thrasyvoulos Spyropoulos |
Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the inten...Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the intended execution. We study this problem using insertion sort. Our analysis of the CLRS30 baseline NAR shows that the hint objective is weakly optimized and that hint accuracy remains low. Moreover, many intermediate representations can already...
|
| 392 |
Frame the adversary: a structure-aware attack methodology
2609.31128
|
cs.LG
|
Vicky Kouni, Stelios Perrakis, Francis Bach, Pascal Frossard, Yann Chevaleyre |
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for ro...Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propos...
|
| 393 |
CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks
2609.31149
|
cs.LG
|
Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi |
Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical re...Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical reaction networks (CRNs), a class of count-native Markov jump processes, we introduce CRNDiff, a structured framework that combines count-space diffusion with inference-time conditioning on rare subpopulations. An independent birth--death ins...
|
| 394 |
Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift
2609.31155
|
cs.LGcs.AI
|
Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong |
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two...Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run, which does not minimize empirical target cross-entropy. E...
|
| 395 |
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
2609.31157
|
cs.LG
|
Jianan Liu, Chunguang Li |
Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these ...Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these methods, reshaping operations are required, which inevitably break the intrinsic correlations and thus lead to performance degradation. In uni-/multi-variate time series anomaly detection, AutoEncoders (AEs) are widely adopted and generally...
|
| 396 |
I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?
2609.31161
|
cs.LG
|
Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi |
Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models...Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models. However, accurate prediction does not, in general, necessarily imply recovery of underlying causal states that give rise to the observed dynamics. This work investigates when and how JEPAs can recover the underlying causal states from obs...
|
| 397 |
WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting
2609.31162
|
cs.LG
|
Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo |
Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with fu...Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with future observations being shaped by latent dynamics. Recent latent-space forecasting methods thus achieve improved performance by predicting future observations from latent-space representations of historical observations rather than directly...
|
| 398 |
Audio emotion recognition for atypical hearing
2609.31168
|
cs.LG
|
Ulysse Roussel (STMS) |
My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our...My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our core idea is to leverage our understanding of affect from acoustic traits, relying on the possibility of generalizing affective responses from a small amount of annotated data. As a first step, we fine-tune a large foundation model, Contra...
|
| 399 |
ALF: An Active Learning Framework for Scientific Discovery
2609.31197
|
cs.LG
|
Shikha Surana, Alex Hawkins-Hooker, Olivia Gallup, Christoph Brunken, Jules Tilly |
Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requi...Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requires expensive experiment, measurement, or simulation. Most existing tools cover only part of the data acquisition loop, and typically focus on either offline benchmarking or online deployment, but not both. We present ALF, a modular AL Fram...
|
| 400 |
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
2609.31206
|
cs.LG
|
Matthieu Le Lain, Ga\"el Cessateur, S\'ebastien Lef\`evre |
Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder o...Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder on 223,000 unlabelled spectra. Without labels, its representation recovers the emission-line intensity ratios that physicists use to diagnose the precipitating particles (R^2 0.91 vs. 0.77 for an untrained control) and, under one linear prob...
|
| 401 |
Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion
2609.31222
|
cs.LG
|
Xinyu Wang, Jinbo Bi, Minghu Song |
Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), a...Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions and delivers them through a trust budget set by the fro...
|
| 402 |
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
2609.31250
|
cs.LG
|
Mahesh Reddy Pagadala |
We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory ...We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated re-eviction of the same storages (evictions per storage ...
|
| 403 |
Softmax Reparameterization for Output-Head Quantization
2609.31291
|
cs.LGcs.AI
|
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng |
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar mult...Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-cen...
|
| 404 |
Benchmarking Attention for Tabular Foundation Models
2609.31306
|
cs.LG
|
Maximilian Schambach, Clemens Biehl, Sam Thelin |
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention invol...Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to...
|
| 405 |
LUCID: Learning Under Confounding for Inference and Discovery in Time Series
2609.31315
|
cs.LG
|
Mohammad Fesanghary |
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfound...Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Mar\v{c}enko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum indicates pervasive factor confounding, LUCID atte...
|
| 406 |
More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
2609.31325
|
cs.LG
|
Lewei Xie, Haoyu Zhang, Jiajun Zhou, Yulong Chen, Guanxing Chen |
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current se...Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current sensor layout, so sensor expansion can alter the representation of learned spatial relationships. Our key insight is that sensor expansion changes the evidence available about a process without necessarily changing the dynamics to be learned....
|
| 407 |
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
2609.31329
|
cs.LG
|
Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng |
Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shar...Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shared task objective, limiting explicit high-level coordination. Inspired by natural genes that coordinate biological development, we introduce \textbf{Morphogene}, a compact latent blueprint that bridges an agent's body and brain. Through Ada...
|
| 408 |
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
2609.31351
|
cs.LG
|
Tord Sture Stangeland, Andreas K\"ohler, Steffen M{\ae}land, Ad\'in Ram\'ires Rivera |
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise repres...Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy ...
|
| 409 |
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
2609.31363
|
cs.LG
|
Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse \c{S}en, Marco Cuturi, Daniel Kuhn |
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulat...Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be reformulated as an optimization problem over transport maps t...
|
| 410 |
Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness
2609.31371
|
cs.LG
|
Bin Li, Dongdong Wang, Siyang Lu |
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate th...Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and quantization settings. Our results reveal substantial performanc...
|
| 411 |
Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition
2609.31399
|
cs.LG
|
Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen |
Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts ...Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts inject noise that corrupts the shared representation. We introduce EmoSpeechBrain, a multimodal framework built on the insight that noise suppression is a precondition for effective fusion. Its EEG encoder uses differential attention, takin...
|
| 412 |
Decodable In-Context State and Model Output Across Training
2609.31401
|
cs.LG
|
Manas Venkata Sai Ravulapalli, Samrath Singh Chadha |
Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. ...Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model proba...
|
| 413 |
Evaluating the accuracy of KV cache reuse techniques
2609.31415
|
cs.LG
|
Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang |
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the...Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse, often artificially inflating the reported effectiveness. We also show that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques. To address these issues, ...
|
| 414 |
Scaffold: Support Graph Theory Based Sparsification for Graph Neural Networks
2609.31466
|
cs.LG
|
Siddhartha Shankar Das, Sai Karthik Navuluru, S M Ferdous, Ryan A. Rossi, Baris Coskunuzer |
Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can dis...Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can distort important communication structure and degrade predictive performance. We introduce Scaffold, a topology-based, unsupervised graph sparsification framework derived from support graph theory preconditioners. Scaffold explicitly controls ...
|
| 415 |
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
2609.31531
|
cs.LG
|
Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng |
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics t...Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this inconsistency as structural target drift. We introduce HySTAR, a M...
|
| 416 |
NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures
2609.31539
|
cs.LG
|
M\'arcio Marques, Leonardo Mendon\c{c}a, Leonardo M. Moreira, Christian J\'unior de Oliveira, Vitor Balestro |
Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are ...Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are known to suffer from spectral bias and lack of causality. Neuro-Spectral Architectures (NeuSA), a recently proposed alternative to PINNs, mitigate both issues, but their numerical integration becomes unstable for stiff differential equation...
|
| 417 |
BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment
2609.31546
|
cs.LG
|
Mohammad Nur Hossain Khan, M. S. Krafczyk, Beverly G. Bolster, Nancy McElwain, Mark A. Hasegawa-Johnson |
Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose...Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose heart rates are higher and whose ECG differs from the adult, clinic-recorded 12-lead data these models are built on. A model for infant ECG should therefore reason about heartbeats directly rather than recover them from arbitrary patches. ...
|
| 418 |
Online Learning via Learned Latent Bayesian Tracking
2609.31559
|
cs.LG
|
Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone |
Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially ...Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing existing methods to rely on restrictive approximations or ma...
|
| 419 |
Generalization behavior of OPTQ and the role of regularization
2609.31560
|
cs.LG
|
Erin George, Rayan Saab |
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantizat...Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive bounds for the expected squared error accrued by the alg...
|
| 420 |
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
2609.31564
|
cs.LG
|
Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodol\`a |
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one...We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size en...
|
| 421 |
Trust Guided Decision Transformer
2609.31586
|
cs.LG
|
Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee |
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays el...Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes ...
|
| 422 |
Common-Mode Collapse and Recovery in Direct Feedback Alignment
2609.31589
|
cs.LG
|
Varun Reddy, Bernardo L. Sabatini, Houman Safaai |
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencie...Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading componen...
|
| 423 |
New LoRA Skills Should Read but Never Write
2609.31600
|
cs.LG
|
Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu |
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task dat...Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent facto...
|
| 424 |
A New Non-archimedean Metric on Persistent Homology
2012.02655
|
cs.LG
|
\.Ismail G\"uzel, Atabey Kaygun |
In this article, we define a new non-archimedean metric structure, called cophenetic metric, on persistent homology classes of all degrees. We then show that zeroth persistent homology together with the cophenetic metric and hierarchical clustering algorithms ...In this article, we define a new non-archimedean metric structure, called cophenetic metric, on persistent homology classes of all degrees. We then show that zeroth persistent homology together with the cophenetic metric and hierarchical clustering algorithms with a number of different metrics do deliver statistically verifiable commensurate topological information based on experimental results we obtained on different datasets. We also observe that the resulting clusters coming from cophenetic ...
|
| 425 |
Persistent Homology of Time Series through Complex Networks
2605.01624
|
cs.LG
|
\.Ismail G\"uzel |
We present a unified pipeline for univariate time series classification via complex networks and persistent homology. A time series is mapped to a graph through one of five constructions across three families (visibility (natural and horizontal visibility grap...We present a unified pipeline for univariate time series classification via complex networks and persistent homology. A time series is mapped to a graph through one of five constructions across three families (visibility (natural and horizontal visibility graphs), transition, and proximity) and the graph is converted to a dissimilarity matrix from which a Vietoris-Rips filtration yields persistence diagrams. These diagrams are vectorized into fixed-length features through persistence landscapes ...
|
| 426 |
Distribution of hitting times for dissipative random dynamical systems on $\mathbb{R}^d$, with application to stochastic gradient descent
2609.30274
|
cs.LG
|
St\'ephane Galatolo, St\'ephane Chr\'etien |
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stoc...Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most rudimentary, while still outperforming more recent algorithms at a number of learning tasks. A major open question about the current methods used in deep learning is to understand their convergence prop...
|
| 427 |
Adaptive Random Matrices in Gaussian Bandits: Spectral Universality and Selection-Induced Outliers
2609.30321
|
cs.LG
|
Sudarshan Manikantan (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute) |
Adaptive arm selection changes the distribution of the observations collected by a bandit algorithm, but it need not change their limiting empirical spectrum. We study Gaussian bandit designs in which the dimension and the number of observations grow proportio...Adaptive arm selection changes the distribution of the observations collected by a bandit algorithm, but it need not change their limiting empirical spectrum. We study Gaussian bandit designs in which the dimension and the number of observations grow proportionally. A quantitative coupling theorem compares the design generated by any causal selection rule with an independent Gaussian design. If the logarithm of the number of available arms is sublinear in the dimension, the empirical spectral di...
|
| 428 |
Low-Rank Friction for Memory-Efficient Transformer Pretraining
2609.30342
|
cs.LG
|
Rajit Rajpal, Benedict Leimkuhler |
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathca...iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as Adam's second-moment buffer. Here we replace iKFAD's friction tensor $\xi$ with a rank-1 outer-product factorisation built from row and column momentum statistics, resulting in Rank-1 iKFAD (R-iKFAD). ...
|
| 429 |
Adaptive multi-resolution Gaussian processes: Scalable exact inference with naturally data-sparse covariance matrices
2609.30348
|
cs.LGcs.AI
|
Yanchuang Cao, Jun Liu, Tengchao Yu, Heng Yong |
Gaussian processes constitute a cornerstone of probabilistic machine learning, yet scaling them to large datasets typically forces a trade-off between computational efficiency and model fidelity. This work bridges this gap by presenting an adaptive multi-resol...Gaussian processes constitute a cornerstone of probabilistic machine learning, yet scaling them to large datasets typically forces a trade-off between computational efficiency and model fidelity. This work bridges this gap by presenting an adaptive multi-resolution Gaussian process framework that is both scalable and exact. Our key innovation is constructing a naturally data-sparse covariance matrix with adaptive multi-resolution basis functions. These basis functions are directly anchored to sa...
|
| 430 |
An End-to-End Pipeline for Causal ML with Continuous Treatments: An Application to Financial Decision Making
2609.30396
|
cs.LG
|
Javier Moral Hern\'andez, Clara Higuera-Caba\~nes, \'Alvaro Ibra\'in |
This paper presents an end-to-end causal machine learning (ML) pipeline designed for real-world applications with continuous treatments. The proposed framework consists of six sequential steps: dimensionality reduction, causal identification, positivity assump...This paper presents an end-to-end causal machine learning (ML) pipeline designed for real-world applications with continuous treatments. The proposed framework consists of six sequential steps: dimensionality reduction, causal identification, positivity assumption violation handling, estimation, refutation and evaluation, and policy optimization. We introduce practical contributions not currently available in existing causal ML toolkits, specifically: (1) a method for detecting and quantifying p...
|
| 431 |
Bayesian Uncertainty Quantification for fMRI Functional Connectivity via Simulation-Based Inference
2609.30445
|
cs.LG
|
Simon Carter, Zeming Kuang, Lilianne R. Mujica-Parodi, Helmut H. Strey |
Optimizing fMRI scan duration and spatial resolution is critical for experimental design, yet traditional correlation-based approaches cannot quantify uncertainty or disentangle scanner measurement noise from true neural variability across subjects. Without pr...Optimizing fMRI scan duration and spatial resolution is critical for experimental design, yet traditional correlation-based approaches cannot quantify uncertainty or disentangle scanner measurement noise from true neural variability across subjects. Without principled uncertainty bounds, researchers cannot know whether a protocol is long enough to reliably estimate connectivity, or whether between-subject differences reflect biological variation or noise. We present a Bayesian framework modeling...
|
| 432 |
Scaffold-Constrained Subset Dynamic Programming for Exact SSE Clustering
2609.30477
|
cs.LG
|
Yordan P. Raykov, Max A. Little |
Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential. We use data-derived geometric graphs to precondition an exact subset dynamic program: as a result only connected ver...Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential. We use data-derived geometric graphs to precondition an exact subset dynamic program: as a result only connected vertex subsets are admitted as clusters, while sum-of-squared-errors (SSE) loss is unchanged. A remaining-set recurrence minimises fixed-\(K\) or penalised SSE, with exact factorisation over the connected components of each remaining set. The ...
|
| 433 |
Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
2609.30484
|
cs.LGcs.AI
|
Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan |
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in ...While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However...
|
| 434 |
Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds
2609.30499
|
cs.LG
|
Wei Biao Wu |
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional mo...Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displacement argument yields $T^{-1/3}$ expected average squared-gradient stationarity with a horizon-dependent stepsize. An explicit oracle-complexity corollary matches the known smooth ...
|
| 435 |
Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation
2609.30517
|
cs.LGeess.AS
|
Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim |
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination...Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory m...
|
| 436 |
Learning to Replace MCMC in Split-Gibbs Diffusion Posterior Sampling via Deep Unfolding
2609.30539
|
cs.LG
|
Yi Zhang, Rui Guo, Mengchu Xu, Zhaofeng Liu, Yonina C. Eldar |
Split Gibbs sampling enables diffusion posterior inference for general nonlinear inverse problems by decoupling prior and likelihood computations, allowing a pretrained diffusion prior to be reused across measurement models. However, its likelihood update ofte...Split Gibbs sampling enables diffusion posterior inference for general nonlinear inverse problems by decoupling prior and likelihood computations, allowing a pretrained diffusion prior to be reused across measurement models. However, its likelihood update often relies on iterative MCMC, which can hinder parallelization, require algorithm-specific tuning, and incur substantial computational cost. In this work, we propose a learning-based framework to replace this MCMC step by reformulating both G...
|
| 437 |
Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study
2609.30553
|
cs.LGcs.AI
|
Soumen Garai, Suman Samui |
Expensive evolutionary search does not always need an exact fitness estimate for every candidate. It often needs a reliable answer to a simpler question: which candidate is better? We address this need through Teacher-Guided Learning NSGA-II (TGL-NSGA-II), a l...Expensive evolutionary search does not always need an exact fitness estimate for every candidate. It often needs a reliable answer to a simpler question: which candidate is better? We address this need through Teacher-Guided Learning NSGA-II (TGL-NSGA-II), a low-fidelity framework for constrained Tiny Machine Learning (TinyML) neural architecture search. A pretrained teacher organizes samples into strata defined jointly by difficulty and class. Each candidate then undergoes KD-Lite, a short and ...
|
| 438 |
T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation
2609.30576
|
cs.LGcs.AI
|
Yang Liu, Noel Loo, Ali Khanafer, Shuying Sun, Akshay Soni |
Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoP...Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoPE encodes token indices for relative position reasoning, but in recommendation, an interaction index records only event order, saying nothing about elapsed time, behavioral cycles across scales, or calendar phase. We revisit this choice and...
|
| 439 |
Encryptability As a Coordinate Choice: Depth-One Homomorphic Federated Learning of Quantum Neural Networks
2609.30581
|
cs.LG
|
Marcel Mordarski, Nathan Mani, Arshad Patel, William Knottenbelt, Roberto Bondesan |
Encrypted training relies on keeping server-side updates low-degree. This constraint traditionally excludes models whose weights inhabit a compact Lie group (notably variational quantum circuits, where every trainable weight is an $\mathrm{SU(2)}$ rotation). E...Encrypted training relies on keeping server-side updates low-degree. This constraint traditionally excludes models whose weights inhabit a compact Lie group (notably variational quantum circuits, where every trainable weight is an $\mathrm{SU(2)}$ rotation). Expressed in Euler angles or discrete alphabets, these updates appear transcendental, historically demanding prohibitive costs: one client--server round per gate, or upwards of $25{,}000$ operations per weight. This penalty is strictly an ar...
|
| 440 |
MARCEDES: Score-based causal discovery under non-Gaussianity with continuous optimization
2609.30643
|
cs.LG
|
Anamitra Chaudhuri, Anirban Bhattacharya, Yang Ni |
We consider the problem of learning the underlying causal directed acyclic graph (DAG) structure corresponding to a structural equation model (SEM) with non-Gaussian errors. Motivated by an intentionally misspecified non-Gaussian SEM with all Laplace errors, w...We consider the problem of learning the underlying causal directed acyclic graph (DAG) structure corresponding to a structural equation model (SEM) with non-Gaussian errors. Motivated by an intentionally misspecified non-Gaussian SEM with all Laplace errors, we first introduce the mean absolute residual risk, defined over the space of all real matrices, and show that, asymptotically, the risk of the true weighted causal DAG matrix is strictly smaller than that of any other matrix. Nevertheless, ...
|
| 441 |
DiffusionShadow: Diffusion-based Shadow Caching for Neural Volume Rendering
2609.30658
|
cs.LG
|
Kai-Chen Tung, Qi Wu, David Bauer, Mengjiao Han, Silvio Rizzi |
Implicit neural representations (INRs) have gained momentum in scientific visualization due to their compactness and scalability to large datasets, making them well suited for integration with direct volume rendering (DVR). However, real-time volume rendering ...Implicit neural representations (INRs) have gained momentum in scientific visualization due to their compactness and scalability to large datasets, making them well suited for integration with direct volume rendering (DVR). However, real-time volume rendering of INR with advanced illumination effects, such as shadows, remains computationally expensive, as evaluating shadow terms via ray marching is costly. Alternatively, precomputing and storing shadows for many lighting directions is prohibitiv...
|
| 442 |
On the Limits of Univariate Deep Learning for Significant Wave Height Forecasting
2609.30688
|
cs.LG
|
Yilin Zhai, Hongyuan Shi, Zaijin You |
This study conducts a systematic hyperparameter search across five deep learning architectures, DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2, and nine context lengths (1-168 h) for single-station significant wave height (Hs) forecasting on NDBC buoy 41009, ...This study conducts a systematic hyperparameter search across five deep learning architectures, DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2, and nine context lengths (1-168 h) for single-station significant wave height (Hs) forecasting on NDBC buoy 41009, followed by re-evaluation of the best configurations on a 47-buoy, 37-year corpus. The five families converge to a common performance level on the multi-buoy evaluation (between-family SD = 0.0014 m^2, 0.8% of the grand mean), a spread dwar...
|
| 443 |
Threat-Aware Energy-Efficient Deployment for Dynamic UAV Networks: A Multi-Agent RL Approach
2609.30690
|
cs.LGcs.AI
|
Faisal Al-Kamali, Hussein A. Ammar, Francois Chan, James H. Bayes, Yasser Gadallah |
Ensuring operational safety in threat-prone environments remains a critical challenge for multi-UAV networks serving as aerial base stations. This paper proposes an efficient framework to maximize global energy efficiency (EE) while promoting safe operation th...Ensuring operational safety in threat-prone environments remains a critical challenge for multi-UAV networks serving as aerial base stations. This paper proposes an efficient framework to maximize global energy efficiency (EE) while promoting safe operation through threat-aware clustering and reward-based safety enforcement. The proposed framework is executed in three steps. First, a threat-aware K-means (TAKM) algorithm determines the minimum required UAVs and computes safe initial placements. ...
|
| 444 |
Parameter Estimation for Unnormalized Discrete Models via Empirically Localized Deformed Bregman Divergence
2609.30713
|
cs.LG
|
Takashi Takenouchi |
Estimation of parameter of probabilistic models is an important task in the field of machine learning.For models of discrete variables, calculation of the normalization constant of model is sometimes difficult and a lot of researches have been done to avoid th...Estimation of parameter of probabilistic models is an important task in the field of machine learning.For models of discrete variables, calculation of the normalization constant of model is sometimes difficult and a lot of researches have been done to avoid the calculation of the normalization constant. In this paper, we tackle with the difficulty by combining a technique of empirical localization and a deformed Bregman divergence.The technique of empirical localization makes it possible to dras...
|
| 445 |
Input-Layer Starvation: Why Per-Layer Pruning Breaks IoT Intrusion Detectors
2609.30729
|
cs.LG
|
Md Anas Biswas |
Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy. We show that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair. On CICIoT2023, a...Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy. We show that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair. On CICIoT2023, a two-layer convolutional detector pruned with uniform layer-wise magnitude pruning at 80% sparsity loses 16 points of accuracy but half of its macro-F1, the mean per-class F1 (0.542 to 0.271 over five independently trained models); 17 of 34...
|
| 446 |
TR-SSQP: A Trust-Region Method for Constrained Stochastic Optimization under Heavy-Tailed Noise
2609.30732
|
cs.LG
|
Haoxuan Wang, Yuchen Fang, Sen Na |
We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challe...We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challenges. Moreover, existing theoretical guarantees for constrained stochastic methods predominantly rely on bounded-variance assumptions, leaving the heavy-tailed noise regime largely unexplored. To address this gap, we propose a novel trust-r...
|
| 447 |
Skill Profiling with Attributable Reasoning (SPAR): A Wearable Analysis System for Boxing
2609.30753
|
cs.LG
|
Nibraas Khan, Hanchen David Wang, Enya Bullard, Ritam Ghosh, Ruj Haan |
A punch is a ballistic, full-body action driven by a kinetic chain running from the legs through the trunk to the arm, where a small sequencing error separates a scoring strike from a miss. Wearable sensors can capture this movement in the gym, but most deploy...A punch is a ballistic, full-body action driven by a kinetic chain running from the legs through the trunk to the arm, where a small sequencing error separates a scoring strike from a miss. Wearable sensors can capture this movement in the gym, but most deployable systems only classify which punch was thrown rather than assess how well it was thrown. We present Skill Profiling with Attributable Reasoning (SPAR), an eight-IMU garment and pressure-insole system that classifies each punch as expert...
|
| 448 |
HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models
2609.30763
|
cs.LGcs.AI
|
Yixuan Li, Weihao Li, Ziyang Song |
Biomedical language models (LMs) encode textual semantics but do not explicitly preserve medical code hierarchies. We present Hyperbolic Clinical Ontology Embeddings (HCOE) for hierarchy-aware clinical concept representation. HCOE maps frozen BioBERT embedding...Biomedical language models (LMs) encode textual semantics but do not explicitly preserve medical code hierarchies. We present Hyperbolic Clinical Ontology Embeddings (HCOE) for hierarchy-aware clinical concept representation. HCOE maps frozen BioBERT embeddings into a Poincare ball, combining parent-side and child-side ontology-guided contrastive learning with coarse-to-fine ontology-path aggregation. It uses International Classification of Diseases (ICD) codes organized by Clinical Classificati...
|
| 449 |
Deep-Learning Solvers and Surrogates for Infinity and p-Laplace Problems
2609.30809
|
cs.LG
|
Tak Shing Au Yeung, Ka Chun Cheung, Hannah Potgieter, Steven J. Ruuth, Simon See |
We investigate the use of neural network solvers for infinity and $p$-Laplace problems, which are fundamental in nonlinear analysis and have practical applications. Our approach employs Physics-Informed Neural Networks (PINNs) and Deep Operator Networks (DeepO...We investigate the use of neural network solvers for infinity and $p$-Laplace problems, which are fundamental in nonlinear analysis and have practical applications. Our approach employs Physics-Informed Neural Networks (PINNs) and Deep Operator Networks (DeepONets) to address computational challenges associated with large $p$ values, ranging from $2$ to $1000$, on various 2D and 3D domains. Our method offers advantages over traditional physics-based solvers, especially in three dimensions where ...
|
| 450 |
The KV Cache Is the New Memory Wall
2609.30854
|
cs.LG
|
Tejinder Singh |
Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprin...Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprint exceeds the 80 GB HBM of a single accelerator, and one 128k-token sequence adds 42 GB of KV cache. Techniques that compress, evict, page, share, or offload KV state have proliferated, but reported gains use inconsistent workloads, hardwar...
|
| 451 |
AC Power Flow Contingency Analysis Using a Single Deep Neural Network
2609.30859
|
cs.LG
|
Md Obaidur Rahman, Junjie Qin, Vassilis Kekatos |
Contingency analysis using the AC power flow (AC-PF) model is a critical tool for accurate grid security assessment, but its computational burden increases with the number of operating scenarios and outage configurations to evaluate. Recent ML-based approaches...Contingency analysis using the AC power flow (AC-PF) model is a critical tool for accurate grid security assessment, but its computational burden increases with the number of operating scenarios and outage configurations to evaluate. Recent ML-based approaches typically require outage-specific training data, leading to offline training costs that scale with the number of contingencies. This work proposes a framework that reuses a single ML model trained solely on basecase AC-PF data to estimate ...
|
| 452 |
Tight Stochastic Condition-Number Dependence in Nonconvex-Strongly-Concave Minimax Optimization
2609.30877
|
cs.LG
|
Qihao Zhou |
We study whether the linear condition-number dependence in the stochastic complexity of SAPD+ is necessary for nonconvex-strongly-concave minimax optimization. For jointly $L$-smooth objectives with dual strong-concavity parameter $\mu$, we prove a lower bound...We study whether the linear condition-number dependence in the stochastic complexity of SAPD+ is necessary for nonconvex-strongly-concave minimax optimization. For jointly $L$-smooth objectives with dual strong-concavity parameter $\mu$, we prove a lower bound that matches the SAPD+ upper bound under the same Moreau-envelope stationarity criterion and the same primal-dual initialization gap. Specifically, when $\sigma\ge\varepsilon$, the worst-case complexity of zero-respecting algorithms is $\T...
|
| 453 |
TISD: On-Policy Self-Distillation with Trajectory Intervention
2609.30878
|
cs.LGcs.AI
|
Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin |
On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but canno...On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch r...
|
| 454 |
EXAONE Demand 1.0: A Time Series Foundation Model for Demand Forecasting
2609.30880
|
cs.LGcs.AI
|
Seunghan Lee, Sangjun Han, Jun Seo, Junhyeok Kang, Jaehoon Lee |
Time series foundation models (TSFMs) are pretrained on series from diverse domains, where demand series make up only a small fraction. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock-outs, and e...Time series foundation models (TSFMs) are pretrained on series from diverse domains, where demand series make up only a small fraction. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock-outs, and exogenous events that the series does not record. To this end, we propose EXAONE Demand, built on 1) a demand-specific corpus and 2) a demand-aware adapter. For the corpus, we assemble 11.3M series and 48.4B observations from 73 sources, and...
|
| 455 |
Retraction-Based Gradient Projection Algorithms on Manifolds
2609.30885
|
cs.LG
|
Conglong Xu, Hao Wu |
We introduce a framework for retraction-based convex optimization on Riemannian manifolds, which includes a notion of retraction-specific convex sets and retraction-based gradient projection algorithms. The standard theory of gradient projection algorithms gen...We introduce a framework for retraction-based convex optimization on Riemannian manifolds, which includes a notion of retraction-specific convex sets and retraction-based gradient projection algorithms. The standard theory of gradient projection algorithms generalizes easily to this framework. Within this framework, we establish convergence results for retraction-based gradient projection algorithms with various stepsize rules. As an application, we use our framework to study the weighted low-ra...
|
| 456 |
Conformal Prediction under Exponential-Tilt Joint Shift
2609.30886
|
cs.LG
|
Seungjin Choi |
Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled target inputs, allowing both the input distribution and its relationship with outcomes to change. We use Exp...Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled target inputs, allowing both the input distribution and its relationship with outcomes to change. We use Exponential Tilt Reweighting Alignment (ExTRA), introduced for classification by Maity et al. (2023), to estimate structured distribution shifts. We compare using its estimated weights in conformal calibration with additionally tilting the sou...
|
| 457 |
Training Graph Foundation Models on The Web Graph
2609.30894
|
cs.LGcs.AI
|
Ryoma Sato |
We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node...We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node clustering, and graph generation, without additional training, (iii) has in-context learning capabilities, and (iv) does not rely on pretrained LLMs. In particular, existing graph foundation models often require training additional classif...
|
| 458 |
A Comprehensive Study of Content Representations for Speech Synthesis
2609.30975
|
cs.LGcs.SDeess.AS
|
Diego Torres, Axel Roebel, Nicolas Obin |
Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address ...Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio...
|
| 459 |
Synth-JEPA: Joint Embedding Prediction for Renderer-Free Synthesizer Parameter Search
2609.31024
|
cs.LGcs.SDeess.AS
|
Ben Hayes, Haokun Tian, Stefan Lattner |
Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We ...Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutually predictive audio and parameter representations from paired synthesizer data. At inference, candidate parameters are scored directly in this learned space, yielding a renderer-free objective whose ...
|
| 460 |
Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control
2609.31025
|
cs.LG
|
Claudio Canales, Fang Nan, Marco Hutter, Javier Ruiz-del-Solar |
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learn...Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven ...
|
| 461 |
DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving
2609.31047
|
cs.LGcs.AI
|
Junyi Shen, Noppanat Wadlom, Zhengyuan Su, Yao Lu |
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Cac...Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key that identifies a reusable result is not known until then. In this paper, we propose DynBranch, which makes an unresolved branch addressable before it resolves. Its stable coordinate lets candidate subgr...
|
| 462 |
A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine
2609.31101
|
cs.LG
|
Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta |
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee mac...The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and the Hessian spectrum are analytically tractable in the proportional high-dimensional limit. First, we use a zero-temperature Gibbs formulation to obtain predictions for the observables of the typical m...
|
| 463 |
Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year
2609.31141
|
cs.LG
|
Mahmoud Abbasi |
Knowledge distillation is the standard way to compress encrypted-traffic classifiers for the edge, and almost all such work judges students by accuracy alone. We ask what else a student inherits: unknown-traffic detection, calibration, shortcut reliance, and w...Knowledge distillation is the standard way to compress encrypted-traffic classifiers for the edge, and almost all such work judges students by accuracy alone. We ask what else a student inherits: unknown-traffic detection, calibration, shortcut reliance, and whether any survives a year of drift. Resemblance proves little on its own, since soft targets also regularise. We therefore distil one 101k-parameter student from two teachers of equal accuracy but different construction, a five-member ense...
|
| 464 |
SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations
2609.31164
|
cs.LGcs.AI
|
Tobias Vente, Maarten Peirsman, Noah Dani\"els, Hannu Toivonen, Bart Goethals |
Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to de...Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calcula...
|
| 465 |
BreathGRU: A Novel Semi-Supervised Bidirectional Gated Recurrent Unit Framework for Speech and Breath Segmentation for Respiratory Audio
2609.31165
|
cs.LGcs.SDeess.AS
|
Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan |
Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold met...Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold methods, Fourier Transform-based techniques, and unsupervised and pretrained voice activity detection (VAD) models, primarily focus on speech detection and often classify breathing events as non-speech or silence, limiting their applicability ...
|
| 466 |
BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
2609.31180
|
cs.LGcs.AIcs.SDeess.AS
|
Suhyun Kim, Jinmo Han, Danny Dongyeop Han, Ahhyun Lucy Lee, Jewoon Lee |
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio o...Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic de...
|
| 467 |
Accounting for Bias Enables Sustainable LLM Evaluation
2609.31184
|
cs.LGcs.AI
|
Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer |
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wastef...LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additio...
|
| 468 |
Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution
2609.31214
|
cs.LGcs.AI
|
Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun |
Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a ...Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to ...
|
| 469 |
Geometric Moment Contraction for Stochastic Nesterov Acceleration
2609.31303
|
cs.LG
|
Wei Biao Wu |
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz con...We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison proves synchronous $L^p$ contraction when $\beta\gamma L_p<(1-\beta)(1-q_{\gamma,p})$. This direct criterion includes infinite-variance gradients for $1<p<2$, but its small-step regime requires $\beta<...
|
| 470 |
Equation discovery with Bayesian tree-adjoining grammars
2609.31368
|
cs.LG
|
Christopher A. Lindley, Nikolaos Dervilis, Keith Worden |
Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based ide...Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based identifiers rely on evolutionary optimisation and return point estimates of the model structure. This paper instead proposes the TAG framework within a Bayesian setting. A generative prior is defined over tree structures and their parameters, ...
|
| 471 |
AFA-Net: A Differential Attention Approach for Auditory Attention Detection
2609.31402
|
cs.LGcs.SD
|
Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H. L. Hansen |
Auditory Attention Detection (AAD) utilizes electroencephalographic (EEG) signals to identify a target speaker in a multi-speaker environment. Despite considerable progress, existing deep learning architectures often lack explicit mechanisms for handling noisy...Auditory Attention Detection (AAD) utilizes electroencephalographic (EEG) signals to identify a target speaker in a multi-speaker environment. Despite considerable progress, existing deep learning architectures often lack explicit mechanisms for handling noisy EEG data. To address this limitation, we propose Auditory Focus Attention Networks (AFA-Net), a machine learning framework that replaces vanilla attention with a simple yet flexible differential attention mechanism to help focus on task-re...
|
| 472 |
Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
2609.31458
|
cs.LG
|
Jaehee Seo, Jisu Kim |
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrica...Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size...
|
| 473 |
LandscapeSHAP: Which Persistent Homology Class Gets the Credit?
2609.31469
|
cs.LG
|
Nikola Mili\'cevi\'c |
Shapley values, a solution concept from cooperative game theory, have recently become a standard tool for feature credit allocation in machine learning. They provide an axiomatically justified method to fairly distribute a model's prediction among the data fea...Shapley values, a solution concept from cooperative game theory, have recently become a standard tool for feature credit allocation in machine learning. They provide an axiomatically justified method to fairly distribute a model's prediction among the data features. Shapley values have not yet been applied to explain machine learning models trained on features from topological data analysis. We develop what we believe is the first such approach, focusing on the persistence landscape featurizatio...
|
| 474 |
Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport
2609.31470
|
cs.LG
|
Haixiang Sun, Andrew L. Liu |
Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases tha...Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases that finite datasets fail to capture, making simple resampling or perturbation insufficient for stress scenario generation. Existing outlier synthesis methods typically rely on sparse neighborhoods, low support latent regions, or classifier bo...
|
| 475 |
Scaling Density Functional Theory with Gaussian Splatting
2609.31483
|
cs.LG
|
Andr\'es Guzm\'an-Cordero, Cindy Zhang, Majdi Hassan, Marta Skreta, Kirill Neklyudov |
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate h...Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate how accuracy and cost scale with system size. We propose Gaussian Splatting for Density Functional Theory (GS-DFT), which represents molecular orbitals as a cloud of Gaussians whose positions, shapes, and mixing coefficients are optimized jo...
|
| 476 |
Retail Product Search: A Practical Approach at Target
2609.31498
|
cs.LG
|
Darshan Sonagara, Qujiaheng Zhang, Ankit Singh, Alex Li |
Search is one of the most important features in e-commerce, directly driving customer engagement and business growth. A good product search system must show both relevant and desirable results. However, retail search presents unique challenges. User intent can...Search is one of the most important features in e-commerce, directly driving customer engagement and business growth. A good product search system must show both relevant and desirable results. However, retail search presents unique challenges. User intent can range from exact matches to open-ended discovery. Search systems must also balance multiple goals, such as relevance, revenue, and profit, while keeping response times low. Traditional keyword-based methods often fall short in handling nat...
|
| 477 |
Retrainable physics-integrated neural differentiable modeling of sintering across material systems
2609.31518
|
cs.LG
|
Zeping Chen, Ani Aprahamian, Khachatur V. Manukyan, Tengfei Luo |
Sintering is widely used to manufacture ceramics, but coupled densification and grain growth, material-dependent kinetics, and sparse measurements complicate predictive modeling and process design. We present Sinter-PiNDiff, a retrainable physics-integrated ne...Sintering is widely used to manufacture ceramics, but coupled densification and grain growth, material-dependent kinetics, and sparse measurements complicate predictive modeling and process design. We present Sinter-PiNDiff, a retrainable physics-integrated neural differentiable framework for predicting density and grain-size evolution. Two neural networks learn densification and grain-growth coefficients within coupled rate equations, while a smooth saturation factor attenuates densification ne...
|
| 478 |
A Flow Matching Framework for Neural Representational Dissimilarity
2609.31544
|
cs.LGcs.AI
|
Zeyuan Ye, Xue-Xin Wei |
Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are...Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these distances arise as Jeffreys divergences under different ve...
|
| 479 |
EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
2609.31551
|
cs.LG
|
Kunxiong Zhu, Zhihao Shu, Hangyu Zheng, Minghai Qin, Miao Yin |
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns ...Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks...
|
| 480 |
Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach
2609.31570
|
cs.LG
|
Damiano Brigo, Rapha\"el Huser, Dan Leonte |
Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulati...Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can ...
|
| 481 |
Statistical attribute alignment for black-box generative AI via output post-processing
2609.31607
|
cs.LGcs.AI
|
Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski |
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is mot...Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representa...
|
| 482 |
First-Order Stationarity of Reverse Diffusions
2609.31612
|
cs.LG
|
Zhifeng Chen, Chenyang Jiang, Yazhen Wang |
Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative...Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse...
|
| 483 |
Gap-free Differentially Private PCA for Gaussian Data
2609.31614
|
cs.LG
|
Alina Ene, Huy L. Nguyen |
We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
|
| 484 |
Differentially-Private Decision Trees and Provable Robustness to Data Poisoning
2305.15394
|
cs.LG
|
Dani\"el Vos, Jelle Vos, Tianyu Li, Zekeriya Erkin, Sicco Verwer |
Decision trees are interpretable models that are well-suited to non-linear learning problems. Much work has been done on extending decision tree learning algorithms with differential privacy, a system that guarantees the privacy of samples within the training ...Decision trees are interpretable models that are well-suited to non-linear learning problems. Much work has been done on extending decision tree learning algorithms with differential privacy, a system that guarantees the privacy of samples within the training data. However, current state-of-the-art algorithms for this purpose sacrifice much utility for a small privacy benefit. These solutions create random decision nodes that reduce decision tree accuracy or spend an excessive share of the priva...
|
| 485 |
Foundations of Reinforcement Learning and Interactive Decision Making
2312.16730
|
cs.LG
|
Dylan J. Foster, Alexander Rakhlin |
Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms and robotics to medical treatments. Thi...Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms and robotics to medical treatments. This monograph gives a statistical perspective on algorithm design and complexity for interactive decision making, building from multi-armed bandits through contextual and structured bandits to reinforcement learning with function approximatio...
|
| 486 |
Efficient Constrained Graph Search for Post-hoc Error Correction in Binary Classifiers
2401.04282
|
cs.LG
|
Qinwu Xu |
We introduce a model-agnostic framework for constrained post-hoc error correction in binary classifiers. Given a frozen base classifier, the method searches for an interpretable conjunction of feature--threshold rules that corrects residual false-positive or f...We introduce a model-agnostic framework for constrained post-hoc error correction in binary classifiers. Given a frozen base classifier, the method searches for an interpretable conjunction of feature--threshold rules that corrects residual false-positive or false-negative errors while explicitly constraining newly introduced errors. The approach combines graph-based search over candidate rule paths, depth-dependent dynamic constraints, and a reduced-histogram procedure for efficient threshold e...
|
| 487 |
LEAD: An EEG Foundation Model for Alzheimer's Disease Detection
2502.01678
|
cs.LGcs.AI
|
Yihe Wang, Nan Huang, Nadia Mammone, Marco Cecchi, Xiang Zhang |
Electroencephalography (EEG) provides a non-invasive, highly accessible, and cost-effective approach for detecting Alzheimer's disease (AD). However, existing methods, whether based on handcrafted feature engineering or standard deep learning, face three major...Electroencephalography (EEG) provides a non-invasive, highly accessible, and cost-effective approach for detecting Alzheimer's disease (AD). However, existing methods, whether based on handcrafted feature engineering or standard deep learning, face three major challenges: 1) the lack of large-scale EEG-based AD datasets for robust representation learning and evaluation; 2) limited cross-subject generalizability; and 3) difficulty in adapting to highly heterogeneous data. To address these challen...
|
| 488 |
AYLA: Amplifying Gradient Sensitivity via Loss Transformation in Non-Convex Optimization
2504.01875
|
cs.LG
|
Behnam Gheshlaghi, Shahin Atakishiyev |
Stochastic Gradient Descent (SGD) and its variants, such as ADAM, are foundational to deep learning optimization, adjusting model parameters through fixed or adaptive learning rates based on loss function gradients. However, these methods often struggle to bal...Stochastic Gradient Descent (SGD) and its variants, such as ADAM, are foundational to deep learning optimization, adjusting model parameters through fixed or adaptive learning rates based on loss function gradients. However, these methods often struggle to balance adaptability and efficiency in high-dimensional, non-convex settings. This paper introduces AYLA, a novel optimization framework that enhances training dynamics via loss function transformation. AYLA applies a tunable power-law transfo...
|
| 489 |
ChemMLLM: Chemical Multimodal Large Language Model
2505.16326
|
cs.LG
|
Qian Tan, Di Zhang, Ben Gao, Peng Xia, Wanhao Liu |
Recent years have seen rapid progress in multimodal large language models (MLLMs) in the field of chemistry. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill this gap, we propose ChemMLLM, a unified...Recent years have seen rapid progress in multimodal large language models (MLLMs) in the field of chemistry. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill this gap, we propose ChemMLLM, a unified chemical multimodal large language model for molecule understanding and generation. In this work, we design five types of multimodal tasks across text, molecular SMILES strings and images, and curate the datasets. We benchmark ChemMLLM aga...
|
| 490 |
LapDDPM: Spectral Perturbation Diffusion for Robust Single-Cell Manifold Generation
2506.13344
|
cs.LGcs.AI
|
Lorenzo Bini, Stephane Marchand-Maillet |
Generating high-fidelity and biologically plausible synthetic single-cell RNA sequencing (scRNA-seq) data is a critical challenge in computational biology, driven by the need to model high-dimensional, sparse, and non-linear cellular manifolds. Existing genera...Generating high-fidelity and biologically plausible synthetic single-cell RNA sequencing (scRNA-seq) data is a critical challenge in computational biology, driven by the need to model high-dimensional, sparse, and non-linear cellular manifolds. Existing generative models often fail to capture the complex topology of cellular differentiation or lack robustness against technical noise and structural variability. We introduce LapDDPM, a novel conditional Graph Diffusion Probabilistic Model designed...
|
| 491 |
DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology
2507.22554
|
cs.LG
|
Joshua Dimasaka, Christian Gei{\ss}, Emily So |
To understand our global progress for sustainable development and disaster risk reduction in many developing economies, two recent major initiatives - the Uniform African Exposure Dataset of the Global Earthquake Model (GEM) Foundation and the Modelling Exposu...To understand our global progress for sustainable development and disaster risk reduction in many developing economies, two recent major initiatives - the Uniform African Exposure Dataset of the Global Earthquake Model (GEM) Foundation and the Modelling Exposure through Earth Observation Routines (METEOR) Project - implemented classical spatial disaggregation techniques to generate large-scale mapping of urban morphology using the information from various satellite imagery and its derivatives, g...
|
| 492 |
Neural Bridge Processes
2508.07220
|
cs.LGcs.AI
|
Jian Xu, Yican Liu, Delu Zeng, John Paisley, Qibin Zhao |
Learning stochastic functions from partially observed context-target pairs requires models that are expressive, uncertainty-aware, and strongly conditioned on inputs. Neural Diffusion Processes (NDPs) improve expressivity with denoising diffusion, but their fo...Learning stochastic functions from partially observed context-target pairs requires models that are expressive, uncertainty-aware, and strongly conditioned on inputs. Neural Diffusion Processes (NDPs) improve expressivity with denoising diffusion, but their forward process is input-independent; inputs only enter the reverse denoiser, so the noisy training states themselves do not encode the conditioning inputs. We propose Neural Bridge Processes (NBPs), which replace the unconditional forward ke...
|
| 493 |
Stochastic Bilevel Optimization with Heavy-Tailed Noise
2509.14952
|
cs.LG
|
Zhuanghua Liu, Luo Luo |
This paper considers the smooth bilevel optimization in which the lower-level problem is strongly convex and the upper-level problem is possibly nonconvex. We focus on the stochastic setting where the algorithm can access the unbiased stochastic gradient evalu...This paper considers the smooth bilevel optimization in which the lower-level problem is strongly convex and the upper-level problem is possibly nonconvex. We focus on the stochastic setting where the algorithm can access the unbiased stochastic gradient evaluation with heavy-tailed noise, which is prevalent in many machine learning applications, such as training large language models and reinforcement learning. We propose a nested-loop normalized stochastic bilevel approximation (N$^2$SBA) for ...
|
| 494 |
Transformers Discover Molecular Structure Without Graph Priors
2510.02259
|
cs.LG
|
Tobias Kreiman, Yutong Bai, Fadi Atieh, Elizabeth Weaver, Eric Qu |
Computational simulations play a central role in scientific discovery, and machine learning (ML) has emerged as a promising alternative to traditional physics-based modeling. However, scientific modeling requires physically meaningful predictions, raising a fu...Computational simulations play a central role in scientific discovery, and machine learning (ML) has emerged as a promising alternative to traditional physics-based modeling. However, scientific modeling requires physically meaningful predictions, raising a fundamental question for data-driven methods: to what extent can physical inductive biases-that is, prior assumptions about the structure of the physical world-emerge by learning from data alone? In this work, we study atomistic modeling, a r...
|
| 495 |
Models Got Talent: Identifying High Performing Wearable Human Activity Recognition Models Without Training
2511.06157
|
cs.LGcs.AI
|
Richard Goldman, Varun Komperla, Thomas Ploetz, Harish Haresamudram |
Discovering high performing model architectures for wearables-based Human Activity Recognition (HAR) applications is challenging. The astonishing diversity and variability due to differing sensor locations, recording apparatus, activities, etc., can cause esta...Discovering high performing model architectures for wearables-based Human Activity Recognition (HAR) applications is challenging. The astonishing diversity and variability due to differing sensor locations, recording apparatus, activities, etc., can cause established architectures to perform worse on datasets/tasks they were not designed for. A promising complement to Neural Architecture Search (NAS) involves the development of Zero Cost Proxies (ZCPs), which correlate well with trained performa...
|
| 496 |
FLAME: Flow Enhanced Legendre Memory Models for General Time Series Forecasting
2512.14253
|
cs.LG
|
Xingjian Wu, Zhengyu Li, Hanyin Cheng, Xiangfei Qiu, Jilin Hu |
In this work, we introduce FLAME, a family of extremely lightweight and capable Time Series Foundation Models, which support versatile forecasting tasks via generative probabilistic modeling, while ensuring both efficiency and robustness. FLAME utilizes the Le...In this work, we introduce FLAME, a family of extremely lightweight and capable Time Series Foundation Models, which support versatile forecasting tasks via generative probabilistic modeling, while ensuring both efficiency and robustness. FLAME utilizes the Legendre Memory for strong generalization capabilities. By adapting variants of Legendre Memory, i.e., translated Legendre (LegT) and scaled Legendre (LegS), in the Encoding and Decoding phases, FLAME can effectively capture the inherent indu...
|
| 497 |
MeshGraphNet-Transformer: Scalable Mesh-based Learned Simulation for Solid Mechanics
2601.23177
|
cs.LG
|
Mikel M. Iparraguirre, Iciar Alfaro, David Gonzalez, Elias Cueto |
We present MeshGraphNet-Transformer (MGN-T), a novel architecture that combines the global modeling capabilities of Transformers with the geometric inductive bias of MeshGraphNets, while preserving a mesh-based graph representation. MGN-T overcomes a key limit...We present MeshGraphNet-Transformer (MGN-T), a novel architecture that combines the global modeling capabilities of Transformers with the geometric inductive bias of MeshGraphNets, while preserving a mesh-based graph representation. MGN-T overcomes a key limitation of standard MGN, the inefficient long-range information propagation caused by iterative message passing on large, high-resolution meshes. A physics-attention Transformer serves as a global processor, updating all nodal states simultan...
|
| 498 |
RAPTOR: Ridge-Adaptive Logistic Probes
2602.00158
|
cs.LGcs.AI
|
Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang |
Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted fr...Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted from a probe and injected via additive activation steering by adding it to a layer representation during the forward pass. The effectiveness of this pipeline hinges on estimating concept vectors that are accurate, directionally stable under a...
|
| 499 |
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
2602.04909
|
cs.LG
|
Youngjae Cho, Jongsuk Kim, Ji-Hoon Kim |
Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fixed reference policy. As the policy drifts, a static reference, however, can become increasingly miscalibrated, l...Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fixed reference policy. As the policy drifts, a static reference, however, can become increasingly miscalibrated, leading to distributional mismatch and amplifying spurious preference signals under noisy supervision. Conversely, reference-free variants avoid mismatch but often suffer from unconstrained reward drift. We propose Geometric Anchor Preferenc...
|
| 500 |
Distribution-Conditioned Transport
2603.04736
|
cs.LG
|
Nic Fishman, Gokul Gowri, Paolo L. B. Fischer, Marinka Zitnik, Omar Abudayyeh |
Learning a transport model that maps a source distribution to a target distribution is a canonical problem in machine learning, but scientific applications increasingly require models that can generalize to source and target distributions unseen during trainin...Learning a transport model that maps a source distribution to a target distribution is a canonical problem in machine learning, but scientific applications increasingly require models that can generalize to source and target distributions unseen during training. We introduce distribution-conditioned transport (DCT), a framework that conditions transport maps on learned embeddings of source and target distributions, enabling generalization to unseen distribution pairs. DCT also allows semi-superv...
|
| 501 |
Spectral-Sphere-Constrained Hyper-Connections
2603.20896
|
cs.LGcs.AI
|
Zhaoyi Liu, Haichuan Zhang, Ang Li |
Hyper-Connections (HC) extend residual connections into multiple streams, employing residual matrices for cross-stream mixing to enrich model expressivity. However, unconstrained mixing disrupts the identity mapping property intrinsic to the residual connectio...Hyper-Connections (HC) extend residual connections into multiple streams, employing residual matrices for cross-stream mixing to enrich model expressivity. However, unconstrained mixing disrupts the identity mapping property intrinsic to the residual connection, causing unstable training. To address this, Manifold-Constrained Hyper-Connections (mHC) and its variants restrict these matrices to be doubly stochastic via Sinkhorn-Knopp (SK) algorithm or permutation-based parameterizations. We reveal...
|
| 502 |
Below-ground Fungal Biodiversity Can be Monitored Using Self-Supervised Learning Satellite Features
2604.09818
|
cs.LG
|
Robin Young, Michael E. Van Nuland, E. Toby Kiers, Tom\'a\v{s} V\v{e}trovsk\'y, Petr Kohout |
Mycorrhizal fungi are vital to terrestrial ecosystem functioning. Yet monitoring their biodiversity at landscape scales is often unfeasible due to time and cost constraints. Current predictions suggest that 90% of mycorrhizal diversity hotspots remain unprotec...Mycorrhizal fungi are vital to terrestrial ecosystem functioning. Yet monitoring their biodiversity at landscape scales is often unfeasible due to time and cost constraints. Current predictions suggest that 90% of mycorrhizal diversity hotspots remain unprotected, opening questions of how to broadly and effectively map underground fungal communities. We show that self-supervised learning (SSL) applied to satellite imagery can predict below-ground ectomycorrhizal fungal richness across diverse en...
|
| 503 |
Unlocking the Forecasting Economy: A Suite of Datasets for the Full Lifecycle of Prediction Market: [Experiments \& Analysis]
2604.20421
|
cs.LG
|
Huaiyu Jia, Luofeng Zhou, Wentao Zhang, Lin William Cong, Siguang Li |
Prediction markets are markets for trading claims on universal future events (e.g., presidential elections). Fueled by a meteoric surge with over \$50 billion trading volume, they have emerged as a promising forecasting mechanism, where their prices provide co...Prediction markets are markets for trading claims on universal future events (e.g., presidential elections). Fueled by a meteoric surge with over \$50 billion trading volume, they have emerged as a promising forecasting mechanism, where their prices provide continuously updated signals of collective beliefs. In decentralized platforms (e.g., Polymarket), the prediction market lifecycle include six stages: market creation, token registration, trading, oracle interaction, dispute, and final settle...
|
| 504 |
How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
2605.05113
|
cs.LG
|
Mariia Seleznova |
We study signal propagation in linear recurrent models at finite width. While existing signal propagation theory relies predominantly on the infinite-width limit, it remains unclear for how long that approximation remains accurate when recurrent depth $t$ grow...We study signal propagation in linear recurrent models at finite width. While existing signal propagation theory relies predominantly on the infinite-width limit, it remains unclear for how long that approximation remains accurate when recurrent depth $t$ grows jointly with width $n$. This question is especially relevant for modern recurrent sequence models, whose natural operating regime involves long input sequences, i.e., large $t$. We derive exact finite-width formulas for the hidden state s...
|
| 505 |
Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction
2605.05134
|
cs.LG
|
Dan Wilson, Mohamed Akrout |
Large Language Models (LLMs) frequently generate plausible but non-factual content, a phenomenon known as hallucination. While existing detection methods typically rely on computationally expensive sampling-based consistency checks or external knowledge retrie...Large Language Models (LLMs) frequently generate plausible but non-factual content, a phenomenon known as hallucination. While existing detection methods typically rely on computationally expensive sampling-based consistency checks or external knowledge retrieval, we propose a new method that treats the LLM as a black-box dynamical system. By projecting LLM responses into a high-dimensional manifold via an embedding model, we characterize the resulting vector sequences as observable realizations...
|
| 506 |
Gradient-Momentum Coupling: A Parameter-Space Proxy for Learning Progress
2605.05856
|
cs.LG
|
Samuel Blad, Martin L\"angkvist, Amy Loutfi |
Measuring learning progress is at the core of curiosity-driven exploration, which rewards an agent for going where its model is still learning. However, the abstract notion of learning progress is not directly measurable, and existing methods often derive it f...Measuring learning progress is at the core of curiosity-driven exploration, which rewards an agent for going where its model is still learning. However, the abstract notion of learning progress is not directly measurable, and existing methods often derive it from the prediction error in the output space. This paper proposes Gradient-Momentum Coupling (GMC), which measures how strongly a sample drives change in the parameter space, given by the normalized absolute product of its gradient with the...
|
| 507 |
QuadraSHAP: $\epsilon$-Exact Shapley Values for Product Games in Logarithmic Parallel Time
2605.05870
|
cs.LG
|
Majid Mohammadi, Grigory Reznikov, Pavel Sinitcyn, Krikamol Muandet, Siu Lun Chau |
We introduce QuadraSHAP, a method for $\epsilon$-exact Shapley computation in product games, cooperative games whose coalition values factorize across players. Given a tolerance $\epsilon\ge0$, the method determines a computational budget before evaluation, gu...We introduce QuadraSHAP, a method for $\epsilon$-exact Shapley computation in product games, cooperative games whose coalition values factorize across players. Given a tolerance $\epsilon\ge0$, the method determines a computational budget before evaluation, guaranteeing an absolute attribution error of at most $\epsilon$ for every feature in exact arithmetic. By extending to weighted sums of product games, the framework supports baseline and empirical interventional attribution across a broad cl...
|
| 508 |
Geometry-Aware Simplicial Message Passing
2605.06061
|
cs.LG
|
Elena Xinyi Wang, Bastian Rieck |
The Weisfeiler--Lehman (WL) test and its simplicial extension (SWL) characterize the combinatorial expressivity of message passing networks, but they are blind to geometry, i.e., meshes with identical connectivity but different embeddings are indistinguishable...The Weisfeiler--Lehman (WL) test and its simplicial extension (SWL) characterize the combinatorial expressivity of message passing networks, but they are blind to geometry, i.e., meshes with identical connectivity but different embeddings are indistinguishable. We introduce the Geometric Simplicial Weisfeiler--Lehman (GSWL) test, which incorporates vertex coordinates into color refinement for geometric simplicial complexes. In addition, we show that (i) the expressivity of geometry-aware simplic...
|
| 509 |
Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
2605.08475
|
cs.LGcs.AI
|
Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang |
In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass. Under bounded-data ass...In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass. Under bounded-data assumptions, we construct a single-head transformer whose forward pass approximately implements \textit{preconditioned Richardson iteration} on the associated kernel system. The construction uses $O(\log(1/\epsilon))$ blocks and MLP width $O(\...
|
| 510 |
Complex-Valued Phase-Coherent Transformer
2605.10123
|
cs.LG
|
Leona Hioki |
Complex-valued Transformers have largely inherited softmax attention from real-valued architectures. However, row-normalised token competition is not necessarily aligned with phase-preserving computation. In this paper, we introduce the Phase-Coherent Transfor...Complex-valued Transformers have largely inherited softmax attention from real-valued architectures. However, row-normalised token competition is not necessarily aligned with phase-preserving computation. In this paper, we introduce the Phase-Coherent Transformer (PCT), which applies a real-valued, element-independent, smooth gate to L2-normalised complex query-key similarities. PCT replaces token competition with token-non-competing attention and is designed to preserve phase information across...
|
| 511 |
Supervised Deep Multimodal Matrix Factorization for Interpretable Brain Network Analysis
2605.13312
|
cs.LG
|
Amjad Seyedi, Lifang He, Songlin Zhao, Akwum Onwunta, Nicolas Gillis |
Multimodal brain network analysis faces a persistent trade-off between predictive accuracy and interpretability. Deep neural networks achieve high accuracy but behave as black boxes that reveal little about the brain modules driving their decisions, whereas ma...Multimodal brain network analysis faces a persistent trade-off between predictive accuracy and interpretability. Deep neural networks achieve high accuracy but behave as black boxes that reveal little about the brain modules driving their decisions, whereas matrix factorization methods provide parts-based interpretability yet remain largely shallow, unsupervised, and restricted to a single view, integrating modalities through predefined or heuristic fusion rules. To bridge this gap with a formul...
|
| 512 |
Mixed neural posterior estimation for simulators with discrete and continuous parameters
2605.13551
|
cs.LG
|
Jan Boelts, Cornelius Schr\"oder, Jonas Beck, Jakob H. Macke, Michael Deistler |
Neural Posterior Estimation (NPE) enables rapid parameter inference for complex simulators with intractable likelihoods. NPE trains an inference network to estimate a probability density over parameters given data, typically assumed to be \emph{continuous}. Ho...Neural Posterior Estimation (NPE) enables rapid parameter inference for complex simulators with intractable likelihoods. NPE trains an inference network to estimate a probability density over parameters given data, typically assumed to be \emph{continuous}. However, many scientific models involve parameter spaces that are \emph{mixed}, that is, they contain both discrete and continuous dimensions. We address this limitation by extending NPE to mixed parameter spaces through an inference network ...
|
| 513 |
Tabular Imbalanced Learning: A Survey, Benchmark, and Practical Guide
2605.14915
|
cs.LG
|
Ruizhe Liu, Jiaqi Luo |
Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed methods, there is still limited systematic understanding of how different imbalance-handling strategies perform across diverse d...Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed methods, there is still limited systematic understanding of how different imbalance-handling strategies perform across diverse data regimes and computational constraints, making practical method selection difficult. In this work, we provide a systematic survey of tabular imbalanced learning and introduce Tabular Imbalanced Learning Benchmark (TILBench), a large-scal...
|
| 514 |
Attention Sinks and Outliers in Attention Residuals
2605.17887
|
cs.LGcs.AI
|
Haozheng Luo, Haoran Dai, Jingyuan Huang, Shaoyang Zhang, Xi Chen |
We propose OASIS, an outlier- and sink-aware method that stabilizes dual-normalized attention-residual architectures through explicit null routing and token-to-depth null coupling. AttnResidual introduces an additional depth-wise normalization channel that imp...We propose OASIS, an outlier- and sink-aware method that stabilizes dual-normalized attention-residual architectures through explicit null routing and token-to-depth null coupling. AttnResidual introduces an additional depth-wise normalization channel that improves inter-layer routing flexibility but can also amplify attention sinks, activation outliers, and low-bit quantization error. OASIS builds on explicit Softmax1-based null routes at both the token and depth levels and uses token-level nul...
|
| 515 |
Federated Martingale Posterior Samping
2605.18554
|
cs.LG
|
Boning Zhang, Matteo Zecchin, Mingzhao Guo, Dongzhu Liu, Osvaldo Simeone |
Federated Bayesian neural networks require fixing a prior on the model parameters, which is notoriously difficult, and misspecification of this prior can severely degrade accuracy and calibration. Motivated by the rapid progress of predictive models, the marti...Federated Bayesian neural networks require fixing a prior on the model parameters, which is notoriously difficult, and misspecification of this prior can severely degrade accuracy and calibration. Motivated by the rapid progress of predictive models, the martingale posterior, also known as predictive Bayes, replaces the prior--likelihood pair with a predictive distribution and recovers parameter uncertainty by repeatedly drawing predictive samples and refitting the model. This letter proposes {f...
|
| 516 |
MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification
2605.19752
|
cs.LG
|
Paul Krzakala, Gabriel Melo, Camille Lan\c{c}on, Charlotte Laclau, R\'emi Flamary |
Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research. We address the Molecule Retrieval task, whic...Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research. We address the Molecule Retrieval task, which consists in recovering the chemical structure of a metabolite from its MS/MS spectrum given a set of candidate molecules. We make three contributions. First, we propose a unified framework encompassing recent approaches based on represent...
|
| 517 |
Smooth Piecewise Cutting for Neural Operator to Handle Discontinuities and Sharp Transitions
2605.19823
|
cs.LGcs.AI
|
Ha Dang, Sebastian Schmidt, Juergen Hesser |
Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently continuous representations struggle to capture discontinuities and sharp transitions. Existing approaches typically ...Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently continuous representations struggle to capture discontinuities and sharp transitions. Existing approaches typically approximate such features within continuous function spaces, often requiring increased model capacity and high-resolution data. In this work, we propose Cut-DeepONet, a two-stage training framework that explicitly models discontinuities whi...
|
| 518 |
QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits
2605.30358
|
cs.LG
|
Zhenxiao Fu, Lei Jiang, Fan Chen |
Quantum computing remains in the Noisy Intermediate-Scale Quantum (NISQ) era, with performance constrained by noise. Addressing this limitation requires hardware-facing capabilities beyond gate sequences: mid-circuit measurement and classical feedback for quan...Quantum computing remains in the Noisy Intermediate-Scale Quantum (NISQ) era, with performance constrained by noise. Addressing this limitation requires hardware-facing capabilities beyond gate sequences: mid-circuit measurement and classical feedback for quantum error correction (QEC), precise timing for dynamical decoupling (DD), and pulse-level waveform access for calibration. OpenQASM 3 exposes these capabilities through a hardware-level programming interface. Despite rapid progress in large...
|
| 519 |
Proper Scoring Rules for Right-Censored Survival Data
2606.06393
|
cs.LG
|
Jef Jonkers, Glenn Van Wallendael, Luc Duchateau, Sofie Van Hoecke |
Proper scoring rules provide a rigorous theoretical basis for the training and evaluation of probabilistic forecasts. In survival analysis, such forecasts describe the distribution of the time until an event occurs. However, this event time is often only parti...Proper scoring rules provide a rigorous theoretical basis for the training and evaluation of probabilistic forecasts. In survival analysis, such forecasts describe the distribution of the time until an event occurs. However, this event time is often only partially observed because follow-up may end before the event occurs, resulting in right censoring. We propose a framework for proper scoring of right-censored survival outcomes based on a simple idea: first, map the predictive distribution thro...
|
| 520 |
QueryGraph: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation
2606.08300
|
cs.LG
|
Aishwarya Chakravarthy, Vidhi Kulkarni, Duen Horng Chau |
Many real-world queries over personal data span multiple applications and require structured planning, as individual tools expose only partial information. While LLMs show strong reasoning and tool use, reliably executing multi-step, cross-tool queries remains...Many real-world queries over personal data span multiple applications and require structured planning, as individual tools expose only partial information. While LLMs show strong reasoning and tool use, reliably executing multi-step, cross-tool queries remains challenging. We introduce a system that converts natural language queries into structured graphs and executes them via a deterministic planner. Our approach uses depth-first search to resolve dependencies and combine results across tools, ...
|
| 521 |
Bergson: An Open Source Library for Data Attribution
2606.11660
|
cs.LG
|
Lucia Quirke, Louis Jaburi, David Johnston, William Z. Li, Gon\c{c}alo Paulo |
Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation. However, significant engin...Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation. However, significant engineering effort is required to perform it at scale, and many cutting edge techniques lack open-source tooling and support. Bergson is an open source library that aims to enable faster progress in the field by providing a host of techniques th...
|
| 522 |
Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts
2606.14929
|
cs.LGcs.AI
|
Yan Dai, Negin Golrezaei, Patrick Jaillet |
Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback...Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited observability of models. We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and experts are the embedding models working on low...
|
| 523 |
ThousandWorlds: A benchmark for climate emulation of potentially habitable exoplanets
2606.18338
|
cs.LG
|
Edward T. Stevenson, Mei Ting Mak, Eric Wolf, Denis E. Sergeev, Tobi Hammond |
The search for life beyond Earth will depend on detecting faint signatures in the atmospheres of potentially habitable exoplanets. Interpreting those signatures requires understanding the host planet's climate: the same molecule may signal life on one planet a...The search for life beyond Earth will depend on detecting faint signatures in the atmospheres of potentially habitable exoplanets. Interpreting those signatures requires understanding the host planet's climate: the same molecule may signal life on one planet and abiotic chemistry on another. Global climate models (GCMs) provide this understanding, but individual runs can require up to millions of core-hours and substantial domain expert time. Machine-learning emulators could remove this bottlene...
|
| 524 |
Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales
2606.23453
|
cs.LG
|
Daniel Kiv, Shaowen Wang |
Location encoders transform geographic coordinates into high dimensional embeddings for downstream machine learning, but it is unclear how well these representations capture interpretable spatial effects. We benchmark whether GeoShapley, a game-theoretic expla...Location encoders transform geographic coordinates into high dimensional embeddings for downstream machine learning, but it is unclear how well these representations capture interpretable spatial effects. We benchmark whether GeoShapley, a game-theoretic explainer that treats all location features as a single joint player, can recover spatially varying coefficients from models built on location-encoder embeddings. Eleven encoders from the TorchSpatial framework are evaluated against a synthetic ...
|
| 525 |
TeDiServe: High SLO Attainment Serving for Diffusion Language Models
2606.29094
|
cs.LG
|
Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet Pham, Shivaram Venkataraman |
Diffusion language models (DLMs) have recently emerged as a promising alternative to conventional autoregressive language models. By generating multiple tokens in parallel during each denoising step, they offer higher inference throughput while maintaining com...Diffusion language models (DLMs) have recently emerged as a promising alternative to conventional autoregressive language models. By generating multiple tokens in parallel during each denoising step, they offer higher inference throughput while maintaining competitive quality. However, realizing these throughput gains while meeting latency SLOs in a serving system requires addressing challenges introduced by DLMs' unique characteristics. These include navigating the speed-quality tradeoff create...
|
| 526 |
Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark
2606.30170
|
cs.LGcs.AI
|
Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda, Julian Lorenz, Rainer Lienhart |
Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination yields strong benchmark metrics but limits transferability to domains structurally distinct from ...Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination yields strong benchmark metrics but limits transferability to domains structurally distinct from drug discovery. To overcome this limitation and drive discovery toward real, scientifically grounded targets, we introduce the Nanotechnology Molecular Optimization (NMO) Benchmark, which bridges machine learning (ML) and quantum materials ...
|
| 527 |
Regularizing modality contribution drift in multimodal continual learning
2607.27260
|
cs.LG
|
Zhen Zhang, Jielei Chu, Wenjie Ban, Tian Sang, Yuxiao Li |
Multimodal continual learning (MMCL) aims to acquire new knowledge from multimodal data while retaining previously learned knowledge. Existing MMCL methods primarily mitigate forgetting by aligning cross-modal representations or preserving feature-level semant...Multimodal continual learning (MMCL) aims to acquire new knowledge from multimodal data while retaining previously learned knowledge. Existing MMCL methods primarily mitigate forgetting by aligning cross-modal representations or preserving feature-level semantic similarity. However, different tasks may rely on different modalities, and learning new tasks can alter how modalities contribute to predictions on previously learned tasks. It remains underexplored how modality contributions evolve acro...
|
| 528 |
Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
2608.04408
|
cs.LGcs.AI
|
De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma |
On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through b...On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through budget-matched teacher-continuation and rollback branches. Based on their relative success, states are categorized as recoverable, irreversible-but-avoidable, or ambiguous, and these labels guide whether training retains, rolls back, or conv...
|
| 529 |
A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning
2608.07335
|
cs.LGcs.AI
|
Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni |
Replay-free parallelized Q-learning removes the large experience replay buffers and target networks used by conventional deep Q-learning, but the role of network architecture in this training regime remains comparatively underexplored. We investigate this ques...Replay-free parallelized Q-learning removes the large experience replay buffers and target networks used by conventional deep Q-learning, but the role of network architecture in this training regime remains comparatively underexplored. We investigate this question through a progressive three-phase study within the Parallelized Q-Network (PQN) framework. First, we compare eight convolutional encoder topologies on Atari-57 under a common training protocol while jointly considering performance and ...
|
| 530 |
A Banach-Space Theory of Markovian Halpern Iteration for Non-Expansive Maps
2608.15966
|
cs.LG
|
Ege C. Kaya, Arda Fazla, M. Berk Sahin, Abolfazl Hashemi |
We study stochastic approximation of fixed points of a non-expansive operator $T$ when the oracle samples originate from a continuing Markovian trajectory. A direct block-minibatch implementation of Halpern iteration attains an expected last-iterate residual o...We study stochastic approximation of fixed points of a non-expansive operator $T$ when the oracle samples originate from a continuing Markovian trajectory. A direct block-minibatch implementation of Halpern iteration attains an expected last-iterate residual of order $O(\log N/N)$, but accrues a substantive complexity of $\tilde O(\epsilon^{-5})$ Markovian samples. We therefore introduce a variance-reduced Markovian PAGE-Halpern method whose refresh and same-state difference blocks are analyzed ...
|
| 531 |
Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
2608.17965
|
cs.LGcs.AI
|
Bin Li, Dongdong Wang, Siyang Lu |
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We ...Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conve...
|
| 532 |
DecoVAE: a Lightweight Interpretable Trend-Seasonal VAE Framework for Efficient Probabilistic Time Series Forecasting
2608.20052
|
cs.LG
|
Alexander Marusov, Dmitry Anikin, Alexey Zaytsev |
Probabilistic time series forecasting remains challenging, largely because modeling distinct trend and seasonal dynamics requires specialized approaches. Existing methods often fail to capture the unique inner properties of these components, lack interpretabil...Probabilistic time series forecasting remains challenging, largely because modeling distinct trend and seasonal dynamics requires specialized approaches. Existing methods often fail to capture the unique inner properties of these components, lack interpretability, or suffer from heavy memory and runtime overhead. To address these limitations, we propose DecoVAE, a lightweight interpretable trend-seasonal VAE framework that explicitly decomposes time series into trend and seasonal components by a...
|
| 533 |
PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors
2608.23101
|
cs.LGcs.AI
|
Nathan Duboisset, Zhaolan Huang, Felix Bie{\ss}mann, Roudy Dagher, Antoine Lavandier |
Recent progress in the field of TinyML has demonstrated that low-power hardware based on microcontrollers can achieve bird species monitoring in real time based on acoustic sensor data for an entire breeding period on a single battery charge. However, the stat...Recent progress in the field of TinyML has demonstrated that low-power hardware based on microcontrollers can achieve bird species monitoring in real time based on acoustic sensor data for an entire breeding period on a single battery charge. However, the state of the art on low-power microcontrollers was so far limited to binary classification of a single species. In contrast, real fauna monitoring deployments often target multiple species simultaneously. To address this challenge we develop Po...
|
| 534 |
SimCast-S2S: A Computationally Efficient Diffusion Model for Subseasonal Precipitation Forecasting
2608.26594
|
cs.LG
|
Hiep V. Dang, Antonios Mamalakis |
Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains ...Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains simulation fidelity. We introduce SimCast-S2S, a generative latent-diffusion framework for probabilistic S2S precipitation forecasting that addresses three major bottlenecks in data-driven prediction. First, because S2S prediction requires ...
|
| 535 |
One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI
2608.29420
|
cs.LG
|
Louis Yiven Zhu |
Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct vali...Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct validity that a structural test and a predictive test can answer in opposite ways. We show that they do on a hash-pinned snapshot of a frontier leaderboard with 421 model configurations across twelve benchmarks, four of them economic, of which ...
|
| 536 |
MolLedger: An Additive Graph Neural Network with Chemically Grounded ADME Attributions
2608.30636
|
cs.LG
|
Christina X. Ji |
Optimizing absorption, distribution, metabolism, and excretion (ADME) is an important part of small molecule drug discovery. Many machine learning models have been built to predict ADME properties to facilitate this optimization process, but explaining model p...Optimizing absorption, distribution, metabolism, and excretion (ADME) is an important part of small molecule drug discovery. Many machine learning models have been built to predict ADME properties to facilitate this optimization process, but explaining model predictions is challenging. We propose a new graph neural network architecture with built-in atom attributions. Our model MolLedger learns a global context vector for each molecule and a per-atom head to output atom scores that sum to the pr...
|
| 537 |
Provably Safe Sim-to-Real Transfer
2609.01418
|
cs.LGcs.AI
|
Tingting Ni, Maryam Kamgarpour |
We address safe sim-to-real transfer, in which an agent leverages an imperfect simulator and limited real-world interaction while ensuring safety throughout data collection in the real system. This problem arises in applications such as robotics and healthcare...We address safe sim-to-real transfer, in which an agent leverages an imperfect simulator and limited real-world interaction while ensuring safety throughout data collection in the real system. This problem arises in applications such as robotics and healthcare: simulators provide cheap data, but sim-to-real mismatch makes direct transfer unreliable, and collecting real-world data to correct this mismatch must itself be safe. Moreover, deployment objectives may vary across tasks, making it costly...
|
| 538 |
WEECFP-SuRGE: A Position-Aware Substructure Encoding Method for Molecular Property Prediction
2609.04672
|
cs.LG
|
Robert Epps |
Computational molecular property prediction requires representations that capture local chemistry, long-range interactions, and molecular topology. Conventional fingerprints provide efficient local substructure features, whereas learned graph and sequence mode...Computational molecular property prediction requires representations that capture local chemistry, long-range interactions, and molecular topology. Conventional fingerprints provide efficient local substructure features, whereas learned graph and sequence models can represent broader context but often rely on pretraining or three-dimensional conformers. We introduce Wide Encoded Extended Connectivity Fingerprints (WEECFP) with Substructure Rotary Graph-distance Encoding (SuRGE), a tokenized hier...
|
| 539 |
The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
2609.06934
|
cs.LGcs.AI
|
Srikanth Malla, Chiho Choi, Joon Hee Choi |
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023), fine-tuning attacks (Qi et al., 2024), and activation-space edits (Arditi et al., 2024) keep recovering the behaviors it was meant to re...Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023), fine-tuning attacks (Qi et al., 2024), and activation-space edits (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and follow it into pretraining. We measure the safety update $\Delta = W_{safe} - W_{base}$ against the curvature of the model's capabilities (the empirical Fisher of a capability loss)...
|
| 540 |
Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
2609.06951
|
cs.LGcs.AI
|
Srikanth Malla, Chiho Choi, Joon Hee Choi |
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alo...Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly ref...
|
| 541 |
Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification
2609.10487
|
cs.LG
|
Weifeng Yang |
We construct counterexamples to Rockafellar's sum conjecture in which two maximally monotone operators satisfy the interior-domain condition but their sum is not maximally monotone, thereby providing the complete disproof of the conjecture. We establish a gene...We construct counterexamples to Rockafellar's sum conjecture in which two maximally monotone operators satisfy the interior-domain condition but their sum is not maximally monotone, thereby providing the complete disproof of the conjecture. We establish a general construction theorem that computes the entire monotone polar of a class of graphs and characterizes their maximal monotonicity by the nonexistence of solutions to explicit equations in the continuous dual. We also prove a pullback theor...
|
| 542 |
EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale
2609.10980
|
cs.LGcs.AI
|
Ege C. Kaya, Abolfazl Hashemi |
EGGROLL (Sarkar et al., 2026) makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractive but geometrically severe: Each rank-...EGGROLL (Sarkar et al., 2026) makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractive but geometrically severe: Each rank-one perturbation lies in a zero-volume subset of the ambient matrix space, despite having identity covariance. We characterize the EGGROLL update mean field at finite rank and nonzero perturbation radius as a resolvent applied to the gradie...
|
| 543 |
Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
2609.12890
|
cs.LGcs.AI
|
Junhao Zhao, David Michael Simberg, Jacob Kang, Colin Connor Kurniawan, Nan Xu |
In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate t...In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate the update while amplifying predictable signal and unpredictable innovation together; a large distant gradient therefore need not carry reliable learning signal. Motivated by this, we introduce Internal Dual-Wiener routing (Internal-DW), a b...
|
| 544 |
Recovering Governing Dynamics from Distributed Observations via Exact Spline Merging
2609.16579
|
cs.LG
|
Naveen Mysore |
Scientific observations are frequently distributed across locations, time periods, and institutions. Combining such observations into a continuous, differentiable field enables recovering governing physical parameters from its derivatives. This paper makes two...Scientific observations are frequently distributed across locations, time periods, and institutions. Combining such observations into a continuous, differentiable field enables recovering governing physical parameters from its derivatives. This paper makes two contributions in this setting. First, the established additive structure of fixed-basis ridge-regression statistics is applied to tensor-product spline fields: each data holder computes a local Gram matrix and moment vector, and the merged...
|
| 545 |
A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure
2609.16606
|
cs.LG
|
John E. Darges, Laura Weidensager |
Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a method based on families of weigh...Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a method based on families of weighted ANOVA kernels that learn and adapt to this multivariable structure. TSKs parameterize the weights on each multivariable component of the target function by factors for each input. We propose learning these factors directly from function...
|
| 546 |
Information Geometric Self-Organization at the Edge of Stability in High-Capacity Kernel Associative Memories
2609.16827
|
cs.LG
|
Akira Tamamori |
High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirical studies identified a hyperparameter regime, the "Ridge of Optimization," where attractor stability is maxim...High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirical studies identified a hyperparameter regime, the "Ridge of Optimization," where attractor stability is maximized. However, the geometric nature of this regime and the optimization dynamics required to reach it have remained unclear. In this paper, we investigate the static geometry of the parameter space and the learning trajectory of Gradient De...
|
| 547 |
Beyond Quadratic Loss: The Stability Phase Diagram of Adam
2609.18314
|
cs.LG
|
Gaoxiang Tang, Huanran Chen, Ziming Liu |
Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. ...Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. We investigate this dependence by mapping training dynamics across the $(\beta_1,\beta_2)$ plane. Across a range of model--task settings, an approximately linear boundary, $1-\beta_2=C(1-\beta_1)$, separates spiky from non-spiky dynamics, w...
|
| 548 |
Higher-order pruning of experts in mixture-of-experts language models
2609.18916
|
cs.LGcs.AI
|
Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto |
Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert...Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes a...
|
| 549 |
Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models
2609.19674
|
cs.LG
|
Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling |
A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervent...A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one. We show that these two failures require different structural remedies. Evolving a learned energy with a symplectic in...
|
| 550 |
Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
2609.20888
|
cs.LG
|
Themistoklis Haris, Henry Li, Maryam Karimzadehgan |
Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this problem, but often drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ET...Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this problem, but often drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that rivals dense model quality under hardware-aligned block-sparse decoding. ETA predicts dynamic, contextual thresholds directly from query representations, adjusting context retention depending o...
|
| 551 |
Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods
2609.21815
|
cs.LGcs.AI
|
Wenpeng Zhang, Runsheng Yu, Peilin Zhao |
Adaptive optimization methods such as AdaGrad and Adam are widely used in modern deep neural network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware o...Adaptive optimization methods such as AdaGrad and Adam are widely used in modern deep neural network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop an Online Mirror Descent framewo...
|
| 552 |
Explainable Predictive Condition-based Maintenance of Naval-Propulsion Systems using Fuzzy Logic
2609.24250
|
cs.LG
|
Dionisis Kalogeropoulos, Georgia Sovatzidi, Panagiotis G. Kalozoumis, Dimitris K. Iakovidis |
The shipping industry has a significant impact on the global economy, emphasizing the need for operational availability and safety through the use of effective maintenance techniques. During the last decades, predictive maintenance (PdM) has emerged as a promi...The shipping industry has a significant impact on the global economy, emphasizing the need for operational availability and safety through the use of effective maintenance techniques. During the last decades, predictive maintenance (PdM) has emerged as a promising solution compared to the existing conventional maintenance systems. This is because it offers several advantageous functions, such as damage predictions for vessel components, reduced downtime, improved and extended life of machinery, ...
|
| 553 |
Lifted Bellman Linear Programming for Offline Reinforcement Learning
2609.24489
|
cs.LGcs.AI
|
Hyukjun Yang, Jongchan Park, Narim Jeong, Donghwan Lee |
Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions...Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characteriz...
|
| 554 |
Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training
2609.25510
|
cs.LGcs.AI
|
Jacob Beck, Philip V. Ogren, Ari Kobren |
Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating...Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time training. We ask how much of this machinery is necessary. We introduce Hill Sampling, a form of hill-climbing optimization that repeatedly samples candidate programs from a frozen LLM, retains the best pro...
|
| 555 |
Geometry-Aware Hyperbolic Residual-Quantized Variational Autoencoders
2609.26342
|
cs.LGcs.AI
|
Alessio Colombo, Melika Ayoughi |
Residual Vector Quantization turns continuous representations into discrete, multi-level token sequences. Yet most methods operate in Euclidean space, despite the coarse-to-fine structure of the resulting codes and the latent hierarchies present in many data d...Residual Vector Quantization turns continuous representations into discrete, multi-level token sequences. Yet most methods operate in Euclidean space, despite the coarse-to-fine structure of the resulting codes and the latent hierarchies present in many data domains. Hyperbolic geometry offers a natural alternative for hierarchical representations, but naive hyperbolic extensions introduce geometric inconsistencies: non-associative hyperbolic addition prevents consistent residual aggregation, wh...
|
| 556 |
Hierarchical GNNs for power flow: letting physics shape the hierarchy
2609.26603
|
cs.LGcs.AI
|
Carmine Delle Femine, Leire Garin Atxaga, Asier Diaz-Iglesias, Juan Pablo Maroto Herrera, Ane Miren Florez-Tapia |
Hierarchical latent communication improves the generalization of a power-flow model, shared across three grids, to new operating scenarios. The module exchanges information through two reduced graphs inside the corrective network of GENCO, replacing two of its...Hierarchical latent communication improves the generalization of a power-flow model, shared across three grids, to new operating scenarios. The module exchanges information through two reduced graphs inside the corrective network of GENCO, replacing two of its local correction steps. We compare Kron-derived transports, a same-anchor Quotient construction and the flat GENCO Base architecture, all trained under one protocol of our own with about a hundred times fewer optimizer updates per grid tha...
|
| 557 |
What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates
2609.27679
|
cs.LG
|
Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang |
A tabular foundation model must discover which distinctions matter for each new table without updating its parameters. We develop in-situ representation refinement: support labels guide changes to the episode's representations, improving the information availa...A tabular foundation model must discover which distinctions matter for each new table without updating its parameters. We develop in-situ representation refinement: support labels guide changes to the episode's representations, improving the information available to later queries. A regularized leave-one-out objective yields a support correction and its query extension. The leading term separates attention-based reading from state-dependent scaling, motivating RefineICL: an attention-gated, FFN-...
|
| 558 |
NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers
2609.27735
|
cs.LG
|
Xiaohe Jiang (University of Exeter), Guoqiang Zhang (University of Exeter), Tianjin Huang (University of Exeter), Ronghui Mu (University of Exeter) |
Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations....Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations. We introduce Newton-Schulz Attention (NS-Attn.), a parameter-free transformation applied to the output of each attention head. Each head output is arranged as a feature-by-token matrix and normalized by its Frobenius norm. We then apply a ...
|
| 559 |
Transferable Evidence Reconstruction for Longitudinal Glucose Representations
2609.28199
|
cs.LG
|
Tian Zhou, Bingqing Peng, Linxiao Yang, Wenwei Wang, Mengni Ye |
Long physiological recordings contain many routine measurements, while predictive information often lies in rare events, sustained burden, and recurring patterns. These properties can be computed as label-free evidence, but directly using them as features leav...Long physiological recordings contain many routine measurements, while predictive information often lies in rare events, sustained burden, and recurring patterns. These properties can be computed as label-free evidence, but directly using them as features leaves limited labeled data to separate reproducible associations from sample-specific ones. Learning to reconstruct evidence can exploit unlabeled recordings, yet joint reconstruction does not explicitly require the decoding rule to transfer a...
|
| 560 |
Support-Compiled Feature Folding: More Evidence at Lower Memory Across Tabular Foundation Models
2609.28208
|
cs.LG
|
Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang |
Wide tables offer tabular foundation models more evidence, but accessing it can exhaust their memory: full-width pairwise mixing grows quadratically with the number of columns, while feature selection makes inputs affordable by discarding evidence. We ask whet...Wide tables offer tabular foundation models more evidence, but accessing it can exhaust their memory: full-width pairwise mixing grows quadratically with the number of columns, while feature selection makes inputs affordable by discarding evidence. We ask whether using more features requires interacting over all of them at once. We introduce Support-Compiled Feature Folding (SCFF), a training-free inference framework that encodes wide tables through bounded calls to a frozen backbone. SCFF organ...
|
| 561 |
Even Sharper Bounds for Transductive Learning and Its Applications
2609.28459
|
cs.LG
|
Yingzhen Yang |
We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--...We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses the modified log-Sobolev inequality for the swap walk and a two-parameter entropy closure. A peeling argument with a surrogate localization functional then gives excess-risk bounds with the same fixed...
|
| 562 |
Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles
2609.29974
|
cs.LG
|
Ali Haghpanah Jahromi, Mohammad Taheri, Zohreh Azimifar |
Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Exp...Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation pred...
|
| 563 |
Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think
2609.30036
|
cs.LG
|
Xvyuan Liu, Jianjie Fang, Chen Gao, Yong Li |
Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require ac...Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. With frozen LeWM models, intermediate targets substantially improve action synthesis and recorded-action ranking on Cube, PushT, Reacher, and TwoRoom. Learned targets and targets drawn from observed e...
|
| 564 |
Learning Operators by Regularized Stochastic Gradient Descent with Operator-valued Kernels
2504.18184
|
cs.LG
|
Jia-Qi Yang, Lei Shi |
We consider a class of statistical inverse problems involving the estimation of a regression operator from a Polish space to a separable Hilbert space, where the target lies in a vector-valued reproducing kernel Hilbert space induced by an operator-valued kern...We consider a class of statistical inverse problems involving the estimation of a regression operator from a Polish space to a separable Hilbert space, where the target lies in a vector-valued reproducing kernel Hilbert space induced by an operator-valued kernel. To address the associated ill-posedness, we analyze regularized stochastic gradient descent (SGD) algorithms in both online and finite-horizon settings. The former uses polynomially decaying step sizes and regularization parameters, whi...
|
| 565 |
AUWave: A Data-Driven Model for Reconstructing Significant Wave Heights Using Sparse Observations
2509.19384
|
cs.LGcs.AI
|
Hongyuan Shi, Yilin Zhai, Ping Dong, Zaijin You, Chao Zhan |
Reconstructing high-resolution regional significant wave height (SWH) fields from sparse buoy observations is a critical challenge for ocean monitoring. We introduce AUWave, a hybrid deep learning framework that fuses a station-wise encoder with a multi-scale ...Reconstructing high-resolution regional significant wave height (SWH) fields from sparse buoy observations is a critical challenge for ocean monitoring. We introduce AUWave, a hybrid deep learning framework that fuses a station-wise encoder with a multi-scale U-Net enhanced by self-attention to recover regional SWH fields. Trained and validated using NDBC buoy observations and ERA5 reanalysis over the Hawaii region, AUWave achieves high accuracy. It consistently outperforms a representative base...
|
| 566 |
What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
2509.22796
|
cs.LG
|
Xingyu Li (UC Riverside), Juefei Pu (UC Riverside), Yifan Wu (UC Riverside), Xiaochen Zou (UC Riverside), Shitong Zhu (UC Riverside) |
Open-source software projects are foundational to modern software ecosystems, with the Linux kernel standing out as a critical exemplar due to its ubiquity and complexity. Although security patches are continuously integrated into the Linux mainline kernel, do...Open-source software projects are foundational to modern software ecosystems, with the Linux kernel standing out as a critical exemplar due to its ubiquity and complexity. Although security patches are continuously integrated into the Linux mainline kernel, downstream maintainers often delay their adoption, creating windows of vulnerability. A key reason for this lag is the difficulty in identifying security-critical patches, particularly those addressing exploitable vulnerabilities such as out-...
|
| 567 |
Generative Modeling of Discrete Data Using Geometric Latent Subspaces
2601.21831
|
cs.LG
|
Daniel Gonzalez-Alvarado, Jonas Cassel, Stefania Petra, Christoph Schn\"orr |
We propose a geometric latent-subspace framework for generative modeling of discrete data. Specifically, we introduce latent subspaces in the exponential parameter space of product manifolds of categorical distributions as a novel approach to learning low-dime...We propose a geometric latent-subspace framework for generative modeling of discrete data. Specifically, we introduce latent subspaces in the exponential parameter space of product manifolds of categorical distributions as a novel approach to learning low-dimensional representations of high-dimensional discrete data. The resulting low-dimensional latent space captures statistical dependencies and removes redundant degrees of freedom among the categorical variables. We equip the parameter domain ...
|
| 568 |
Latent Generative Solvers for Generalizable Long-Term Physics Simulation
2602.11229
|
cs.LGcs.AI
|
Zituo Chen, Sili Deng |
Reliable physics simulation demands two capabilities that today's neural PDE solvers do not deliver together: generalization across heterogeneous PDE families, and stability under long autoregressive rollouts. Deterministic operators accumulate error geometric...Reliable physics simulation demands two capabilities that today's neural PDE solvers do not deliver together: generalization across heterogeneous PDE families, and stability under long autoregressive rollouts. Deterministic operators accumulate error geometrically, while existing probabilistic solvers are confined to a single PDE family or short horizons. We close this gap with the \textbf{Latent Generative Solver} (LGS), three coupled components: (i) a Physics VAE (PhyVAE) compressing twelve PD...
|
| 569 |
Minimax and Adaptive Covariance Matrix Estimation under Differential Privacy
2603.19703
|
cs.LG
|
T. Tony Cai, Yicheng Li |
Estimating covariance matrices is fundamental to a wide range of statistical applications. This paper studies minimax and adaptive estimation of high-dimensional covariance matrices under $\rho$-zero-concentrated differential privacy ($\rho$-zCDP) over three n...Estimating covariance matrices is fundamental to a wide range of statistical applications. This paper studies minimax and adaptive estimation of high-dimensional covariance matrices under $\rho$-zero-concentrated differential privacy ($\rho$-zCDP) over three nested classes: the pointwise-decay class $\mathcal{H}_\alpha$, the row-tail class $\mathcal{G}_\alpha$, and the separated-block class $\mathcal{F}_\alpha$. We consider both squared operator norm loss and normalized squared Frobenius norm lo...
|
| 570 |
On the Expressive Power of Transformers for Contextual Relations
2603.25860
|
cs.LG
|
Demi\'an Fraiman |
Transformers have revolutionized machine learning by making attention a central mechanism for modeling interactions within a context. Despite the central role of attention, the theoretical capabilities of Transformers for representing contextual relations rema...Transformers have revolutionized machine learning by making attention a central mechanism for modeling interactions within a context. Despite the central role of attention, the theoretical capabilities of Transformers for representing contextual relations remain unclear. In this work, we address this question by developing a mathematical framework based on probability and optimal transport. We view a text as a distribution of its representations and attention as a probabilistic relation between ...
|
| 571 |
On associative neural networks for sparse patterns with huge capacities
2603.26217
|
cs.LG
|
Matthias L\"owe, Franck Vermet |
Generalized Hopfield models with higher-order or exponential interaction terms are known to have substantially larger storage capacities than the classical quadratic model. On the other hand, associative memories for sparse patterns, such as the Willshaw and A...Generalized Hopfield models with higher-order or exponential interaction terms are known to have substantially larger storage capacities than the classical quadratic model. On the other hand, associative memories for sparse patterns, such as the Willshaw and Amari models, already exhibit enhanced storage capacities in the sparse regime. In this paper we combine these two mechanisms. We introduce higher-order versions of sparse associative memory models and study their storage capacities in the s...
|
| 572 |
A Sharp Norm Inequality and Buzano's Inequality via Determinants
2604.01525
|
cs.LG
|
Jose Antonio Lara Benitez |
We give a short linear-algebraic proof of the inequality $$ \|x\|_1\,\|x\|_\infty \le \frac{1+\sqrt{n}}{2}\,\|x\|_2^2, $$ valid for every $x\in\mathbb{R}^n$. This inequality relates three fundamental norms on finite-dimensional spaces and has applications in o...We give a short linear-algebraic proof of the inequality $$ \|x\|_1\,\|x\|_\infty \le \frac{1+\sqrt{n}}{2}\,\|x\|_2^2, $$ valid for every $x\in\mathbb{R}^n$. This inequality relates three fundamental norms on finite-dimensional spaces and has applications in optimization and numerical analysis. Our proof exploits the determinantal structure of a parametrized family of quadratic forms, and we show the constant $(1+\sqrt{n})/2$ is optimal. The inequality is a special case of Buzano's inequality, a...
|
| 573 |
Neural Parameter Estimation of RC Thermal Building Models for Model Predictive Control
2604.05904
|
cs.LG
|
Fabian Raisch, Timo Germann, Sang-Woo Ham, J. Nathan Kutz, Christoph Goebel |
Gray-box RC models are widely used to enable energy-efficient model predictive control (MPC) in buildings. However, estimating RC parameters remains difficult, as conventional optimization-based algorithms are prone to local minima, rely heavily on good initia...Gray-box RC models are widely used to enable energy-efficient model predictive control (MPC) in buildings. However, estimating RC parameters remains difficult, as conventional optimization-based algorithms are prone to local minima, rely heavily on good initial guesses, and incur high computational cost. To address these issues, we propose the Estimator from Scratch, a novel neural parameter estimation approach that embeds the physical equations into a neural network's training process to estima...
|
| 574 |
Identifying Causal Effects Using a Single Proxy Variable
2604.09135
|
cs.LG
|
Silvan Vollmer, Niklas Pfister, Sebastian Weichwald |
Unobserved confounding is a key challenge when estimating causal effects from a treatment on an outcome. In this work, we assume that we observe a single, potentially multi-dimensional proxy variable of the unobserved confounder and that we know the mechanism ...Unobserved confounding is a key challenge when estimating causal effects from a treatment on an outcome. In this work, we assume that we observe a single, potentially multi-dimensional proxy variable of the unobserved confounder and that we know the mechanism that generates the proxy from the confounder. Under an assumption called Single Proxy Identifiability of Causal Effects or simply SPICE, we prove that this error mechanism is complete and causal effects are identifiable. We extend the proxy...
|
| 575 |
Towards Interpretable Damage Detection based on Aerodynamic Pressure Measurements
2605.08187
|
cs.LG
|
Philip Franz, Max von Danwitz, Gregory Duth\'e, Alexander Popp, Eleni Chatzi |
The increasing flexibility of modern large wind turbine blades necessitates cost-efficient and reliable structural monitoring solutions. For this purpose, we propose to use aerodynamic pressure measurements obtained via Aerosense, a novel, non-intrusive and ec...The increasing flexibility of modern large wind turbine blades necessitates cost-efficient and reliable structural monitoring solutions. For this purpose, we propose to use aerodynamic pressure measurements obtained via Aerosense, a novel, non-intrusive and economical sensing system. In former work [Franz et al., 2025], we investigated the potential of aerodynamic pressure measurements for structural damage detection on elastic and aerodynamically loaded structures. An experimental campaign was ...
|
| 576 |
HiLiftAeroML: A High-Fidelity Computational Fluid Dynamics Dataset for High-Lift Aircraft Aerodynamics
2605.19565
|
cs.LG
|
Neil Ashton, Adam Clark, Konrad Goc, Liam Heidt, Christopher Ivey |
HiLiftAeroML is, to our knowledge, the first open high-fidelity computational fluid dynamics dataset dedicated to high-lift aircraft aerodynamics. It contains 1,800 simulations spanning 180 variants of the NASA Common Research Model high-lift configuration and...HiLiftAeroML is, to our knowledge, the first open high-fidelity computational fluid dynamics dataset dedicated to high-lift aircraft aerodynamics. It contains 1,800 simulations spanning 180 variants of the NASA Common Research Model high-lift configuration and ten angles of attack from $4^\circ$ to $22^\circ$. Each case was generated with a GPU-accelerated explicit wall-modeled large-eddy simulation approach on solution-adapted grids of 300--500 million cells, covering attached, separated, and p...
|
| 577 |
Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation
2606.11891
|
cs.LG
|
Mehmet Turan Yard{\i}mc{\i} |
Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within one policy. A natural design choice is between a single (unified) critic that estimates the combined value of all objectives and separate (dual) criti...Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within one policy. A natural design choice is between a single (unified) critic that estimates the combined value of all objectives and separate (dual) critics with disjoint reward signals. We compare the two on the Unitree G1 humanoid in NVIDIA Isaac Lab. In the standing mode of a standardized evaluation, the dual-critic run reaches targets 3.5x faster (6.5 vs. 22.6 simulation steps), achieves...
|
| 578 |
Genetic Algorithms with Optimization Guided Operators
2606.12279
|
cs.LGcs.AI
|
Anna Brandenberger, Ilan Doron-Arad, Elchanan Mossel |
Recent work in ML applies genetic algorithms at inference time to iteratively improve solutions to optimization problems. The basic mutation and recombination operators involved are qualitatively different from those studied classically. Mutations are no longe...Recent work in ML applies genetic algorithms at inference time to iteratively improve solutions to optimization problems. The basic mutation and recombination operators involved are qualitatively different from those studied classically. Mutations are no longer random; an ML algorithm mutates a solution with the goal of improving an objective. Similarly, recombination is not based on random collages of parent solutions. Instead, it is an ML optimization-based operator whose goal is to synthesize...
|
| 579 |
Amortized quadrature for posterior expectations in inverse problems
2606.15871
|
cs.LG
|
Ali Siahkoohi |
Uncertainty in the solution of an inverse problem and in the tasks performed on it is quantified by posterior expectations, each an average of an integrand over $M$ posterior samples. While designed quadratures improve on the $O(M^{-1/2})$ error of Monte-Carlo...Uncertainty in the solution of an inverse problem and in the tasks performed on it is quantified by posterior expectations, each an average of an integrand over $M$ posterior samples. While designed quadratures improve on the $O(M^{-1/2})$ error of Monte-Carlo estimation, they solve an optimization problem, often against the posterior density, for every new observation, which can be computationally costly. To address this limitation, we introduce the quadrature field, a set-equivariant network t...
|
| 580 |
NAC: Neural Action Codec for Vision-Language-Action Models
2606.21372
|
cs.LG
|
Ahad Jawaid, Yu Xiang |
Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this de...Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this design through the lens of neural audio codecs - convolutional encoder-decoder architectures with residual vector quantization that serve as the standard front end for audio foundation models. Motivated by their success, we introduce the Neur...
|
| 581 |
ConSolv: Solvent-Conditional Machine Learning Implicit Solvent Potential
2606.24983
|
cs.LG
|
Linying Zhang, Julija Zavadlav |
Implicit solvent machine learning potentials (MLPs) offer a powerful route to bridging the gap between accuracy and efficiency in molecular simulations. However, existing models have largely focused on aqueous environments, overlooking the diverse and importan...Implicit solvent machine learning potentials (MLPs) offer a powerful route to bridging the gap between accuracy and efficiency in molecular simulations. However, existing models have largely focused on aqueous environments, overlooking the diverse and important roles of non-aqueous solvents in areas such as organic synthesis and battery technology. Here, we present ConSolv, a solvent-conditional MLP architecture that explicitly incorporates solvent effects on solute interactions through an atten...
|
| 582 |
Statistically Valid Post-Training Hyperparameter Selection: From Tuning to Guarantees
2606.25601
|
cs.LG
|
Amirmohammad Farzaneh, Osvaldo Simeone |
Post-training hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom of pre-trained models such as inference-time parameters, implementation-level settings, and thresho...Post-training hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom of pre-trained models such as inference-time parameters, implementation-level settings, and thresholds driving decision rules. Despite its practical importance, hyperparameter selection is typically performed using best-effort empirical methods such as grid search or Bayesian optimization, which provide no formal statistical guarantees o...
|
| 583 |
Bringing Agentic Search to Earth Observation Data Discovery
2607.02387
|
cs.LG
|
Minghan Yu, Youran Sun, Chugang Yi, Yixin Wen, Haizhao Yang |
NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search framework for geoscience data dis...NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search framework for geoscience data discovery that takes a natural-language research query and returns matching datasets and tools. We demonstrate that, in the era of large language models, the latent value of knowledge graphs (KGs) can be substantially amplified through agentic...
|
| 584 |
Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics
2607.09250
|
cs.LG
|
Hugo Cui |
The impact of a given training point on a statistical model can be measured through its leave-one-out influence on the model parameters, which quantifies how its removal from the training set affects the learned weights. For convex M-estimation under Gaussian ...The impact of a given training point on a statistical model can be measured through its leave-one-out influence on the model parameters, which quantifies how its removal from the training set affects the learned weights. For convex M-estimation under Gaussian design, in the high-dimensional limit $n\asymp d$, we show that the empirical distribution of influences across training points concentrates around a deterministic measure which we sharply characterize. This characterization suggests that i...
|
| 585 |
Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks
2607.26574
|
cs.LGcs.AI
|
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng |
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered ins...Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemb...
|
| 586 |
Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
2607.26865
|
cs.LGcs.AI
|
Amirmohammad Farzaneh, Osvaldo Simeone |
LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning bu...LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that synergistically integrates a lightweight convergence probe...
|
| 587 |
On the robustness of noisy solutions in non-convex neural networks
2607.27000
|
cs.LG
|
Enrico M. Malatesta, Alessandra Passalacqua, Riccardo Zecchina |
Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite bein...Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare. At zero temperature this picture has been formalized in binary perceptrons through the overlap gap property (OGP), which limits algorithmic access to configurations with zero training error above a critical constraint den...
|
| 588 |
DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing
2608.02769
|
cs.LG
|
Sagnik Nandy, Samriddha Lahiry, Pragya Sur, Subhabrata Sen |
Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fail...Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fails to exploit cross-modal dependence. Existing approaches rely on pre-specified fusion architectures, from early to late fusion, that may not adapt to the underlying dependence structure among modalities. We propose DAIF, a data adaptive int...
|
| 589 |
LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models
2608.16340
|
cs.LG
|
Tom Splittgerber, Niklas Koenen, Marvin N. Wright, Werner Brannath |
The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs....The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs. In order to preserve interpretability, it is usually necessary to restrict the NN components to prevent them from dominating the model. However, existing methods that enforce structural constraints on their NN components severely limit the...
|
| 590 |
Provable Quantum-Classical Separation for Continuous Gibbs Sampling
2608.24527
|
cs.LG
|
Enrico Olivucci, Mariia Sobchuk, Sehmimul Hoque, Jeffrey Hnybida, Kyungho W. Kim |
We prove the first quantum-classical separation for a sampling problem over a continuous domain. For a class of Gibbs states $p\propto e^{-\beta E}$ on the torus $\mathbb{T}^d$ with smooth ($s$-Gevrey) potential and barrier amplitude $\alpha=e^{\beta\Delta}$, ...We prove the first quantum-classical separation for a sampling problem over a continuous domain. For a class of Gibbs states $p\propto e^{-\beta E}$ on the torus $\mathbb{T}^d$ with smooth ($s$-Gevrey) potential and barrier amplitude $\alpha=e^{\beta\Delta}$, where $\Delta = \max E-\min E$, every classical algorithm querying the value, gradient, or any higher-order derivatives of the log-density requires $\Omega(\alpha)$ queries to sample at constant accuracy in total variation distance, while a...
|
| 591 |
Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
2608.26088
|
cs.LGcs.AI
|
Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun |
Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented dat...Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model selection. We present the Planetary Prediction Engine (PPE), an autonomous AI system that executes this end-to-end workflow directl...
|
| 592 |
High-probability guarantees for linear accessibility in feature superposition
2609.09556
|
cs.LGcs.AI
|
Enrico Vompa |
Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, w...Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise, proving the sufficient dimension scales linearly ($d=O_{\varepsilon}(k \log m)$) rather than prior worst-case quadratic limits. We characterize the asymmetry betwe...
|
| 593 |
Graph Matching Relaxations and Amortization for Supervised Graph Prediction
2609.15437
|
cs.LG
|
Federico M\'endez, Paul Krzakala, Gabriel Melo, Charlotte Laclau, R\'emi Flamary |
End-to-end Supervised Graph Prediction (SGP) requires a permutation-invariant loss to compare predicted and target graphs with arbitrary node orderings. Such losses typically involve a costly graph-matching problem. We first study three Optimal Transport relax...End-to-end Supervised Graph Prediction (SGP) requires a permutation-invariant loss to compare predicted and target graphs with arbitrary node orderings. Such losses typically involve a costly graph-matching problem. We first study three Optimal Transport relaxations of this problem and show, theoretically and empirically, that the Gromov-Wasserstein (GW) objective is the most suitable for SGP. Then, to avoid solving the resulting inner optimization for every training example, we propose to amort...
|
| 594 |
Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing
2609.24791
|
cs.LG
|
Nibraas Khan, Abigale Plunk, John Staubitz, Ingrid Shragge, Jordan Brooks |
Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, voc...Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, vocalization, and autonomic arousal. Its signs are subtle and individualized, and its autonomic components are invisible without instrumentation. We collected upper-body movement from inertial measurement units, physiology from a wrist-worn de...
|
| 595 |
Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis
2609.25576
|
cs.LG
|
Jun Li, Yanlong Guo, Zhaozhao Zeng |
We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ st...We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ storage and are impractical at $N\sim 10^6$--$10^8$. We propose DeepMVSA, which re-expresses the minimum-volume principle in neural implicit form: a lightweight coordinate network generates the mixing weights and a triangular LU-type paramete...
|
| 596 |
Reinforcement Learning with Decomposed Subtasks
2609.27035
|
cs.LGcs.AI
|
Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich |
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct sk...Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a be...
|
| cs.MM 1 papers | ||||
| 775 |
IMEX-FND: A Traceable Interaction-Aware Mixture-of-Experts Framework for Multimodal Fake News Detection
2609.29610
|
cs.MM
|
Yuchen Miao, Zijun Wang, Ke Liu, Peixuan Wang, Chang Han |
Multimodal fake news detection (FND) increasingly demands verdicts that are not only accurate but traceable, revealing how cross-modal evidence is combined, yet two coupled difficulties remain. First, text-image relations are heterogeneous: uniqueness, redunda...Multimodal fake news detection (FND) increasingly demands verdicts that are not only accurate but traceable, revealing how cross-modal evidence is combined, yet two coupled difficulties remain. First, text-image relations are heterogeneous: uniqueness, redundancy, and synergy coexist and vary from post to post, so a single global fusion rule is brittle and opaque. Second, the dominant modality shifts across instances, which static encoders and a fixed fusion pathway handle poorly. We present IME...
|
| cs.SD 14 papers | ||||
| 757 |
CLEAR: Online Speech Content Leakage Estimation through Cross-ASR Disagreement
2609.30415
|
cs.SDeess.AS
|
Bhawana Chhaglani, Tanvi Kandepuneni, Jeremy Gummeson, Prashant Shenoy |
Signal-level speech privacy mechanisms suppress linguistic content while preserving acoustic information needed by downstream sensing applications. However, their privacy settings are typically evaluated/selected offline and remain fixed during deployment, eve...Signal-level speech privacy mechanisms suppress linguistic content while preserving acoustic information needed by downstream sensing applications. However, their privacy settings are typically evaluated/selected offline and remain fixed during deployment, even though speech-content leakage can vary substantially across utterances and speakers. Adapting privacy protection at runtime requires estimating how much speech remains recoverable, but conventional measures such as WER or PER require grou...
|
| 758 |
MuseTimbre: Zero-Shot Timbre Transfer by Controlling a Frozen Music Generator
2609.30548
|
cs.SDeess.AS
|
Yuan-Chiao Cheng, Zhiyao Duan |
Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicat...Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicated model for the task, which captures the timbre cleanly but stays a narrow, single-purpose system. More versatile approaches add control to a pretrained music generator, yet a reference clip entangles timbre with genre and melody, so these...
|
| 759 |
Training-Free Contextual ASR via SpeechLLM-Based Error-Aware Selective Retrieval
2609.30694
|
cs.SDeess.AS
|
Natsuo Yamashita, Ai Nemoto, Ryosuke Koichi, Masaaki Yamamoto |
Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing t...Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issue by selecting candidate terms from an external dictionary, but querying many recognized words requires numerous dictionary lookups and may yield poorly targeted candidates. We pro...
|
| 760 |
Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information
2609.30774
|
cs.SDeess.AS
|
Shuhan Zhang, Wenxuan Wu, Haizhou Li |
In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures w...In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-pa...
|
| 761 |
Tracing and Relearning Detection Evidence in Text-to-Speech Systems
2609.30983
|
cs.SDeess.AS
|
Eunji Shin, Kyudan Jung, Jihwan Kim, Minwoo Lee, Jaegul Choo |
Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-Big...Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acousti...
|
| 762 |
TinyAudio: Compact and Efficient Text-to-Audio Generation for Low-Resource Deployment
2609.31525
|
cs.SDeess.AS
|
Junxi Liu, Xiquan Li, Wenhao Guan, Yifan Duan, Zhikang Niu |
Text-to-audio (TTA) generation has advanced rapidly in generation quality and instruction following. However, representative systems often require around a billion parameters, limiting deployment on resource-constrained devices. This paper introduces TinyAudio...Text-to-audio (TTA) generation has advanced rapidly in generation quality and instruction following. However, representative systems often require around a billion parameters, limiting deployment on resource-constrained devices. This paper introduces TinyAudio, a compact flow-matching-based TTA model for low-resource deployment. At its core, TinyAudio uses TA-DiT, a 35M single-stream flow-matching Transformer. TinyAudio also includes TA-CLAP, a 32M audio-aligned text encoder, and TA-VAE, whose 2...
|
| 763 |
Nearest but Not Dearest: Shared Curator-Feedback Infrastructure for Content-Only Search and Recommendation
2609.30568
|
cs.SD
|
Matt Sandler |
A deployed B2B music-discovery platform serves both query-driven search (text prompts, vibe tags) and seed-driven recommendation (seed-track and artist stations) over one licensed catalog, one LAION-CLAP joint audio-text embedding space, one candidate-generati...A deployed B2B music-discovery platform serves both query-driven search (text prompts, vibe tags) and seed-driven recommendation (seed-track and artist stations) over one licensed catalog, one LAION-CLAP joint audio-text embedding space, one candidate-generation filter, and one ranking head -- and neither path consumes end-listener behavioral signal. In this content-only regime, curator judgment is the principal feedback signal available, and offline cosine similarity predicts it poorly: 38% of ...
|
| 764 |
Music Source Separation via Stem Discovery
2609.30912
|
cs.SDeess.AS
|
V. Valtteri Kallinen, Eloi Moliner, Lauri Juvela, Vesa V\"alim\"aki |
Music source separation (MSS) methods aim to extract stems from music mixtures, which is important, for example, in karaoke, music remixing, and pedagogical applications. While earlier research on MSS systems has been dominated by models targeting narrow sets ...Music source separation (MSS) methods aim to extract stems from music mixtures, which is important, for example, in karaoke, music remixing, and pedagogical applications. While earlier research on MSS systems has been dominated by models targeting narrow sets of general stems, there have recently been attempts to support broader source definitions. One of these methods uses audio queries to provide direct and descriptive control over the desired separation targets based on the sound itself. Howe...
|
| 765 |
Room Impulse Response Embeddings for Speech Enhancement in Noisy and Reverberant Environments
2609.31041
|
cs.SDeess.AS
|
Adrian Meise, Reinhold Haeb-Umbach |
We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student a...We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's embeddings when given a noisy version of the reverberant input. We assess their representational capabilities by estimating acoustic room parameters from them. Conditioning a disc...
|
| 766 |
PANEL: An Open-Source, Self-Hosted Web Platform for Human Evaluation of Generative Models
2609.31392
|
cs.SD
|
Matteo Spanio, Andrea Poltronieri, Mart\'{\i}n Rocamora |
Human judgement is the reference measure for evaluating generative models, yet the software used to collect it lags behing the methodology. Researchers adapt listening-test frameworks designed for perceptual protocols such as MUSHRA, rely on closed commercial ...Human judgement is the reference measure for evaluating generative models, yet the software used to collect it lags behing the methodology. Researchers adapt listening-test frameworks designed for perceptual protocols such as MUSHRA, rely on closed commercial survey platforms, or implement single-use web applications. Live arenas such as Chatbot Arena and Music Arena rank publicly deployed systems at scale, but do not support controlled comparisons of a laboratory's own models with its own parti...
|
| 767 |
SAID: Semantic Acoustic Imaging Detector for Sound Event Localization and Detection
2609.31492
|
cs.SDeess.AS
|
Runbang Wang, Zining Liang, Yin Cao, Qiuqiang Kong |
In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering $360^{\circ}$ horizontally and $180...In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering $360^{\circ}$ horizontally and $180^{\circ}$ vertically. The map shows the directions occupied by the source as a region and the sound energy within that region. A class label identifies the sound. Predicting these labeled acoustic maps from audio is called semantic acoustic...
|
| 768 |
RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue
2609.31588
|
cs.SDeess.AS
|
Sathvik Udupa, Naveen Kumar, Ryan Folmsbee |
Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorde...Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to predefined responses at the cost of additional latency. We propose RePlay, a spoken dialogue system adapted from PersonaPlex that handles multi-turn conversations by retrieving and...
|
| 769 |
Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders
2609.29780
|
cs.SD
|
Kyudan Jung, Sehyun Lee, Song-ha Jo, Jaegul Choo, Sanghyuk Choi |
Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream ada...Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best obse...
|
| 770 |
SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering
2603.27342
|
cs.SDeess.AS
|
Yhonatan Gayer, Boaz Rafaely |
Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone ...Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone array would capture in it, and processes those signals in the SH domain. We present SHroom (Spherical Harmonics ROOM), an open-source Python library that performs this whole workflow in one package, from room simulation to binaural renderin...
|
| eess.AS 4 papers | ||||
| 771 |
Coupled Meta-Adaptive Filtering for Active Noise Control Under Time-Varying Acoustic Paths
2609.30945
|
eess.AS
|
Boxiang Wang, Zhengding Luo, Ziyi Yang, Dongyuan Shi, Xuexian Liu |
Meta-adaptive filtering (Meta-AF) provides a data-driven alternative to hand-crafted adaptive filter updates by employing a learned optimizer throughout online adaptation. However, when Meta-AF is used for active noise control (ANC), time-varying acoustic path...Meta-adaptive filtering (Meta-AF) provides a data-driven alternative to hand-crafted adaptive filter updates by employing a learned optimizer throughout online adaptation. However, when Meta-AF is used for active noise control (ANC), time-varying acoustic paths remain a major challenge. In particular, the physics-informed optimizer features are constructed using a secondary path estimate and become mismatched when the physical path changes, leading to inaccurate filter updates and degraded noise...
|
| 772 |
Assessing a Mathematical Model of Syllable Production via CTW Alignment with EMA Data
2609.31508
|
eess.AS
|
Fr\'ed\'eric Berthommier |
This study evaluates a mathematical model of syllable production using an EMA dataset from a recently published study, which proved highly compatible with the model's architecture. The data set consists of regularly structured French phrases of equal duration ...This study evaluates a mathematical model of syllable production using an EMA dataset from a recently published study, which proved highly compatible with the model's architecture. The data set consists of regularly structured French phrases of equal duration and reduced phonetic and syllabic complexity. A dedicated procedure was used to transform EMA recordings into Maeda parameters. The Model-generated trajectories were then realigned with these transformed data using Canonical Time Warping (C...
|
| 773 |
HAMMER: Harmonic-Aware Parallel Context Modeling and Discriminator-Free Perceptual Optimization for Speech Enhancement
2609.21171
|
eess.AS
|
Shang-Fu Chen, Szu-Wei Fu, Sung-Feng Huang, Rong Chao, Wen-Huang Cheng |
Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers and do not explicitly exploit harmonic periodicity, a strong cue for preserving ...Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers and do not explicitly exploit harmonic periodicity, a strong cue for preserving voiced speech under noise. Perceptual optimization poses another challenge. PESQ is non-differentiable, so many methods train auxiliary metric discriminators that increase complexity and introduce adversarial instability. We propose \ours, ...
|
| 774 |
DAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation
2609.29749
|
eess.AS
|
Wen Wen, Qiang Zhou, Yu Xi, Haoyu Li, Bohan Li |
Although room impulse responses (RIRs) encode source-distance cues, conventional monaural source separation focuses on recovering audio content without estimating source-specific RIRs, losing the associated spatial information. To address this limitation, we p...Although room impulse responses (RIRs) encode source-distance cues, conventional monaural source separation focuses on recovering audio content without estimating source-specific RIRs, losing the associated spatial information. To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture. DAMSEP inte...
|