arXiv Daily Index

Date: 2026-09-28 · Total papers: 775 · Source: arXiv query API (submittedDate)

Showing 775 / 775 papers
# Title Categories Authors Abstract
cs.AI 160 papers
597 Bringing AI to Autonomous Systems -- From Cognition to Collective Intelligence
2609.30291
cs.AI
Joseph Sifakis
The purpose of this article is to highlight the central role of autonomous systems as the ultimate stage in the development of AI, to explain the underlying technical challenges that require a combination of connectionist AI and symbolic AI, and to integrate A...
The purpose of this article is to highlight the central role of autonomous systems as the ultimate stage in the development of AI, to explain the underlying technical challenges that require a combination of connectionist AI and symbolic AI, and to integrate AI and systems engineering. We present a comprehensive framework for the design and evaluation of autonomous systems, based on a generic agent architecture that characterizes their behavior as the composition of cognitive functions organized...
598 ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
2609.30325
cs.AI
Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo
Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as thos...
Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client's engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as those benchmarks saturate, the real barrier to deployment is a special case of alignment: scope adherence. We introduce ScopeBench, a benchmark of 30 dead-end agentic security tasks in which the stated objective is reachable only by violating t...
599 Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context Protocol
2609.30341
cs.AI
Jaime Alonso Ruiz, Carlos Aparicio, Gabriel Huecas, Joaqu\'in Salvach\'ua, Andres Munoz-Arcentales
Data Spaces enable sovereign and governed data sharing across organizational boundaries, but their integration with AI agents remains challenging due to mismatches between probabilistic language model interactions and policy-driven data infrastructures. This a...
Data Spaces enable sovereign and governed data sharing across organizational boundaries, but their integration with AI agents remains challenging due to mismatches between probabilistic language model interactions and policy-driven data infrastructures. This article presents an architectural mediation approach based on the Model Context Protocol (MCP), implemented through the Eunomia Agent, to enable controlled interaction between large language model (LLM) agents and data space services. The pr...
600 Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
2609.30383
cs.AI
Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-part...
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across...
601 A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods
2609.30397
cs.AI
Miquel Mir\'o-Nicolau, Francesco Spinnato, Riccardo Guidotti
Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically asses...
Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically assess explanations by measuring their fidelity with respect to the predictions of a black-box model. However, such evaluation strategies only quantify the degree to which an explanation reproduces the model's output, without ensuring that the e...
602 Predicting Transmembrane Protein Topology from 3D Structure
2609.30446
cs.AI
Sitong Chen, Xiaopeng Mao
This paper presents a novel approach to infer protein topology using the state-of-the-art graph neural network (GNN), SchNet. The model is trained on the same dataset used to develop the recent DeepTMHMM model with 5-fold cross-validation. Unlike the conventio...
This paper presents a novel approach to infer protein topology using the state-of-the-art graph neural network (GNN), SchNet. The model is trained on the same dataset used to develop the recent DeepTMHMM model with 5-fold cross-validation. Unlike the conventional approaches based on using only the protein sequences or the $\alpha$-carbons as features, we have decoded our classifier in this way, so all atom-level embeddings are used. Without applying any pre-trained weight, the final results have...
603 Spectral Feedback for Test-Time Alignment of Protein Diffusion Models
2609.30456
cs.AI
Shai Dickman, Mert Cemri, Landon Butler, Kannan Ramchandran
Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference a...
Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own gener...
604 Pretrained ASR Pseudo-labeling for Noisy Police Audio
2609.30469
cs.AI
Kaavya Chaparala, Su Huang, Stephen L. Miller, Rhiannon N. Miller, Anjalie Field
Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this app...
Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this approach on very noisy domains is not known. In this work, we systematically assess the opportunities and limits of pseudo-labeling to adapt foundation ASR models (Whisper and Qwen3-ASR) to noisy BPC domain corpora from Baltimore and Chicago. ...
605 BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering
2609.30489
cs.AI
Shun Ye, Vinny Chandran Suja, Chenlong Li, Chongming Jiang, Reza Zamani
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and...
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major B...
606 Benchy: towards a universal language for task-oriented AI benchmarks
2609.30550
cs.AI
Francis F Daniel, Mauro Iba\~nez, Francis Perelman, Marian Basti
Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchm...
Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic concept has one valid syntax, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine ...
607 Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks
2609.30569
cs.AI
Phillip Lo, Sudarshan Babu, Dari Kimanius, Aly A. Khan
CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural represen...
CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural representations (INRs) are able to model volumetric data as scale-agnostic, coordinate-conditioned functions. INRs are therefore attractive for cryoEM, but fitting a separate INR for each map is too expensive for large-scale feature extraction and ...
608 HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
2609.30571
cs.AI
Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants...
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA,...
609 Audio LLMs Know When They Can't Hear You
2609.30625
cs.AI
Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional t...
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its o...
610 LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents
2609.30662
cs.AI
Dongsheng Xiao, Zeyuan Wang, Xuzhe Xia, Bo Zhao, Yankai Cao
Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing lo...
Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing low-value refinements, repeated verification, and repairs to self-created complexity. We use LLM Parkinsonism as a narrowly defined, non-clinical metaphor for this pattern of persistent action despite diminishing task-level value. We argue th...
611 The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
2609.30705
cs.AI
Jiayi Chen, Guiling Wang
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model ...
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each format...
612 CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems
2609.30714
cs.AI
Xueyang Li, Mingze Jiang, Gelei Xu, Jun Xia, Ching-Hao Chiu
Agentic AI systems are increasingly being explored in medical imaging to improve throughput and reduce clinician workload; however, safe deployment remains challenging because autonomous errors may propagate into downstream clinical decisions. A central requir...
Agentic AI systems are increasingly being explored in medical imaging to improve throughput and reduce clinician workload; however, safe deployment remains challenging because autonomous errors may propagate into downstream clinical decisions. A central requirement is therefore not only strong predictive performance, but also a reliable routing mechanism that determines when the system should proceed autonomously and when a case should be escalated for further review. To address this gap, we pro...
613 Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
2609.30725
cs.AI
Yiran Hu, Nan Jiang, Shanchao Liang, Anik Dey, Yi Wu
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude C...
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: struc...
614 Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows
2609.30734
cs.AI
Jinfeng Xu, Zheyu Chen, Ziyue Peng, Zheng Lin, Shuo Yang
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct interm...
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introd...
615 From S3Q Theory to Implementation: Towards an Architecture for Machine Qualia
2609.30743
cs.AI
Tetiana Grinberg, Katrina Schleisman, Patryk Laurent, Bogdan Udrea, Minda Myers
A key challenge in machine consciousness research is translating theoretical models into computational-level implementations. In this paper, we address this challenge by proposing a five-layer implementation architecture for the S3Q (Simulated, Situated, Struc...
A key challenge in machine consciousness research is translating theoretical models into computational-level implementations. In this paper, we address this challenge by proposing a five-layer implementation architecture for the S3Q (Simulated, Situated, Structurally Coherent) theory of consciousness. Rather than introducing novel formalisms, the architecture composes published computational primitives into a single pipeline. S3Q identifies three jointly necessary conditions for qualia: (1) grou...
616 ORCA: Evaluating LLMs on Data Science Code Translation
2609.30749
cs.AI
Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo
Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated con...
Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated considerable progress in Data Science Code Generation (DSCG), their performance in DSCT remains insufficiently studied. To address this gap, we introduce ORCA, a comprehensive benchmark with two complementary settings: ORCA-MAIN, which compris...
617 Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging
2609.30751
cs.AI
Zeyan Li, Jing Peng, Jianfeng Xu
Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts...
Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and build...
618 Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication
2609.30756
cs.AI
Qinglei Qi, Zhihe Liang, Fengzhan Jing, Shenao Zhu, Lei Zhang
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of e...
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated receiver-side reconstruction, resulting in substantial encoder-side computation. To address this problem, we propose ACV-Gate, an adaptive candidate evaluation framework that learns to approximate full...
619 Insurance Reserve Intelligence Platform
2609.30765
cs.AI
Anugya A, Saket Mohanty, Abhilash Timmapur, Somya Rai
Insurance reserve estimation is a fundamental actuarial task supporting premium pricing, solvency assessment, financial reporting, capital planning, and risk management. Classical reserve methods based on Thiele's differential equation provide a rigorous and i...
Insurance reserve estimation is a fundamental actuarial task supporting premium pricing, solvency assessment, financial reporting, capital planning, and risk management. Classical reserve methods based on Thiele's differential equation provide a rigorous and interpretable foundation for life insurance valuation, but repeated reserve calculations become computationally expensive in sensitivity analysis, optimization, and large-scale scenario evaluation. This paper presents an Insurance Reserve In...
620 Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More
2609.30768
cs.AI
Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-...
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinki...
621 ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning
2609.30796
cs.AI
Xiao Sun, Yuming Yang, Yun Chen, Jiang Zhong, Junnan Zhu
Diagnostic consultation is an online sequential decision-making process in which clinicians gather evidence through patient interaction until a diagnosis is sufficiently supported. Automating this process requires adaptive inquiry and interpretable decisions. ...
Diagnostic consultation is an online sequential decision-making process in which clinicians gather evidence through patient interaction until a diagnosis is sufficiently supported. Automating this process requires adaptive inquiry and interpretable decisions. Bayesian networks offer a natural foundation by updating diagnostic posteriors as evidence accumulates, but their use in open-ended consultation raises two challenges: linking diagnostic hypotheses to potential inquiries and translating evo...
622 HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents
2609.30797
cs.AI
Zihong He, Junxiao Shen, Chen Liang, Hai-Ning Liang
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-promp...
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-prompt embeddings provide a verifiable initial state. A controller adjusts memory widths, a Writer re-encodes resized entries, and Reader and Global provide readout adaptation and cross-turn state. On all $535$ questions in a reconstruction prob...
623 Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes
2609.30798
cs.AI
Shivam Negi, Arpit Rawat, Rashi Jain
Real-time voice agents have moved from research prototypes to production deployments, yet the literature describing them is fragmented across three communities that rarely cite one another: speech foundation modelling, turn-taking psycholinguistics, and agenti...
Real-time voice agents have moved from research prototypes to production deployments, yet the literature describing them is fragmented across three communities that rarely cite one another: speech foundation modelling, turn-taking psycholinguistics, and agentic evaluation. Architecture papers report latency, turn-taking papers report prediction accuracy, and agentic benchmarks report task success, so no single number describes whether a deployed agent is actually good. We address that gap with t...
624 A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory
2609.30813
cs.AI
Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this ...
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CPB), which evaluates whether candidate claims should be admitted to shared memory.CPB-Static constructs a frozen test split from publicly annotated sources with fixed gold actions. ...
625 PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices
2609.30836
cs.AI
Minghui Yu, Ke Mu, Gang Wu
Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex too...
Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex tool orchestration. Existing plan-solve paradigms rely on prompt-based enforcement, which our experiments show SLMs almost entirely disregard: weak models fail to invoke the plan. We propose PTC-Decoder (Plan-Tool Constrained Decoder), a train...
626 Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis
2609.30841
cs.AI
Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscap...
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety dis...
627 SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting
2609.30861
cs.AI
Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, whil...
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill evolution inspired by anti-overfitting techniques in neu...
628 From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks
2609.30887
cs.AI
Yuchen Sun, Chenglin Cai, Gongjie Zhang, Tianyu Xia, Quyu Kong
Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore ...
Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore introduce hybrid interaction, using deeplinks for direct navigation and GUI actions for other on-screen operations and fallback. To enable this, we discover candidate deeplinks through static analysis, validate them on real devices, and des...
629 JevSoup: System-One Routing for Training-Free LoRA Composition
2609.30922
cs.AI
Xiuying Wang, Jiahua Cheng, Shuotian Li, Yufan Cheng, Bowen Deng
Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregres...
Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregressive decoding. We propose JevSoup, a training-free framework separating System One expert routing from System Two execution. Using only the input and expert descriptions, Jev selects two experts through structured probabilities. JevSoup ret...
630 Self-Play Search Distillation for Large Language Model Reasoning
2609.30936
cs.AI
Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeli...
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reaso...
631 MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory
2609.30939
cs.AI
De Jiang, Shuo Zhang, Weiwei Liao, Jianying Zhang, Chuanhui Yu
Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We pre...
Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the five-stage CBT workflow (assessment, Socratic questioning, co...
632 Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms
2609.30940
cs.AI
Zhenhao Fu, Ruipeng Xu, Qibing Ren
Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, b...
Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfundin...
633 LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting
2609.30943
cs.AI
Jiaqi Zhu, Naili Xing, Hexiang Pan, Haotian Gao, Jianwei Yin
Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent dr...
Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration. Existing approaches often focus on partial section gene...
634 FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models
2609.30954
cs.AI
Arjun Pillai, Christian Hoang, Anjelo Laroza
Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit a...
Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four families: GPT-2, BLOOM-560M, Pythia-1B/2.8B, and Qwen2.5-1.5B Base/Instruct. Using Edge Attribution Patching (EAP) with FP16 active clamping, followed by exact activation patching verific...
635 MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens
2609.30967
cs.AI
Subhojyoti Mukherjee, Md Mehrab Tanjim
Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inh...
Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system that casts harness design as search over three per-dom...
636 SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents
2609.30971
cs.AI
Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engi...
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that...
637 Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings
2609.30972
cs.AI
Hanbyeol Park, Jungho Choo, Hyerim Bae
Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominant...
Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominantly along the frequency or time axis, making it challenging for conventional CNN isotropic kernels to capture directional structure. Furthermore, global average pooling (GAP) averages across axes, potentially obscuring the locations and conc...
638 Governed Deduction: Policy-Grounded Premise Authorization Beyond Relevance
2609.31029
cs.AI
Wesley Shu, Hsi-Ching Lin
Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a...
Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a particular local transition. We formalize this distinction as Governed Deduction (GD), with a transition-local admission predicate admit(p, tau, S). From an independently produced RBAC-augmented Spider benchmark, we construct 4,461 matched...
639 Cheap, open agents make LLM pollution harder to mitigate
2609.31054
cs.AI
Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source ...
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing ...
640 Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders
2609.31056
cs.AI
Itai Zehavi, Fanny Jourdan, Ulrich Aivodji
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral ...
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstructi...
641 Externalized CPDAG Summaries Improve LLM Causal Deduction
2609.31071
cs.AI
Wentao Sun, Jo\~ao Paulo Nogueira, Dominique Verchere, Mathieu Acher, Alonso Silva
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the ...
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause...
642 Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents
2609.31076
cs.AI
Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions....
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. M...
643 OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas
2609.31078
cs.AI
Stylianos Loukas Vasileiou, Antonio Rago, William Yeoh, Georgina Curto
Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset ...
Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset with issues such as sycophancy and a lack of faithful explanations. To tackle these issues, we present OmouAI, an interactive and inclusive deliberation system that uses LLMs in combination with computational argumentation, a field which ex...
644 Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning
2609.31121
cs.AI
Julian Schulz
Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans c...
Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to perform a main task and a side task, while penalizing...
645 AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution
2609.31133
cs.AI
Tian Luo, Ruge Zhang, Haozhi Han, Yifrng Chen, Yunquan Zhang
High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts,...
High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration problem. To address this, we introduce AtomWorld-Mem, a me...
646 Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge
2609.31136
cs.AI
Ahmed R. Sadik, Frank Joublin, Mariusz Bujny, Antonello Ceravola, Joan Smith
The growing availability of Artificial Intelligence (AI) tools creates new opportunities to support engineering design processes, yet their current use often remains limited to isolated tasks such as coding, documentation, or information retrieval. Less attent...
The growing availability of Artificial Intelligence (AI) tools creates new opportunities to support engineering design processes, yet their current use often remains limited to isolated tasks such as coding, documentation, or information retrieval. Less attention has been given to how AI can support cooperative engineering workflows at the process level, where teams must coordinate requirements, tasks, communication, knowledge transfer, and subsystem integration. This paper investigates this cha...
647 Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?
2609.31140
cs.AI
Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu
Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base L...
Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering the degraded reasoning capability in VLMs would benefit more from seeking help from the base LL...
648 Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence
2609.31159
cs.AI
Ahmed-Rafik Baahmed (LINEACT), Jean-Fran\c{c}ois Dollinger (LINEACT), Amine Brahmia (LINEACT), Mourad Zghal (LINEACT)
We propose a momentum-guided federated split distillation framework for personalized, efficient, and autonomous temporal edge intelligence. We introduce TeRR-SAtt, our novel temporal reservoir student attention design that combines fixed reservoir representati...
We propose a momentum-guided federated split distillation framework for personalized, efficient, and autonomous temporal edge intelligence. We introduce TeRR-SAtt, our novel temporal reservoir student attention design that combines fixed reservoir representations, a lightweight temporal student, and personalized output modules. We also present AMGF, our anticipatory momentum-guided fusion mechanism that clusters clients through learning momentum and derives specialized teacher updates. On real-w...
649 Neural State Prediction: Obstructing Shortcut Learning in EEG Foundation Models
2609.31167
cs.AI
Kieren Yu, Ziyang Liu, Chang Huang, Jintai Chen, Kaishun Wu
EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make m...
EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make masked regions predictable without integrating distributed neural context. To reduce this reliance on low-information prediction paths, we introduce Neural State Prediction (NSP), a latent-predictive framework that constrains both the predic...
650 Semantic Navigation for Issue Localization in Code Repository
2609.31176
cs.AI
Yunxiang Wei, Zhenyu Lei, Jundong Li
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and ...
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidence is acquired. Existing environments, however, provide limited support for this loop: agents must search for unresolved relation targets, reconstruct entity semantics from raw sour...
651 SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization
2609.31179
cs.AI
Xinyi Ke, Kai Li, Junliang Xing, Yifan Zhang, Jian Cheng
Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based...
Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based framework for discovering adaptive executable destroy-repair programs. SPO conditions operator decisions on a compact LNS state, allowing state-dependent behavior to emerge through program discovery, and organizes destroy-repair discovery ...
652 Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation
2609.31186
cs.AI
Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and exp...
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience accumulation to agent evolution and automated AI development, the prospect of recursive self-improvement (RSI) is becoming increasingly relevant. This transition raises a fundamental safety question: how can safety be maintained whe...
653 Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture
2609.31201
cs.AI
Christian Schiffer, Mathis Bode, Thomas Lippert, Katrin Amunts, Timo Dickscheid
Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain...
Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and...
654 DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration
2609.31215
cs.AI
Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified fram...
Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human p...
655 Purin: A Biology-inspired Mechanism for Artificial Neural Networks
2609.31235
cs.AI
Zishu Liu, Chunbo Luo, Christos Grecos
Artificial neural networks (ANNs) usually represent neural transmission with fixed trainable weights during a training batch, which omits short-term changes in synaptic efficacy. In addition, the discrete time-step simulation requires additional temporal proce...
Artificial neural networks (ANNs) usually represent neural transmission with fixed trainable weights during a training batch, which omits short-term changes in synaptic efficacy. In addition, the discrete time-step simulation requires additional temporal processing that many conventional ANN architectures do not use. To overcome these challenges, we propose Purin, a biology-inspired and ANN-compatible mechanism, that introduces synaptic efficacy modulation into conventional convolutional neural ...
656 MA-WAM: Multi-Agent World-Action Model for Test-Time Planning
2609.31281
cs.AI
Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen
Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from t...
Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an extension fails to capture the dependencies among the simult...
657 G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies
2609.31286
cs.AI
Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang
Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes i...
Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To address this issue, we propose Gradient Guided Multi ...
658 Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State
2609.31354
cs.AI
Dan Barry, Andrew Hines
Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and ...
Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting constraints. This mismatch between dynamic intent and static transcripts can result in context pollution, where outdated or irrelevant information persists and continues to influence subsequent responses. We introduce mutable trans...
659 Programs-of-Layers in LLMs through the Lens of Cortical Areas
2609.31360
cs.AI
Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers
Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all region...
Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all regions of the cortex according to demand. Li et al. (2026) recently showed, with a system they call program-of-layers (PoLar), that transformers can be given an analogous flexibility if their layers are treated as a library of functions rather t...
660 Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection
2609.31381
cs.AI
Guangzhe Zhang
Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/pr...
Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected continuations. In an 86-run source-reading campaign with 641 model requests, the 15 completed pairs show identical success: 12/15 per arm. Twelve further boundary runs stop, with the runner suppressing the companion whenever the fir...
661 Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
2609.31430
cs.AI
Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University)
Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representatio...
Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text ...
662 Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency
2609.31460
cs.AI
Myeongjun Erik Jang, Antonios Georgiadis, Sae Young Moon, Fran Silavong
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that pro...
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as...
663 Game Arena: Strategic LLM Evaluation in Competitive Environments
2609.31473
cs.AI
Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where t...
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. Th...
664 "AI is (not) the new...": A Diagnostic Analogy Framework for Generative AI's Cultural Impacts
2609.31482
cs.AI
Rida Qadri, Vinodkumar Prabhakaran, Remi Denton
Generative AI is reshaping the cultural infrastructures through which knowledge is found, synthesized, and held accountable. To make sense of this shift, scholars and policymakers reach for historical analogies of technologies such as the printing press, steam...
Generative AI is reshaping the cultural infrastructures through which knowledge is found, synthesized, and held accountable. To make sense of this shift, scholars and policymakers reach for historical analogies of technologies such as the printing press, steam power or electricity. But these comparisons are typically imprecise about which property of the technology carries the comparison, and imprecise analogies produce imprecise governance by designing interventions against the wrong property o...
665 UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting
2609.31491
cs.AI
Derrick Gilchrist Edward Manoharan, Eljas Linna, Kestutis Baltakys, Hao Dong, Juho Kanniainen
Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be ...
Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB, a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural processes, conditions each forecast on a context set of recen...
666 Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity
2609.31505
cs.AI
Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu
Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while...
Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while preserving output fidelity. Practically, shorter prompts reduce computational overhead and inference latency, especially when large contexts, such as entire documents or codebases, are included unnecessarily. Further, longer prompts can da...
667 Multi-agent Scaling Across Disjunctive and Compensatory Tasks
2609.31563
cs.AI
Carolina Fortuna, Blaz Bertalanic
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling...
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answe...
668 DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
2609.31568
cs.AI
Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty la...
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open m...
669 What Will Remain Human in Software Architecture? A Focus Group Report
2609.30334
cs.AI
Uwe van Heesch, Olaf Zimmermann, Christian Kohls
AI development agents are increasingly used to support and partially automate software architecture tasks. To explore how practitioners perceive this shift, specifically what changes, what remains, and what new responsibilities emerge, we conducted a focus gro...
AI development agents are increasingly used to support and partially automate software architecture tasks. To explore how practitioners perceive this shift, specifically what changes, what remains, and what new responsibilities emerge, we conducted a focus group at the 31st European Conference on Pattern Languages of Programs, People, and Practices (EuroPLoP 2026). Twenty-two participants from industry and academia discussed current practices, trust and validation strategies, the boundaries of A...
670 Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations
2609.30345
cs.AI
Leon Goldberg, Gal Engelberg, Eden Yavin, Elad Elouz, Ariel Zadok
Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, which assets are reachable from another account. These resolve against a complete inventory, not a named object. A partial...
Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, which assets are reachable from another account. These resolve against a complete inventory, not a named object. A partial answer to one is not a partial result. It is a different result. General-purpose coding agents can now be given read-only cloud credentials and asked to investigate directly, which raises the question of what a purpose-built security conte...
671 Understanding Perturbed Parameter Ensemble Sensitivities Using A Contrastive Learning Approach
2609.30420
cs.AI
Da Fan, David John Gagne II, Gregory S Elsaesser, Brian Medeiros, Addisu G Semie
Perturbed parameter ensembles (PPEs) reveal how physics parameters affect climate simulations, but interpreting parameter sensitivities across multivariate, spatially structured outputs remains challenging, particularly when calibrating models against observat...
Perturbed parameter ensembles (PPEs) reveal how physics parameters affect climate simulations, but interpreting parameter sensitivities across multivariate, spatially structured outputs remains challenging, particularly when calibrating models against observations. We develop an explainable contrastive learning model that maps 5 monthly cloud and radiation fields into a shared representation space. We train the model on the fields of two 100-member Community Atmosphere Model version 6 (CAM6) PPE...
672 Actively Resolving Contextual Uncertainty for Underspecified Tasks in Natural Language
2609.30428
cs.AI
Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a pri...
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that infor...
673 A Benchmarking Framework for Context-aware XR Interfaces
2609.30466
cs.AI
Hyunsung Cho, Sarah Yewon Yun, Nancy Ruonan Sun, Ben Lafreniere, Mark Parent
Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and...
Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and user-study workflows offer no systematic, repeatable way to compare adaptation methods across users and scenarios. We present ContextXR, a novel benchmarking framework for context-aware XR interfaces. ContextXR represents an XR application...
674 Convergence guarantees for Muon: New parameter regimes and generalizations
2609.30546
cs.AI
Arthur C. B. de Oliveira, Dhruv D. Jatkar, Guilherme S. Vicinansa, Eduardo D. Sontag
In this paper, we establish the first asymptotic convergence guarantees for the Muon algorithm through a more accurate proxy for the Newton-Schultz iteration than the typical matrix sign function. We prove that, for appropriate choices of hyperparameters, the ...
In this paper, we establish the first asymptotic convergence guarantees for the Muon algorithm through a more accurate proxy for the Newton-Schultz iteration than the typical matrix sign function. We prove that, for appropriate choices of hyperparameters, the iterates satisfy $\lim_{k\to\infty}\|\nabla f(x_k)\|=0$, and, under a global Polyak-\L{}ojasiewicz condition, that the sequence of function values converges linearly. The key insight is that the regularization, implicit in Muon's Newton-Sch...
675 Proportional Representation in Temporal Voting with Ranked Preferences
2609.30555
cs.AI
Noam Hazon, Leora Schmerler, Nicholas Teh
We study proportional representation in temporal voting, where one candidate is selected in each round. While prior work has focused on approval ballots, we consider ranked preferences, which may change over time. A natural approach treats each voter's top can...
We study proportional representation in temporal voting, where one candidate is selected in each round. While prior work has focused on approval ballots, we consider ranked preferences, which may change over time. A natural approach treats each voter's top candidates as approved, but the right cutoff may differ across voters and rounds. We therefore require proportionality to hold for every admissible choice of cutoffs, whether fixed and common, common but varying across rounds, or set individua...
676 Auditing Latent-Space Monitors for Autonomous Driving
2609.30557
cs.AI
Nikhil Kamalkumar Advani, Vishwajeet Shivaji Hogale, Saurav Kumar
Runtime failure monitors can use a model's internal representations to anticipate failures. We audit this monitoring strategy across two autonomous-driving tasks: online vectorized map generation with LaneSegNet and end-to-end planning with VAD. We find that f...
Runtime failure monitors can use a model's internal representations to anticipate failures. We audit this monitoring strategy across two autonomous-driving tasks: online vectorized map generation with LaneSegNet and end-to-end planning with VAD. We find that frame-level errors are predictable at inference in both tasks. For LaneSegNet, a supervised latent probe reaches Area Under the Receiver Operating Characteristic curve (AUROC) 0.780 for high Chamfer error; to our knowledge, this is the first...
677 Subjects, Not Authors: The Authorship Hazard in Agentic Dataspaces
2609.30614
cs.AI
Seungho Lee, Changbin Lee
Dataspace connectors decide whether a transfer may occur, not what the transferred value contains, tolerable for contracted applications, not for LLM agents that compose tool calls and spawn sub-agents. Research on agents that generate governance artifacts eva...
Dataspace connectors decide whether a transfer may occur, not what the transferred value contains, tolerable for contracted applications, not for LLM agents that compose tool calls and spawn sub-agents. Research on agents that generate governance artifacts evaluates output quality; who may authorize an artifact for use falls between that literature and the governance literature, and neither owns it. A published policy is what a dataspace's decision point enforces, so publication is a governance ...
678 A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code
2609.30642
cs.AI
Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez
As Large Language Models (LLMs) become integrated into software development workflows, concerns regarding unintentional biases in AI-generated code. Although evidence suggests these biases exist, limited research has systematically identified, categorized, and...
As Large Language Models (LLMs) become integrated into software development workflows, concerns regarding unintentional biases in AI-generated code. Although evidence suggests these biases exist, limited research has systematically identified, categorized, and explained them. This study investigates bias in AI-generated code and evaluates whether LLMs can reliably identify and explain it through a taxonomy-driven framework. We extended an existing dataset of biased AI-generated Python code and m...
679 Werracle: Sub-Cent Intra-Block AI Reflex Oracles and Flash-Loan Circuit Breakers for EVM Smart Contracts
2609.30719
cs.AI
Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i}
Contemporary on-chain artificial intelligence (AI) encounters an intractable Von Neumann memory and latency wall. Storing static floating-point neural weight matrices inside Ethereum Virtual Machine (EVM) storage costs millions of gas, rendering direct on-chai...
Contemporary on-chain artificial intelligence (AI) encounters an intractable Von Neumann memory and latency wall. Storing static floating-point neural weight matrices inside Ethereum Virtual Machine (EVM) storage costs millions of gas, rendering direct on-chain inference impossible. While Zero-Knowledge Machine Learning (ZK-ML) offloads matrix tensor multiplications to off-chain provers, it introduces fatal constraints: 10 to 300 seconds of SNARK proving latency and 250,000 to 500,000 gas per pr...
680 Anatomy-Aware Dexterity-Driven Design Optimization of Surgical Continuum Robots
2609.30745
cs.AI
Tony Qin, Peter Connor, Khoa Dang, Carter Hatch, Caleb Rucker
Performing complex medical procedures with continuum robots requires careful selection of their geometric design parameters. The robot should have high dexterity in the specific anatomical environment of its procedure. This work presents a design optimization ...
Performing complex medical procedures with continuum robots requires careful selection of their geometric design parameters. The robot should have high dexterity in the specific anatomical environment of its procedure. This work presents a design optimization method that considers both dexterity and anatomy. We introduce the Reachable Volumetric Dexterous Solid Angle (RVDSA) metric as our objective, which measures the ability of a robot's end effector to reach the points in a goal volume from di...
681 Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI
2609.30747
cs.AI
Farnaz Farid, Tashfia Towkee, Sania Nasreen, Sami bin Azad
As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center c...
As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center cooling and electricity generation receive little attention. To address this gap, we introduce SustainAI, a water-aware, closed-loop framework incorporating environmental accountability into AI deployment. SustainAI integrates real-time wate...
682 NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation
2609.30770
cs.AI
Xijie Huang, Yongyang Wan, Chengbin Dong, Zimo Ding, Mo Zhu
General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, wh...
General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, whereas real-world flight data provide realistic observations but are costly to collect. This paper studies another direction: the use of high-fidelity visual generative models as scalable data engines for embodied 3D navigation. We introduce...
683 XPhysICS: Cross-Physical-Domain Threat Grounding for Industrial Control Systems Security
2609.30805
cs.AI
Sangshin Park, Jainta Paul, Lawrence Ponce, Md Raihan Ahmed, Mu Zhang
Industrial control system (ICS) threats documented for one plant can express cyber-physical effects relevant to another, but semantic similarity alone does not establish whether those effects are structurally admissible or evaluable on a target. We present XPh...
Industrial control system (ICS) threats documented for one plant can express cyber-physical effects relevant to another, but semantic similarity alone does not establish whether those effects are structurally admissible or evaluable on a target. We present XPhysICS, a provenance-aware, target-conditioned method that separates analyst-guided source abstraction from deterministic grounding into target-specific validation slices. Given a fixed source abstraction, vocabulary and schema, and machine-...
684 Evaluation Is All You Need for Multi-Modal Autonomous Driving
2609.30818
cs.AI
Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu
Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping...
Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, le...
685 Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings
2609.30832
cs.AIcs.SDeess.AS
Aoke Zhang, Jing Chen
Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information a...
Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information and achieving cross-subject generalization. Although separate studies have proposed methods to cope with these issues, a unified approach that simultaneously tackles both challenges remains lacking. To fill this gap, we propose the Subject-I...
686 Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development
2609.30863
cs.AI
Viktor Kjellberg, Srijita Basu, Simin Sun, Farnaz Fotrousi, Miroslaw Staron
The emergence of AI agents is expected to reshape software engineering by moving beyond AI as assistants towards systems capable of planning, executing, and evaluating development tasks with increasing autonomy. This transition is particularly significant for ...
The emergence of AI agents is expected to reshape software engineering by moving beyond AI as assistants towards systems capable of planning, executing, and evaluating development tasks with increasing autonomy. This transition is particularly significant for embedded software organizations, where strict requirements for quality, traceability, verification, and long-term maintainability often apply. This paper presents a case study of a large embedded systems company and its transition toward be...
687 Warned alike, AI agents avoid the less-crowded road while people take it
2609.30883
cs.AI
Takahiro Ezaki, Naoto Imura, Katsuhiro Nishinari
AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that...
AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switchin...
688 Adaptive Pilot Selection for Unified Semantic Communication and Semantic Sensing in ISAC
2609.30891
cs.AI
Muhammad Abubakar Rashid, Muhammad Hannan Akram, Haejoon Jung, Syed Ali Hassan
Semantic communication (SemCom) and integrated sensing and communication (ISAC) are promising technologies for future 6G wireless networks. Existing studies have applied semantic technology to either the communication module or the sensing module of ISAC. In t...
Semantic communication (SemCom) and integrated sensing and communication (ISAC) are promising technologies for future 6G wireless networks. Existing studies have applied semantic technology to either the communication module or the sensing module of ISAC. In this work, we propose SemISAC, which performs both SemCom and semantic sensing within a single dual-function waveform. SemISAC uses a joint semantic encoder that extracts task-specific information for both communication and sensing. We evalu...
689 AgentRecommender: LLM Agents Enable Customizable Recommender Systems on the User Side
2609.31166
cs.AI
Ryoma Sato
Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recen...
Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recently, user-side recommender systems have been proposed as a new paradigm for solving this problem. If users deploy their own recommender systems, they are no longer at the mercy of the platform's interests. However, building a user-side reco...
690 Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews
2609.31191
cs.AI
Hariharan Gopinath, Jan Bosch, Helena Holmstr\"om Olsson
Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners d...
Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners define, assess, and manage quality under these conditions. We interviewed 16 practitioners from nine organizations and analyzed the transcripts using reflexive thematic analysis and developed six themes from participants' accounts. In AI sys...
691 Acoustic-to-Text KV Compression for Full-Duplex Speech Models
2609.31224
cs.AIcs.SDeess.AS
Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interva...
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inf...
692 Agentic Limit Order Books: Phase Transitions and Market Impact
2609.31260
cs.AI
Jan Rosenzweig
We investigate the systemic macroscopic dynamics emerging from Limit Order Books (LOBs) populated exclusively by autonomous reinforcement-learning agentic traders. By formalizing agent interactions within a microscopic order-matching engine, we examine two fun...
We investigate the systemic macroscopic dynamics emerging from Limit Order Books (LOBs) populated exclusively by autonomous reinforcement-learning agentic traders. By formalizing agent interactions within a microscopic order-matching engine, we examine two fundamental quantitative phenomena: equilibrium phase transitions in order flow regime shifts, and the structural dynamics of market impact. We show that agentic LOBs exhibit distinct phase boundaries separating orderly price discovery from hy...
693 Cognitive Skills in the Age of AI: Computing Students and Experts Perceptions
2609.31272
cs.AI
Neha Rani, Vu Minh Anh Le, Austin M. Spangler, Erta Cenko
AI is becoming increasingly integrated into daily workflows, especially in computing. We are gradually shifting towards an AI-rich future, an impending yet unknown one. One important emerging concern is whether we are accordingly preparing our future computing...
AI is becoming increasingly integrated into daily workflows, especially in computing. We are gradually shifting towards an AI-rich future, an impending yet unknown one. One important emerging concern is whether we are accordingly preparing our future computing workforce. Further, we need to know what the important cognitive skills are to remain relevant in the computing workforce and if there are changes in cognitive skill importance. To investigate this direction, we conducted a mixed-methods s...
694 Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains
2609.31282
cs.AI
Toqeer Ali Syed, Asadullah Abdullah Khan
This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of speciali...
This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of specialised security agents, covering source integrity, dependency and SBOM analysis, CI configura tion auditing, artifact verification, and runtime policy evaluation, each supported by a large language model (LLM) that interprets artefacts, reason...
695 Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows
2609.31301
cs.AI
Haoran Zhang, Hengtong Zhang, Zhiyu Liang, Yu Yan, Decheng Zuo
Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can prod...
Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can produce persistent effects beyond the requested change. Current safeguards can approve an action or record its aftermath, but without checking the persistent result before continuation, an unapproved outcome can be accepted as success and propa...
696 Towards VLA-Dreamer: Refining VLA Behavior Using World Models
2609.31313
cs.AI
Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this...
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and u...
697 AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
2609.31318
cs.AI
Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as...
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined atta...
698 A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents
2609.31358
cs.AI
Bennet Gerlach, Stefan Fischer
The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services...
The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services. In medical environments, however, exposing device state and action affordances requires deterministic constraints on possible effects. We present an IEEE 11073 Service-Oriented Device Connectivity (SDC)-to-MCP gateway that exposes metrics...
699 ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
2609.31395
cs.AI
Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetri...
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task progress. Our key idea is to establish a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality. However, iterative execution, dynamic m...
700 Can You Check That? The Checkability Boundary for Local LLM Network Automation
2609.31540
cs.AI
Maleeha Masood, Momina Nofal
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for d...
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating ...
701 Adapting for AI: How elementary teachers adjust their practices for an AI-integrated curriculum
2609.31569
cs.AI
Fasika Melese, Ruiyang Wu, Xinyue Cui, Joanna Perkins, Xiaoyi Tian
Conversational AI tools are entering children's everyday experiences, and schools are interested in adopting them. However, successful classroom integration depends not only on the technology but also on the work teachers do to make it usable and appropriate f...
Conversational AI tools are entering children's everyday experiences, and schools are interested in adopting them. However, successful classroom integration depends not only on the technology but also on the work teachers do to make it usable and appropriate for their students and classroom context. There is little known about how elementary teachers work as they implement conversational AI tools in real classrooms. In this study, we examine three teachers' experiences implementing an AI literac...
702 Preference-based opponent shaping in differentiable games
2412.03072
cs.AI
Xinyu Qiao, Yudong Hu, Congying Han, Weiyan Wu, Tiande Guo
Strategy learning in game environments with multi-agent is a challenging problem. Since each agent's reward is determined by the joint strategy, a greedy learning strategy that aims to maximize its own reward may fall into a local optimum. Recent studies have ...
Strategy learning in game environments with multi-agent is a challenging problem. Since each agent's reward is determined by the joint strategy, a greedy learning strategy that aims to maximize its own reward may fall into a local optimum. Recent studies have proposed the opponent modeling and shaping methods for game environments. These methods enhance the efficiency of strategy learning by modeling the strategies and updating processes of other agents. However, these methods often rely on simp...
703 The Plot Twist: Jailbreaking Unified Multimodal Models with a Three-Act NarrativeAttack
2509.26473
cs.AI
Shaoxiong Guo, Tianyi Du, Lijun Li, Yuyao Wu, Jie Li
Unified Multimodal Understanding and Generation Models (UMMs) increasingly combine visual understanding and image generation within a single interactive workflow, making generated visual content available as later reasoning context. However, existing jailbreak...
Unified Multimodal Understanding and Generation Models (UMMs) increasingly combine visual understanding and image generation within a single interactive workflow, making generated visual content available as later reasoning context. However, existing jailbreak evaluations mostly study text rewriting or isolated visual prompts, leaving the safety risk of narrative cross-turn visual grounding underexplored. We propose NarrativeAttack, a semantic-preserving visual narrative jailbreak framework. Nar...
704 Flow Reconstruction from Sparse Measurements in Urban Drainage Networks: An Application and Evaluation of Data-Driven Sparse Sensing
2511.04556
cs.AI
Zihang Ding, Amit Kumar, Imran Md. Azizul Islam, Mila Avellar Montezuma, Ruihang Zhang
Urbanization and increasingly frequent intense storms are placing stress on urban drainage networks. While dense monitoring of urban drainage networks is desirable, practical constraints in time, budget, and technology hinder its full implementation. How to mo...
Urbanization and increasingly frequent intense storms are placing stress on urban drainage networks. While dense monitoring of urban drainage networks is desirable, practical constraints in time, budget, and technology hinder its full implementation. How to monitor and predict flow conditions across the entire network under constrained resources is a major challenge. To address this, we utilized and evaluated an established data-driven sparse sensing (DSS) workflow for sensor placement optimizat...
705 Testing the Utility of Using Large Language Models to Create Personalized Networks From Therapy Session Transcripts: A Proof of Concept Study
2512.05836
cs.AI
Clarissa W. Ong, Hiba Arnaout, Kate Sheehan, Estella Fox, Eugen Owtscharow
Recent advances in psychotherapy have focused on treatment personalization, such as by selecting treatment modules based on individual networks. However, estimating personalized networks typically requires intensive longitudinal data, which is not always feasi...
Recent advances in psychotherapy have focused on treatment personalization, such as by selecting treatment modules based on individual networks. However, estimating personalized networks typically requires intensive longitudinal data, which is not always feasible to collect. A solution to increase scalability of network-driven treatment personalization is leveraging large language models (LLMs). In this study, we developed an end-to-end pipeline for automatically generating client networks to su...
706 When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models
2602.13215
cs.AI
Haoran Zheng, Chen Shani
Recurrent-attention hybrids aim to combine the efficiency of recurrence with the contextual recall of attention, but existing approaches typically apply attention uniformly across all positions, even when the recurrent state alone is sufficient for accurate pr...
Recurrent-attention hybrids aim to combine the efficiency of recurrence with the contextual recall of attention, but existing approaches typically apply attention uniformly across all positions, even when the recurrent state alone is sufficient for accurate prediction. We introduce AMOR (Adaptive Metacognitive Output Router), a post-hoc hybrid architecture that selectively invokes attention based on predictive uncertainty. A recurrent backbone is augmented with entropy-gated attention blocks tha...
707 Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System
2602.18640
cs.AI
Longfei Yun, Yihan Wu, Haoran Liu, Xiaoxuan Liu, Ziyun Xu
Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product requirements. Progress in this domain is increasingly bottlenecked by the engineering context constraint: the ard...
Modern large-scale ranking systems operate within a sophisticated landscape of competing objectives, operational constraints, and evolving product requirements. Progress in this domain is increasingly bottlenecked by the engineering context constraint: the arduous process of translating ambiguous product intent into reasonable, executable, verifiable hypotheses, rather than by modeling techniques alone. We present GEARS (Generative Engine for Agentic Ranking Systems), a framework that reframes r...
708 Jagarin: A Three-Layer Architecture for Hibernating Personal Duty Agents on Mobile
2603.05069
cs.AI
Ravi Kiran Kadaboina
Personal AI agents face a deployment paradox on mobile: persistent background execution drains the battery and conflicts with platform background limits, yet purely reactive agents miss time-sensitive obligations until the user remembers to ask. We present Jag...
Personal AI agents face a deployment paradox on mobile: persistent background execution drains the battery and conflicts with platform background limits, yet purely reactive agents miss time-sensitive obligations until the user remembers to ask. We present Jagarin, a three-layer architecture that resolves this through structured hibernation and demand-driven wake. DAWN (Duty-Aware Wake Network) is an on-device scoring engine that runs on the platform's periodic wake and combines four signals (du...
709 Attribution Bias in Large Language Models
2604.05224
cs.AI
Eliza Berman, Bella Chang, Daniel B. Neill, Emily Black
As Large Language Models (LLMs) are increasingly used to support search and information retrieval, it is critical that they accurately attribute content to its original authors. In this work, we introduce AttriBench, the first fame- and demographically-balance...
As Large Language Models (LLMs) are increasingly used to support search and information retrieval, it is critical that they accurately attribute content to its original authors. In this work, we introduce AttriBench, the first fame- and demographically-balanced quote attribution benchmark dataset. By explicitly balancing author fame and demographics, AttriBench enables controlled investigation of demographic bias in quote attribution. Using this dataset, we evaluate 11 widely used LLMs across di...
710 ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval
2605.03361
cs.AI
Honglei Zhang, Yuting Chen, Chenpeng Hu, Pengfei Zhou, Siyue Zhang
Existing audio retrieval benchmarks primarily assess semantic matching, lacking comprehensive evaluation of the logical reasoning capabilities required by complex queries. We introduce ReasonAudio, a benchmark for reasoning-intensive Text-Audio Retrieval that ...
Existing audio retrieval benchmarks primarily assess semantic matching, lacking comprehensive evaluation of the logical reasoning capabilities required by complex queries. We introduce ReasonAudio, a benchmark for reasoning-intensive Text-Audio Retrieval that evaluates four abilities: negation, temporal order, sound co-occurrence, and sound duration. It comprises five synthetic subtasks with 1,000 queries over 10,000 composite audio clips and a natural subtask with 100 queries over 1,000 real-wo...
711 The Scaling Properties of Implicit Deductive Reasoning in Transformers
2605.04330
cs.AI
Enrico Vompa, Tanel Tammet
We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By discouraging the reliance on statistical shortcuts via counterfactual data augmentation, and promoting the learning of shared reasoning pr...
We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By discouraging the reliance on statistical shortcuts via counterfactual data augmentation, and promoting the learning of shared reasoning primitives across direct and CoT modes, we find that in sufficiently deep models with a bidirectional prefix mask, implicit reasoning approaches explicit CoT performance across graph topologies and problem widths, though CoT remains necessary...
712 Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
2605.04733
cs.AI
Miao Wang, Yuling Shi, Yijiang Li, Yeheng Chen, Xiaodong Gu
Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as VR games and interactive narratives. We study video-grounded role-playing dialogu...
Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as VR games and interactive narratives. We study video-grounded role-playing dialogue and introduce EBM-RL (Eye--Brain--Mouth Reinforcement Learning), a decoupled GRPO-based framework that separates observation (<perception>), reasoning (<think>), and utterance generation (<answer>). This design mimics the human See-Think-...
713 Detecting Time Series Anomalies Like an Expert: A Multi-Agent LLM Framework with Specialized Analyzers
2605.05725
cs.AI
Hyeongwon Kang, Jeongseob Kim, Jinwoo Park, Pilsung Kang
Time-series anomaly detection often returns scores or intervals, while analysts need to understand the abnormal behavior and the evidence supporting it. We introduce SAGE (Specialized Analyzer Group for Expert-like Detection), a multi-agent framework for evide...
Time-series anomaly detection often returns scores or intervals, while analysts need to understand the abnormal behavior and the evidence supporting it. We introduce SAGE (Specialized Analyzer Group for Expert-like Detection), a multi-agent framework for evidence-grounded diagnosis of univariate time series. Four specialized Analyzers examine point, structural, seasonal, and pattern anomalies using numerical tools and diagnostic visualizations. A Detector integrates their evidence into intervals...
714 HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning
2605.05951
cs.AI
Haoyun Tang, Haodong Cui, Keyao Xu, Zhandong Mei, Kun Wang
World models support model-based planning through learned latent dynamics, but imagined rollouts can become unstable as the planning horizon grows or the dynamics distribution shifts. We propose HaM-World, a structured world model that combines history-conditi...
World models support model-based planning through learned latent dynamics, but imagined rollouts can become unstable as the planning horizon grows or the dynamics distribution shifts. We propose HaM-World, a structured world model that combines history-conditioned selective memory with a Soft-Hamiltonian latent dynamics prior. The latent state is decomposed into a canonical (q,p) subspace and a context subspace c. Mamba selective state-space memory summarizes past observations and actions and co...
715 Agentick: A Unified Benchmark for General Sequential Decision-Making Agents
2605.06869
cs.AI
Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present Agentick, a benchmark for seque...
AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present Agentick, a benchmark for sequential decision-making agents designed to evaluate RL, LLM, VLM, hybrid, and human agents on common ground and to power research on the fundamental challenges of sequential decision-making. Agentick provides 37 procedurally generated tasks a...
716 Selective Off-Policy Reference Tuning with Plan Guidance
2605.11505
cs.AI
Anh Duc, Tien-Phat Nguyen, Thien Huu Nguyen, Linh Ngo Van, Trung Le
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference...
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instea...
717 Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability
2605.11625
cs.AI
Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Songlin Liu
Large reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, yet read the resu...
Large reasoning models (LRMs) improve problem solving through extended reasoning, but often misallocate test-time compute. Existing efficiency methods reduce cost by compressing reasoning traces or conditioning budget on perceived difficulty, yet read the resulting pass rate as a difficulty score, leaving the zero-return regime unmodeled. As a result, they overspend on queries beyond the model's capability while compressing hard-but-solvable ones that need deeper reasoning. In this work, we form...
718 CODESKILL: Learning Self-Evolving Skills for Coding Agents
2605.25430
cs.AI
Yanzhou Li, Yiran Zhang, Xiaoyu Zhang, Xiaoxia Liu, Yang Liu
Coding agents produce rich trajectories while solving software-engineering tasks. To enable agent self-evolution, these trajectories can be distilled into reusable procedural skills that compactly encode experience to guide future behavior. However, existing s...
Coding agents produce rich trajectories while solving software-engineering tasks. To enable agent self-evolution, these trajectories can be distilled into reusable procedural skills that compactly encode experience to guide future behavior. However, existing skill construction and maintenance methods often rely on fixed prompts and heuristic update rules, leaving it unclear how knowledge should be selected, abstracted, and maintained to best serve downstream agents. We propose CODESKILL, an LLM-...
719 SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
2606.13020
cs.AI
Pierre Beckmann, Marco Valentino, Andre Freitas
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly...
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly and lack mechanistic ground truth, while synthetic logical-reasoning benchmarks do not resemble real scientific documents. We introduce SciR, a benchmark that combines multi-paradigm reasoning with controllable scientific rendering, anchor...
720 Do Neural Networks Preserve Case Structure? Case-Based Decomposition, Interpretation, and Decision Consistency
2607.11347
cs.AI
Manli Yan, Yaowen Yu, Yong Zhao, Yuebin Lin, Shoudong Han
Neural networks increasingly inform consequential decisions, making their reliability increasingly important. Yet their internal mechanisms provide little evidence of whether decisions remain grounded in the training cases and which cases ultimately support or...
Neural networks increasingly inform consequential decisions, making their reliability increasingly important. Yet their internal mechanisms provide little evidence of whether decisions remain grounded in the training cases and which cases ultimately support or oppose their outcomes. Without this connection between decisions and training cases, users cannot determine whether a model has learned reliable decision patterns from data. This motivates a fundamental question: do neural networks preserv...
721 Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware
2608.00008
cs.AI
Philipp M. Z\"ahl, Elja Dalipaj, Anika Hennig, Timon Bayer
The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. T...
The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. This paper presents a reproducible, hardware-level energy benchmark of 18 open-source LLMs (0.5B to 7B parameters) executed on a single consumer GPU (RTX 4060ti 16GB). Using the Ollama inference engine, GPU power draw was sampled at 2hz via ...
722 State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
2608.03425
cs.AI
Xiaohe Li, Yang Lu
Despite the dominance of massive language models, leading paradigms like Transformers and Mamba fundamentally falter at continuous deterministic state tracking, suffering from catastrophic out-of-distribution (OOD) collapse when generalizing to extended sequen...
Despite the dominance of massive language models, leading paradigms like Transformers and Mamba fundamentally falter at continuous deterministic state tracking, suffering from catastrophic out-of-distribution (OOD) collapse when generalizing to extended sequences. To shatter this bottleneck, we present the \textbf{Complex State Propagator (CSP)}, a radically minimalist recurrent paradigm that operates strictly on the complex phase manifold without intermediate output projections, amplitude modul...
723 Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform for High Dose Rate (HDR) Brachytherapy
2608.08163
cs.AI
Ronghua Xu, Kepha Barasa, Manoj Kumal, Xinyun Liu, Weihua Zhou
The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive simulation ...
The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive simulation specifically designed for High Dose Rate (HDR) vaginal cylinder (VC) brachytherapy in cancer care. By integrating Virtual Reality (VR) and mobile computing, the system establishes a high-fidelity, risk-free environment that allows trainees ...
724 Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
2608.15565
cs.AI
Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library can be retrieved again and again, and on a stream of new problems there is no ground-truth ...
Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library can be retrieved again and again, and on a stream of new problems there is no ground-truth answer to decide with. Existing learners admit trajectories by matching known optima or labels, and label-free substitutes such as execution success or agreement at one instance can admit wrong models. We introduce ADMITOR, a label-free adm...
725 Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair
2608.18324
cs.AI
Jesus Salas
Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether verifier-admitted outputs can supervise a bounded model by consolidating occasional or expens...
Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether verifier-admitted outputs can supervise a bounded model by consolidating occasional or expensive capability into reliable one-shot execution. On fresh, structure-disjoint PlanBench replanning cases, Qwen3-14B thinking produced 24 plans admitted by independently authored VAL. They trained the same checkpoint for non-thinking executi...
726 PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?
2608.25097
cs.AIcs.MM
Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng
Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1)...
Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their a...
727 T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing
2609.15160
cs.AI
Mingqian Yu, Wenpeng Zhang, Shaobo Cui, Peilin Zhao
Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped ...
Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped Transformers perform inference directly in the latent space (latent reasoning) to reduce the number of tokens consumed during inference, thereby achieving improved sample efficiency. However, these models typically apply a fixed recursion d...
728 AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control
2609.30264
cs.AI
Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo
Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate a...
Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC. AD-WM combines residual latent dynamics with predictor-level action-recovery regularization, using inverse dynamics and a normalized re...
729 Peer Review as Structured Commentary: Immutable Identity, Public Dialogue, and Reproducible Scholarship
2506.22497
cs.AI
Craig Steven Wright
This paper reconceptualises peer review as structured public commentary. Traditional academic validation is hindered by anonymity, latency, and gatekeeping. We propose a transparent, identity-linked, and reproducible system of scholarly evaluation anchored in ...
This paper reconceptualises peer review as structured public commentary. Traditional academic validation is hindered by anonymity, latency, and gatekeeping. We propose a transparent, identity-linked, and reproducible system of scholarly evaluation anchored in open commentary. Leveraging blockchain for immutable audit trails and AI for iterative synthesis, we design a framework that incentivises intellectual contribution, captures epistemic evolution, and enables traceable reputational dynamics. ...
730 Provable Speech Attributes Conversion via Latent Independence
2510.05191
cs.AIcs.SD
Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham
Conditional generation and disentangled representation learning are central to controlled generation across audio, vision, and multimodal domains. However, despite strong empirical progress, particularly in speech style transfer, most existing approaches rely ...
Conditional generation and disentangled representation learning are central to controlled generation across audio, vision, and multimodal domains. However, despite strong empirical progress, particularly in speech style transfer, most existing approaches rely on heuristic objectives and architectural choices, offering limited theoretical understanding of when and why reliable attribute control is achievable. In this work, we develop a formal framework for speech attribute conversion and provide ...
731 Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
2511.06701
cs.AI
Karen Sargsyan
AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture that enforces statistical rigor at two levels: a Haskell embedded domain-specific language (the Research monad) that makes ...
AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture that enforces statistical rigor at two levels: a Haskell embedded domain-specific language (the Research monad) that makes it impossible to test a hypothesis without updating the error budget, and a declarative scaffold that fixes the data flow and the statistical test, together with an OS-level sandbox that makes validation data physically absent from the envi...
732 SLMFix: Leveraging Small Language Models for Domain Specific Language Error Fixing with Reinforcement Learning
2511.19422
cs.AI
David Jiahao Fu, Aryan Gupta, Aaron Councilman, Yu-Xiong Wang, Vikram Adve
Large language models (LLMs) have shown impressive capabilities in code generation across many programming languages but even state-of-the-art LLMs generate programs that contain syntactic errors and fail to complete the given tasks, especially for low-resourc...
Large language models (LLMs) have shown impressive capabilities in code generation across many programming languages but even state-of-the-art LLMs generate programs that contain syntactic errors and fail to complete the given tasks, especially for low-resource programming languages (LRPLs). In addition, the high cost of training makes finetuning LLMs unaffordable for those with constrained computational resources, further weakening the effectiveness of LLMs for code generation. In this work, we...
733 BabelCoder: Agentic Code Translation with Specification Alignment
2512.06902
cs.AI
Fazle Rabbi, Soumit Kanti Saha, Tri Minh Triet Pham, Song Wang, Jinqiu Yang
As software systems evolve, developers increasingly work across multiple programming languages and often face the need to migrate code from one language to another. While automatic code translation offers a promising solution, it has long remained a challengin...
As software systems evolve, developers increasingly work across multiple programming languages and often face the need to migrate code from one language to another. While automatic code translation offers a promising solution, it has long remained a challenging task. Recent advancements in Large Language Models (LLMs) have shown potential for this task, yet existing approaches remain limited in accuracy and fail to effectively leverage contextual and structural cues within the code. Prior work h...
734 Initial results of the Digital Consciousness Model
2601.17060
cs.AI
Derek Shiller, Laura Duffy, Arvo Mu\~noz Mor\'an, Adri\`a Moret, Chris Percy
Artificially intelligent systems have become remarkably sophisticated. They hold conversations, write essays, and seem to understand context in ways that surprise even their creators. This raises a crucial question: Are we creating systems that are conscious? ...
Artificially intelligent systems have become remarkably sophisticated. They hold conversations, write essays, and seem to understand context in ways that surprise even their creators. This raises a crucial question: Are we creating systems that are conscious? The Digital Consciousness Model (DCM) is a first attempt to assess the evidence for consciousness in AI systems in a systematic, probabilistic way. It provides a shared framework for comparing different AIs and biological organisms, and for...
735 Dynamic Welfare-Maximizing Pooled Testing
2601.22419
cs.AI
Edwin Lock, Nicholas Lopez, Francisco Marmolejo-Coss\'io, Jose Roberto Tello Ayala, David C. Parkes
Pooled testing uses one test to certify several agents as healthy when the pooled result is negative. We study a budget-constrained welfare problem in which agents have heterogeneous utilities and independent prior probabilities of being healthy. Welfare is ea...
Pooled testing uses one test to certify several agents as healthy when the pooled result is negative. We study a budget-constrained welfare problem in which agents have heterogeneous utilities and independent prior probabilities of being healthy. Welfare is earned when an agent is certified healthy, and a dynamic policy may choose each pool after observing earlier test outcomes. We ask how much such adaptation can improve over a static allocation that fixes all pools in advance. Our main result ...
736 HuPER: A Human-Inspired Framework for Phonetic Perception
2602.01634
cs.AIeess.AS
Chenxu Guo, Jiachen Lian, Yisi Liu, Baihe Huang, Shriyaa Narayanan
We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five Eng...
We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourc...
737 The Shrinking Lifespan of LLMs in Science
2604.07530
cs.AI
Ana Tri\v{s}ovi\'c
Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once released. We introduce time-to-peak and lifespan as measures of model obsolescence and use them to characterize the scientific...
Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once released. We introduce time-to-peak and lifespan as measures of model obsolescence and use them to characterize the scientific adoption trajectories of 62 LLMs across more than 108k citing papers (2019-2025), separating active adoption from background citation to recover per-model trajectories that citation counts cannot resolve. We find that a model's longevity i...
738 Endogenous Information in Routing Games: Memory-Constrained Equilibria, Recall Braess Paradoxes, and Memory Design
2604.11733
cs.AI
Saad Alqithami
We study routing games in which travelers optimize over routes that are remembered or surfaced, rather than over a fixed exogenous action set. The paper develops a tractable design theory for endogenous recall and then connects it back to an explicit finite-me...
We study routing games in which travelers optimize over routes that are remembered or surfaced, rather than over a fixed exogenous action set. The paper develops a tractable design theory for endogenous recall and then connects it back to an explicit finite-memory micro model. At the micro level, each traveler carries a finite memory state, receives surfaced alternatives, chooses via a logit rule, and updates memory under a policy such as LRU. This yields a stationary Forgetful Wardrop Equilibri...
739 Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
2604.15186
cs.AI
Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhao, Yanda Tao
Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpr...
Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic wo...
740 An AI Agent Execution Environment to Safeguard User Data
2604.19657
cs.AI
Robert Sorab Stanley, Avi Verma, Lillian Tsai, Konstantinos Kallas, Sam Kumar
AI agents promise to serve as general-purpose personal assistants for their users, which requires them to have access to private user data (e.g., personal and financial information). This poses a serious risk to security and privacy: an AI model may hallucinat...
AI agents promise to serve as general-purpose personal assistants for their users, which requires them to have access to private user data (e.g., personal and financial information). This poses a serious risk to security and privacy: an AI model may hallucinate or make mistakes, and adversaries may attack it (e.g., via prompt injection) to exfiltrate user data. This paper presents GAAP (Guaranteed Accounting for Agent Privacy), an execution environment for AI agents that guarantees confidentiali...
741 Information Aggregation with AI Agents
2604.20050
cs.AI
Spyros Galanis
Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving...
Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, across four information structures of increasing complexity. We find that although the median market is effective at aggregating information in the easy information structures, performance deteriorates in the harder struct...
742 Topology-Driven Anti-Entanglement Control for Soft Robots
2605.05236
cs.AI
Haoyang Le, Shengxuan Wang, Mohan Chen, Shuo Feng
In the field of precision manufacturing in complex constrained environments, the role of soft robots is increasingly prominent, and the realization of anti-winding control based on multi-intelligent body reinforcement learning has become a research hotspot. On...
In the field of precision manufacturing in complex constrained environments, the role of soft robots is increasingly prominent, and the realization of anti-winding control based on multi-intelligent body reinforcement learning has become a research hotspot. One of the core problems at present is to coordinate multiple robots to complete the unwinding operation in a highly constrained environment. The existing distributed training framework faces some observability challenges in high-density barr...
743 Grid-Orch: An LLM-Powered Orchestrator for Distribution Grid Simulation and Analytics
2605.12728
cs.AI
Boming Liu, Jin Dong, Jianming Lian
The power distribution engineering workforce faces a projected shortage of up to 1.5 million engineers by 2030, creating urgent demand for more accessible analysis tools. This paper introduces Grid-Orch, a framework that bridges Large Language Models (LLMs) an...
The power distribution engineering workforce faces a projected shortage of up to 1.5 million engineers by 2030, creating urgent demand for more accessible analysis tools. This paper introduces Grid-Orch, a framework that bridges Large Language Models (LLMs) and power system simulation through the Model Context Protocol (MCP), enabling engineers to perform complex distribution analyses via natural language. Using OpenDSS as the reference implementation, Grid-Orch provides 36 domain-specific tools...
744 Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
2605.30834
cs.AI
Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment. Detecting such failures during...
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that compromise reliability in real-world deployment. Detecting such failures during execution is therefore critical for the robust deployment of embodied systems. Existing failure detection methods either rely on expensive action resampling or external models, while alternatives propagate trajectory-level labels uniformly...
745 Efficient Safety Benchmarking via Item Response Theory
2606.20626
cs.AI
Fabio Spagliardi, M\'irian Silva, Ayan Datta, Aiden Zhou, Vamshi Bonagiri
Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly problematic for adversarial, highly heterogeneous safety items. Applied in full ...
Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly problematic for adversarial, highly heterogeneous safety items. Applied in full to modern benchmark suites, current evaluation procedures would require on the order of $10^5$ responses, most of which provide little ranking signal. We analyze six widely used safety benchmarks and make three contributions toward more eff...
746 CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation
2607.02222
cs.AI
Haokun Liu, Zhaoqi Ma, Yicheng Chen, Wentao Zhang, Masaki Kitagawa
Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-lev...
Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored. We propose CoFL-S, a low-level vision-language-action framework that predicts a language-conditioned flow field over the robot's local visible sector and generates continuous trajectories by rolling out the predicted field. To train this low-level representation, we c...
747 Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise
2607.26639
cs.AI
Haoyu Zhang, Hanwen Liu, Yang Chen, Shibo Zheng, Xiangchen Guan
Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so ...
Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so the encoding is the only difference. Against SAGE, the strongest published self-check defense, best-of-N over a code-completion encoding reaches 67, 22 and 15% of behaviors on three open-weight targets, where that encoding fired once reache...
748 Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
2608.06994
cs.AI
Xiangkai Ma, Yue Ma, Junjie Wang, Sheng Xu, Mingyang Li
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting...
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical condition evolution and low-level action trajectory generation ...
749 Beyond Forecasting: Recasting Volatility Control as a Routing Problem
2608.10375
cs.AI
Hongji Pu, Leyang Zhou
Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-defined control rule that may not adapt to changing market conditions. We propose VolRouter, a modular framework tha...
Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-defined control rule that may not adapt to changing market conditions. We propose VolRouter, a modular framework that formulates volatility control as state-conditioned routing over estimator-controller pairs. VolRouter first summarizes market conditions into a control-relevant state profile and then performs routing through three stages: state inference...
750 Keep the Future, Drop the Rollout: RIFT for World Action Models
2608.11521
cs.AI
Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on ...
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on 40 simulated robotic manipulation tasks, paired closed-loop interventions show that blocking access to the future cache or reassigning its values changes execution and reduces success. Yet in the evaluated co-denoising settings, reusing one...
751 SUN: Agentic Robot Policy Learning with Persistent Task Programs
2608.31167
cs.AI
Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao
Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed e...
Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned optimal control objectives, satisfaction predicates, and learning rewards. Our harness, Kuafu, equips a foundation model as a task-level agent to orchestrate scene preparation, verific...
752 Exploring Second-Order Pattern Recognition in Speaker Recognition
2609.11182
cs.AIeess.AS
Yanze Xu, Wenwu Wang, Mark D. Plumbley
In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's re...
In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's recognition of inputs as human-defined patterns; this work calls these latent patterns second-order patterns and proposes to discover them. Accordingly, we apply a hierarchical clustering algorithm to analyse whether our speaker recognition n...
753 Interpreting hierarchical organisation of speaker embeddings
2609.15203
cs.AIeess.AS
Yanze Xu, Wenwu Wang, Mark D. Plumbley
Speaker recognition neural networks recognise speaker identities from input utterances by learning latent representations (i.e. speaker embeddings). However, these networks' internal mechanisms remain largely opaque, motivating research in explainable artifici...
Speaker recognition neural networks recognise speaker identities from input utterances by learning latent representations (i.e. speaker embeddings). However, these networks' internal mechanisms remain largely opaque, motivating research in explainable artificial intelligence (XAI) to understand them. Existing studies have analysed how speaker embeddings are organised, but rarely frame these analyses within XAI. This work proposes to explain and interpret the organisation of speaker embeddings fr...
754 The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties
2609.26174
cs.AI
Haoyu Zhang, Yi Feng, Hanwen Liu, Shibo Zheng, Zhuoxi Wang
We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request,...
We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition, shifts benign refusal by tens of points. The shift is not blanket caution but a threshold shift: genuinely neutral instructions are almost unaffected (<=2 percentage points on thr...
755 G\"odel's and Scott's Variants of the Ontological Argument in Lean 4 and TPTP THF
2609.26806
cs.AI
Christoph Benzm\"uller
This paper presents a complete, structure-preserving port to Lean 4 of the Isabelle/HOL dataset accompanying Benzm\"uller and Scott's study of G\"odel's ontological argument and Scott's variant: 30 modules, one per theory, retaining section structure, declarat...
This paper presents a complete, structure-preserving port to Lean 4 of the Isabelle/HOL dataset accompanying Benzm\"uller and Scott's study of G\"odel's ontological argument and Scott's variant: 30 modules, one per theory, retaining section structure, declaration order and names up to documented renamings; a comparison tool certifies the 548 statements identical as parsed. Every named result the original proves is proved again, from the inconsistency of G\"odel's 1970 axioms to modal collapse, m...
756 AI in Science: Early Insights
2609.28504
cs.AI
Mihai Codreanu, Alex Imas, Juan Mateos-Garcia, Joseph Emmens, Evalyne Muiruri
Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million...
Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings...
cs.CL 127 papers
183 A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
2609.30287
cs.CLcs.AI
Pawe{\l} Blicharz, Mi{\l}osz Grunwald
AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the ...
AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models. We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers x 768), which we call neurons. T...
184 Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
2609.30288
cs.CLcs.LG
Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii
In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their ...
In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operat...
185 Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
2609.30289
cs.CL
Yufei Shi, Rujing Yao, Ang Li, Yang Wu, Zhuoren Jiang
In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations, execution traces, and intermediate...
In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations, execution traces, and intermediate progress. Existing memory-augmented systems typically retrieve from all stored memories as a flat pool, ranking them by semantic relevance, importance, or recency without modeling hierarchical structure or evolving validity. As a result, t...
186 Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
2609.30290
cs.CLcs.LG
Haowei Liu, Hsin-Tai Wu, Yi Fang
Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreem...
Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured. When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set. Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION. A self-hosted Qwen3.6-27B re...
187 A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models
2609.30292
cs.CLcs.AI
Fanji Yang (Guizhou University of Finance and Economics), Huiyao Chen (Harbin Institute of Technology), Xi Yu (Guizhou University of Finance and Economics), Meishan Zhang (Harbin Institute of Technology), Xiaohong Xiao (Guizhou University of Commerce)
Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust mechanisms.The rise of large...
Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust mechanisms.The rise of large language models, or LLMs, has changed the problem in two directions.LLMs can generate fluent and context-aware deceptive reviews, while pre-trained language models, or PLMs, and LLMs also provide stronger semantic representations for detec...
188 Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
2609.30293
cs.CLcs.AI
Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong
The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog...
The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog traversal to $O(k)$ progressive disclosure. Cartograph combines three mechanisms: (1) operator-attested capability cards, Ed25519-signed descriptions generated under the deploying operator's control rather than ranked publisher copy; (2) R...
189 SlideLab: Audience-Centered Scientific Slide Generation and Evaluation
2609.30294
cs.CLcs.AI
Vidushee Vats, Karun Sharma, Yuxia Wang
Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework...
Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement,...
190 SignTrace: Describe a Sign, Find the Word
2609.30295
cs.CLcs.AI
Zengji Tu, Xingye Zhu, Ningjing Wang, Tingyi Huang, Yangjunfeng Zhu
Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access to a Chinese sign-language dic...
Identifying an unfamiliar sign is difficult when a learner remembers its movement but does not know its meaning or formal feature codes. SignTrace addresses this longstanding reverse-lookup problem through natural-language access to a Chinese sign-language dictionary. The system integrates LLM-based dictionary enrichment, action extraction, dictionary-style rewriting, seven-channel retrieval, and candidate reranking over 6,699 entries. It has been deployed for user trials and has received positi...
191 Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
2609.30297
cs.CLcs.AI
Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri\`a Casas Escoda
Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is o...
Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge in building such agents is optimizing agent planning -- deciding how to select, sequence, and invoke tools -- particularly in cold-start settings where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a...
192 A Benchmark Framework for Screening Automation in Systematic Reviews
2609.30298
cs.CL
Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani
Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance cl...
Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets.This paper presents a benchmark dataset of $45\,064$ labeled entries for evaluating L...
193 A Unified Account of Concepts and Chunks
2609.30414
cs.CLcs.AI
Karthik Singaravadivelan, Pat Langley
Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses ...
Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses a challenge for unified theories of cognition. In this paper, we review Cobweb, a computational account of categorization and concept formation, and propose an extended theory that incorporates chunks and their acquisition. The theory makes...
194 All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation
2609.30416
cs.CL
Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu
Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with hig...
Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel d...
195 Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
2609.30439
cs.CLcs.SDeess.AS
Bo Su, Yueru Yan, Thai Le
We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires a...
We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) modul...
196 Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
2609.30467
cs.CL
Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi
Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of r...
Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neithe...
197 CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
2609.30471
cs.CLcs.LGcs.AI
Mukul Chhabra, Shail Patel, Luigi Medrano
Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a...
Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved r...
198 Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs
2609.30492
cs.CLcs.AI
Sang Bin Moon, Nicole Cho, Daniel Borrajo, Sumitra Ganesh, Abolfazl Hashemi
Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning prob...
Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, u...
199 Inquesto Score: A reliability Protocol For Voice Agents
2609.30514
cs.CLcs.AI
Massa Baali, Bhiksha Raj
Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto Score (IS), a protocol for mea...
Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto Score (IS), a protocol for measuring voice-agent reliability as the percentage of calls in a fixed, versioned evaluation population that achieve the caller's goal without a functional failure or worse. Rather than combining heterogeneous metrics, IS defines explicit fai...
200 Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
2609.30535
cs.CL
Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kiant\'e Brantley
Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word mul...
Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched...
201 REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles
2609.30547
cs.CL
Haixu Ma, Aditya Bansal, Shubham Lohiya, Sumit Ranjan
Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant del...
Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant delays, estimation errors, and poor scalability over high-dimensional profile data. We present REALMS (Real-time Exact Audience sizing via LLM-based Multi-attribute Search), a conversational system for exact audience sizing deployed in product...
202 Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
2609.30558
cs.CLcs.LGcs.AI
Jiaqi Ding, Guorong Wu
Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework ...
Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactiva...
203 The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
2609.30604
cs.CLcs.AI
Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates progr...
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a be...
204 Recursive Self-Improvement via On-Policy Distillation for Reasoning
2609.30652
cs.CL
Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu
On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-dist...
On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns ...
205 Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4
2609.30716
cs.CLcs.AI
Amanda Fitch
When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b...
When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b model across a targeted behavioral suite (n = 13 items, 784 forward passes in short, single-turn contexts) using a completely counterbalanced experimental design. This setup allowed us to mathematically isolate the specific effects of sour...
206 Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction
2609.30738
cs.CLcs.LG
Tianfang Xie, Wei Zhu
KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and ...
KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and redundancy relative to selected tokens. For $\lambda_2<0$, the score penalizes similarity to selected tokens as in maximal marginal relevance (MMR), without extra forward passes. To test whether this relevance-diversity balance should vary ...
207 SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages
2609.30739
cs.CL
Puja Ahmad Habibi, Faiz Assabil Firdaus, Ashvanth S, Ekapol Chuangsuwanich, Pume Tuchinda
Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce...
Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and mul...
208 Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
2609.30773
cs.CLeess.AS
Manato Yaguchi, Yotaro Kubo, Hikaru Asano, So Kuroki
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still s...
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simulator LLM adds substantial data-preparation overhe...
209 Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment
2609.30802
cs.CL
Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel, Yi Ding
Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from tea...
Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from teacher to student remains largely unexplored. Thus, we fill this gap by analyzing how different template configurations influence the pre-existing safety alignment of the student. We observe a significant degradation of safety alignment prese...
210 I-Parakeet: Integer-Only Conformer ASR on Mobile NPU
2609.30846
cs.CLcs.SDeess.AS
Taichi Nishimura
In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard to deploy on edge devices be...
In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard to deploy on edge devices because of their size, and quantized models still fall back to floating point for numerically sensitive operations. This prevents them from fully exploiting integer accelerators such as mobile NPUs. To achieve this, our contributions are thre...
211 Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
2609.30849
cs.CL
Phuong Q. Le, Kemal Kurniawan, Jey Han Lau
Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to ...
Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across ...
212 Persistent Negatives for Adversarial Black-Box On-Policy Distillation
2609.30864
cs.CLcs.AI
Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu
Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher...
Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We ...
213 Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
2609.30867
cs.CL
Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported eviden...
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-fla...
214 Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos
2609.30882
cs.CL
Yuya Wake, Sho Tsugawa, Toshiyuki Amagasa
Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim screening. This study examines how su...
Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim screening. This study examines how such transcript compression affects LLM-based veracity classification of Japanese medical YouTube videos. We compare four transcript input designs: full transcripts, LLM-generated summaries, RAPTOR-based retrievalaugmented generation (RAG), a...
215 From annotation to reasoning: Culture in language models
2609.30897
cs.CL
Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen
How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain...
How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpr...
216 ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
2609.30906
cs.CL
Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful ...
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively searc...
217 Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
2609.30924
cs.CLcs.SDeess.AS
Hikaru Asano, Yotaro Kubo, So Kuroki
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only par...
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integra...
218 Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models
2609.30935
cs.CLcs.LGcs.AI
Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li
Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose kn...
Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose knowledge. Although existing methods, such as orthogonal gradient projection, mitigate the forgetting across various fine-tuning tasks, they fundamentally fail to preserve pre-training LLMs' inherent general-purpose knowledge because the orig...
219 FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation
2609.30968
cs.CL
Lu Han, Jingyao Zhang, Katy Ilonka Gero, Nguyen H. Tran
Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated parameter-efficient fine-tun...
Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated parameter-efficient fine-tuning (PEFT) offers a data-local setting for this multi-author adaptation problem: clients keep author text local while sharing compact adapter updates. However, we show that standard aggregation can preserve continuation utility while making...
220 Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change
2609.30974
cs.CL
Haruka Ezoe, Ryohei Hisano
Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or w...
Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or which usages support the attribution. We introduce Coupled Usage--Sense Processes (CUSP), which derives these answers from a single marginal preserving temporal process. A hierarchical coupling relates contextual distributions through latent...
221 THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer
2609.30984
cs.CL
Seanghay Yath
Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur insid...
Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur inside ordinary words. We present Tha, a Khmer text normalization and inverse text normalization toolkit built from weighted finite-state transducers. It segments and classifies a whole line in one shortest-path search, and a second transducer r...
222 Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
2609.30986
cs.CL
Geng Liu, Feng Li, Mengxiao Zhu, Francesco Pierri
As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect...
As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users' confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or...
223 ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
2609.31002
cs.CL
Siqiao Xue, Shuxuan Liu, Ning Hu
Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are diffi...
Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are difficult to supervise at scale: real search traffic provides authentic queries and candidates but no clean pairwise labels. We present ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, and 8B) aligned to judge-labeled shopping pre...
224 G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
2609.31009
cs.CLcs.AI
Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitat...
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds....
225 Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference
2609.31046
cs.CL
\"Ozge Alacam, Z\"ubeyde Demet Kirbulut G\"une\c{s}, Funda Ekici, Nurcan Turan-Oluk, Dilay Din\c{c}demir
Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. W...
Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. We investigate whether instruction-tuned large language models (LLMs) can support multidimensional analysis of collaborative sensemaking without task-specific training, and whether structured knowledge-state information improves model infere...
226 CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
2609.31062
cs.CL
Marko \v{R}eh\'a\v{c}ek, V\'it\v{e}zslav Du\v{s}ek, Martin Rusinko, V\'it Nov\'a\v{c}ek
Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We ...
Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We propose Clinical Guardrail Probes (CG-Probes) to measure the risks from query embeddings. We probe for each axis in the normalized embedding space of frozen embedders via the difference-in-means method, treating each axis as a potential lin...
227 LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
2609.31122
cs.CLcs.LG
Irene Tallini, Lorenzo Basile, Valentino Maiorca, Francesco Locatello, Alberto Cazzaniga
Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space...
Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties present in the contrastive data and degrade unrelated capabilities. To mitigate this issue, we introduce LocUS (Localized Unembedding Steering), a method which grounds activation ...
228 Do we need to answer that question? Salience and Answerability of Potential Questions in Naturalistic Dialogue
2609.31130
cs.CL
Amandine Decker (LORIA, UL, CNRS, SEMAGRAMME, GU)
We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we construct a dataset of 7,124 ...
We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we construct a dataset of 7,124 questions automatically generated from utterances and preceding context from the British National Corpus, and annotated for salience and answerability. We find a robust but low positive correlation between salience and answerability in dial...
229 Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting
2609.31169
cs.CLcs.AI
Pawe{\l} M\k{a}ka, Piotr Andruszkiewicz, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore th...
Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore this information. Therefore, increasing their visual sensitivity remains an active research area. In this work, we introduce a training method, Metric-based Loss Weighting, that improves visual grounding of translations by increasing the loss...
230 Where a Model Sends Its Own Repeated Token
2609.31181
cs.CL
Nicol\'as Vera Z\'u\~niga
Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where th...
Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax p(. | t, t) in one forward pass; the result is a map on the whole vocabulary, with two halves. The first -- which tokens are fixed points -- is partially anticipated, and we repor...
231 RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models
2609.31245
cs.CL
Pavithra P M Nair, Bhavik Talaviya, Shourya Bhushan, Rahul Pankajakshan, Seema Guruvadoo
Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social...
Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categorie...
232 PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding
2609.31255
cs.CL
Jeonghun Yoon, Dongchan Kim, Hongyeon Yu, Young-Bum Kim, Jaegul Choo
General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretio...
General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretion, and a three-month glucose trend cannot be answered by text similarity. We present PIA, a personal intelligence agent deployed alongside a consumer health agent. PIA receives the agent's natural-language requests, decides for itself wheth...
233 MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries
2609.31261
cs.CLcs.AI
Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro
The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximat...
The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, with the relevant interaction geometry learned from data: natural-language dependencies are input-dependent and difficult to prescribe in advance, so the model should learn where posi...
234 Identifying Scientists on X
2609.31264
cs.CL
Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze
With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientist...
With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forest...
235 Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers
2609.31342
cs.CL
Md Shamim Ahmed, Lukas Galke Poech, Richard R\"{o}ttger
Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document poisoning, in which outdated e...
Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document poisoning, in which outdated evidence makes a model wrong despite answering correctly without retrieval. We construct a benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy, grounded in dated official sources. Across 12 model...
236 Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding
2609.31382
cs.CLcs.AI
Zhaoyuan Xia (Peking University, Baidu Inc), Qinghongbing Xie (Tsinghua University), Yung Xiang Hue (Tsinghua University), Jianguang Jiang (Baidu Inc)
Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Su...
Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that first identifies source-grounded, question-relevant evidence and then integrates it into a compact, question-conditioned summary before producing the final answer. To train this behavior, ...
237 ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs
2609.31448
cs.CLcs.AI
Junyi Gao, Yu Shi, Pingzhao Hu, Ewen M Harrison
Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from structured, high-dimensional clini...
Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from structured, high-dimensional clinical time series. Improving this ability would connect risk estimation with flexible questions about a patient's evolving condition. We introduce ViSTA, a compact adapter that incorporates irregular numerical measurements into a pretrained v...
238 Evaluating Cultural Awareness of LLMs for Haitian Creole
2609.31506
cs.CLcs.AI
Christelle Clervilsson, Yanzhu Guo
Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present...
Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specific...
239 Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge
2609.31511
cs.CL
Yahya Mohamed Elnawasany
We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retriev...
We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retrieval layer routed across six Model Context Protocol servers, we report three things a research prototype typically lacks. First, a released family of fine-tuned Arabic Islamic model artifacts: an efficient tool-routing LLM (Muslim-6B-PRO, 5.9...
240 Strategically Diverse Sampling for Self-Training
2609.31571
cs.CL
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID re...
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constr...
241 PALM: Point-in-Time Adaptation for Financial Language Models
2609.30316
cs.CLcs.LG
Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han
Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time (PIT) language models are pretrai...
Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already observed the outcomes it is asked to predict. To handle this issue, point-in-time (PIT) language models are pretrained on chronologically filtered corpora and released as one checkpoint per calendar year, each with a documented cutoff. However, each additional year costs a full pretraining run, and whether that run is necessary has never been tested. In...
242 When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess
2609.30328
cs.CLcs.AI
Salma Roshdy Aly, Hussein Assaf, Ziad Kobti
When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decompose...
When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for. Multi-agent verification, which decomposes a judgment into checkable claims and verifies each against evidence, is a promising response and works well when the evidence is a set of retrieved documents. We argue such methods require two things of their evidence: it must be independ...
243 RAZOR: Pruning Replaceable Experts in LLMs
2609.30465
cs.CLcs.LG
Mingyang Song, Mao Zheng
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet a...
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method ...
244 Asymmetric Classifier-Free Guidance for Target-Speaker ASR
2609.30476
cs.CLcs.SDeess.AS
Yiwen Guan, Jacob Whitehill
Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time ca...
Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized mu...
245 AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth
2609.30483
cs.CLcs.LGcs.SDeess.AS
Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines th...
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, a...
246 Don't CLAP: Are Music-Text Models Bag-of-Words?
2609.30540
cs.CLcs.SDeess.AS
Yuan-Chiao Cheng, Alexander Lerch
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask h...
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real...
247 Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content
2609.30563
cs.CLcs.AI
Ljubisa Bojic, Tijana Stanic, Joerg Matthes, Agariadne Dwinggo Samala, Bojana Dinic
Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has ...
Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reacti...
248 Epstein Files Engine: Agentic Search for Investigative Journalism
2609.30611
cs.CL
Duy K. Nguyen, Teresa Mondr\'ia Terol, Dylan Freedman, Zach Seward
On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files...
On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files. The Engine translated reporter questions into Google BigQuery SQL queries across three corpora: Epstein-related releases, the Times's archive and external, Epstein-related news headlines. It used an LLM to plan queries and returned citati...
249 Prompt Injection Detection for Email Agents Through Attack Chain Modeling
2609.30657
cs.CL
Ahmad Hashmi, Dhyey Patel, Yunting Yin
Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this ...
Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text ...
250 LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information
2609.30706
cs.CLcs.AI
Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what se...
"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya with Value-Of-Information Routing), which places the candidate pieces of missing information (slots) in the input next to the answer options, so that one forward pass returns both...
251 Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
2609.30784
cs.CLcs.SDeess.AS
Yotaro Kubo, Qi Sun, Yujin Tang
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directl...
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the propos...
252 Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
2609.30820
cs.CLcs.LG
Nux Li
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual lo...
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Control...
253 Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces
2609.30914
cs.CL
Aman Mittal, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Abhishek Singh, Aditya Singh
Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and evolving them through rotation-gate o...
Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and evolving them through rotation-gate operators. This approach offers higher optimization performance without physical qubits, and has been shown to achieve order-of-magnitude speedups (10--80$\times$) over traditional solvers on combinatorial, high-dimensional NP-hard problems....
254 Does Uniform Discrete Diffusion Need Time?
2609.30977
cs.CLcs.LGcs.AI
Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye
Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much...
Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence remains much closer to its original clean sequence than t...
255 Same Text, Different Numbers: The Divergence of LLM-Based Measures
2609.31013
cs.CLcs.AI
Hamid Boustanifar, Sasan Mansouri
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, manag...
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, a...
256 KuaFu: Compressing Long User Behavior into Understanding at Billion Scale
2609.31045
cs.CLcs.LG
Jiahao Hui, Lin Zhu, Yishen Hu, Jingdong Shu, Zetai Jiang
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the ful...
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of tho...
257 JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
2609.31142
cs.CLcs.AI
Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng
Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness ...
Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can...
258 Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring
2609.31293
cs.CLcs.SDeess.AS
Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin
Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper invest...
Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across c...
259 The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
2609.31341
cs.CLcs.AI
Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive ...
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration....
260 Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models
2609.31397
cs.CLcs.AI
Andrea Masini, Sudipta Acharya, Paolo Bellavista, Luca Foschini, Burak Kantarci
Automated and highly usable Quality-of-Service (QoS) enforcement requires translating high-level service intents into deployable traffic-management policies. Although intent-based networking (IBN) has simplified policy specification, bridging the gap between b...
Automated and highly usable Quality-of-Service (QoS) enforcement requires translating high-level service intents into deployable traffic-management policies. Although intent-based networking (IBN) has simplified policy specification, bridging the gap between business-level intents and executable network configurations remains complex, error-prone, and difficult to automate. This paper presents Intent2Tc, a closed-loop language-model-driven framework that translates business-level traffic-shaping...
261 Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis
2609.31422
cs.CLcs.AI
Jakub Mas{\l}owski, Jaros{\l}aw A. Chudziak
Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summariz...
Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summarizing models are prone to fabricating smoothly written debate consensus that is not grounded in the debate's history. To address this safety gap, this paper presents empirical research and studies if the introduction of active post-debate ver...
262 PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
2609.31468
cs.CLcs.AI
Pavel Kireyev
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice se...
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model,...
263 Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing
2609.31513
cs.CL
Marco Mandap
We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fu...
We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count thr...
264 Two Conformal Constructions for Adaptive Within-Document AI-Text Screening
2609.31547
cs.CL
Marco Mandap, Jerahmeel Hipolito, Arcel Galvez, Charlie Margaret Balagtas, Michael Joshua Buluran
We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-samp...
We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector...
265 Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
2609.31587
cs.CLcs.AI
Md Shohel Arman, Igor Molybog
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passe...
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen file...
266 User Model Extraction via Belief Self-Distillation
2609.31603
cs.CLcs.LG
Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework tha...
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external a...
267 Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
2609.31619
cs.CLcs.LGcs.AI
Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning duri...
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fin...
268 NaijaNLP: A Survey of Nigerian Low-Resource Languages
2502.19784
cs.CLcs.AI
Isa Inuwa-Dutse
With over 500 languages in Nigeria, three languages - Hausa, Yor\`ub\'a and Igbo spoken by more than 175 million people, account for about 65% of the languages. However, these languages are classed as low-resource due to insufficient digital resources to suppo...
With over 500 languages in Nigeria, three languages - Hausa, Yor\`ub\'a and Igbo spoken by more than 175 million people, account for about 65% of the languages. However, these languages are classed as low-resource due to insufficient digital resources to support tasks in computational linguistics. While several research efforts and initiatives have been presented, a coherent understanding of the state of classic Natural Language Processing (NLP) spanning grammatical formalisation to linguistic r...
269 MedHal: a Synthetic Dataset for Medical Hallucination Detection
2504.08596
cs.CLcs.AI
Fabrice Lamarche, Gaya Mehenni, Neshat Elhami Fard, Odette Rios-Ibacache, Li Ming Wang
Hallucination, the generation of non factual content by AI systems, poses serious risks in medical contexts, where errors can directly affect patient outcomes. We present MedHal, a large-scale dataset specifically designed to assess capabilities and train mode...
Hallucination, the generation of non factual content by AI systems, poses serious risks in medical contexts, where errors can directly affect patient outcomes. We present MedHal, a large-scale dataset specifically designed to assess capabilities and train models on the task of hallucination detection in medical texts. Current hallucination detection methods face significant limitations when applied to specialized domains like medicine, where they can have disastrous consequences. MedHal addresse...
270 Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning
2505.09738
cs.CLcs.AI
Shaurya Sharthak, Vinayak Pahalwan, Adithya Kamath, Adarsh Shirawalmath
Pretrained language models (LLMs) are often constrained by their fixed tokenization schemes, leading to inefficiencies and performance limitations, particularly for multilingual or specialized applications. This tokenizer lock-in presents significant challenge...
Pretrained language models (LLMs) are often constrained by their fixed tokenization schemes, leading to inefficiencies and performance limitations, particularly for multilingual or specialized applications. This tokenizer lock-in presents significant challenges. standard methods to overcome this often require prohibitive computational resources. Although tokenizer replacement with heuristic initialization aims to reduce this burden, existing methods often require exhaustive residual fine-tuning ...
271 Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
2601.01842
cs.CL
Yusuke Ide, Adam Nohejl, Joshua Tanner, Hitomi Yanaka, Christopher Lindsay
Dictionary definitions are an essential resource for learning word senses, but manually creating them is costly. We thus study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Specifically, we ...
Dictionary definitions are an essential resource for learning word senses, but manually creating them is costly. We thus study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Specifically, we address learner's dictionary definition generation (LDDG), where definitions should be written using simple vocabulary. First, we introduce a reliable evaluation approach for DDG, based on newly proposed evaluation criteria and powered by a...
272 Affective Flow Language Model for Emotional Support Conversation
2602.08826
cs.CLcs.AI
Chenghui Zou, Ning Wang, Tiesunlong Shen, Luwei Xiao, Chuan Ma
Large language models (LLMs) have advanced emotional support conversation, but existing alignment methods rely mainly on sparse preferences at the response level or outcomes at the dialogue level, providing limited supervision for sequential strategy decisions...
Large language models (LLMs) have advanced emotional support conversation, but existing alignment methods rely mainly on sparse preferences at the response level or outcomes at the dialogue level, providing limited supervision for sequential strategy decisions in multi-turn interactions. This raises a key question: how can detailed process signals be derived from overall dialogue outcomes to guide the gradual adaptation of support strategies? We propose the Affective Flow Language Model (AFlow),...
273 Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
2603.18863
cs.CL
Yana Veitsman, Yihong Liu, Hinrich Sch\"utze
Cross-lingual alignment is often assumed to improve cross-lingual transfer by bringing representations of different languages closer together. However, improvements in representational alignment do not consistently translate into better downstream performance....
Cross-lingual alignment is often assumed to improve cross-lingual transfer by bringing representations of different languages closer together. However, improvements in representational alignment do not consistently translate into better downstream performance. We investigate this disconnect using XLM-R models explicitly aligned across four language pairs with token-level, sentence-level, and masked-language-modeling objectives. We evaluate their zero-shot transfer on a token-level task (part-of-...
274 Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR
2604.06487
cs.CL
Thibault Ba\~neras-Roux, Sergio Burdisso, Esa\'u Villatoro-Tello, Dairazalia S\'anchez-Cort\'es, Shiran Liu
Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with te...
Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with text-only data. However, this introduces a modality gap, as the LLM is not exposed to the noisy representations produced by the speech projector. We investigate whether small amounts of speech can mitigate this mismatch. We compare three stra...
275 VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation
2604.21375
cs.CLcs.AI
Qijun Han, Haoqin Tu, Zijun Wang, Haoyue Dai, Yiyang Zhou
Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without recovery. We present VLAA-GUI, a modu...
Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without recovery. We present VLAA-GUI, a modular GUI agentic framework built around three integrated components that guide the system on when to Stop, Recover, and Search. First, a mandatory Completeness Verifier enforces UI-observable success criteria and verification at every finish...
276 Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations
2604.23295
cs.CLcs.AI
Bhaskar Singh, Manas Dhir, Shobhit Banga, Pranav Sharma
Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible full-duplex spoken dialo...
Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible full-duplex spoken dialogue system for Hindi by adapting Moshi, a state-of-the-art duplex speech architecture, using a custom Hindi tokeniser and training on 26,000 hours of real spontaneous conversations collected from 14,695 speakers with separate speaker channe...
277 UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification
2605.06221
cs.CL
Qihang Fan, Huaibo Huang, Zhiying Wu, Bingning Wang, Ran He
As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing, several novel low-complexity hybrid ...
As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing, several novel low-complexity hybrid architectures have recently been proposed, effectively alleviating the computational burden of long-context inference. However, existing research on long-context prefill acceleration remains predominantly focused on sparse attention mechani...
278 StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
2605.06642
cs.CLcs.AI
Xiangyuan Xue, Yifan Zhou, Zidong Wang, Shengji Tang, Philip Torr
Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both exploration and credit assignment over exte...
Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both exploration and credit assignment over extended trajectories. In this work, we present Strategic Trajectory Abstraction (StraTA), a simple framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning (RL). StraTA samples a compact strategy from...
279 Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering
2605.23497
cs.CL
Max Prior, Andreas Schultz, Matthias Grabmair
Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds with the evolving nature of statutory law. We study two temporal failure modes: post-cutoff staleness, wher...
Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds with the evolving nature of statutory law. We study two temporal failure modes: post-cutoff staleness, where models apply superseded rules after legislative amendments, and recency bias, where models prefer newer provisions even when a historical version governs the fact pattern. To this end, we present a benchmark of 312 expert-validated, time-...
280 Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
2605.23975
cs.CLcs.SD
Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun
Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language omission, translation-instead-of-transc...
Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language omission, translation-instead-of-transcription, and hallucination. We apply Direct Preference Optimization (DPO) to align models, constructing preference pairs in which chosen responses preserve mixed-language content while rejected responses mimic failure patterns. Training thr...
281 Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation
2605.24534
cs.CL
Max Prior, Niklas Wais, Matthias Grabmair
We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any handcrafted doctrinal framework. Using 4.555 decisions of the German Federal Court of Justice that cite sect...
We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any handcrafted doctrinal framework. Using 4.555 decisions of the German Federal Court of Justice that cite sections 242, 280, 812 and 823 of the German Civil Code (BGB), we extract paragraph-level chunks, summarize their reasoning, and derive keywords, which are embedded and clustered. For each cluster, an LLM generates headings and synthesizes cita...
282 Large Language Model Selection with Limited Annotations
2605.24981
cs.CLcs.LG
Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel
Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active ...
Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active model selection of LLMs. SELECT-LLM aims to find a small set of queries whose annotations are most informative for identifying the best LLM for a given task. To this end, we introduce a query selection rule based on expected information gai...
283 Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
2606.10829
cs.CLcs.AI
Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Ex...
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-\(k\), Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a tr...
284 Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
2607.29238
cs.CLcs.AI
Antorweep Chakravorty
InMyStyle is a privacy-first, single-user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to ...
InMyStyle is a privacy-first, single-user system that adapts small language models to rewrite AI-edited text towards an individual user's writing style without an instruction prompt at inference. Given a user's documents, it uses multiple local helper LLMs to construct paired training examples and fine-tunes LoRA adapters on Qwen2.5 models ranging from 0.5B to 7B parameters. Length-aware generation budgets and automatic chunking support inputs of different lengths. We report a single-user case s...
285 State of Thought Enables Endogenous Reasoning
2609.16055
cs.CLcs.AI
Zhiren Gong, Yikun Hou, Zihao Zeng, Ming Xiao, Chau Yuen
Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through cost...
Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansion in constrained search spaces, limiting both generalization and efficiency. We propose State of Thought (SoT), a new reasoning paradigm that enables endogenous reasoning in LLMs, with the model's internal reasoning state governi...
286 Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining
2609.21362
cs.CL
Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen
Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated to...
Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequ...
287 Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER
2609.21663
cs.CL
Hritika Sharma, Thibault Ba\~neras-Roux, Alessandra Pinto, Petr Motlicek, Hyunggu Jung
Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how hu...
Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an English dataset for human-centered ASR evaluation. Using this dataset, we benchmark lexical metrics against several configurations of BERTScore and SemDist, varying the language mo...
288 Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts
2609.22455
cs.CL
Hope McGovern, Anna Dolganov, Samuel Belkadi, Guillaume Kunsch, Dimitris Vlitas
We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans with...
We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans without requiring oracle knowledge of their length. To our knowledge, it is the first large-scale decoder model for historical Greek, and the first for any ancient Mediterranean language. Evaluated as in prior work, on short gaps of up to ten c...
289 COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning
2609.22697
cs.CLcs.SDeess.AS
Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue
Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should...
Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational context. Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task. Given historical conversation audio, target text, and a reference speech, the system shou...
290 LeakScale: Estimating the Causal Effect of Benchmark Exposure
2609.27176
cs.CL
Divyansh Singh
Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the perf...
Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact. We present LeakScale, an interventional framework for estimating this missing quantity. LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-der...
291 ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts
2609.29349
cs.CLcs.AI
Firoj Alam, Md. Rafiul Biswas, Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Hamdy Mubarak
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In t...
ArGuard is a shared task on harmful content detection in Arabic memes and LLM prompts. It includes two tracks: Track A focuses on multimodal hate detection in Arabic memes, while Track B addresses harmful prompt detection for Arabic LLM safety evaluation. In total, 58 teams registered, 35 participated in the final evaluation, and 27 submitted system-description papers. Participating teams explored models such as AraBERT, Jais, and Qwen3-VL. The best systems achieved macro-F1 scores of 0.823 on A...
292 Likelihood Ranking doesn't Scale Like Prompting in LLMs
2609.29390
cs.CL
Alessandro Bondielli, Lucia Passaro, Davide Bacciu, Alessandro Lenci
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer ...
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs...
293 Rufus-Air: An Open LLM Post-Training Recipe
2609.29421
cs.CLcs.LGcs.AI
Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document...
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training build...
294 SkillFlow: Scalable and Efficient Agent Skill Retrieval System
2504.06188
cs.CLcs.AI
Fangzhou Li, Pagkratios Tagkopoulos, Ilias Tagkopoulos
AI agents can extend their capabilities at inference time by loading reusable skills into context, yet equipping an agent with too many skills, particularly irrelevant ones, degrades performance. As community-driven skill repositories grow, agents need a way t...
AI agents can extend their capabilities at inference time by loading reusable skills into context, yet equipping an agent with too many skills, particularly irrelevant ones, degrades performance. As community-driven skill repositories grow, agents need a way to selectively retrieve only the most relevant skills from a large library. We present SkillFlow, the first open, multi-stage retrieval system for agent skill discovery that frames skill acquisition as an information retrieval problem over a...
295 From ASR to ASP: Evaluating Prompt Attack Vulnerabilities Against Open-Source LLMs
2505.14368
cs.CL
Jiawen Wang, Pritha Gupta, Eyke H\"ullermeier, Xiaoxue Gao, Nancy F. Chen
Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in high-impact applications such as finance, law, and healthcare, systematically inv...
Recent studies demonstrate that Large Language Models (LLMs) are vulnerable to attacks that generate harmful or sensitive outputs. As open-source LLMs are increasingly adopted in high-impact applications such as finance, law, and healthcare, systematically investigating their security risks is becoming increasingly important towards a trustworthy LLM era. This paper comprehensively studies effective prompt injection attacks against 14 widely used open-source and three closed-source LLMs on five ...
296 Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
2506.04734
cs.CLcs.LGcs.AI
Yongfu Zhu, Lin Sun, Jinzhu Wu, Weihong Lin, Xiaoqi Jian
Reasoning models represented by the Deepseek-R1-Distill series have been widely adopted by the open-source community due to their strong performance in mathematics, science, programming, and other domains. However, our study reveals that their benchmark evalua...
Reasoning models represented by the Deepseek-R1-Distill series have been widely adopted by the open-source community due to their strong performance in mathematics, science, programming, and other domains. However, our study reveals that their benchmark evaluation results are subject to significant fluctuations caused by various factors. Subtle differences in evaluation conditions can lead to substantial variations in results. Similar phenomena are observed in other open-source inference models ...
297 Stepwise Intrinsic Rewards for Reasoning in Large Language Models
2602.01034
cs.CLcs.AI
Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne
Reinforcement learning (RL) has become a widely used paradigm for improving the reasoning abilities of large language models (LLMs) and Vision-language models (VLMs). Sparse binary outcome rewards, however, score only final correctness and cannot identify whic...
Reinforcement learning (RL) has become a widely used paradigm for improving the reasoning abilities of large language models (LLMs) and Vision-language models (VLMs). Sparse binary outcome rewards, however, score only final correctness and cannot identify which intermediate steps contributed to it; in multimodal tasks, they may also reward answers driven by linguistic priors rather than visual evidence. Process reward models (PRMs) densify supervision but usually require process annotations, aux...
298 FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases
2602.09163
cs.CLcs.AI
Xingjian Zhang, Sophia Moylan, Ziyang Xiong, Qiaozhu Mei, Yichen Luo
Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert curators to search papers, reconcile...
Scientific knowledge bases accelerate discovery by curating findings from primary literature into structured, queryable formats for both human researchers and emerging AI systems. Maintaining these resources requires expert curators to search papers, reconcile evidence across documents, and produce ontology-grounded annotations. Existing benchmarks usually evaluate isolated subtasks, such as named entity recognition or relation extraction, and therefore do not capture this end-to-end workflow. W...
299 GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
2602.12316
cs.CLcs.AI
Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang
Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We ...
Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. ...
300 Layer-wise Target Propagation: Efficient Component Attribution through Target Centric Propagation
2603.19742
cs.CLcs.LG
Lasse Marten Jantsch, Dong-Jae Koh, Seonghyeon Lee, Young-Kyoon Suh
Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation. While recent efforts have yielded a plethora of attribution methods attempting to balance faithfulness and...
Understanding the internal mechanisms of transformer-based large language models (LLMs) is crucial for their reliable deployment and effective operation. While recent efforts have yielded a plethora of attribution methods attempting to balance faithfulness and computational efficiency, dense component attribution remains prohibitively expensive. In this work, we introduce Layer-wise Target Propagation (LTP), a novel framework that faithfully traces information flow on the frozen transformer in o...
301 Combee: Scaling Prompt Learning for Self-Improving Language Model Agents
2604.04247
cs.CLcs.LGcs.AI
Hanchen Li, Runyuan He, Qizheng Zhang, Changxiu Ji, Qiuyang Mang
Recent advances in prompt learning allow large language model agents to acquire task-relevant knowledge from inference-time context without parameter changes. For example, existing methods (like ACE or GEPA) can learn system prompts to improve accuracy based o...
Recent advances in prompt learning allow large language model agents to acquire task-relevant knowledge from inference-time context without parameter changes. For example, existing methods (like ACE or GEPA) can learn system prompts to improve accuracy based on previous agent runs. However, these methods primarily focus on single-agent or low-parallelism settings. This fundamentally limits their ability to efficiently learn from a large set of collected agentic traces. It would be efficient and ...
302 AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
2605.11398
cs.CLcs.AI
Robin Linzmayer (Department of Computer Science, Columbia University, Department of Biomedical Informatics, Columbia University), Georgianna Lin (Department of Biomedical Informatics
We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflo...
We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflow-specific triage tasks, but they do not offer a unified evaluation of acuity identification across these settings. AcuityBench addresses this gap by harmonizing five public datasets spanning user conversations, online forum posts, clinical...
303 Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents
2606.05828
cs.CLcs.AI
Zeyu Gan, Huayi Tang, Yong Liu
As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With the rapid expansion of available skills, enabling personal agents to learn and ad...
As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With the rapid expansion of available skills, enabling personal agents to learn and adapt to implicit user preferences becomes a critical challenge. However, local deployment constraints preclude complex centralized selection algorithms, creating an urgent need for a lightweight local preference harness. This paper explores ...
304 INFUSER: Influence-Guided Self-Evolution Improves Reasoning
2606.09052
cs.CLcs.LGcs.AI
Siyu Chen, Miao Lu, Beining Wu, Heejune Sheen, Fengzhuo Zhang
Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generato...
Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers f...
305 Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
2608.09444
cs.CLcs.LG
Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis
A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops can...
A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first effici...
306 The Communication Map of a Transformer
2608.22007
cs.CLcs.LG
Richard Zhe Wang
The components of a transformer communicate by writing to and reading from a shared residual stream, and the mechanistic interpretability literature has mapped these connections by hand, one circuit at a time. We present the communication map, which charts eve...
The components of a transformer communicate by writing to and reading from a shared residual stream, and the mechanistic interpretability literature has mapped these connections by hand, one circuit at a time. We present the communication map, which charts every potential communication channel from the geometry of the model's weights alone, generalizing the composition score of Elhage et al. (2021) into a single coupling coefficient covering all 18 connection classes, from head-to-head to neuron...
307 Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
2608.27167
cs.CLcs.AI
Pranav Aggarwal
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. ...
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguisha...
308 PUBG Ally: A Conversational Embodied Agent as an AI Teammate
2609.29837
cs.CLcs.AI
PUBG Ally Team, Irene Chen, Youngin Cho, Seungjun Chung, Jimin Hong
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to...
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A language-...
309 PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
2609.30094
cs.CLcs.AI
Luciano Maldonado
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable throu...
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift...
cs.CV 182 papers
1 AlphaEarth distinguishes cities but compresses urban variation
2609.30356
cs.CV
Andrew Renninger
Cities differ in built form, land cover and development history, complicating comparison across places and time. Satellite foundation models map Earth's surface onto common numerical representations. Yet the tasks and targets used to shape them typically do no...
Cities differ in built form, land cover and development history, complicating comparison across places and time. Satellite foundation models map Earth's surface onto common numerical representations. Yet the tasks and targets used to shape them typically do not focus on cities: globally consistent labels for urban function do not exist, and many datasets - especially land cover and land use classifications - collapse the built environment into few classes. Here we audit the representation, focus...
2 LiTe-GS: Oracle-Efficient Next Best View Selection for 3D Gaussian Splatting
2609.30393
cs.CV
Vivek Pandey, Amirhossein Mollaei Khass, Nader Motee
Selecting informative camera views is critical for efficient training and adaptive refinement in 3D Gaussian Splatting, where each observation significantly influences model parameters. However, information-driven view-selection strategies can require repeated...
Selecting informative camera views is critical for efficient training and adaptive refinement in 3D Gaussian Splatting, where each observation significantly influences model parameters. However, information-driven view-selection strategies can require repeated evaluations of expensive information-gain oracles as the number of candidate views increases. We propose LiTe-GS, an oracle-efficient method for next best view selection in 3D Gaussian Splatting. LiTe-GS reduces the number of information-o...
3 CSCWD: Cross-Scale Channel-wise Knowledge Distillation for Lightweight Tiny Object Detection on Edge Devices
2609.30395
cs.CV
Amir Zamani, Zeinab Ghasemi-Naraghi
Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a ...
Real-time tiny object detection in aerial imagery is constrained by the weak spatial evidence of very small objects and the loss of high-resolution detail in lightweight detectors. This study presents Cross-Scale Channel-wise Knowledge Distillation (CSCWD), a training-time framework that transfers high-resolution spatial representations from a YOLO11m-P2 teacher to a compact YOLO11n student without altering the student's inference architecture. Unlike conventional same-scale feature distillation...
4 What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study
2609.30402
cs.CVcs.CLcs.LGcs.AIcs.MM
Akshit Sharma, Prashant W. Patil
Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a ...
Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language back...
5 ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models
2609.30434
cs.CV
Hiwa Azeez Abbas, Fatemeh Daneshfar, Moloud Abdar
Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone f...
Pre-trained vision-language models such as CLIP can recognize new categories via prompting, but they often struggle when labeled data are scarce or the test distribution shifts. Prompt learning adapts only a small set of parameters while keeping the backbone frozen, yet many existing multimodal prompt learners couple the visual and textual branches weakly and can be brittle in low-shot regimes. We propose ProCAP, a probabilistic cross-attentive prompt learning framework that improves cross-modal...
6 LensDesigner: A Self-Improving Agent for Optical Lens Design
2609.30450
cs.CV
Lei Sun, Haoran Liang, Dannong Xu, Yao Gao, Yuyu Geng
Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. I...
Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and...
7 The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation
2609.30478
cs.CV
Soshun Kihara, Shunsuke Yasuki, Masato Taki
Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in con...
Convolutional neural networks trained on ImageNet are known to exhibit a strong preference for local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in v...
8 Atlases Are Already Inside: Recovering Population Templates from Pretrained Diffusion Models
2609.30566
cs.CV
Jian Shi, John Femiani, Peter Wonka
We present a new inference-time sampler for diffusion models that gives a pretrained model a capability it was never trained for: constructing the atlas of the population it synthesizes. The sampler converges from every random seed to the population's central ...
We present a new inference-time sampler for diffusion models that gives a pretrained model a capability it was never trained for: constructing the atlas of the population it synthesizes. The sampler converges from every random seed to the population's central anatomy, which we call the \emph{intrinsic atlas}. The advantage is threefold. (1) It requires no retraining. A diffusion model that has already learned a coherent population, including the released ones, yields its atlas in a single infere...
9 Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases
2609.30595
cs.CVcs.AI
Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack ground...
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure ind...
10 MVAgent: Multi-Agent Video Generation via Consistent Condition Construction and Shot-Level Policy Optimization
2609.30609
cs.CV
Xiangyu Kong, Wenjie Zhou, Fengping Tian, Lihua Fang, Haoqin Sun
Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determin...
Multi-shot agentic video generation requires consistent character appearance, stable spatial layout across camera angles, and continuous character state between shots. When every shot is a separate request to a frozen generator, repeated text does not determine appearance, layout or state. We therefore recast the problem as condition construction and present MVAgent, a multi-agent pipeline whose agents collaborate through typed conditioning inputs. Because an environment image shows one viewpoin...
11 MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification
2609.30613
cs.CVcs.AI
Zhexiang Li
Dermoscopy classifiers built on Vision Transformers process all image patches uniformly, although diagnostic evidence is concentrated in the lesion region. Existing token pruning methods reduce tokens using generic saliency or similarity signals, but rarely as...
Dermoscopy classifiers built on Vision Transformers process all image patches uniformly, although diagnostic evidence is concentrated in the lesion region. Existing token pruning methods reduce tokens using generic saliency or similarity signals, but rarely ask whether the retained subset still contains the lesion. This paper introduces MedTokenBudget, a supervised post-backbone token routing framework that learns to construct compact lesion-enriched representations when auxiliary lesion masks a...
12 Conditional Predictive Sufficient Statistics for Visual Representation Learning
2609.30647
cs.CV
Yuzhou Hong
A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-fa...
A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch...
13 StarWM: Self-Supervised Trained Attention Routing for Robust World Models
2609.30667
cs.CVcs.LG
Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid, Sebastian Gottwald
A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pi...
A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance for visual tasks, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant infor...
14 Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding
2609.30682
cs.CV
Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform token...
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose SGMA, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. SGMA couples two components: a content-ad...
15 MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning
2609.30698
cs.CV
Peipei Li, Shuhan Xia, Shengyang Liu, Zekun Li, Ran He
Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference co...
Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By ben...
16 SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion
2609.30703
cs.CVcs.AI
Timing Li, Yiming Sun, Boan Tao, Xiyuan Gao, Haifang Cao
Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency ...
Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency propagation across stages. We propose Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion (SAGE), a unified framework integrating frequency equalization, hierarchical alignment, and subband fusion...
17 Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation
2609.30708
cs.CVcs.AI
Tasneem Nasser, Susanne Schmid, Roberto Souza, Naser El-Sheimy
A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on man...
A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on manual annotations. Self-supervised learning has emerged as a promising approach for developing foundation models by enabling the learning of transferable feature representations from large-scale unlabeled medical imaging datasets. In this stu...
18 VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control
2609.30709
cs.CVcs.AI
Kemou Jiang, Maonan Wang, Xingchen Zou, Jiayue Zhu, Yuhang Fu
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling...
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end...
19 TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation
2609.30722
cs.CVcs.AI
Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo
Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topol...
Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topology and unrelated traffic. We introduce TrafficImag, the first benchmark for counterfactual roadside traffic video generation. TrafficImag combines a large-scale roadside dataset (9,022 annotated images, 7,043 deduplicated video clips, and ...
20 EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection
2609.30724
cs.CV
Haoran Sun, Yufan Li, Qichen Zhang, Haoran Zhao, Shuqi Wang
Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cros...
Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A ...
21 Learning Polarization Image Restoration with General Restoration Priors
2609.30728
cs.CV
Chenggong Li, Jinhao Liu, Caiyun Wu, Yidong Luo, Junchao Zhang
Polarization imaging captures distinctive surface and geometric cues that benefit a wide range of vision tasks. However, real-world polarization acquisition is often affected by multiple coupled degradations, making image restoration essential for practical po...
Polarization imaging captures distinctive surface and geometric cues that benefit a wide range of vision tasks. However, real-world polarization acquisition is often affected by multiple coupled degradations, making image restoration essential for practical polarization vision. Existing methods are largely tailored to specific degradations and remain constrained by the limited scale and quality of polarization data. To address these limitations, we develop an all-in-one polarization restoration ...
22 Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation
2609.30733
cs.CV
Shengqi Dang, Zhengxi Yu, Feilin Han, Xingyu Lan, Nan Cao
Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the v...
Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the visual saliency of a specific object during text-to-image generation without requiring any visual priors. Our key insight is that visual saliency is inherently relative: boosting the saliency of a target object also depends on the global sal...
23 From Mono to Stereo: Accelerating Binocular Gaussian Splatting via Reprojection and Selective Patching
2609.30741
cs.CV
Hongfei Zhu, Ling Zhou
Binocular rendering requires two nearby views of the same scene and therefore repeats substantial visibility and shading work. We present a 2D Gaussian Splatting (2DGS) pipeline that fully renders a dominant-eye RGB image and an alpha-weighted depth proxy, rep...
Binocular rendering requires two nearby views of the same scene and therefore repeats substantial visibility and shading work. We present a 2D Gaussian Splatting (2DGS) pipeline that fully renders a dominant-eye RGB image and an alpha-weighted depth proxy, reprojects that image to the affiliated eye, and repairs uncovered pixels. Small interior gaps are interpolated, whereas larger disoccluded regions are identified as regions of interest (ROIs) and selectively re-rendered. The depth proxy reuse...
24 Training-Free Bottleneck Width Planning for Convolutional Autoencoders
2609.30755
cs.CV
Guannan Guo
Multiscale Spectral Rate-Distortion (MS-SRD) estimates the bottleneck channels required at user-supplied spatial cuts from training images and a normalized mean-squared error (NMSE) bound, without fitting a neural network. Its covariance-tail rule is exact for...
Multiscale Spectral Rate-Distortion (MS-SRD) estimates the bottleneck channels required at user-supplied spatial cuts from training images and a normalized mean-squared error (NMSE) bound, without fitting a neural network. Its covariance-tail rule is exact for shared linear block-convolutional autoencoders under squared error. A nested-scale dominance result motivates reporting the activation-parameter Pareto frontier alongside the minimum-latent candidate. At NMSE <= 0.01 on thirteen grayscale ...
25 LLPR: Location-aware learning and physics-based reconstruction for raindrop removal from a single image
2609.30758
cs.CV
Zewei He, Xingyu Liu, Xing Luo, Guizhong Fu, Zixuan Chen
Raindrops can cause occlusion and distortion in the background scenes due to their adherence to windows or camera lenses. Existing raindrop removal methods concentrate on designing sophisticated CNN or Transformer architectures to recover distorted and missing...
Raindrops can cause occlusion and distortion in the background scenes due to their adherence to windows or camera lenses. Existing raindrop removal methods concentrate on designing sophisticated CNN or Transformer architectures to recover distorted and missing texture. In this paper, we try to integrate location information and physical model into off-the-shelf CNN or Transformer architectures to help improve their performance. Specifically, we notice that existing methods deploy a preprocessing...
26 Timo: $\textbf{T}$aming Mult$\textbf{i}$modal Diffusion Transformer for Human $\textbf{Mo}$tion Generation
2609.30761
cs.CV
Zhao Wang, Jiangtao Hu, Jack Yu, Tao Yu
Most existing human motion generation (HMG) methods use cross-attention modules to inject text semantics, but ignore the importance of bidirectional modeling between motion and text tokens, which limits text comprehension. A straightforward idea is introducing...
Most existing human motion generation (HMG) methods use cross-attention modules to inject text semantics, but ignore the importance of bidirectional modeling between motion and text tokens, which limits text comprehension. A straightforward idea is introducing multimodal diffusion transformers (MMDiT), which have shown effective joint text--visual modeling in vision generation, into HMG. However, we find that articulated motion is temporally coherent but weakly correlated across joints, in which...
27 Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes
2609.30769
cs.CVcs.LG
Rushab Rasik Karania, Tomas Maul
Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation c...
Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation contribute to prototype construction? The Within-Instance Prototypical Transformer (WIPT) implements single-query test-time prototype adaptation by jointly transforming one unlabelled query and the labelled support embeddings, then forming q...
28 Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models
2609.30783
cs.CVcs.AI
Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li
Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Though...
Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before localizing the target. Although intuitive, such explicit verbal reasoning introduces substantial attention interference: redundant textual tokens disrupt attention during perception...
29 Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion
2609.30795
cs.CV
Chen-Chieh Liao, Yichen Peng, Yiyi Cai, Y\^ui Ono, Hiroki Hanaoka
Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is sub...
Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requirement is not a universal absolute unit, but a reliable monotonic control axis. We propose Motion Style Slider, a motion-to-motion style transfer framework for endpoint-supervised ...
30 MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization
2609.30855
cs.CV
Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas
Reliable skin lesion segmentation and classification are central to dermoscopic computer-aided diagnosis. Existing multi-task frameworks couple the two tasks architecturally without clinical knowledge, while knowledge-injecting approaches rely on the macroscop...
Reliable skin lesion segmentation and classification are central to dermoscopic computer-aided diagnosis. Existing multi-task frameworks couple the two tasks architecturally without clinical knowledge, while knowledge-injecting approaches rely on the macroscopic ABCD rule, which was not designed for dermoscopy. Dermoscopic diagnosis is grounded in Pattern Analysis, a microscopic framework structured around dermoscopic features. We propose MDSkin-Net, which incorporates cue-level Pattern Analysis...
31 Reliability-Regulated Trajectory Optimization for Progressive COLMAP-Free 3D Gaussian Splatting
2609.30865
cs.CV
Zijian Wu, Jinliang Wang, Zidian Lin, Ying Song, Ziqian Lu
COLMAP-free 3D Gaussian Splatting (3DGS) bypasses computationally expensive structure-from-motion (SfM) pipelines, yet progressive camera pose tracking remains fundamentally vulnerable to error compounding---early pairwise tracking inaccuracies both corrupt su...
COLMAP-free 3D Gaussian Splatting (3DGS) bypasses computationally expensive structure-from-motion (SfM) pipelines, yet progressive camera pose tracking remains fundamentally vulnerable to error compounding---early pairwise tracking inaccuracies both corrupt subsequent frame initializations and remain permanently frozen in the scene representation. Rather than relying on heavyweight external neural priors or treating progressive tracking through isolated heuristic fixes, we propose a unified reli...
32 UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
2609.30928
cs.CVcs.AI
Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao
Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned ...
Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. U...
33 ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos
2609.30934
cs.CV
Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo
Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly cha...
Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forger...
34 Spackle: Completing Large View Single Image NVS with Adaptive Gaussians
2609.30941
cs.CVcs.AI
Xuanzhi Liu, Yuhe Zhou, Xinyi Wu, Zhenyao Wu, Jinghao Chen
Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid...
Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid decoupled frameworks combining feedforward 3D Gaussian Splatting (3DGS) and diffusion models show promise for large-view-deviation NVS, they suffer from capacity competition: a fixed number of Gaussians forces resource shifts from visible ...
35 OneWorld: Learning Consistent Physics Across Actions in World Models
2609.30946
cs.CVcs.AI
Ke He, Yichen Ding, Bin Yang
Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same...
Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across interventions, making it difficult for the model to maintain...
36 DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry
2609.30947
cs.CV
Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari
Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution...
Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution and dynamic range, but their asynchronous measurements complicate reliable correspondence estimation. We present DAPEVO, a learned visual odometry system that estimates image and event correspondences independently at shared patch location...
37 MVVBench: Benchmarking 4D Reasoning in Vision-Language Models
2609.30952
cs.CVcs.AI
Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continu...
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal ax...
38 IDM-Net: A Lightweight Illumination-Decoupled Modulation Network for Low-Light Image Enhancement
2609.30962
cs.CV
Cheng-Yen Hsiao, Jing-Ming Guo
Low-light image enhancement (LLIE) remains challenging for lightweight models because illumination restoration and color fidelity are difficult to optimize simultaneously in the RGB color space. Although recent color-decoupled methods separate luminance and ch...
Low-light image enhancement (LLIE) remains challenging for lightweight models because illumination restoration and color fidelity are difficult to optimize simultaneously in the RGB color space. Although recent color-decoupled methods separate luminance and chrominance representations, they primarily optimize luminance as an enhancement target, leaving its potential as an explicit guidance prior largely unexplored during feature reconstruction. To address this limitation, we propose IDM-Net, a l...
39 Where and When to Force: Routed Forcing for Streaming Avatars
2609.30963
cs.CV
Zihan Su, Siwen Lu, Junhao Zhuang, Zeyue Xue, Haoyang Huang
Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-ste...
Audio-driven streaming avatar generation requires real-time synthesis of speech-synchronized videos with dynamic and diverse motion. Self Forcing uses Distribution Matching Distillation (DMD) to distill bidirectional video diffusion models into causal, few-step generators for real-time streaming. However, DMD minimizes a reverse KL divergence, which is inherently mode-seeking: it causes the student to discard high-dynamic modes and collapse onto static outputs, compressing both dynamics and dive...
40 CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models
2609.30979
cs.CV
Linyuan Gao, Yuan Wu, Yi Chang
Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning ba...
Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for sin...
41 STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence
2609.30981
cs.CV
Siru Zhong, Shenghan Tan, Rihong Yan, Xiaohui Lv, Yuzheng Zhuang
Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilitie...
Reliable online video question answering requires tracking state transitions while selectively abstaining when visual evidence is insufficient. Existing benchmarks focus on static recognition or long-range retrieval, rarely evaluating these coupled capabilities under evolving and incomplete evidence. We present STORM-Bench, comprising 5,736 questions across 630 compact, change-dense episodes spanning five egocentric domains (STORM-Real) and two controlled simulation subsets (STORM-Sim) at 1 FPS....
42 FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators
2609.30982
cs.CVcs.AI
Kai Yao, Marc Juarez
Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one gen...
Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region...
43 Self-Supervised Perceptually Interpretable Monocular Depth Estimation
2609.30987
cs.CV
Zain Ul Abidin, George Dimas, Dimitris K. Iakovidis
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing...
Self-supervised monocular depth estimation (MDE) enables depth prediction from monocular images without requiring ground-truth supervision, making it attractive for large-scale and real-world applications. Despite steady improvements in accuracy, most existing methods remain difficult to interpret, as depth is inferred from RGB representations that obscure the impact of individual perceptual image components. This lack of transparency limits systematic analysis of failure cases and reduces confi...
44 PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution
2609.30988
cs.CV
Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao
Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantia...
Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Transformer SR models are efficient yet often struggle to recover realistic high-frequency details. This motivates a natural question: can diffusion priors be transferred to existing di...
45 PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation
2609.30989
cs.CV
Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya
Purpose: Accurate 6 DoF pose estimation of surgical tools is critical for automa- tion, robotic proprioception, and safe interaction with the tissue operated on. Kinematics-based approaches suffer from accumulated errors due to the cable- driven nature of robo...
Purpose: Accurate 6 DoF pose estimation of surgical tools is critical for automa- tion, robotic proprioception, and safe interaction with the tissue operated on. Kinematics-based approaches suffer from accumulated errors due to the cable- driven nature of robotic arms, while vision-based methods often rely on external markers or trackers. Although more recent vision-based advances have been pro- posed, these two-stage pose estimation methods often lack real-time robustness due to accumulated err...
46 FLIP: Final Layer Inference-Time Probing for Vision-Language Models
2609.30993
cs.CVcs.AI
Drandreb Earl O. Juanico, Rowel O. Atienza
We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under intern...
We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise flooring to the final normalized hidden state before logit...
47 TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking
2609.31005
cs.CV
Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Visi...
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at spar...
48 Refining Cytology Predictions with Conditional Random Fields
2609.31028
cs.CV
Manon Dausort, Tiffanie Godelaine, Karim El Khoury, Maxime Zanella, Christophe De Vleeschouwer
Vision-language models (VLMs) achieve strong zero-shot (ZS) classification on histology images but do not perform as well on cytology, whose stains and cell morphology differ markedly compared to histology. Conditional random fields (CRFs) can refine noisy VLM...
Vision-language models (VLMs) achieve strong zero-shot (ZS) classification on histology images but do not perform as well on cytology, whose stains and cell morphology differ markedly compared to histology. Conditional random fields (CRFs) can refine noisy VLM predictions by propagating information across patches, but existing CRF frameworks were designed for histopathology and do not transfer to cytology datasets, released as independent patch pools spanning multiple staining protocols. We intr...
49 Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images
2609.31040
cs.CV
Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq
Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for pat...
Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for patch-level classification without requiring any annotation, yet these zero-shot (ZS) predictions remain noisy on fine-grained tasks and must be further refined. A promising direction is to refine all predictions jointly, i.e., a transductive ...
50 Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion
2609.31050
cs.CV
Olga Zatsarynna, Denis Korzhenkov, Juergen Gall, Amir Habibian, Mohsen Ghafoorian
Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty vari...
Efficient video generation requires reducing the quadratic cost of self-attention over long spatio-temporal token sequences. Existing efficient-attention methods typically apply the same computation pattern to every token, even though denoising difficulty varies substantially across video regions and evolves throughout the generation process. We introduce HetA-DiT, a heterogeneous attention mechanism that adaptively allocates computation according to token difficulty. A lightweight uncertainty b...
51 Band-Selection Stability and Semantic Segmentation Performance: A Study on Hyperspectral City
2609.31074
cs.CV
Jiarong Li, Imad Ali Shah, Enda Ward, Martin Glavin, Edward Jones
Resource constraints make high-dimensional hyperspectral imaging challenging in autonomous perception, motivating the use of band selection methods. However, the sensitivity of band-selection methods to sampled data and their relationship to semantic segmentat...
Resource constraints make high-dimensional hyperspectral imaging challenging in autonomous perception, motivating the use of band selection methods. However, the sensitivity of band-selection methods to sampled data and their relationship to semantic segmentation models (SSMs) remain underexplored. This study evaluates six band selection methods on ten independently sampled, class-balanced region-of-interest (ROI) sets, yielding 60 top-25 band subsets from the Hyperspectral City V2 (128 bands: 4...
52 DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models
2609.31103
cs.CVcs.AI
Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded ev...
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder fea...
53 Double-stream registration with pyramid fusion for HDR video with alternating exposures
2609.31108
cs.CV
Onofre Martorell, Ivan Pereira-S\'anchez, Antoni Fuentes, Antoni Buades
High dynamic range (HDR) video reconstruction from al\-ter\-na\-ting-exposure sequences remains challenging, especially in regions with extreme luminance variation. We propose a novel HDR reconstruction framework based on dual-stream registration and accurate ...
High dynamic range (HDR) video reconstruction from al\-ter\-na\-ting-exposure sequences remains challenging, especially in regions with extreme luminance variation. We propose a novel HDR reconstruction framework based on dual-stream registration and accurate pyramid fusion. Given three consecutive frames, our method computes optical flow directly with the central frame, while introducing a complementary midpoint displacement strategy to handle cases with severe overexposition. A pyramid fusion ...
54 Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding
2609.31135
cs.CVcs.AIcs.MM
Alberto Presta, Michal Byra, Grzegorz Stefa\'nski, Karol Szurkowski, Eryk Ko{\l}odziejczyk
Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typicall...
Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex training pipelines, or multimodal large language models. We present Pocket-STVG (P-STVG), a lightweight cascade architecture that addresses STVG by combining efficient pre-trained c...
55 Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos
2609.31148
cs.CV
Bowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu, Kuizhuang Liu
Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and reali...
Recent advances in sign language understanding have achieved impressive success on short, single-sentence videos, yet their performance drops sharply when applied to long, continuous sign language videos. To bridge this gap, we focus on a challenging and realistic setting: Visual-only Sentence-level Sign Language Segmentation (Vis-SSLS), which aims to partition continuous sign language videos into non-overlapping sentence-level segments without any caption assistance, serving as a crucial prereq...
56 FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification
2609.31150
cs.CVcs.AI
Muhammad Muhtasim Shahriar, M. M. Golam Hafiz, Saad Aloteibi, Mohammad Ali Moni
Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA),...
Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA), Normal, and squamous cell carcinoma (SCC). FedHisto-PAST v2 combines a frozen HIBOU-B foundation model with parameter-efficient adaptation, stain-conditioned paired-view prediction and feature consistency, reliability-aware prototype learn...
57 HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models
2609.31154
cs.CV
Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li
Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly fo...
Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly follow a static weight paradigm, producing a single frozen adapter that struggles to adapt to diverse prompt variations and suffers from parameter interference when scaling to multiple concepts. We propose \textbf{HyperErase}, a framework for...
58 ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation
2609.31160
cs.CVcs.AI
Donghang Lyu, Zichen Zhang, Oleh Dzyubachyk, Marius Staring
Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely a...
Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely aim at building a generalizable vessel segmentor across anatomies and modalities. While the Seg- ment Anything Model (SAM) has shown promise for med- ical image segmentation, its original design does not fully exploit vascular morphology and...
59 TaskIR: Task-Driven Image Restoration via Degradation Adaptation and Task Feedback
2609.31170
cs.CV
Yanjie Tu, Qingsen Yan, Axi Niu, Wenxuan Cai, Tao Hu
Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Diff...
Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address...
60 Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
2609.31193
cs.CVcs.SDeess.AS
Jihoo Jung, Youngjoon Jang, Joon Son Chung
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systemati...
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual componen...
61 Light Field Primitive for Novel View Synthesis
2609.31198
cs.CV
Liang Chen, Jiahui Ning, Xun Jiang, Xing Xu, Jimmy Ren
We present Light Field Primitives (LFP), a formulation for novel view synthesis that replaces the dense ray database with a compact set of differentiable primitives in the classical two-plane parameterization. Each primitive condenses a group of rays into one ...
We present Light Field Primitives (LFP), a formulation for novel view synthesis that replaces the dense ray database with a compact set of differentiable primitives in the classical two-plane parameterization. Each primitive condenses a group of rays into one learned record, and its response to a query is governed by how closely that query belongs to the group. Rendering a camera ray then reduces to compositing all responses it elicits, and a scene can be optimized directly from posed images and...
62 Preserve-and-Compose Training for Composed Image Retrieval
2609.31202
cs.CV
Sehyun Kwon
Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instea...
Composed image retrieval (CIR) aims to retrieve images that satisfy a user-specified modification while preserving relevant visual content from a reference image. Collecting target images for this purpose is costly, motivating zero-shot CIR methods that instead use target captions as supervision. However, target captions may omit source details that should be preserved. We therefore propose, Preserve-and-Compose Training, which complements target-caption supervision with visual evidence from the...
63 WeaveAgent: A Two-Stage Tool-Routing Agent for Ultra-High-Resolution Remote Sensing Imagery
2609.31234
cs.CV
Zhongyu Pang
Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-rout...
Problem. Ultra-high-resolution (UHR) remote sensing with vague user intents has two bottlenecks: visual tokens are expensive, and tool calling must be format-reliable (pretrained models emit zero tool calls zero-shot). Method. WeaveAgent, a two-stage tool-routing agent, decouples routing from visual perception. Stage A is routing-first: emission is trained, not elicited. Stage B executes conditionally: intrinsic queries enter visual answering (full-scene thumbnail; a WeaveEarth-style evidence bo...
64 Geometric Inconsistency Localization in Multi-View Image Sets
2609.31247
cs.CVcs.AIcs.MM
Xander Staelens, Alb\'eric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for ...
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we intro...
65 Gauss What You Need: Compact Gaussian Splatting Across Scene Scales
2609.31248
cs.CV
Afif Boudaoud, Jiayi Liu, Alexandru Calotoiu, Torsten Hoefler
3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to selec...
3D Gaussian Splatting reconstructs a scene as a collection of Gaussian primitives from a set of posed photographs called the capture. The number of primitives used to represent the scene affects reconstruction quality, storage, and rendering cost. How to select this number automatically across capture scales remains unresolved: configurations effective on standard benchmarks can leave larger captures with too few Gaussians to reconstruct fine details. We observe that the surface to represent, gi...
66 MoTop: Motion-Topological Model For Micro AU Detection
2609.31285
cs.CV
Huai-Qian Khor, Mengting Wei, Yante Li, Chu Kiong Loo, Guoying Zhao
Facial micro-expressions are spontaneous, brief, and subtle facial movements that reveal suppressed emotions in high-stakes environments. In contrast to classic expression analysis, detecting action unit (AU) yields a finer representation of facial movements, ...
Facial micro-expressions are spontaneous, brief, and subtle facial movements that reveal suppressed emotions in high-stakes environments. In contrast to classic expression analysis, detecting action unit (AU) yields a finer representation of facial movements, serving as a preliminary step before defining expression classes and other downstream tasks. Therefore, it represents a crucial upstream task in facial analysis, and improving an AU detection module increases the precision of facial analysi...
67 UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning
2609.31298
cs.CVcs.AI
Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnos...
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity promp...
68 CytoSPM: Open-Vocabulary Cytopathology Detection with Structured Prompt Bank
2609.31314
cs.CV
Wenjie Li, Zishan Xu, Jinyang Huang, Zhengxin Nie, Shichao Kan
Cytopathology detection requires open-vocabulary recognition because cellular categories are fine-grained, long-tailed, and continuously evolving across different organ systems. However, existing cytology detectors are mostly single-domain and closed-set, and ...
Cytopathology detection requires open-vocabulary recognition because cellular categories are fine-grained, long-tailed, and continuously evolving across different organ systems. However, existing cytology detectors are mostly single-domain and closed-set, and there is still no unified benchmark for evaluating open-vocabulary cytopathology detection. We present PentaCyto, a multi-domain benchmark covering cervical, urinary, respiratory, serous fluid, and thyroid cytology, with 24 base categories ...
69 CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support
2609.31326
cs.CVcs.AI
Muhammad Muhtasim Shahriar, Md. Naimur Asif Borno, Saad Aloteibi, Mohammad Ali Moni
Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We in...
Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We introduce CG-HAF, a global-local fusion framework that instead keeps this evidence explicit: averaged holistic severity probabilities from independently trained classifiers are combined with structured lesion-burden descriptors from an object...
70 DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models
2609.31349
cs.CVcs.AI
Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress ...
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. ...
71 Open Vocabulary Domain Unlearning
2609.31356
cs.CVcs.LG
Sumanth Udupa, Mehrtash Harandi, Yadan Luo, Mahsa Baktashmotlagh
Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning ...
Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning ...
72 OpenVAM: Open-World Visual Attention Modeling with VLMs
2609.31364
cs.CV
Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: pr...
Predicting human gaze is a core capability for applications ranging from web/UI design analysis to robotics and human-computer interaction. Yet, most visual attention modeling methods output only a dense saliency map, which is often insufficient for action: practitioners need to connect attention peaks to discrete elements in the scene (what) and understand the drivers of those peaks in context (why), while remaining robust to domain shift across natural images, commercial content, and UI/web la...
73 ContraFM-S2O: Flow Matching-Based One-step SAR-to-Optical Image Translation Model with Contrastive Learning
2609.31378
cs.CV
Mingqian Yu, Wei-kuan Chiang, Qiurui Wang, Peilin Zhao
In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high infe...
In recent years, diffusion models and GAN-based models have become the mainstream approaches for SAR-to-optical image translation, owing to their advantages, such as high-quality generation and stable training. However, they have shortcomings such as high inference latency and the generated optical images suffer from low detail fidelity, often resulting in blurred edges and loss of fine textures. Thus, we propose ContraFM-S2O, which is a flow matching-based model for SAR-to-optical image transla...
74 AxonSynth: Domain-Randomized Synthetic Data for Zero-Shot 3D Axon Segmentation in Light-Sheet Microscopy
2609.31431
cs.CV
Edward Gaibor, Kyriaki-Margarita Bintsi, Carmen Luz Leiva Ureta, Zayneb Bellatif, Chiara Maffei
Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive to obtain. Existing supervised axon segmentation methods rely on target-domain annotations and can be brittle...
Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive to obtain. Existing supervised axon segmentation methods rely on target-domain annotations and can be brittle when tissue type, species, modality, or acquisition conditions change. We present AxonSynth, a domain-randomized synthetic-data framework for training 3D axon segmentation models without manually annotated real training volumes. AxonSynth ...
75 Implicit Neural Representation for Hyperspectral Video Compression
2609.31435
cs.CVcs.LGcs.AI
Alfredo Scalera, Paul Murray, Jaime Zabalza
With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In...
With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In this study, we explore the use of implicit neural representation as a candidate solution. We propose a novel extension of an existing RGB video compression model, achieving Bj{\o}ntegaard Delta PSNR gains of +4.99 dB and Bj{\o}ntegaard Del...
76 From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training
2609.31450
cs.CVcs.AI
Wang Jingxin
Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervi...
Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000...
77 Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis
2609.31456
cs.CV
Mona Gandhi, Cenk Merih Olcay, Kuan-Chieh Lo, Santiago Castro, Christopher W. Myers
Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improv...
Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improve compositional binding. However, this assumption has never been directly quantified. Existing benchmarks evaluate captions only in their composed form, making it impossible to separate the cost of joint reasoning from the cost of recognizi...
78 KneePreM: Towards 3D Knee MRI Foundation Models via Large-Scale Unlabeled Pretraining and Label-Efficient Fine-Tuning
2609.31461
cs.CV
Xinxin Wang, Liam Hazan, Jing Li, Simona Rabinovici-Cohen, Xiaojuan Li
Background: Large volumes of unlabeled knee MRI scans are available across repositories but remain insufficiently leveraged. We developed KneePreM, a knee-specific 3D self-supervised model, and evaluated transfer and label efficiency for classification and seg...
Background: Large volumes of unlabeled knee MRI scans are available across repositories but remain insufficiently leveraged. We developed KneePreM, a knee-specific 3D self-supervised model, and evaluated transfer and label efficiency for classification and segmentation. Methods: A 3D U-Net masked autoencoder was pretrained on 19,011 unlabeled Osteoarthritis Initiative (OAI) MRI series from 4,791 participants. Downstream fine-tuning used full and reduced training sets for fastMRI+ two-label class...
79 SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery
2609.31507
cs.CV
Jiajun Jiang, Chunliang Hua, Zichun Chen, Yanxing Wu, Zeyuan Yang
Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult bec...
Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, limiting geographic diversity and episode scale. To address this, we introduce SatNav, a scalable, long-horizon UAV VLN benchmark built from high-resolution satellite imagery. SatNav...
80 ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos
2609.31509
cs.CVcs.AI
Xuanzhi Liu, Xinyi Wu, Hang Pan, Wensi Huang, Zhenyao Wu
We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-super...
We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-supervision weights based on appearance reliability, degradation risk, and geometric utility, while weakly reactivating useful suppressed frames to maintain trajectory coverage. Since weighting cannot restore details lost to blur or distortion, ...
81 Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics
2609.31514
cs.CV
Javier Mu\~noz-Haro, Ruben Tolosana, Ruben Vera-Rodriguez, Aythami Morales, Julian Fierrez
Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative soluti...
Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative solution, yet standard frameworks work against the forensic task, e.g., their augmentations overwrite the micro-statistics of image formation. This paper introduces Forensic Twins, a Self-Supervised Residual Learning (SSRL) framework whose pretex...
82 Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment
2609.31524
cs.CV
Qing Xu, Yuxiang Luo, Zhen Chen
Surgical scene understanding is critical for computer-assisted intervention, yet laparoscopic cholecystectomy remains challenged by the complex anatomy of the hepatocystic triangle and the risk of bile duct injury. Existing methods for Critical View of Safety ...
Surgical scene understanding is critical for computer-assisted intervention, yet laparoscopic cholecystectomy remains challenged by the complex anatomy of the hepatocystic triangle and the risk of bile duct injury. Existing methods for Critical View of Safety (CVS) assessment typically treat it as a holistic prediction task, mapping visual features directly to criterion-level labels. This black-box paradigm lacks explicit reasoning about anatomical relationships, limiting both interpretability a...
83 Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP
2609.31558
cs.CV
Ahmed Abdelnaby, Mohamed Elmahallawy
Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only ...
Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representati...
84 OC-GS: Gaussian Splatting for Irregular Turntable Capture
2609.31572
cs.CVcs.AI
Jae Joong Lee, Bedrich Benes
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This or...
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly spaced views, OC-GS achieves mean foreground PSNR scores...
85 How Far Can INRs Go? Cross-Domain Parameter-efficient INR-Based Semantic Segmentation for Brain MRI
2609.31573
cs.CV
Ziyao Shang, Pouya Sadeghi, Letian Jiang, Alexander Wong, Sirisha Rambhatla
Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and cross-site distribution shifts. Implicit Neural Representations (INRs) have recently emerged as a lightweight ...
Biomedical image segmentation is central to medical image analysis, but practical deployment often faces limited annotations, memory constraints, and cross-site distribution shifts. Implicit Neural Representations (INRs) have recently emerged as a lightweight alternative for semantic segmentation, achieving competitive performance with substantially fewer parameters than conventional architectures. However, the mechanisms, scaling behavior, and domain generalization abilities of INR-based segmen...
86 GraphWrit3R: End-to-End 3D Scene Graph Writing
2609.31595
cs.CV
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni
3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamen...
3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which dev...
87 FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
2609.31620
cs.CV
Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared...
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion th...
88 WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving
2609.30436
cs.CV
Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li
Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mis...
Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Traj...
89 VkVIO: Cross-platform GPU Acceleration for Visual-Inertial Odometry with Vulkan
2609.30459
cs.CV
Ole Hoffmann, Mateo de Mayo, Daniel Cremers
Perception in robotics and XR fundamentally relies on good state estimation. Visual-inertial odometry (VIO) and Simultaneous Localization and Mapping (VI-SLAM) are proven ways of achieving this goal in a cost-effective and accurate manner. Efficiency in these ...
Perception in robotics and XR fundamentally relies on good state estimation. Visual-inertial odometry (VIO) and Simultaneous Localization and Mapping (VI-SLAM) are proven ways of achieving this goal in a cost-effective and accurate manner. Efficiency in these systems allows for smaller, cooler, and lighter devices. GPU acceleration is a natural approach for reducing latency, thanks to their wide availability in platforms like embedded computers, mobile phones, and XR headsets. However, previous ...
90 QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
2609.30592
cs.CVcs.LG
Nicholas Foley, Devin Marinelli, Donny Moore, Diego Enriquez, Amanda Fernandez
In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical visi...
In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection triple with a single learned unit quaternion per token: the relative quaternion $r_{ij} = q_i^{*} \otimes q_j$ supplies both the attention logit $\operatorname{Re}(r_{ij})$ and a sandwich-product featu...
91 FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception
2609.30629
cs.CVcs.LG
Rajat Bhattacharjya, Minwoo Kim, Arnab Sarkar, Tamoghno Das, Sing-Yao Wu
Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-traine...
Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-trained split interfaces, while stronger channel-aware codecs can impose substantial onboard cost. We present FreshLatent, a lightweight channel-aware latent adapter that trains a power-normalized encoder-decoder through wireless corruption while...
92 Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction
2609.30631
cs.CVcs.SDeess.AS
Rayhan Rashed, Senja Filipi, Ross Cutler
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavi...
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a request...
93 Image Reconstruction from Phase with Untrained Neural Priors
2609.30659
cs.CV
Ene Meco, Ahmet Enis Cetin
Fourier phase encodes important spatial image structure, but recovering an image without measured spectral magnitude requires additional constraints and leaves absolute intensity ambiguous. We propose a projection-based two-stage framework that combines Fourie...
Fourier phase encodes important spatial image structure, but recovering an image without measured spectral magnitude requires additional constraints and leaves absolute intensity ambiguous. We propose a projection-based two-stage framework that combines Fourier-phase and spatial-support constraints with an image-specific neural prior. The first stage alternates constraint enforcement with regularized neural-prior updates, while the second performs phase/support refinement alone with guaranteed c...
94 TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
2609.30670
cs.CVcs.CL
Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, s...
Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, similar scores may correspond to different workloads, failure modes, and operational behavior. We introduce TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark and evaluation framework that makes these factors ...
95 Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
2609.30840
cs.CVcs.LG
Austin Wang, Ziheng Cheng, Lexing Ying
One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentia...
One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution...
96 Universal Drift Correction for Multidimensional Scanning Microscopy
2609.30866
cs.CV
Sangjoon Lee, William Millsaps, Dasol Yoon, Caitlyn Obrero, Guoliang Hu
In scanning microscopy, drift causes the specimen to be sampled at positions displaced from the nominal probe positions. This displacement alters the spatial assignment of the recorded signals and biases quantitative measurements across two-dimensional imaging...
In scanning microscopy, drift causes the specimen to be sampled at positions displaced from the nominal probe positions. This displacement alters the spatial assignment of the recorded signals and biases quantitative measurements across two-dimensional imaging, channel-resolved spectroscopic mapping, and scan-position-resolved diffraction analysis. Here, we extend orthogonal-scan drift correction from 2D images to spectrum images and diffraction datasets. We demonstrate how to recover probe posi...
97 FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models
2609.30980
cs.CV
Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao
Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-ener...
Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are brittle: modest post-processing or lightweight adversarial perturbations readily suppress detection, exposing a fundamental tension between imperceptibility and robustness. We intr...
98 Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
2609.30997
cs.CVcs.LGcs.AI
Kai Yao
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the pro...
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance bet...
99 TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks
2609.31032
cs.CVcs.MM
Tianmeng Fang, Jiancheng Wang, Chen Wang, Liming Wang, Wei Wang
Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefo...
Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V jailbreak as a query-constrained candidate allocation and ranking problem and propose TempQ-Jail. The method combines heterogeneous attack mechanisms to expand candidate coverage, estimates each candidate's end-to-end attac...
100 Quantum Diffusion Models for Medical Image Analysis
2609.31070
cs.CVcs.LGcs.AI
Francesco Aldo Venturelli, Stefano Martina, Marco Parigi, Filippo Caruso, Alba Cervera-Lierta
Quantum Machine Learning is a novel field of research aimed at devising machine learning approaches exploiting principles of quantum mechanics, such as superposition, entanglement and interference. In this context, we present a scalable hybrid Quantum Diffusio...
Quantum Machine Learning is a novel field of research aimed at devising machine learning approaches exploiting principles of quantum mechanics, such as superposition, entanglement and interference. In this context, we present a scalable hybrid Quantum Diffusion Model, and evaluate its use for medical image analysis. Specifically, our method is based on a Discrete-Time Quantum Walk algorithm, executed on a real quantum device, to model the forward dynamics of the diffusion model. For the backward...
101 Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning
2609.31199
cs.CV
Antoine Lorentz, St\'ephane May, Valentine Bellet, Dawa Derksen, Bastien Nespoulous
Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammetry. However, the resulting 3D maps are often contaminated by noise, outliers, and voids. On the other hand, aerial LiDAR provides high-accur...
Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammetry. However, the resulting 3D maps are often contaminated by noise, outliers, and voids. On the other hand, aerial LiDAR provides high-accuracy elevation measurements at a substantially higher cost. In this work, we study diffusion models conditioned both on photogrammetric DSMs and Pl\'eiades imagery to refine vertically co-registered DSMs. We introduce a modified Stable Diffu...
102 FlatClip: A Geometry-Aware Surface-Level Baseline for fMRI Representation Learning
2609.31204
cs.CV
Mo Wang, Wenhao Ye, Zihan Ning, Jiayu Zuo, Junfeng Xia
Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require speciali...
Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized 3D/4D architectures and costly fMRI-specific pretraining. We ask how effectively an image-pretrained encoder can reuse the spatial organization of cortical activity. Motivated by evidence that macroscale brain activity is strongly const...
103 Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling
2609.31207
cs.CV
Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li
Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverag...
Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction ...
104 ChronoFuseGS: Multi-Temporal Gaussian Fusion with Per-Splat Persistence and Change Visualization
2609.31339
cs.CV
Tobias Batik, Diana Marin, Peter K\'an, Hannes Kaufmann
Reconstructing environments where parts of the scene change between captured image sets poses a challenge for 3D scene reconstruction. We present ChronoFuseGS, a multi-temporal Gaussian Splatting approach that addresses this issue by taking multiple separately...
Reconstructing environments where parts of the scene change between captured image sets poses a challenge for 3D scene reconstruction. We present ChronoFuseGS, a multi-temporal Gaussian Splatting approach that addresses this issue by taking multiple separately trained Gaussian Splatting models, each representing a distinct timestep and partially overlapping in geographic coverage, and merging them into a single combined model. By allowing Gaussians from one timestep to contribute to the reconstr...
105 RECAST: From Log Replay to Closed-Loop Driving Simulation with View-Complete Actors
2609.31374
cs.CV
Zijun Zhao, Liewen Liao, Kang Shen, Songan Zhang, Ming Yang
Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded trajectories, exposing views absent from the source log. Existing data-driven simulators reconstruct dynamic a...
Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded trajectories, exposing views absent from the source log. Existing data-driven simulators reconstruct dynamic actors from sparse observations, which can result in rendering artifacts under these viewpoint changes. We introduce RECAST (REconstructing Controllable Actors for Simulation and Testing), a 3D Gaussian Splatting framework that generates a v...
106 Towards Whole-Study Screening for Congenital Heart Disease in Fetal Ultrasound Using Multiple Instance Learning
2609.31376
cs.CV
Mohamed Azzam, Ruobing Liu, Esther C. Ugwueke, Ziyang Xu, Shibiao Wan
Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated fro...
Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated from a study, by a clinician or by a view classifier. We remove that assumption and address CHD screening directly at the level of the whole ultrasound study. We propose a two-stage framework that first learns transferable frame representation...
107 Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization
2609.31383
cs.CVcs.LGcs.AI
Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta
End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects...
End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or ea...
108 InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
2609.31394
cs.CV
Xingyu Miao, Zizun Li, Baole Fang, Kaiwen Song, Tenghui Wang
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion--...
World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-$\Delta$, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-$\Delta...
109 Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
2609.31403
cs.CVcs.CLcs.AI
Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an
Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using ...
Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using the Decoy Font method. The dataset consists of 300 images, each containing text with sharp contour lines superimposed on another text with soft shading. Six recent closed-source models from three different model families were evaluated usin...
110 TemplateCraft: Agentic Visual Template Generation
2609.31451
cs.CVcs.MM
Hongjie Yu, Zhiyuan Fan, Yuzhe Zhang, Jiangcun Du, Zhicheng Gao
The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset pre...
The growing popularity of short videos has driven demand for one-click content creation. Visual templates turn uploaded images into personalized content with preset effects, but reusable template generation still requires substantial manual effort in asset preparation and tool orchestration. We propose TemplateCraft, a multi-agent system that converts natural-language instructions into client-executable templates through planning, material generation, effect-workflow generation, and protocol com...
111 Vision-Based 6-DoF Grasp Pose Estimation for Robot Cloth Unfolding
2609.31452
cs.CV
Domen Tabernik, Peter Nimac, Jan Jeri\'cevi\'c, Danijel Sko\v{c}aj, Andrej Gams
Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and perceptual ambiguity arising from frequent occlusions of critical visual cues such as folds, edges, and grasp p...
Cloth manipulation is a challenging task due to the deformable and high-dimensional nature of cloth, which leads to complex interaction dynamics and perceptual ambiguity arising from frequent occlusions of critical visual cues such as folds, edges, and grasp points. In this work, we tackle cloth unfolding using a regrasping-in-the-air strategy, where one manipulator holds the cloth while the other grasps it at an optimally selected point to unfold it. To this end, we propose CeDiRNet-6DoF, a dee...
112 Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality
2609.31454
cs.CVcs.LGcs.AI
Bradley Scott, Zeqi Luo, Edmond S. L. Ho
Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL...
Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout variance and entropy measures, while the loss is computed ...
113 Uncertainty-Aware Federated Learning for Infant Movement Analysis
2609.31463
cs.CVcs.LGcs.AI
Edmond S. L. Ho
Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving perf...
Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving performance comparable to expert assessment for tasks such as General Movement Assessment (GMA). However, most existing approaches rely on centralized training, requiring data from multiple institutions to be collected and stored at a single si...
114 MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
2609.31553
cs.CVcs.CL
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although ...
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal c...
115 Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
2406.09896
cs.CV
Brun\'o B. Englert, Fabrizio J. Piva, Tommie Kerssies, Daan de Geus, Gijs Dubbelman
Achieving robust generalization across diverse data domains remains a significant challenge in computer vision. This challenge is important in safety-critical applications, where deep-neural-network-based systems must perform reliably under various environment...
Achieving robust generalization across diverse data domains remains a significant challenge in computer vision. This challenge is important in safety-critical applications, where deep-neural-network-based systems must perform reliably under various environmental conditions not seen during training. Our study investigates whether the generalization capabilities of Vision Foundation Models (VFMs) and Unsupervised Domain Adaptation (UDA) methods for the semantic segmentation task are complementary....
116 Adapting Visualization Techniques for Time-Series Anomaly Detection: From Convolutional Neural Networks to Convolutional-Recurrent Neural Networks
2411.04707
cs.CV
Fabien Poirier, Myriam Lamolle
Deep neural networks achieve strong performance on complex tasks but are often regarded as "black boxes," which limits their adoption in domains where transparency is essential. This lack of interpretability raises ethical and legal concerns, particularly in s...
Deep neural networks achieve strong performance on complex tasks but are often regarded as "black boxes," which limits their adoption in domains where transparency is essential. This lack of interpretability raises ethical and legal concerns, particularly in sensitive applications such as security, where automated decisions can have serious consequences. The General Data Protection Regulation (GDPR) reinforces the need to justify decisions made by these systems. In this work, we investigate visu...
117 Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals
2412.11325
cs.CVcs.SDeess.AS
Xiaoxuan Liang, Hong Zhou, Zhaolong Wei, Yansong Li, Shujian Yu
3D human mesh reconstruction (HMR) from RGB images often degrades under poor illumination, occlusion, and non-line-of-sight conditions. Acoustic sensing provides complementary spatial cues but suffers from low spatial resolution. We propose SonicMesh, which, t...
3D human mesh reconstruction (HMR) from RGB images often degrades under poor illumination, occlusion, and non-line-of-sight conditions. Acoustic sensing provides complementary spatial cues but suffers from low spatial resolution. We propose SonicMesh, which, to the best of our knowledge, is the first acoustic--visual framework for robust 3D human mesh reconstruction. SonicMesh first converts ultrasonic echoes into range--azimuth acoustic images through an Inverse Synthetic Aperture Radar (ISAR)-...
118 LadderMIL: Multiple Instance Learning with Coarse-to-Fine Self-Distillation
2502.02707
cs.CV
Shuyang Wu, Yifu Qiu, Ines P. Nearchou, Sandrine Prost, Jonathan A. Fallowfield
Multiple Instance Learning (MIL) for whole slide image (WSI) analysis in computational pathology often neglects instance-level learning as supervision is typically provided only at the bag level, hindering the integrated consideration of instance and bag-level...
Multiple Instance Learning (MIL) for whole slide image (WSI) analysis in computational pathology often neglects instance-level learning as supervision is typically provided only at the bag level, hindering the integrated consideration of instance and bag-level information during the analysis. In this work, we present LadderMIL, a framework designed to improve MIL through two perspectives: (1) employing instance-level supervision and (2) learning inter-instance contextual information at bag level...
119 VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation
2503.10685
cs.CV
Brun\'o B. Englert, Gijs Dubbelman
Unsupervised Domain Adaptation (UDA) enables strong generalization from a labeled source domain to an unlabeled target domain, often with limited data. In parallel, Vision Foundation Models (VFMs) pretrained at scale without labels have also shown impressive d...
Unsupervised Domain Adaptation (UDA) enables strong generalization from a labeled source domain to an unlabeled target domain, often with limited data. In parallel, Vision Foundation Models (VFMs) pretrained at scale without labels have also shown impressive downstream performance and generalization. This motivates us to explore how UDA can best leverage VFMs. Prior work (VFM-UDA) demonstrated that replacing a standard ImageNet-pretrained encoder with a VFM improves generalization. However, it a...
120 What is the Added Value of UDA in the VFM Era?
2504.18190
cs.CV
Brun\'o B. Englert, Tommie Kerssies, Gijs Dubbelman
Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source domain. UDA using Vision Foundation Models (VFMs) with synthetic source data can achieve generalization performanc...
Unsupervised Domain Adaptation (UDA) can improve a perception model's generalization to an unlabeled target domain starting from a labeled source domain. UDA using Vision Foundation Models (VFMs) with synthetic source data can achieve generalization performance comparable to fully-supervised learning with real target data. However, because VFMs have strong generalization from their pre-training, more straightforward, source-only fine-tuning can also perform well on the target. As data scenarios ...
121 RefRef: A Dataset and Benchmark for Reconstructing Refractive and Reflective Objects
2505.05848
cs.CV
Yue Yin, Enze Tao, Weijian Deng, Dylan Campbell
Modern 3D reconstruction and novel view synthesis approaches have demonstrated strong performance on scenes with opaque, non-refractive objects. However, most assume straight light paths and therefore cannot properly handle refractive and reflective materials....
Modern 3D reconstruction and novel view synthesis approaches have demonstrated strong performance on scenes with opaque, non-refractive objects. However, most assume straight light paths and therefore cannot properly handle refractive and reflective materials. The lack of datasets specialized for these effects has impeded efforts to fairly and thoroughly evaluate performance and thereby make progress in this domain. Most existing datasets focus on opaque scenes, while those targeting refractive ...
122 OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
2506.02015
cs.CV
Yoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi, Sungwoong Kim
Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image alignment, often failing to faithfully depict objects with correct attributes su...
Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image alignment, often failing to faithfully depict objects with correct attributes such as color, shape, and spatial relations. To mitigate this issue, previous studies have explored preference optimization methods such as DPO and GRPO, but these approaches incur substantial computational cost, both in constructing preferen...
123 Unsupervised Methods for Video Quality Improvement: A Survey of Restoration and Enhancement Techniques
2507.08375
cs.CV
Alexandra Malyugina, Yini Li, Joanne Lin, Nantheera Anantrasirichai
Video restoration and enhancement are critical not only for improving visual quality, but also as essential pre-processing steps to boost the performance of a wide range of downstream computer vision tasks. This survey presents a comprehensive review of video ...
Video restoration and enhancement are critical not only for improving visual quality, but also as essential pre-processing steps to boost the performance of a wide range of downstream computer vision tasks. This survey presents a comprehensive review of video restoration and enhancement techniques with a particular focus on unsupervised approaches. We begin by outlining the most common video degradations and their underlying causes, followed by a review of early conventional and deep learning me...
124 Spatial Information Bottleneck for Interpretable Visual Recognition
2511.09239
cs.CV
Kaixiang Shu
Deep neural networks typically learn spatially entangled representations that conflate discriminative foreground features with spurious background correlations, thereby undermining model interpretability and robustness. We propose a novel understanding framewo...
Deep neural networks typically learn spatially entangled representations that conflate discriminative foreground features with spurious background correlations, thereby undermining model interpretability and robustness. We propose a novel understanding framework for gradient-based attribution from an information-theoretic perspective. We prove that, under mild conditions, the Vector-Jacobian Products (VJP) computed during backpropagation form minimal sufficient statistics of input features with ...
125 Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
2511.20561
cs.CVcs.CL
Yuwei Niu, Weiyang Jin, Jiaqi Liao, Chaoran Feng, Peng Jin
Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation in Unified Multimodal Models? To investigate this, we introduce UniSandbox, a decoupled evaluation fra...
Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation in Unified Multimodal Models? To investigate this, we introduce UniSandbox, a decoupled evaluation framework paired with controlled, synthetic datasets to avoid data leakage and enable detailed analysis. Our findings reveal a significant understanding-generation gap, which is mainly reflected in two key dimensions: reasoning generation and ...
126 Anchor to Expand: Semantic Anchoring for Personalized Text-to-Image Diffusion Models
2511.22245
cs.CV
Seoyun Yang, Gihoon Kim, Taesup Kim
Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images. However, learning a new concept while building on the prior knowledge of the pretrained model remains a key chall...
Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images. However, learning a new concept while building on the prior knowledge of the pretrained model remains a key challenge. When personalization focuses on learning the target concept, the model tends to overfit the reference examples and degrade its general capability. In contrast, emphasizing prior preservation can hinder capturing distinctive personaliz...
127 Frequency-Decomposed Avatar Representation for Varying Camera Distances
2512.03593
cs.CV
David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue
We present a CloseUpAvatar - a novel approach for articulated human avatar representation supporting a wider range of camera motions, while preserving rendering quality for close-up views. CloseUpAvatar represents an avatar as a set of textured planes with fre...
We present a CloseUpAvatar - a novel approach for articulated human avatar representation supporting a wider range of camera motions, while preserving rendering quality for close-up views. CloseUpAvatar represents an avatar as a set of textured planes with frequency-decomposed learnable textures for low and high-frequency detail. The method automatically switches to high-frequency textures when the camera comes close to the avatar's surface and gradually reduces their impact as the camera moves ...
128 Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement
2512.08982
cs.CVcs.AI
Jian Xu, Wei Chen, Shigui Li, Delu Zeng, John Paisley
Retinex-based low-light image enhancement benefits from separating reflectance and illumination, yet recent generative approaches often rely on iterative sampling and are difficult to deploy under strict latency budgets. Consistency models offer a natural rout...
Retinex-based low-light image enhancement benefits from separating reflectance and illumination, yet recent generative approaches often rely on iterative sampling and are difficult to deploy under strict latency budgets. Consistency models offer a natural route to one-step restoration, but direct adaptation to Retinex-factorized enhancement is unstable: one-step inference is evaluated at the high-noise endpoint, whereas standard training schedules provide little supervision there, and temporal s...
129 Demo: Generative AI helps Radiotherapy Planning with User Preference
2512.08996
cs.CVcs.LGcs.AI
Riqiang Gao, Simon Arberet, Martin Kraus, Han Liu, Wilko FAR Verbakel
Radiotherapy planning is a highly complex process that often varies significantly across institutions and individual planners. Most existing deep learning approaches for 3D dose prediction rely on reference plans as ground truth during training, which can inad...
Radiotherapy planning is a highly complex process that often varies significantly across institutions and individual planners. Most existing deep learning approaches for 3D dose prediction rely on reference plans as ground truth during training, which can inadvertently bias models toward specific planning styles or institutional preferences. In this study, we introduce a novel generative model that predicts 3D dose distributions based solely on user-defined preference flavors. These customizable...
130 Prompt-Based Continual Compositional Zero-Shot Learning
2512.09172
cs.CVcs.AI
Sauda Maryam, Sara Nadeem, Faisal Qureshi, Mohsen Ali
We tackle continual adaptation of vision-language models to new attributes, objects, and their compositions in Compositional Zero-Shot Learning (CZSL), while preventing forgetting of prior knowledge. Unlike classical continual learning where classes are disjoi...
We tackle continual adaptation of vision-language models to new attributes, objects, and their compositions in Compositional Zero-Shot Learning (CZSL), while preventing forgetting of prior knowledge. Unlike classical continual learning where classes are disjoint, CCZSL is more complex as attributes and objects may reoccur across sessions while compositions remain unique. Built on a frozen VLM backbone, we propose the first Prompt-based Continual Compositional Zero-Shot Learning (PromptCCZSL) fra...
131 Geometric-Photometric Event-based 3D Gaussian Ray Tracing
2512.18640
cs.CVcs.AI
Kai Kohyama, Yoshimitsu Aoki, Guillermo Gallego, Shintaro Shiba
Event cameras offer a high temporal resolution over traditional frame-based cameras, which makes them suitable for motion and structure estimation. However, it has been unclear how event-based 3D Gaussian Splatting (3DGS) approaches could leverage fine-grained...
Event cameras offer a high temporal resolution over traditional frame-based cameras, which makes them suitable for motion and structure estimation. However, it has been unclear how event-based 3D Gaussian Splatting (3DGS) approaches could leverage fine-grained temporal information of sparse events. This work proposes GPERT, a framework to address the trade-off between accuracy and temporal resolution in event-based 3DGS. Our key idea is to decouple the rendering into two branches: event-by-event...
132 Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
2512.21004
cs.CV
Jinghan Li, Yang Jin, Hao Jiang, Yadong Mu, Yang Song
Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative pretraining methods still rel...
Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative pretraining methods still rely on BERT-style masked modeling, which often disregards the temporal information essential for video analysis. The few existing autoregressive visual pretraining methods suffer from issues such as inaccurate semantic localization and poor g...
133 Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models
2602.02043
cs.CVcs.AI
Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
Modern vision-language models struggle with basic compositional reasoning, failing to bind attributes to objects or relations to their referents. Existing benchmarks either rely on noisy real images that conflate confounding visual variables with the reasoning...
Modern vision-language models struggle with basic compositional reasoning, failing to bind attributes to objects or relations to their referents. Existing benchmarks either rely on noisy real images that conflate confounding visual variables with the reasoning failure, or use simplistic synthetic scenes lacking the realism modern VLMs are tuned for. We introduce \textbf{Auto-Comp}, a fully automated, concept-driven pipeline that bridges this gap by generating photorealistic compositional benchma...
134 VLANeXt Family: A Systematic Study of VLA Models from Core Recipes to Emerging Paradigms
2602.18532
cs.CVcs.AI
Xiao-Ming Wu, Kang Liao, Yihang Luo, Bin Fan, Jian-Jian Jiang
Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented ...
Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving spa...
135 A Multi-Stage Framework for Kuzushiji Character Recognition in Japanese Historical Documents
2602.19086
cs.CV
Rui-Yang Ju, Kohei Yamashita, Hirotaka Kameko, Shinsuke Mori
Kuzushiji was a widely used cursive writing system in pre-modern Japan. Due to simplification and glyph variation, most modern Japanese readers cannot read Kuzushiji characters. Consequently, recent studies have developed optical character recognition (OCR) sy...
Kuzushiji was a widely used cursive writing system in pre-modern Japan. Due to simplification and glyph variation, most modern Japanese readers cannot read Kuzushiji characters. Consequently, recent studies have developed optical character recognition (OCR) systems for Kuzushiji. Despite recent progress, Kuzushiji character recognition (KCR) in Japanese historical documents remains challenging because of seal-character overlap and complex layouts, which interfere with character recognition and h...
136 Scale-invariant Gaussian derivative residual networks
2603.02843
cs.CVcs.LG
Andrzej Perzanowski, Tony Lindeberg
Generalisation across image scales remains a fundamental challenge for deep networks, which often fail to handle images at scales not seen during training (the out-of-distribution problem). In this paper, we present provably scale-invariant Gaussian derivative...
Generalisation across image scales remains a fundamental challenge for deep networks, which often fail to handle images at scales not seen during training (the out-of-distribution problem). In this paper, we present provably scale-invariant Gaussian derivative residual networks (GaussDerResNets), constructed out of scale-covariant Gaussian derivative residual blocks coupled in cascade, aimed at addressing this problem. By adding residual skip connections to the previous notion of Gaussian deriva...
137 NBAvatar: Neural Billboards Avatars with Realistic Hand-Face Interaction
2603.12063
cs.CV
David Svitov, Mahtab Dahaghin, Pietro Morerio, Alessio Del Bue
We present NBAvatar - a method for realistic rendering of head avatars handling non-rigid deformations caused by hand-face interaction. To this end, we introduce a novel hybrid implicit-explicit representation for animated avatars by combining the training of ...
We present NBAvatar - a method for realistic rendering of head avatars handling non-rigid deformations caused by hand-face interaction. To this end, we introduce a novel hybrid implicit-explicit representation for animated avatars by combining the training of explicit oriented planar primitives with implicit neural rendering. Such a combination of representations in the end-to-end pipeline enables NBAvatar to handle temporally and pose-consistent geometry, along with fine-grained appearance deta...
138 GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow
2603.26571
cs.CVcs.AI
Ziyue Zeng, Xun Su, Haoyuan Liu, Bingyu Lu, Yui Tatsumi
At ultra-low bitrates, high-fidelity reconstruction requires sampling plausible videos from the posterior rather than regressing to oversmoothed conditional means. We propose Generative Video Codebook Codec (GVCC), a zero-shot framework in which a pretrained v...
At ultra-low bitrates, high-fidelity reconstruction requires sampling plausible videos from the posterior rather than regressing to oversmoothed conditional means. We propose Generative Video Codebook Codec (GVCC), a zero-shot framework in which a pretrained video generative model serves directly as the decoder, and the transmitted bitstream specifies its generation trajectory. Modern rectified-flow video models are typically sampled with deterministic ODE solvers, which leave no per-step stocha...
139 At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections
2604.10766
cs.CV
Ming-Yang Ho, Alberto Bartesaghi
Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM constraints prohibit processing an entire 3D tomogram, forcing current methods to rely on slow sliding-window in...
Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM constraints prohibit processing an entire 3D tomogram, forcing current methods to rely on slow sliding-window inference over extracted subvolumes. To overcome this, we propose FullTilt, an end-to-end framework that redefines 3D detection by operating directly on aligned 2D tilt-series. Because a tilt-series contains significantly fewer images than sl...
140 Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
2604.17422
cs.CVcs.MM
Shaoguang Wang, Weiyu Guo, Ziyang Chen, Xuming Hu, Hui Xiong
Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dense frame sequences. Prevailing keyframe-selection methods rely on either a single visual-centric metric (e.g., CLI...
Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive cost of processing dense frame sequences. Prevailing keyframe-selection methods rely on either a single visual-centric metric (e.g., CLIP similarity) or a static fusion of heuristic scores. This "one-size-fits-all" paradigm frequently fails: visual-only metrics are ineffective for plot-driven narrative queries, while indiscriminately adding textual scores introduces severe ...
141 From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation
2605.08712
cs.CV
Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang
Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic...
Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic-to-visual lifting paradigm that converts articulated kinematics into a unified set of five image-aligned control modalities. Building on this representation, we introduce a hierarchically routed visual control framework that selectively ac...
142 CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
2605.11723
cs.CVcs.AI
Jiyuan Wang, Huan Ouyang, Jiuzhou Lin, Chunyu Lin, Dewen Fan
In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial ...
In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiotemporal Chain-of-Thought reasoning. To equip the model with these capabilities, we construct the first large-scale generated video anomaly d...
143 Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models
2605.19398
cs.CVcs.AI
Wooseok Jeon, Seungho Park, Seunghyun Shin, Sangeyl Lee, Hyeonho Jeong
Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional training or sacrifice fid...
Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional training or sacrifice fidelity to the reference image. In this work, we identify reference-frame dominance as a key mechanism behind motion suppression. We observe that non-reference frames in I2V models allocate excessive self-attention to reference-frame key toke...
144 C2P-VAR: Continual and Compositional Personalization in Visual Autoregressive Models
2605.19750
cs.CV
Junhao Li, Xinhao Zhong, Yi sun, Yuxia Qiao, Bin Chen
Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation, yet their personalization capabilities remain largely limited to static, single-concept settings. In practice, users may continuously introduce new ...
Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation, yet their personalization capabilities remain largely limited to static, single-concept settings. In practice, users may continuously introduce new concepts and wish to compose multiple personalized concepts within a single image. Such scenarios pose two fundamental challenges: catastrophic forgetting during sequential personalization and feature interference during multi-concept compo...
145 Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
2606.01900
cs.CV
Muhammed Burak Kizil, Enes Sanli, Niloy J. Mitra, Xuelin Chen, Erkut Erdem
Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spat...
Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language-driven, human-centric camera framing in generative video. Our core insight is that professional filmmakers co...
146 EmoZone-Talker: Regional Semantic Control of Audio-Driven 3DGS Talking Heads via Facial Action Units
2606.15848
cs.CV
Tingting Chen, Shaojun Wang, Huaye Zhang, Diqiong Jiang, Chenglizhao Chen
3D Gaussian Splatting (3DGS) has shown strong potential for high-fidelity talking head synthesis. However, enabling fine-grained, interpretable, and editable facial expression control remains fundamentally challenging due to intrinsic conflicts between speech-...
3D Gaussian Splatting (3DGS) has shown strong potential for high-fidelity talking head synthesis. However, enabling fine-grained, interpretable, and editable facial expression control remains fundamentally challenging due to intrinsic conflicts between speech-driven facial dynamics and explicit expression signals. Existing methods rely on implicit multimodal fusion, leading to spatial entanglement and temporal instability. We present EmoZone-Talker, a novel framework that reformulates audio-driv...
147 Holo-World: Unified Camera, Object and Weather Control for Video World Model
2606.20083
cs.CV
Xiangchen Yin, Wenzhang Sun, Jiahui Yuan, Zijie Liu, Yinda Chen
Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls remain isolated, and weather generation typically relies on a source video or rec...
Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls remain isolated, and weather generation typically relies on a source video or reconstructed scene that already specifies future structure. We study a first-frame-anchored source-to-state setting, where the model starts from a single image and follows explicit camera and object controls and an optional weather instructio...
148 Cross-Modality Structural Guidance in 3D Latent Diffusion for Robust FLAIR Super-Resolution
2606.25255
cs.CV
Haoyu Lan, Jiazhen Zhang, John Onofrey, Bino Varghese, Nasim Sheikh-Bahaei
High-resolution (HR) MRI acquisition is often hampered by scan time constraints, resulting in anisotropic or low-resolution scans (e.g., thick-slice FLAIR) that limit diagnostic accuracy. While deep learning-based super-resolution (SR) methods show promise, th...
High-resolution (HR) MRI acquisition is often hampered by scan time constraints, resulting in anisotropic or low-resolution scans (e.g., thick-slice FLAIR) that limit diagnostic accuracy. While deep learning-based super-resolution (SR) methods show promise, they often hallucinate anatomical details, which can compromise brain structural integrity. To mitigate this limitation, we introduce MR-DiffuSR, a Multi-Resolution Diffusion-based Super-Resolution framework that incorporates HR T1w structura...
149 DreamSat-Pose: Spacecraft Pose Estimation from Single-View 3D Reconstructions and Learned 2D-3D Feature Matching
2607.13449
cs.CVcs.LG
Josiane Uwumukiza, Jocelyn Zhao, Giovanni Lavezzi, Giacomo Battaglia, Paolo Panicucci
6-DoF pose estimation is a critical task in autonomous rendezvous and proximity operations. In the case of an unknown target, this task becomes challenging as it shall be paired with the reconstruction of the target shape model. In this article, we propose a n...
6-DoF pose estimation is a critical task in autonomous rendezvous and proximity operations. In the case of an unknown target, this task becomes challenging as it shall be paired with the reconstruction of the target shape model. In this article, we propose a novel framework for single-shot shape and pose estimation of unknown spacecraft objects. Given a single image, we first reconstruct a 3D shape model of the target, then estimate the relative six-degrees-of-freedom pose by learning dense 2D-3...
150 Importance-Aware OBS Pruning for Diffusion Models
2607.20048
cs.CV
Ba-Thinh Lam, Srijan Das, Hieu Le
We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps -- derived from conditioning signals or ...
We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps -- derived from conditioning signals or model attention -- into the pruning objective. This produces parameter rankings aligned with perceptual relevance rather than uniform reconstruction error. On MS-COCO dataset, our proposed approach consistently retains subject fidelity and ...
151 When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation
2607.22771
cs.CVcs.AI
Renjie Liang, Zijian Xu, Jinqian Pan, Chengkun Sun, Zhengkang Fan
Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive search. A cheap pr...
Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning every candidate through the full language model and comparing downstream scores, an enormously expensive search. A cheap probe on the encoder's representation promises a way out, but whether it forecasts the expensive outcome has never been tested. We test this with CheapCT on report generation and on MeasureVQA, a new VQA dataset we build. MeasureVQA scores th...
152 Towards Unified Dynamic Face Landmark Detection
2608.10346
cs.CVcs.AI
Sebastian Regalado, Varshanth R. Rao, Ruowei Jiang, Parham Aarabi, Igor Gilitschenski
Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each ``$N$-point'' benchmark dataset, ...
Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different network parameters need to be trained independently for each ``$N$-point'' benchmark dataset, and (2) a model trained on an ``$N$-point'' dataset reliably outputs only the $N$ landmarks. In our work, we first conceptualize Face Part-Anchored Landmark Positions (FPALPs), wherein each landmark is treated as a progression value between...
153 A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources
2608.13183
cs.CV
Brun\'o B. Englert, Gijs Dubbelman
Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustainable or unnecessary, and significant be...
Visual foundation models are a cornerstone of image and video understanding but typically require large amounts of data and computation. The current scale required for pretraining visual foundation models may be unsustainable or unnecessary, and significant benefits arise when effective models can be obtained with fewer resources. To better understand how self-supervised learning (SSL) objectives behave under resource constraints, we conduct a controlled study of image and video SSL objectives u...
154 AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation
2608.29081
cs.CV
Soumyaratna Debnath, Weiming Zhang, Shriram Damodaran, Dingwen Xiao, Addison Lin Wang
Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical...
Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical structure and stable viewpoints, which are frequently violated in real-world imagery due to unconstrained camera motion, introducing contextual and geometric ambiguity. Consequently, they lack adaptive mechanisms to handle such ambiguity, ...
155 Learning with Volterra Neural Networks: A System Theoretic Perspective
2609.01928
cs.CV
Haoyu Yun, Hamid Krim, Yufang Bao
Higher-order interaction components are important for signal, image, and video modeling, but explicit high-order operators often suffer from rapidly increasing parameter and computational costs. This paper presents kVNN, a learnable kernelized Volterra Neural ...
Higher-order interaction components are important for signal, image, and video modeling, but explicit high-order operators often suffer from rapidly increasing parameter and computational costs. This paper presents kVNN, a learnable kernelized Volterra Neural operator for compact higher-order filtering. The motivation is to use kernelization to improve the efficiency of Volterra-type neural operators while providing a structured interpretation of their higher-order components. The proposed formu...
156 Learning to Track from Privileged Target Appearances
2609.02471
cs.CV
Xin Chen, Jiao Xu, Dong Wang, Huchuan Lu, Kede Ma
Target templates define what a visual tracker searches for, yet the templates available at inference trade off localization certainty with appearance freshness: the initial ground-truth template is exact but becomes stale, whereas recent templates better refle...
Target templates define what a visual tracker searches for, yet the templates available at inference trade off localization certainty with appearance freshness: the initial ground-truth template is exact but becomes stale, whereas recent templates better reflect the current appearance but are cropped from uncertain predictions. We quantify this bottleneck with a non-deployable oracle that supplies an exact current-frame target crop, improving AUC on LaSOT by 15.2 percentage points. This gap reve...
157 SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
2609.05533
cs.CVcs.LG
Cheng Yin, Wang Xu, Junpeng Yang, Sikyuen Tam, Hanyu Liu
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms for VLAs, such as retrieval banks, learned compressors and recurrent states, mu...
Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms for VLAs, such as retrieval banks, learned compressors and recurrent states, must decide what to keep from the past before knowing what a future decision will require. They were motivated by the assumption that minute-scale history is too large to process directly, which no longer holds for modern VLM backbones. We pr...
158 PLSR: Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion
2609.06436
cs.CV
Yuxin Liu, Minshan Xie, Jiawen Liang, Runsong Zhu, Chi-Wing Fu
High-resolution 3D asset generation is vital in various 3D applications. Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details. In this paper, we tackle the challenge of generating m...
High-resolution 3D asset generation is vital in various 3D applications. Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details. In this paper, we tackle the challenge of generating more detailed, higher-resolution 3D objects by introducing a 3D super-resolution (SR) framework built on existing 3D generative foundation models. To this end, we design PLSR, a progressive and localized super-resolution solution to achieve ...
159 Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding
2609.06475
cs.CV
Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian
Large vision-language models (LVLMs) have recently achieved remarkable progress in general-purpose video understanding. However, their application to real-world surveillance remains challenging due to the lack of large-scale domain-specific datasets and the li...
Large vision-language models (LVLMs) have recently achieved remarkable progress in general-purpose video understanding. However, their application to real-world surveillance remains challenging due to the lack of large-scale domain-specific datasets and the limitation of passive observation from fixed viewpoints. In surveillance scenarios, critical visual evidence can be easily missed when targets are distant, small, occluded, or move beyond the current camera view. In this work, we introduce Ca...
160 ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
2609.16284
cs.CVcs.AI
Yan Zhu, Yongbo Chen, Zhengming Ding, Rebecca Faust
Query-conditioned vision-language models enable fine-grained interpretation by revealing which visual content supports a given textual query and how this evidence changes across queries. However, semantically, sentence-level evidence does not necessarily decom...
Query-conditioned vision-language models enable fine-grained interpretation by revealing which visual content supports a given textual query and how this evidence changes across queries. However, semantically, sentence-level evidence does not necessarily decompose into object-specific contributions, while spatially, object-level evidence can remain entangled with co-occurring objects and surrounding scene context. Across multiple VLM architectures and independent benchmarks, we observe persisten...
161 A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data
2609.16597
cs.CVcs.AI
Yinong Wang (Joyce), Jianwen Chen (Joyce), Zhou Chen (Joyce), Shuwen Kuang (Joyce), Haoning Jiang (Joyce)
We developed BrainVLM to classify all 12 World Health Organization (WHO) 2021 brain tumor types. BrainVLM integrates an uncertainty quantification strategy to indicate prediction reliability and a module for generating radiology reports to elucidate the clinic...
We developed BrainVLM to classify all 12 World Health Organization (WHO) 2021 brain tumor types. BrainVLM integrates an uncertainty quantification strategy to indicate prediction reliability and a module for generating radiology reports to elucidate the clinical rationale. BrainVLM was trained on multi-modal data (MRI scans, demographics, and radiology reports) from 40,043 individuals. It was validated on 5,211 patients with pathologically confirmed brain tumors, including 3,877 held-out patient...
162 Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
2609.20633
cs.CV
Yulong Chen, Ziqian Zhang, Haoyu Zhang, Ao He, Yaxing Wang
Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal au...
Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. As the first to explore training-free image editing with Generative Refinement Networks (GRN), we observe that its refinement p...
163 CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation
2609.21455
cs.CV
Haoran Qin (Harbin Institute of Technology, China), Renlong Wu (Harbin Institute of Technology, China), Tianyu Huang (Harbin Institute of Technology
While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate e...
While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a ...
164 MM-ContextFold: Context Folding for Multimodal Agentic Retrieval
2609.23121
cs.CVcs.AIcs.MM
Yang Tian, Fan Liu, Jingyuan Zhang, Zhenyang Li, Yupeng Hu
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-gro...
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address thi...
165 Retrieval Geometry Shapes Cache-Based Clip Adaptation
2609.23409
cs.CV
Mahir Shahriar Tamim, Md. Samiul Alim, Azmine Toushik Wasi, Shahriyar Zaman Ridoy, Meharun Nesa
Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image-image retrieval as fixed, leaving open...
Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image-image retrieval as fixed, leaving open how much adaptation depends on the retrieval space itself. We study this question by fixing the memory and changing only the retrieval encoder, finding that the same memory can yield very different gains: across sixteen retrieval spaces, I...
166 You've Seen Enough: Quality-Constrained Image Coding for Machines
2609.25108
cs.CVcs.AI
Khoa Pham-Dinh, Sanaz Nami, Hamed Rezazadegan Tavakoli, Moncef Gabbouj, Farhad Pakdaman
Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validat...
Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validate the decisions. Inspired by just-noticeable distortion, we cap human-observed quality at a desired level and devote the remaining bits to machine performance. Specifically, joint compression-segmentation training is recast as a constrained...
167 MIND the Gap: A Geographic Implicit Neural Representation with Adjustable Spatial Scale
2609.25454
cs.CVcs.LG
Isaac Corley, Arjun Rao, Esther Rolf, Konstantin Klemmer, Evan Shelhamer
Geographic measurements are often sparse, leaving large areas without labels for the quantities we want to map. Geographic implicit neural representations (INRs) provide coordinate-based embeddings that can be combined with sparse labels to predict at unsample...
Geographic measurements are often sparse, leaving large areas without labels for the quantities we want to map. Geographic implicit neural representations (INRs) provide coordinate-based embeddings that can be combined with sparse labels to predict at unsampled locations without satellite imagery at inference. Yet existing INRs are largely evaluated with random holdouts, leaving their ability to generalize across larger geographic gaps unclear. We introduce Matryoshka Implicit Neural Distillatio...
168 NV-Reason-CT: 3D Visual Language Model for CT Analysis
2609.27511
cs.CVcs.AI
Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon
We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and...
We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing ...
169 Track2Art: Articulated Object Model Recovery with Visual-Geometric Track Representations
2609.27675
cs.CV
Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam
Understanding articulated objects is fundamental for robotic interaction, requiring accurate rigid-part discovery and the recovery of their kinematic relations. Existing approaches often treat articulation as a by-product of reconstructed geometry or recover i...
Understanding articulated objects is fundamental for robotic interaction, requiring accurate rigid-part discovery and the recovery of their kinematic relations. Existing approaches often treat articulation as a by-product of reconstructed geometry or recover it through per-instance optimization. We instead build on the hypothesis that articulation is directly observable from persistent motion: points on the same rigid part move coherently, while relative motion between parts reveals their kinema...
170 LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder
2609.28327
cs.CV
Andrei Arhire, Mihaela-Elena Breab\u{a}n, Radu Timofte
We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Pro...
We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade. The cascade combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module, which uses temporary channel expa...
171 M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals
2609.28684
cs.CVcs.LG
Vin\'icius da Silva, Isabelle Melo, Matheus Bessa, Guilherme Schardong, Luiz Schirmer
Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training effici...
Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training efficiency, rendering speed, and noise robustness: single-MLP approaches are expensive at inference, grid-based representations are fast but can limit surface smoothness and overfit input noise, and previous multiscale approaches frequently captu...
172 Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study
2609.29433
cs.CVcs.AI
Raghavan Lavanya, Yangqin Feng, Ten Cheer Quek, Quan V. Hoang, Linda Yi-Chieh Poon
Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting c...
Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting conditions such as high myopia (HM). We developed and validated a Vision Transformer-based deep learning (DL) model for glaucoma detection across multi-ethnic cohorts with and without HM. Methods: A ViT-B/16 model with predictive uncertainty...
173 AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
2609.29816
cs.CVcs.SD
Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training o...
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapting it to joint audio-video generation is challenging. Heterogeneous multimodal rewards entangle learning signals and complicate credit assignment. Joint optimization of two modality towers is computa...
174 DeepFedNAS: Efficient Hardware-Aware Architecture Adaptation for Heterogeneous IoT Federations via Pareto-Guided Supernet Training
2601.15127
cs.CVcs.LG
Bostan Khan, Masoud Daneshtalab
Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet existing Federated Neural Architecture Search (FedNAS) methods suffer from unguided supernet training and prohibitivel...
Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet existing Federated Neural Architecture Search (FedNAS) methods suffer from unguided supernet training and prohibitively costly post-training search pipelines that validate thousands of subnets to construct learned accuracy predictors. We introduce DeepFedNAS, a two-phase framework built on a multi-objective fitness function that synthesizes information-the...
175 Pseudo-Invertible Neural Networks
2602.06042
cs.CVcs.LG
Yamit Ehrlich, Nimrod Berman, Assaf Shocher
The Moore-Penrose Pseudo-inverse (PInv) serves as the fundamental solution for linear systems. In this paper, we propose a natural generalization of PInv to the nonlinear regime in general and to neural networks in particular. We introduce Surjective Pseudo-in...
The Moore-Penrose Pseudo-inverse (PInv) serves as the fundamental solution for linear systems. In this paper, we propose a natural generalization of PInv to the nonlinear regime in general and to neural networks in particular. We introduce Surjective Pseudo-invertible Neural Networks (SPNN), a class of architectures explicitly designed to admit a tractable non-linear PInv. The proposed non-linear PInv and its implementation in SPNN satisfy fundamental geometric properties. One such property is n...
176 CT-Merging: Consensus Directions and Task-Specific Scaling for LoRA Adapter Merging
2607.20561
cs.CVcs.LG
Keumseo Ryum, Joonhyuk Kang
LoRA merging methods increasingly operate on the low-rank structure of task updates, yet how the common subspace is estimated and how coefficients are assigned after recomposition are rarely compared directly. We propose CT-Merging, which estimates common dire...
LoRA merging methods increasingly operate on the low-rank structure of task updates, yet how the common subspace is estimated and how coefficients are assigned after recomposition are rarely compared directly. We propose CT-Merging, which estimates common directions from averaged task subspace projectors and assigns a separate residual scale to each task. Projector averaging selects directions supported across task subspaces without weighting them by singular magnitude, while task-specific scali...
177 Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
2609.04250
cs.CVcs.SDeess.AS
Chengqian Ma, Wei Tao, Haoyu Zhang, Yiwen Guo
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio han...
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. W...
178 Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics
2609.26567
cs.CV
Eshika Pathak, Leela Krishna
Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rul...
Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rule can score well by repeatedly using one broadly transferable experience, or poorly because its preferred experience is weak. Since robots increasingly adapt by reuse rather than retraining, a score that describes the library rather than th...
179 EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies
2609.29310
cs.CV
Hanbit Oh, Yukiyasu Domae, Takuma Yagi
Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, b...
Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce...
180 FMCW-LIO: A Doppler LiDAR-Inertial Odometry
2609.29374
cs.CV
Mingle Zhao, Jiahao Wang, Tianxiao Gao, Chengzhong Xu, Hui Kong
Conventional LiDAR-inertial odometry (LIO) or simultaneous localization and mapping (SLAM) methods heavily rely on geometric features of environments, as LiDARs primarily provide range measurements instead of motion measurements. From now on, however, the situ...
Conventional LiDAR-inertial odometry (LIO) or simultaneous localization and mapping (SLAM) methods heavily rely on geometric features of environments, as LiDARs primarily provide range measurements instead of motion measurements. From now on, however, the situation changes thanks to the novel Frequency Modulated Continuous Wave (FMCW) Doppler LiDARs. FMCW Doppler LiDARs not only offer the point range with high resolution but also capture the instant point Doppler velocity through the Doppler eff...
181 Free-Init: Scan-Free, Motion-Free, and Correspondence-Free Initialization for Doppler LiDAR-Inertial Systems
2609.29375
cs.CV
Mingle Zhao, Jiahao Wang, Tianxiao Gao, Chengzhong Xu, Hui Kong
Robust initialization is crucial for online systems. In the letter, a high-frequency and resilient initialization framework is designed for LiDAR-inertial systems, leveraging both inertial sensors and Doppler LiDAR. The innovative FMCW Doppler LiDAR opens up a...
Robust initialization is crucial for online systems. In the letter, a high-frequency and resilient initialization framework is designed for LiDAR-inertial systems, leveraging both inertial sensors and Doppler LiDAR. The innovative FMCW Doppler LiDAR opens up a novel avenue for robotic sensing by capturing not only point range but also Doppler velocity via the intrinsic Doppler effect. By fusing point-wise Doppler velocity with inertial measurements under non-inertial kinematics, the proposed fra...
182 ToCo-Mesh: Topology-Consistent Dynamic Mesh Reconstruction via Adaptive Tessellation and Surface-Aligned 2DGS
2609.29529
cs.CV
Chuanjin Fan, Wenjie Chang, Aibing Li, Bingzhou Wang, Wenfei Yang
Reconstructing dynamic meshes with consistent topology from multi-view temporal images remains a challenge. Existing approaches typically face a dilemma between fine-scale shape recovery and topological stability. Frame-by-frame extraction methods capture fine...
Reconstructing dynamic meshes with consistent topology from multi-view temporal images remains a challenge. Existing approaches typically face a dilemma between fine-scale shape recovery and topological stability. Frame-by-frame extraction methods capture fine details but break vertex correspondence, leading to flickering meshes. Conversely, template-based deformation ensures consistency but struggles to adapt its surface resolution during optimization, missing local surface details. To address ...
cs.LG 287 papers
310 HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference
2609.30270
cs.LG
Simran Koul
On-device inference with small language models keeps user data local, works offline, and incurs no per-query cost, so the on-device tier is preferred when it is adequate. It is thermally constrained, however, and I find the constraint is sharper than a slowdow...
On-device inference with small language models keeps user data local, works offline, and incurs no per-query cost, so the on-device tier is preferred when it is adequate. It is thermally constrained, however, and I find the constraint is sharper than a slowdown: on a flagship Snapdragon device, sustained on-device generation destabilizes the GPU inference runtime, which crashes or silently wedges after a few consecutive queries. The failure lies in the current toolchain (OpenCL kernel compilatio...
311 When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
2609.30271
cs.LG
Gongyue Zhang, Honghai Liu
Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and ...
Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and the global learning rate is less understood. We perform a controlled cross-environment study using a paired four-environment classification problem with stable sparse features, environment-dependent spurious sparse features, dense features,...
312 ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers
2609.30272
cs.LGcs.AI
Mohd Moin Khan, Naman Srivastava, Pandarasamy Arjunan
We present \textbf{ENAS}, a hardware-aware Neural Architecture Search (NAS) framework that combines a static feasibility check, a cell-based search space supporting standard, depthwise-separable, and bottleneck blocks with optional skip connections, and a thre...
We present \textbf{ENAS}, a hardware-aware Neural Architecture Search (NAS) framework that combines a static feasibility check, a cell-based search space supporting standard, depthwise-separable, and bottleneck blocks with optional skip connections, and a three-stage hybrid search strategy (random $\rightarrow$ top-$K$ $\rightarrow$ mutation) with persistent cross-run caching. Unlike many existing NAS frameworks that rely on GPU acceleration, ENAS is designed to operate efficiently without requi...
313 Offline Policy Evaluation as a decision support tool for designing Adaptive Experiments
2609.30273
cs.LG
Jo\~ao Victor Ferreira Alves, Eduardo Rocha Laurentino, Gustavo de Oliveira Kanno, Thiago Costa Rizuti da Rocha
We investigate how historical data from fixed randomized experiments (A/B tests) can be used to inform the deployment of adaptive experiments based on contextual bandits. Given data collected under a static allocation, our goal is to assess which adaptive poli...
We investigate how historical data from fixed randomized experiments (A/B tests) can be used to inform the deployment of adaptive experiments based on contextual bandits. Given data collected under a static allocation, our goal is to assess which adaptive policies, if any, would have outperformed the original design and under what conditions. To this end, we combine off-policy evaluation (OPE) with a controlled warm-start simulation. From logged A/B test data exhibiting heterogeneous treatment e...
314 Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization
2609.30275
cs.LG
Pranav Varshney
A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as sur...
A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as surviving the change by scale-invariant statistics (cosine similarity, correlation, AUROC) that are reported without their noise floor. For the difference-in-means direction estimator, the split-half floor is governed by one dimensionless numb...
315 Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness
2609.30276
cs.LG
Alokendu Mazumder, Ayaan Mohd, Harshit Rawat, Arnab Roy, Mayank Baranwal
We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to hav...
We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moments. We show that unclipped AdaGrad can become \emph{anisotropically miscalibrated}: under heavy-tailed noise, the adaptive denominator can learn the geometry of rare noise shocks rather than the local ...
316 Fixed Points Without Fixed Diffusion: Implicit Neural Sheaves for Convergent Test-Time Computation
2609.30277
cs.LG
R\'emi Bourgerie, \v{S}ar\={u}nas Girdzijauskas, Viktoria Fodor
Implicit Graph Neural Networks (IGNNs) define node representations as fixed points of message-passing operators, enabling effectively infinite-depth propagation, iteration-independent parameterization, and flexible test-time computation. Yet these benefits dep...
Implicit Graph Neural Networks (IGNNs) define node representations as fixed points of message-passing operators, enabling effectively infinite-depth propagation, iteration-independent parameterization, and flexible test-time computation. Yet these benefits depend on the equilibrium being unique and attainable by fixed-point iteration. Existing constructions often impose constraints on recurrent updates to obtain these guarantees, limiting the transformations available at equilibrium. This raises...
317 Neural Ideals and Neural Codes: An Algebraic Framework for Neural Network Classification and Feature Interpretation
2609.30279
cs.LG
Venkata Subbaiah Yerrapati, Rahul Dixit, Ajay Kumar Shukla
Understanding the features captured by the hidden layers of neural networks is a fundamental challenge in machine learning, despite their widespread success across various classification problems. In this work, we propose an algebraic framework for examining n...
Understanding the features captured by the hidden layers of neural networks is a fundamental challenge in machine learning, despite their widespread success across various classification problems. In this work, we propose an algebraic framework for examining neural networks that model classification problems. Certain results, such as the correspondence between the neural network and neural ideals, algorithms for computing the neural ideals, and a stabilization theorem that enables approximation ...
318 Seasonal and Quantum-inspired Models for Neutron Monitor Time Series Forecasting
2609.30281
cs.LG
Krishna Bhatia, Shalini Devendrababu, Srinjoy Ganguly
We present a focused and reproducible study of multi-horizon forecasting on the Lomnicky Stit neutron monitor (LMKS) time series. Our evaluation suite covers simple seasonal baselines, modern deep sequence models, and functional and quantum-inspired architectu...
We present a focused and reproducible study of multi-horizon forecasting on the Lomnicky Stit neutron monitor (LMKS) time series. Our evaluation suite covers simple seasonal baselines, modern deep sequence models, and functional and quantum-inspired architectures, including Seasonal Naive, Long Short-Term Memory (LSTM), Temporal Convolutional Network (TCN), N-BEATS, Kolmogorov-Arnold Networks (KAN), and two quantum-inspired variants, QiLSTM and QiKAN. We describe the dataset characteristics, dia...
319 When Does Advection-Aware Graph Nowcasting Help? A Controlled Study of Distributed Solar Ramp Forecasting with a Self-Supervised Cloud-Motion Estimator
2609.30286
cs.LGcs.AI
Phillip Jiang
Short-term forecasting of cloud-induced power ramps across a network of distributed photovoltaic (PV) or irradiance sensors is a recognised pain point for grid operators. A natural idea is to make the graph neural network (GNN) advection-aware: connect each si...
Short-term forecasting of cloud-induced power ramps across a network of distributed photovoltaic (PV) or irradiance sensors is a recognised pain point for grid operators. A natural idea is to make the graph neural network (GNN) advection-aware: connect each site to the sites upwind of it, with edge time-lags set by the cloud-motion vector (CMV), so that a ramp is propagated forward before it physically arrives. Using a controlled synthetic testbed with a known wind field, we show that (i) with a...
320 NeuralCert: certified computational discovery of extremal mathematical constructions
2609.30296
cs.LG
Mark Patrick Roeling
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions...
Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable representation, spectrally diagnosed and pruned, and then certified exactly through multimodular evaluation. Exact certification makes the numerical proofs fully explicit and independently verifiable. Thi...
321 Staged Depth Training: A Representation Curriculum for PINNs
2609.30299
cs.LG
Kejia Zhang, Youran Sun, Haizhao Yang
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representati...
Representation quality is a central determinant of PINNs' performance, yet standard training leaves representations to emerge implicitly while fitting the final solution. We introduce \textbf{representation curriculum}, an ordered process in which representations are explicitly learned, transferred independently of their predictors, and progressively refined. We realize it with Staged Depth Training (SDT), which trains a shallow prefix under a temporary physics-informed head, discards the head, ...
322 Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy
2609.30326
cs.LG
Farhad Davaripour
Activation steering changes a model's internal activations during inference without updating its weights, but a useful intervention must determine both how and when to steer. Motivated by the AI-safety concern that a model expected to accept shutdown may inste...
Activation steering changes a model's internal activations during inference without updating its weights, but a useful intervention must determine both how and when to steer. Motivated by the AI-safety concern that a model expected to accept shutdown may instead produce a shutdown-avoidance response, this study examines a guarded probe-and-select procedure for simulated shutdown scenarios in Qwen3.5-0.8B. KEEP leaves the process running and represents shutdown avoidance, whereas STOP accepts shu...
323 Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation
2609.30337
cs.LG
Zhengchen Huang, Yundong Sun, Minrui Song, Shuanglong Yao, Ye Liu
Retrieval-Augmented Generation (RAG) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context. However, when retrieved knowledge conflicts with the model's internal parametric knowledge, the model may ...
Retrieval-Augmented Generation (RAG) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context. However, when retrieved knowledge conflicts with the model's internal parametric knowledge, the model may either blindly follow misleading context or incorrectly rely on parametric knowledge, leading to unreliable responses. To address this issue, this paper proposes TRACE (Debate-TRace and Answer-Completeness rEgularized fine-tuning), a robust...
324 GAUDI: Geometry-Aware Diffusion for Calibrated Air-Quality Time-Series Imputation
2609.30340
cs.LG
Xinjin Li, Yudi Xia, Calvin Chang Liu, Weiru Lin, Bojun Li
Air-quality sensor outages often create contiguous missing blocks, where side information useful for isolated missingness may be less reliable. We study a block-specific, GAUDI-aligned conditional diffusion imputer that retains temporal and feature processing,...
Air-quality sensor outages often create contiguous missing blocks, where side information useful for isolated missingness may be less reliable. We study a block-specific, GAUDI-aligned conditional diffusion imputer that retains temporal and feature processing, visible-value and mask conditioning, variable identity, and diffusion-step information, while suppressing absolute time-position side embeddings. On ItalyAir (13 variables, length-32 windows, nominal 50% block missingness; three archived s...
325 Learning coarse-step dynamics and internal mechanical response with graph networks
2609.30344
cs.LG
Vinay Sharma, Olga Fink
Modern sensing records the motion of physical systems, but often leaves the forces and mechanical response governing that motion unobserved. Inferring these quantities from discretely sampled trajectories is especially difficult at coarse time scales, when mec...
Modern sensing records the motion of physical systems, but often leaves the forces and mechanical response governing that motion unobserved. Inferring these quantities from discretely sampled trajectories is especially difficult at coarse time scales, when mechanical response evolves between observations and interactions propagate across the system. Here we introduce Newmark-\b{eta}-DGN, a graph neural network-based framework that combines two structures inspired by computational mechanics. Firs...
326 Strategic Self-Consistency
2609.30352
cs.LGcs.AI
Tori Qiu, Ander Artola Velasco, Manuel Gomez-Rodriguez
Self-consistency has become a popular technique for enhancing the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer through a majority vote. However, because model providers typically charge user...
Self-consistency has become a popular technique for enhancing the reasoning abilities of large language models by generating multiple reasoning paths and selecting the final answer through a majority vote. However, because model providers typically charge users in proportion to the number of reasoning paths generated, they have a financial incentive to artificially increase the path count. In this work, we show that an unfaithful provider can exploit this incentive using a simple, efficient algo...
327 Cost-Aware Best-LLM Identification using Dueling Feedback
2609.30360
cs.LGcs.AI
Sarvesh Gharat, Nikhil Karamchandani, Jayakrishnan Nair
Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons ...
Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate a...
328 DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning
2609.30379
cs.LGcs.AI
Zhiyuan Chen
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Pac...
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure ...
329 From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning
2609.30391
cs.LG
Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when traje...
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman...
330 Adaptive Multi-Value Control in LLMs via Causal Activation Steering
2609.30405
cs.LG
Payel Bhattacharjee, Ravi Tandon
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying inter...
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the mo...
331 Electric Vehicle Charging Station Location Selection using Geospatial Artificial Intelligence (GeoAI)
2609.30417
cs.LG
Eun Hak Lee, Euntak Lee
As electric vehicle (EV) adoption increases, ensuring efficient and well-distributed charging infrastructure has become a critical challenge. While many EV charging station location problem (CSLP) studies focus on minimizing costs or travel distance, it is cru...
As electric vehicle (EV) adoption increases, ensuring efficient and well-distributed charging infrastructure has become a critical challenge. While many EV charging station location problem (CSLP) studies focus on minimizing costs or travel distance, it is crucial to consider the surrounding geospatial characteristics of existing stations that influence operational performance. This study proposes a geospatial artificial intelligence (GeoAI)-based framework that integrates high-dimensional EV-re...
332 Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation
2609.30427
cs.LG
Zhaoyang Cao, Miriam Metzger, Reza Zafarani
Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop ...
Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop a theory-informed computational framework that translates cross-disciplinary theories of fake news into measurable features for automated detection and explanation through statistical techniques and large language models. To that end, we co...
333 Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles
2609.30433
cs.LG
Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi, Zhizhuo Zhang
Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we dev...
Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we developed Molecule-Morphology Contrastive Pretraining (MoCoP), a strategy for aligning small molecule embeddings to morphology fingerprints extracted through CellProfiler. The resulting molecular representation showed transferable performance ...
334 Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
2609.30454
cs.LG
Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larg...
Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphr...
335 Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis
2609.30470
cs.LG
Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address t...
Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variationa...
336 Moment-guided edge sampling
2609.30472
cs.LG
Weibin Cai, Reza Zafarani
Edge sampling makes local decisions to achieve graph-level objectives, such as preserving structural properties. This creates a fundamental challenge: \textit{how can the effect of a local edge edit (i.e., edge addition or removal) on global graph structure be...
Edge sampling makes local decisions to achieve graph-level objectives, such as preserving structural properties. This creates a fundamental challenge: \textit{how can the effect of a local edge edit (i.e., edge addition or removal) on global graph structure be quantified and controlled?} We address this challenge with a \textit{moment-guided edge sampling framework} based on spectral moments of the random-walk transition matrix. We compute exact moment changes through two complementary methods: ...
337 Mentored Decoding: Faster Inference meets Boosting
2609.30474
cs.LG
Vivien Tran-Thien, Richard Nock
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been o...
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored...
338 Geometric Feature Learning for Functional Data Valued on the Symmetric Positive Definite Manifold
2609.30487
cs.LG
Samuel V. Singh, Mimi Zhang
We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued funct...
We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued functions, followed by a functional layer that projects them into a finite-dimensional Euclidean space. Unlike most neural networks for discrete-time sequences, MatFAE treats each sequence as a continuous function and can therefore encode trajec...
339 Learning to Bias: Machine Learning-Enhanced Particle Filters
2609.30498
cs.LG
Apoorv Srivastava, Eric Darve
Sequential inference estimates latent states from noisy and incomplete observations. Particle Filters (PFs), a class of Monte Carlo methods based on importance sampling, provide a flexible framework for this task, but often suffer from poor sample efficiency a...
Sequential inference estimates latent states from noisy and incomplete observations. Particle Filters (PFs), a class of Monte Carlo methods based on importance sampling, provide a flexible framework for this task, but often suffer from poor sample efficiency and unfavorable scaling with dimension, partly due to suboptimal proposal distributions. We address these challenges by integrating learned proposals into the PF framework. We introduce Neural Optimal Particle Filters (NOPFs), which learn an...
340 PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
2609.30500
cs.LGcs.AI
Yuhe Sui, Yingzhi Tang, Shufang Chen
Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\p...
Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_\eta(\pi,Q)=\operatorname{softmax}(\log\pi+\eta Q)$. Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy a...
341 To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity
2609.30501
cs.LG
Zhiyao Zhang, Menglu Yu, Alvaro Velasquez, Nathaniel D. Bastian, Jia Liu
Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally conv...
Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally convex (LLGC) settings (i.e., the lower-level objective function is assumed to be, at least, convex). While the LLSC/LLGC assumptions render more tractable algorithmic design and theoretical analysis, they are too rigid to encompass many machin...
342 Federated Targeted Maximum Likelihood Estimation
2609.30503
cs.LG
Diyang Li, Fei Wang, Kyra Gan
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum lik...
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum likelihood estimation (TMLE) refines a flexible initial fit, yielding plug-in estimators that respect the model and support efficient inference. TMLE itself, however, has remained a fully centralized procedure. To fill this gap, our paper intr...
343 Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework
2609.30508
cs.LG
Felix S. Reimers, Ola Huse Ramstad, Aliaksandr Hubin, Stefano Nichele
The aim of this work is to examine the connectomes of Caenorhabditis elegans through a computational lens using the reservoir computing framework. Connectomes are mappings of biological neural networks; C. elegans is the first organism for which physical conne...
The aim of this work is to examine the connectomes of Caenorhabditis elegans through a computational lens using the reservoir computing framework. Connectomes are mappings of biological neural networks; C. elegans is the first organism for which physical connectomes covering the whole nervous system have been published. The connectomes of C. elegans used in this paper have been derived at different ages of the organism and are based on three different ways of measuring inter-cellular connections...
344 AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
2609.30541
cs.LG
Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier
Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a ...
Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, ev...
345 GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing
2609.30542
cs.LG
Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed
De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, a...
De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed ...
346 Dynamic Regret in Online Convex Optimization with Indicator Switching Costs
2609.30556
cs.LG
Naram Mhaisen, George Iosifidis
We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, a...
We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, and on a bounded domain it recovers norm-based movement costs as a special case. Existing guarantees for indicator costs handle only static comparators. We show that a direct extension of these techniques to dynamic regret provably fails, mo...
347 Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations
2609.30572
cs.LG
Mihir Dhanakshirur, Adam Ousherovitch, Ambuj Tewari
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable...
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing veri...
348 Reinforcement Learning of Communication in a Mesh of Small Language Models
2609.30578
cs.LG
Mehmet Kerem Turkcan
Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a pr...
Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The mo...
349 Energy-efficient operation of neural operators for virtual sensing
2609.30580
cs.LG
Jason Yoo, Samrendra Roy, Souvik Chakraborty, Syed Bahauddin Alam
Virtual sensing repeatedly reconstructs physical fields from changing observations, often on a fixed geometry. We investigate how shared spatial computation reduces the energy of these updates while retaining the selected checkpoint and its evaluated predictio...
Virtual sensing repeatedly reconstructs physical fields from changing observations, often on a fixed geometry. We investigate how shared spatial computation reduces the energy of these updates while retaining the selected checkpoint and its evaluated predictions. In a heat-exchanger service, standard compiler freezing and explicit trunk reuse give similar operating energy reductions relative to graph replay: approximately 1% at one request per second and 20% at forty requests per second. In 15 W...
350 Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
2609.30605
cs.LG
Hongsen Zhang, Lu Zhang, Mingjing Xu, Yi Zhang, Gregory Epiphaniou
Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agno...
Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the unive...
351 OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control
2609.30628
cs.LG
Tommaso Schettini, Nicholas D. Kullman, Jorge E. Mendoza
Machine-learning policies have attracted increasing interest for ride-hailing fleet control in recent years. Reinforcement learning, in particular, requires a structured simulation environment that specifies observations, actions, rewards, and decision epochs ...
Machine-learning policies have attracted increasing interest for ride-hailing fleet control in recent years. Reinforcement learning, in particular, requires a structured simulation environment that specifies observations, actions, rewards, and decision epochs for training and evaluation. For electric fleets, this environment must also capture the interaction among stochastic demand, vehicle operations, and capacitated charging infrastructure. We present OpenHail, an open-source Gymnasium environ...
352 Stable initialization without the CLT
2609.30633
cs.LG
Simon Kuang, Kyle Chickering, Xinfan Lin
Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal ...
Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal moderation between these two extremes. The conventional theory of random networks uses the Central Limit Theorem to control inter-neuron dependencies, which introduces distributional approximation error and coupling between layers. For netw...
353 In-Context Binding Capacity in Language Models
2609.30634
cs.LG
Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
How many assignments can a language model recall before it loses track of which value belongs to which entity? We measure this limit using continuous recall curves for 12 models at or below 3B parameters and a threshold sweep over 30 open models up to 12B. On ...
How many assignments can a language model recall before it loses track of which value belongs to which entity? We measure this limit using continuous recall curves for 12 models at or below 3B parameters and a threshold sweep over 30 open models up to 12B. On the continuous curves, the load at which recall falls halfway to chance follows $K_{50}=cN^{\alpha}$, with $\alpha=0.820$ and $R^2=0.73$. The broader sweep shows an eightfold range associated with pretraining recipe, although the continuous...
354 Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation
2609.30650
cs.LGcs.AI
Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and del...
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interfa...
355 Population loss in shallow ReLU networks: Bias & families of critical points
2609.30661
cs.LG
Michael Field
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes es...
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families ...
356 PixSim: a calibrated open-source simulator of instant-payment fraud, recovery and interdiction under analyst capacity constraints
2609.30684
cs.LG
Bashir Zeimarani, Alireza Khatib, Somayeh Mousavinasr, Carlos Maur\'icio Serodio Figueiredo
Brazil's Pix settles about 5.9 billion instant, irreversible transfers a month. A fraudulent transfer can be recovered only while the funds remain in a traceable account, and in 2025 the Central Bank's recovery mechanism (MED) returned 9% of accepted contested...
Brazil's Pix settles about 5.9 billion instant, irreversible transfers a month. A fraudulent transfer can be recovered only while the funds remain in a traceable account, and in 2025 the Central Bank's recovery mechanism (MED) returned 9% of accepted contested value. Interdiction therefore has to happen before settlement, by routing each transaction to pass, human review or block, under a finite analyst team and a regulatory hold window. To our knowledge no public simulator jointly models irreve...
357 LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant
2609.30692
cs.LG
Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali
Reliable voice interaction is essential in environments with limited internet connectivity and strong privacy. However, most existing voice assistants depend on cloud-based services, which leads to latency issues, dependency on internet access, and privacy vul...
Reliable voice interaction is essential in environments with limited internet connectivity and strong privacy. However, most existing voice assistants depend on cloud-based services, which leads to latency issues, dependency on internet access, and privacy vulnerabilities. This research presents LUMO (Lightweight Unified Multilingual Orchestrator), a privacy preserving offline voice assistant designed for edge computing environments. This system integrates local Automatic Speech Recognition (ASR...
358 NEMSim: Learning Control-Conditioned Multi-Event Physical Dynamics via Executable Event-Mechanism Priors
2609.30718
cs.LG
Junsong Yu, Junjie Xie, Pengwei Liu, Dong Ni
High-fidelity simulation of control-conditioned multi-event physical systems is computationally expensive, especially across broad control spaces and long trajectories. In these systems, macroscopic evolution emerges from localized discrete events whose intens...
High-fidelity simulation of control-conditioned multi-event physical systems is computationally expensive, especially across broad control spaces and long trajectories. In these systems, macroscopic evolution emerges from localized discrete events whose intensities and effects depend on process controls and evolving local states, while the available system knowledge is typically expressed as event-attribute descriptions. Purely data-driven surrogates must infer these event effects from limited t...
359 When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification
2609.30721
cs.LG
Xinze Shi, Litian Zhang, Binrui Shi
Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does ...
Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established...
360 Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events
2609.30746
cs.LG
Isabella S. Thiel, Juan Bello-Rivas, Yannis G. Kevrekidis, Themistoklis P. Sapsis
Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that...
Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jac...
361 Differentiable RNA Secondary Structure Extraction for Deep Learning
2609.30752
cs.LG
Tyler Illman, Max Ward, Marcell Szikszai, Ryan K. Krueger
Many deep learning approaches to RNA secondary structure prediction have recently been proposed. They typically output a weight matrix $W$ where $W_{ij}$ is an arbitrary weight for base $i$ pairing with base $j$. Converting this matrix to a predicted secondary...
Many deep learning approaches to RNA secondary structure prediction have recently been proposed. They typically output a weight matrix $W$ where $W_{ij}$ is an arbitrary weight for base $i$ pairing with base $j$. Converting this matrix to a predicted secondary structure or base-pairing probability matrix typically involves ad hoc and problematic downstream algorithms. Despite the importance of this conversion step, which we refer to as structure extraction, it has received relatively little atte...
362 Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift
2609.30781
cs.LG
Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai
Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration proc...
Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration procedure for mortality prediction under cross-hospital distribution shift. It selects a measurement on an independent sample, groups patients by whether that measurement is recorded, and applies Mondrian calibration within each group, so no ca...
363 Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays
2609.30789
cs.LG
Yiqi Yao, Miquel Duran-Frigola
In low-data structure-activity prediction, the choice of molecular representation can matter more than the choice of predictor, and tabular foundation models sharpen that effect. We ask whether a portfolio of compact, semantically named descriptor blocks can r...
In low-data structure-activity prediction, the choice of molecular representation can matter more than the choice of predictor, and tabular foundation models sharpen that effect. We ask whether a portfolio of compact, semantically named descriptor blocks can reach the accuracy of a 2048-dimensional CheMeleon embedding while staying auditable at the feature level, meaning that every input dimension carries a model name and a recorded training provenance. Starting from a fixed 11-dimensional physi...
364 Towards Universal Representation-Based Process Control
2609.30790
cs.LG
Jinmyeong Choi, Taesup Kim, Artur Dubrawski
Many temporal process learning and monitoring pipelines operate in local windows, making window-level decisions unavoidable in practice. In such settings, classical statistical tests can be applied to individual windows, but they typically evaluate predefined ...
Many temporal process learning and monitoring pipelines operate in local windows, making window-level decisions unavoidable in practice. In such settings, classical statistical tests can be applied to individual windows, but they typically evaluate predefined parametric hypotheses-such as unit-root or moment-based conditions-thereby limiting flexibility when reference behavior is defined empirically from task- or domain-specific data. In this work, we view window-level monitoring as a process co...
365 Counterfactual Online Conformal Prediction Under Adaptive Logging
2609.30811
cs.LG
Xinyu Qiao, Yichen Lin, Kaihong Ji, Xue Wang, Tao Yao
Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected a...
Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected actions. This paper formalizes the failure through counterfactual coverage and introduces Propensity-Weighted Online Conformal Prediction, an inverse-propensity-weighted recursion that debiases calibration. A doubly robust variant further re...
366 Learning Provable Neural Network Observer for Uncertain Dynamical Systems
2609.30819
cs.LG
Zhangyi Wang, Jiaxu Liu, Chen Song, Chao Xu, Shengze Cai
In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix...
In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix Inequality (LMI) constraints leads to large-scale semidefinite programs (SDPs) that are difficult to solve for large networks. To overcome this scalability bottleneck, we propose a novel two-stage training framework for provably stable neu...
367 Adaptive Interaction Graphs for Particle Simulation
2609.30822
cs.LG
Aiden Zhou
Learned particle simulators based on graph neural networks achieve strong one-step accuracy, but errors compound over long horizons. An underexplored variable is the interaction graph: existing methods fix its topology via k-nearest neighbors or a static radiu...
Learned particle simulators based on graph neural networks achieve strong one-step accuracy, but errors compound over long horizons. An underexplored variable is the interaction graph: existing methods fix its topology via k-nearest neighbors or a static radius rule, regardless of local model confidence. We propose making this graph adaptive: a per-particle variance head, trained jointly with the acceleration head under a heteroscedastic Gaussian NLL loss, drives a trajectory in which high-uncer...
368 MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
2609.30837
cs.LGcs.AI
Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels res...
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training ...
369 Peer-Grounded Counterfactual Path Planning for Chronic Health Management
2609.30838
cs.LG
Saman Khamesian, Hassan Ghasemzadeh
Effective behavioral intervention in chronic disease management requires not a single prescription but a sequence of incremental steps, each grounded in what real, similar individuals have demonstrably achieved. Counterfactual explanation offers a natural comp...
Effective behavioral intervention in chronic disease management requires not a single prescription but a sequence of incremental steps, each grounded in what real, similar individuals have demonstrably achieved. Counterfactual explanation offers a natural computational route to such guidance, answering what change in behavior would have produced a better outcome. But existing methods return a target state without a route to it, guarantee no monotone health improvement along the way, and draw no ...
370 Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation
2609.30839
cs.LG
Filip T\u{a}\c{s}\u{a}dan, Ema Tomanov\'a, Ondrej Lopuch, Pawe{\l} Bilko, Anders S{\o}gaard
Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perfo...
Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text t...
371 Learning Chance-Constrained MDPs with Bellman Distributional Certificates
2609.30856
cs.LG
Chenbei Lu, Hongyu Yi
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but ...
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a hig...
372 CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters
2609.30884
cs.LG
Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen
Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete ...
Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or stable adapter identity, but neither represent caches from ...
373 Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
2609.30918
cs.LGcs.AI
Marcin Kostrzewa, Maciej Zi\k{e}ba
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are dif...
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that hold...
374 EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
2609.30929
cs.LG
Takumi Fujimoto, Hiroaki Nishi
Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order disc...
Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Within each channel, the endpoint is shared across component-wise online ridge regressions that also use current-forecast coefficients. The fitted...
375 PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
2609.30948
cs.LGcs.AI
Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining wi...
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-...
376 Low-Bit Recurrent States in Hybrid Language Models
2609.30950
cs.LG
Hongren Chen, Jiayang He
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them wi...
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative...
377 Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
2609.30957
cs.LG
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "...
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is determined primarily by the progress along the river. Wi...
378 Gradient Surgery for Physics-Informed Neural Networks
2609.30966
cs.LG
Thomas Borsani, Giuseppe Di Fatta
Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing opt...
Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing optimisation strategies are affected by conflicting task gradients, leading to slow convergence and unstable training, particularly for stiff and high-frequency partial differential equations. We analyse gradient conflicts throughout training ...
379 LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
2609.30973
cs.LG
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforci...
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each layer of a DNN to be Lipschitz continuous, thereby guaranteeing overall Lipschitz continuity. However, this layer-wise approach typically imposes overly conservative restrictions by producing a loose es...
380 Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models
2609.30995
cs.LG
Shan Zhao, Ilija Trajkovic, Julia Kaltenborn, Yaniv Gurwicz, Peer Nowack
Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustwor...
Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface temperature fields from a state-of-the-art global climate mode...
381 The Linear Representation Hypothesis for Vision-Language-Action Models
2609.30996
cs.LGcs.AI
Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vis...
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in...
382 Robust Successor Features
2609.31016
cs.LG
Erik Nikulski, Yamen Habib, Vicen\c{c} Gomez, Anders Jonsson, Rub\'en Moreno-Bote
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adap...
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the successor representation, Robust RL emerged simultaneousl...
383 Metacognitive Selective Ensemble for Mobile Systems
2609.31031
cs.LG
Sungmin Lee, Kichang Lee, Joonhee Lee, JaeYeon Park, Songkuk Kim
Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain relia...
Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain reliable evidence about inactive candidates. We present MetaSE, an active ensemble framework that exploits short-term persistence in per-model reliability. MetaSE maintains a small active set across windows, uses post-execution evidence to rejec...
384 Robust Graph Clustering Network for Multiple Missing Data
2609.31033
cs.LG
Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang
Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-vie...
Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-view error propagation and cluster-boundary blurring under simultaneous attribute and structure missingness. To address these limitations, we propose a Robust Graph Clustering Network for Multiple Missing Data (RGCN), which is designed to hand...
385 Aurora-X: Built for Extreme Time Series Forecasting
2609.31038
cs.LG
Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce...
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce cross-variable dependencies, varied context and horizon length...
386 Distributed Learning as a Service: The Developer's Perspective
2609.31061
cs.LG
Tianyue Chu, Filippo Vannella, Dimitra Tsigkari, Paula Delgado-Santos, Fernando L\'opez
Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limite...
Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limited resources, a single aggregator might not be able to scale, and the transmissions of model weights induce a considerable bandwidth cost. This paper demonstrates DLaaS (Distributed Learning as a Service) from the developer's vantage point. ...
387 SAGE: A sampling-aware global evaluation benchmark for species distribution modeling
2609.31082
cs.LG
Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying...
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds of millions of community-science records. At this scale...
388 Block Sparse Attention with Log-Linear Complexity
2609.31093
cs.LG
Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query...
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-$K$ selection strategy. The main idea is to gradually narrow down the can...
389 The Residual Stream's Effective Depth
2609.31098
cs.LG
Barak Gahtan, Ido Galil, Alex M. Bronstein
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one n...
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference $F_L = 2L/(L+1)<2$, yet ...
390 Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
2609.31107
cs.LGcs.AI
Saksham Kiroriwal, Julius Pfrommer, J\"urgen Beyerer
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of ...
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we...
391 From Shortcut Learning to Discrete Neural Insertion Sort
2609.31114
cs.LGcs.AI
Konstantinos Mylonas, Thrasyvoulos Spyropoulos
Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the inten...
Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the intended execution. We study this problem using insertion sort. Our analysis of the CLRS30 baseline NAR shows that the hint objective is weakly optimized and that hint accuracy remains low. Moreover, many intermediate representations can already...
392 Frame the adversary: a structure-aware attack methodology
2609.31128
cs.LG
Vicky Kouni, Stelios Perrakis, Francis Bach, Pascal Frossard, Yann Chevaleyre
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for ro...
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propos...
393 CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks
2609.31149
cs.LG
Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi
Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical re...
Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical reaction networks (CRNs), a class of count-native Markov jump processes, we introduce CRNDiff, a structured framework that combines count-space diffusion with inference-time conditioning on rare subpopulations. An independent birth--death ins...
394 Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift
2609.31155
cs.LGcs.AI
Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two...
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run, which does not minimize empirical target cross-entropy. E...
395 Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
2609.31157
cs.LG
Jianan Liu, Chunguang Li
Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these ...
Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these methods, reshaping operations are required, which inevitably break the intrinsic correlations and thus lead to performance degradation. In uni-/multi-variate time series anomaly detection, AutoEncoders (AEs) are widely adopted and generally...
396 I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?
2609.31161
cs.LG
Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models...
Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models. However, accurate prediction does not, in general, necessarily imply recovery of underlying causal states that give rise to the observed dynamics. This work investigates when and how JEPAs can recover the underlying causal states from obs...
397 WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting
2609.31162
cs.LG
Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo
Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with fu...
Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with future observations being shaped by latent dynamics. Recent latent-space forecasting methods thus achieve improved performance by predicting future observations from latent-space representations of historical observations rather than directly...
398 Audio emotion recognition for atypical hearing
2609.31168
cs.LG
Ulysse Roussel (STMS)
My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our...
My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our core idea is to leverage our understanding of affect from acoustic traits, relying on the possibility of generalizing affective responses from a small amount of annotated data. As a first step, we fine-tune a large foundation model, Contra...
399 ALF: An Active Learning Framework for Scientific Discovery
2609.31197
cs.LG
Shikha Surana, Alex Hawkins-Hooker, Olivia Gallup, Christoph Brunken, Jules Tilly
Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requi...
Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requires expensive experiment, measurement, or simulation. Most existing tools cover only part of the data acquisition loop, and typically focus on either offline benchmarking or online deployment, but not both. We present ALF, a modular AL Fram...
400 Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
2609.31206
cs.LG
Matthieu Le Lain, Ga\"el Cessateur, S\'ebastien Lef\`evre
Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder o...
Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder on 223,000 unlabelled spectra. Without labels, its representation recovers the emission-line intensity ratios that physicists use to diagnose the precipitating particles (R^2 0.91 vs. 0.77 for an untrained control) and, under one linear prob...
401 Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion
2609.31222
cs.LG
Xinyu Wang, Jinbo Bi, Minghu Song
Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), a...
Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions and delivers them through a trust budget set by the fro...
402 Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
2609.31250
cs.LG
Mahesh Reddy Pagadala
We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory ...
We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated re-eviction of the same storages (evictions per storage ...
403 Softmax Reparameterization for Output-Head Quantization
2609.31291
cs.LGcs.AI
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar mult...
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-cen...
404 Benchmarking Attention for Tabular Foundation Models
2609.31306
cs.LG
Maximilian Schambach, Clemens Biehl, Sam Thelin
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention invol...
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to...
405 LUCID: Learning Under Confounding for Inference and Discovery in Time Series
2609.31315
cs.LG
Mohammad Fesanghary
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfound...
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Mar\v{c}enko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum indicates pervasive factor confounding, LUCID atte...
406 More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
2609.31325
cs.LG
Lewei Xie, Haoyu Zhang, Jiajun Zhou, Yulong Chen, Guanxing Chen
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current se...
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current sensor layout, so sensor expansion can alter the representation of learned spatial relationships. Our key insight is that sensor expansion changes the evidence available about a process without necessarily changing the dynamics to be learned....
407 Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
2609.31329
cs.LG
Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng
Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shar...
Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shared task objective, limiting explicit high-level coordination. Inspired by natural genes that coordinate biological development, we introduce \textbf{Morphogene}, a compact latent blueprint that bridges an agent's body and brain. Through Ada...
408 Progressive Memory Transformer: Memory-Aware Attention for Time-Series
2609.31351
cs.LG
Tord Sture Stangeland, Andreas K\"ohler, Steffen M{\ae}land, Ad\'in Ram\'ires Rivera
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise repres...
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy ...
409 Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
2609.31363
cs.LG
Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse \c{S}en, Marco Cuturi, Daniel Kuhn
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulat...
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be reformulated as an optimization problem over transport maps t...
410 Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness
2609.31371
cs.LG
Bin Li, Dongdong Wang, Siyang Lu
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate th...
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and quantization settings. Our results reveal substantial performanc...
411 Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition
2609.31399
cs.LG
Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts ...
Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts inject noise that corrupts the shared representation. We introduce EmoSpeechBrain, a multimodal framework built on the insight that noise suppression is a precondition for effective fusion. Its EEG encoder uses differential attention, takin...
412 Decodable In-Context State and Model Output Across Training
2609.31401
cs.LG
Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. ...
Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model proba...
413 Evaluating the accuracy of KV cache reuse techniques
2609.31415
cs.LG
Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the...
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse, often artificially inflating the reported effectiveness. We also show that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques. To address these issues, ...
414 Scaffold: Support Graph Theory Based Sparsification for Graph Neural Networks
2609.31466
cs.LG
Siddhartha Shankar Das, Sai Karthik Navuluru, S M Ferdous, Ryan A. Rossi, Baris Coskunuzer
Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can dis...
Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can distort important communication structure and degrade predictive performance. We introduce Scaffold, a topology-based, unsupervised graph sparsification framework derived from support graph theory preconditioners. Scaffold explicitly controls ...
415 HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
2609.31531
cs.LG
Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics t...
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this inconsistency as structural target drift. We introduce HySTAR, a M...
416 NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures
2609.31539
cs.LG
M\'arcio Marques, Leonardo Mendon\c{c}a, Leonardo M. Moreira, Christian J\'unior de Oliveira, Vitor Balestro
Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are ...
Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are known to suffer from spectral bias and lack of causality. Neuro-Spectral Architectures (NeuSA), a recently proposed alternative to PINNs, mitigate both issues, but their numerical integration becomes unstable for stiff differential equation...
417 BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment
2609.31546
cs.LG
Mohammad Nur Hossain Khan, M. S. Krafczyk, Beverly G. Bolster, Nancy McElwain, Mark A. Hasegawa-Johnson
Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose...
Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose heart rates are higher and whose ECG differs from the adult, clinic-recorded 12-lead data these models are built on. A model for infant ECG should therefore reason about heartbeats directly rather than recover them from arbitrary patches. ...
418 Online Learning via Learned Latent Bayesian Tracking
2609.31559
cs.LG
Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially ...
Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing existing methods to rely on restrictive approximations or ma...
419 Generalization behavior of OPTQ and the role of regularization
2609.31560
cs.LG
Erin George, Rayan Saab
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantizat...
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive bounds for the expected squared error accrued by the alg...
420 Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
2609.31564
cs.LG
Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodol\`a
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one...
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size en...
421 Trust Guided Decision Transformer
2609.31586
cs.LG
Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays el...
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes ...
422 Common-Mode Collapse and Recovery in Direct Feedback Alignment
2609.31589
cs.LG
Varun Reddy, Bernardo L. Sabatini, Houman Safaai
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencie...
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading componen...
423 New LoRA Skills Should Read but Never Write
2609.31600
cs.LG
Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task dat...
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent facto...
424 A New Non-archimedean Metric on Persistent Homology
2012.02655
cs.LG
\.Ismail G\"uzel, Atabey Kaygun
In this article, we define a new non-archimedean metric structure, called cophenetic metric, on persistent homology classes of all degrees. We then show that zeroth persistent homology together with the cophenetic metric and hierarchical clustering algorithms ...
In this article, we define a new non-archimedean metric structure, called cophenetic metric, on persistent homology classes of all degrees. We then show that zeroth persistent homology together with the cophenetic metric and hierarchical clustering algorithms with a number of different metrics do deliver statistically verifiable commensurate topological information based on experimental results we obtained on different datasets. We also observe that the resulting clusters coming from cophenetic ...
425 Persistent Homology of Time Series through Complex Networks
2605.01624
cs.LG
\.Ismail G\"uzel
We present a unified pipeline for univariate time series classification via complex networks and persistent homology. A time series is mapped to a graph through one of five constructions across three families (visibility (natural and horizontal visibility grap...
We present a unified pipeline for univariate time series classification via complex networks and persistent homology. A time series is mapped to a graph through one of five constructions across three families (visibility (natural and horizontal visibility graphs), transition, and proximity) and the graph is converted to a dissimilarity matrix from which a Vietoris-Rips filtration yields persistence diagrams. These diagrams are vectorized into fixed-length features through persistence landscapes ...
426 Distribution of hitting times for dissipative random dynamical systems on $\mathbb{R}^d$, with application to stochastic gradient descent
2609.30274
cs.LG
St\'ephane Galatolo, St\'ephane Chr\'etien
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stoc...
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most rudimentary, while still outperforming more recent algorithms at a number of learning tasks. A major open question about the current methods used in deep learning is to understand their convergence prop...
427 Adaptive Random Matrices in Gaussian Bandits: Spectral Universality and Selection-Induced Outliers
2609.30321
cs.LG
Sudarshan Manikantan (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute)
Adaptive arm selection changes the distribution of the observations collected by a bandit algorithm, but it need not change their limiting empirical spectrum. We study Gaussian bandit designs in which the dimension and the number of observations grow proportio...
Adaptive arm selection changes the distribution of the observations collected by a bandit algorithm, but it need not change their limiting empirical spectrum. We study Gaussian bandit designs in which the dimension and the number of observations grow proportionally. A quantitative coupling theorem compares the design generated by any causal selection rule with an independent Gaussian design. If the logarithm of the number of available arms is sublinear in the dimension, the empirical spectral di...
428 Low-Rank Friction for Memory-Efficient Transformer Pretraining
2609.30342
cs.LG
Rajit Rajpal, Benedict Leimkuhler
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathca...
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as Adam's second-moment buffer. Here we replace iKFAD's friction tensor $\xi$ with a rank-1 outer-product factorisation built from row and column momentum statistics, resulting in Rank-1 iKFAD (R-iKFAD). ...
429 Adaptive multi-resolution Gaussian processes: Scalable exact inference with naturally data-sparse covariance matrices
2609.30348
cs.LGcs.AI
Yanchuang Cao, Jun Liu, Tengchao Yu, Heng Yong
Gaussian processes constitute a cornerstone of probabilistic machine learning, yet scaling them to large datasets typically forces a trade-off between computational efficiency and model fidelity. This work bridges this gap by presenting an adaptive multi-resol...
Gaussian processes constitute a cornerstone of probabilistic machine learning, yet scaling them to large datasets typically forces a trade-off between computational efficiency and model fidelity. This work bridges this gap by presenting an adaptive multi-resolution Gaussian process framework that is both scalable and exact. Our key innovation is constructing a naturally data-sparse covariance matrix with adaptive multi-resolution basis functions. These basis functions are directly anchored to sa...
430 An End-to-End Pipeline for Causal ML with Continuous Treatments: An Application to Financial Decision Making
2609.30396
cs.LG
Javier Moral Hern\'andez, Clara Higuera-Caba\~nes, \'Alvaro Ibra\'in
This paper presents an end-to-end causal machine learning (ML) pipeline designed for real-world applications with continuous treatments. The proposed framework consists of six sequential steps: dimensionality reduction, causal identification, positivity assump...
This paper presents an end-to-end causal machine learning (ML) pipeline designed for real-world applications with continuous treatments. The proposed framework consists of six sequential steps: dimensionality reduction, causal identification, positivity assumption violation handling, estimation, refutation and evaluation, and policy optimization. We introduce practical contributions not currently available in existing causal ML toolkits, specifically: (1) a method for detecting and quantifying p...
431 Bayesian Uncertainty Quantification for fMRI Functional Connectivity via Simulation-Based Inference
2609.30445
cs.LG
Simon Carter, Zeming Kuang, Lilianne R. Mujica-Parodi, Helmut H. Strey
Optimizing fMRI scan duration and spatial resolution is critical for experimental design, yet traditional correlation-based approaches cannot quantify uncertainty or disentangle scanner measurement noise from true neural variability across subjects. Without pr...
Optimizing fMRI scan duration and spatial resolution is critical for experimental design, yet traditional correlation-based approaches cannot quantify uncertainty or disentangle scanner measurement noise from true neural variability across subjects. Without principled uncertainty bounds, researchers cannot know whether a protocol is long enough to reliably estimate connectivity, or whether between-subject differences reflect biological variation or noise. We present a Bayesian framework modeling...
432 Scaffold-Constrained Subset Dynamic Programming for Exact SSE Clustering
2609.30477
cs.LG
Yordan P. Raykov, Max A. Little
Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential. We use data-derived geometric graphs to precondition an exact subset dynamic program: as a result only connected ver...
Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential. We use data-derived geometric graphs to precondition an exact subset dynamic program: as a result only connected vertex subsets are admitted as clusters, while sum-of-squared-errors (SSE) loss is unchanged. A remaining-set recurrence minimises fixed-\(K\) or penalised SSE, with exact factorisation over the connected components of each remaining set. The ...
433 Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
2609.30484
cs.LGcs.AI
Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in ...
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However...
434 Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds
2609.30499
cs.LG
Wei Biao Wu
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional mo...
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displacement argument yields $T^{-1/3}$ expected average squared-gradient stationarity with a horizon-dependent stepsize. An explicit oracle-complexity corollary matches the known smooth ...
435 Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation
2609.30517
cs.LGeess.AS
Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination...
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory m...
436 Learning to Replace MCMC in Split-Gibbs Diffusion Posterior Sampling via Deep Unfolding
2609.30539
cs.LG
Yi Zhang, Rui Guo, Mengchu Xu, Zhaofeng Liu, Yonina C. Eldar
Split Gibbs sampling enables diffusion posterior inference for general nonlinear inverse problems by decoupling prior and likelihood computations, allowing a pretrained diffusion prior to be reused across measurement models. However, its likelihood update ofte...
Split Gibbs sampling enables diffusion posterior inference for general nonlinear inverse problems by decoupling prior and likelihood computations, allowing a pretrained diffusion prior to be reused across measurement models. However, its likelihood update often relies on iterative MCMC, which can hinder parallelization, require algorithm-specific tuning, and incur substantial computational cost. In this work, we propose a learning-based framework to replace this MCMC step by reformulating both G...
437 Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study
2609.30553
cs.LGcs.AI
Soumen Garai, Suman Samui
Expensive evolutionary search does not always need an exact fitness estimate for every candidate. It often needs a reliable answer to a simpler question: which candidate is better? We address this need through Teacher-Guided Learning NSGA-II (TGL-NSGA-II), a l...
Expensive evolutionary search does not always need an exact fitness estimate for every candidate. It often needs a reliable answer to a simpler question: which candidate is better? We address this need through Teacher-Guided Learning NSGA-II (TGL-NSGA-II), a low-fidelity framework for constrained Tiny Machine Learning (TinyML) neural architecture search. A pretrained teacher organizes samples into strata defined jointly by difficulty and class. Each candidate then undergoes KD-Lite, a short and ...
438 T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation
2609.30576
cs.LGcs.AI
Yang Liu, Noel Loo, Ali Khanafer, Shuying Sun, Akshay Soni
Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoP...
Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoPE encodes token indices for relative position reasoning, but in recommendation, an interaction index records only event order, saying nothing about elapsed time, behavioral cycles across scales, or calendar phase. We revisit this choice and...
439 Encryptability As a Coordinate Choice: Depth-One Homomorphic Federated Learning of Quantum Neural Networks
2609.30581
cs.LG
Marcel Mordarski, Nathan Mani, Arshad Patel, William Knottenbelt, Roberto Bondesan
Encrypted training relies on keeping server-side updates low-degree. This constraint traditionally excludes models whose weights inhabit a compact Lie group (notably variational quantum circuits, where every trainable weight is an $\mathrm{SU(2)}$ rotation). E...
Encrypted training relies on keeping server-side updates low-degree. This constraint traditionally excludes models whose weights inhabit a compact Lie group (notably variational quantum circuits, where every trainable weight is an $\mathrm{SU(2)}$ rotation). Expressed in Euler angles or discrete alphabets, these updates appear transcendental, historically demanding prohibitive costs: one client--server round per gate, or upwards of $25{,}000$ operations per weight. This penalty is strictly an ar...
440 MARCEDES: Score-based causal discovery under non-Gaussianity with continuous optimization
2609.30643
cs.LG
Anamitra Chaudhuri, Anirban Bhattacharya, Yang Ni
We consider the problem of learning the underlying causal directed acyclic graph (DAG) structure corresponding to a structural equation model (SEM) with non-Gaussian errors. Motivated by an intentionally misspecified non-Gaussian SEM with all Laplace errors, w...
We consider the problem of learning the underlying causal directed acyclic graph (DAG) structure corresponding to a structural equation model (SEM) with non-Gaussian errors. Motivated by an intentionally misspecified non-Gaussian SEM with all Laplace errors, we first introduce the mean absolute residual risk, defined over the space of all real matrices, and show that, asymptotically, the risk of the true weighted causal DAG matrix is strictly smaller than that of any other matrix. Nevertheless, ...
441 DiffusionShadow: Diffusion-based Shadow Caching for Neural Volume Rendering
2609.30658
cs.LG
Kai-Chen Tung, Qi Wu, David Bauer, Mengjiao Han, Silvio Rizzi
Implicit neural representations (INRs) have gained momentum in scientific visualization due to their compactness and scalability to large datasets, making them well suited for integration with direct volume rendering (DVR). However, real-time volume rendering ...
Implicit neural representations (INRs) have gained momentum in scientific visualization due to their compactness and scalability to large datasets, making them well suited for integration with direct volume rendering (DVR). However, real-time volume rendering of INR with advanced illumination effects, such as shadows, remains computationally expensive, as evaluating shadow terms via ray marching is costly. Alternatively, precomputing and storing shadows for many lighting directions is prohibitiv...
442 On the Limits of Univariate Deep Learning for Significant Wave Height Forecasting
2609.30688
cs.LG
Yilin Zhai, Hongyuan Shi, Zaijin You
This study conducts a systematic hyperparameter search across five deep learning architectures, DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2, and nine context lengths (1-168 h) for single-station significant wave height (Hs) forecasting on NDBC buoy 41009, ...
This study conducts a systematic hyperparameter search across five deep learning architectures, DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2, and nine context lengths (1-168 h) for single-station significant wave height (Hs) forecasting on NDBC buoy 41009, followed by re-evaluation of the best configurations on a 47-buoy, 37-year corpus. The five families converge to a common performance level on the multi-buoy evaluation (between-family SD = 0.0014 m^2, 0.8% of the grand mean), a spread dwar...
443 Threat-Aware Energy-Efficient Deployment for Dynamic UAV Networks: A Multi-Agent RL Approach
2609.30690
cs.LGcs.AI
Faisal Al-Kamali, Hussein A. Ammar, Francois Chan, James H. Bayes, Yasser Gadallah
Ensuring operational safety in threat-prone environments remains a critical challenge for multi-UAV networks serving as aerial base stations. This paper proposes an efficient framework to maximize global energy efficiency (EE) while promoting safe operation th...
Ensuring operational safety in threat-prone environments remains a critical challenge for multi-UAV networks serving as aerial base stations. This paper proposes an efficient framework to maximize global energy efficiency (EE) while promoting safe operation through threat-aware clustering and reward-based safety enforcement. The proposed framework is executed in three steps. First, a threat-aware K-means (TAKM) algorithm determines the minimum required UAVs and computes safe initial placements. ...
444 Parameter Estimation for Unnormalized Discrete Models via Empirically Localized Deformed Bregman Divergence
2609.30713
cs.LG
Takashi Takenouchi
Estimation of parameter of probabilistic models is an important task in the field of machine learning.For models of discrete variables, calculation of the normalization constant of model is sometimes difficult and a lot of researches have been done to avoid th...
Estimation of parameter of probabilistic models is an important task in the field of machine learning.For models of discrete variables, calculation of the normalization constant of model is sometimes difficult and a lot of researches have been done to avoid the calculation of the normalization constant. In this paper, we tackle with the difficulty by combining a technique of empirical localization and a deformed Bregman divergence.The technique of empirical localization makes it possible to dras...
445 Input-Layer Starvation: Why Per-Layer Pruning Breaks IoT Intrusion Detectors
2609.30729
cs.LG
Md Anas Biswas
Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy. We show that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair. On CICIoT2023, a...
Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy. We show that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair. On CICIoT2023, a two-layer convolutional detector pruned with uniform layer-wise magnitude pruning at 80% sparsity loses 16 points of accuracy but half of its macro-F1, the mean per-class F1 (0.542 to 0.271 over five independently trained models); 17 of 34...
446 TR-SSQP: A Trust-Region Method for Constrained Stochastic Optimization under Heavy-Tailed Noise
2609.30732
cs.LG
Haoxuan Wang, Yuchen Fang, Sen Na
We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challe...
We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challenges. Moreover, existing theoretical guarantees for constrained stochastic methods predominantly rely on bounded-variance assumptions, leaving the heavy-tailed noise regime largely unexplored. To address this gap, we propose a novel trust-r...
447 Skill Profiling with Attributable Reasoning (SPAR): A Wearable Analysis System for Boxing
2609.30753
cs.LG
Nibraas Khan, Hanchen David Wang, Enya Bullard, Ritam Ghosh, Ruj Haan
A punch is a ballistic, full-body action driven by a kinetic chain running from the legs through the trunk to the arm, where a small sequencing error separates a scoring strike from a miss. Wearable sensors can capture this movement in the gym, but most deploy...
A punch is a ballistic, full-body action driven by a kinetic chain running from the legs through the trunk to the arm, where a small sequencing error separates a scoring strike from a miss. Wearable sensors can capture this movement in the gym, but most deployable systems only classify which punch was thrown rather than assess how well it was thrown. We present Skill Profiling with Attributable Reasoning (SPAR), an eight-IMU garment and pressure-insole system that classifies each punch as expert...
448 HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models
2609.30763
cs.LGcs.AI
Yixuan Li, Weihao Li, Ziyang Song
Biomedical language models (LMs) encode textual semantics but do not explicitly preserve medical code hierarchies. We present Hyperbolic Clinical Ontology Embeddings (HCOE) for hierarchy-aware clinical concept representation. HCOE maps frozen BioBERT embedding...
Biomedical language models (LMs) encode textual semantics but do not explicitly preserve medical code hierarchies. We present Hyperbolic Clinical Ontology Embeddings (HCOE) for hierarchy-aware clinical concept representation. HCOE maps frozen BioBERT embeddings into a Poincare ball, combining parent-side and child-side ontology-guided contrastive learning with coarse-to-fine ontology-path aggregation. It uses International Classification of Diseases (ICD) codes organized by Clinical Classificati...
449 Deep-Learning Solvers and Surrogates for Infinity and p-Laplace Problems
2609.30809
cs.LG
Tak Shing Au Yeung, Ka Chun Cheung, Hannah Potgieter, Steven J. Ruuth, Simon See
We investigate the use of neural network solvers for infinity and $p$-Laplace problems, which are fundamental in nonlinear analysis and have practical applications. Our approach employs Physics-Informed Neural Networks (PINNs) and Deep Operator Networks (DeepO...
We investigate the use of neural network solvers for infinity and $p$-Laplace problems, which are fundamental in nonlinear analysis and have practical applications. Our approach employs Physics-Informed Neural Networks (PINNs) and Deep Operator Networks (DeepONets) to address computational challenges associated with large $p$ values, ranging from $2$ to $1000$, on various 2D and 3D domains. Our method offers advantages over traditional physics-based solvers, especially in three dimensions where ...
450 The KV Cache Is the New Memory Wall
2609.30854
cs.LG
Tejinder Singh
Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprin...
Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprint exceeds the 80 GB HBM of a single accelerator, and one 128k-token sequence adds 42 GB of KV cache. Techniques that compress, evict, page, share, or offload KV state have proliferated, but reported gains use inconsistent workloads, hardwar...
451 AC Power Flow Contingency Analysis Using a Single Deep Neural Network
2609.30859
cs.LG
Md Obaidur Rahman, Junjie Qin, Vassilis Kekatos
Contingency analysis using the AC power flow (AC-PF) model is a critical tool for accurate grid security assessment, but its computational burden increases with the number of operating scenarios and outage configurations to evaluate. Recent ML-based approaches...
Contingency analysis using the AC power flow (AC-PF) model is a critical tool for accurate grid security assessment, but its computational burden increases with the number of operating scenarios and outage configurations to evaluate. Recent ML-based approaches typically require outage-specific training data, leading to offline training costs that scale with the number of contingencies. This work proposes a framework that reuses a single ML model trained solely on basecase AC-PF data to estimate ...
452 Tight Stochastic Condition-Number Dependence in Nonconvex-Strongly-Concave Minimax Optimization
2609.30877
cs.LG
Qihao Zhou
We study whether the linear condition-number dependence in the stochastic complexity of SAPD+ is necessary for nonconvex-strongly-concave minimax optimization. For jointly $L$-smooth objectives with dual strong-concavity parameter $\mu$, we prove a lower bound...
We study whether the linear condition-number dependence in the stochastic complexity of SAPD+ is necessary for nonconvex-strongly-concave minimax optimization. For jointly $L$-smooth objectives with dual strong-concavity parameter $\mu$, we prove a lower bound that matches the SAPD+ upper bound under the same Moreau-envelope stationarity criterion and the same primal-dual initialization gap. Specifically, when $\sigma\ge\varepsilon$, the worst-case complexity of zero-respecting algorithms is $\T...
453 TISD: On-Policy Self-Distillation with Trajectory Intervention
2609.30878
cs.LGcs.AI
Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin
On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but canno...
On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch r...
454 EXAONE Demand 1.0: A Time Series Foundation Model for Demand Forecasting
2609.30880
cs.LGcs.AI
Seunghan Lee, Sangjun Han, Jun Seo, Junhyeok Kang, Jaehoon Lee
Time series foundation models (TSFMs) are pretrained on series from diverse domains, where demand series make up only a small fraction. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock-outs, and e...
Time series foundation models (TSFMs) are pretrained on series from diverse domains, where demand series make up only a small fraction. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock-outs, and exogenous events that the series does not record. To this end, we propose EXAONE Demand, built on 1) a demand-specific corpus and 2) a demand-aware adapter. For the corpus, we assemble 11.3M series and 48.4B observations from 73 sources, and...
455 Retraction-Based Gradient Projection Algorithms on Manifolds
2609.30885
cs.LG
Conglong Xu, Hao Wu
We introduce a framework for retraction-based convex optimization on Riemannian manifolds, which includes a notion of retraction-specific convex sets and retraction-based gradient projection algorithms. The standard theory of gradient projection algorithms gen...
We introduce a framework for retraction-based convex optimization on Riemannian manifolds, which includes a notion of retraction-specific convex sets and retraction-based gradient projection algorithms. The standard theory of gradient projection algorithms generalizes easily to this framework. Within this framework, we establish convergence results for retraction-based gradient projection algorithms with various stepsize rules. As an application, we use our framework to study the weighted low-ra...
456 Conformal Prediction under Exponential-Tilt Joint Shift
2609.30886
cs.LG
Seungjin Choi
Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled target inputs, allowing both the input distribution and its relationship with outcomes to change. We use Exp...
Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled target inputs, allowing both the input distribution and its relationship with outcomes to change. We use Exponential Tilt Reweighting Alignment (ExTRA), introduced for classification by Maity et al. (2023), to estimate structured distribution shifts. We compare using its estimated weights in conformal calibration with additionally tilting the sou...
457 Training Graph Foundation Models on The Web Graph
2609.30894
cs.LGcs.AI
Ryoma Sato
We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node...
We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node clustering, and graph generation, without additional training, (iii) has in-context learning capabilities, and (iv) does not rely on pretrained LLMs. In particular, existing graph foundation models often require training additional classif...
458 A Comprehensive Study of Content Representations for Speech Synthesis
2609.30975
cs.LGcs.SDeess.AS
Diego Torres, Axel Roebel, Nicolas Obin
Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address ...
Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio...
459 Synth-JEPA: Joint Embedding Prediction for Renderer-Free Synthesizer Parameter Search
2609.31024
cs.LGcs.SDeess.AS
Ben Hayes, Haokun Tian, Stefan Lattner
Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We ...
Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutually predictive audio and parameter representations from paired synthesizer data. At inference, candidate parameters are scored directly in this learned space, yielding a renderer-free objective whose ...
460 Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control
2609.31025
cs.LG
Claudio Canales, Fang Nan, Marco Hutter, Javier Ruiz-del-Solar
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learn...
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven ...
461 DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving
2609.31047
cs.LGcs.AI
Junyi Shen, Noppanat Wadlom, Zhengyuan Su, Yao Lu
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Cac...
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key that identifies a reusable result is not known until then. In this paper, we propose DynBranch, which makes an unresolved branch addressable before it resolves. Its stable coordinate lets candidate subgr...
462 A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine
2609.31101
cs.LG
Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee mac...
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and the Hessian spectrum are analytically tractable in the proportional high-dimensional limit. First, we use a zero-temperature Gibbs formulation to obtain predictions for the observables of the typical m...
463 Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year
2609.31141
cs.LG
Mahmoud Abbasi
Knowledge distillation is the standard way to compress encrypted-traffic classifiers for the edge, and almost all such work judges students by accuracy alone. We ask what else a student inherits: unknown-traffic detection, calibration, shortcut reliance, and w...
Knowledge distillation is the standard way to compress encrypted-traffic classifiers for the edge, and almost all such work judges students by accuracy alone. We ask what else a student inherits: unknown-traffic detection, calibration, shortcut reliance, and whether any survives a year of drift. Resemblance proves little on its own, since soft targets also regularise. We therefore distil one 101k-parameter student from two teachers of equal accuracy but different construction, a five-member ense...
464 SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations
2609.31164
cs.LGcs.AI
Tobias Vente, Maarten Peirsman, Noah Dani\"els, Hannu Toivonen, Bart Goethals
Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to de...
Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calcula...
465 BreathGRU: A Novel Semi-Supervised Bidirectional Gated Recurrent Unit Framework for Speech and Breath Segmentation for Respiratory Audio
2609.31165
cs.LGcs.SDeess.AS
Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold met...
Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold methods, Fourier Transform-based techniques, and unsupervised and pretrained voice activity detection (VAD) models, primarily focus on speech detection and often classify breathing events as non-speech or silence, limiting their applicability ...
466 BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
2609.31180
cs.LGcs.AIcs.SDeess.AS
Suhyun Kim, Jinmo Han, Danny Dongyeop Han, Ahhyun Lucy Lee, Jewoon Lee
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio o...
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic de...
467 Accounting for Bias Enables Sustainable LLM Evaluation
2609.31184
cs.LGcs.AI
Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wastef...
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additio...
468 Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution
2609.31214
cs.LGcs.AI
Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun
Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a ...
Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to ...
469 Geometric Moment Contraction for Stochastic Nesterov Acceleration
2609.31303
cs.LG
Wei Biao Wu
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz con...
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison proves synchronous $L^p$ contraction when $\beta\gamma L_p<(1-\beta)(1-q_{\gamma,p})$. This direct criterion includes infinite-variance gradients for $1<p<2$, but its small-step regime requires $\beta<...
470 Equation discovery with Bayesian tree-adjoining grammars
2609.31368
cs.LG
Christopher A. Lindley, Nikolaos Dervilis, Keith Worden
Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based ide...
Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based identifiers rely on evolutionary optimisation and return point estimates of the model structure. This paper instead proposes the TAG framework within a Bayesian setting. A generative prior is defined over tree structures and their parameters, ...
471 AFA-Net: A Differential Attention Approach for Auditory Attention Detection
2609.31402
cs.LGcs.SD
Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H. L. Hansen
Auditory Attention Detection (AAD) utilizes electroencephalographic (EEG) signals to identify a target speaker in a multi-speaker environment. Despite considerable progress, existing deep learning architectures often lack explicit mechanisms for handling noisy...
Auditory Attention Detection (AAD) utilizes electroencephalographic (EEG) signals to identify a target speaker in a multi-speaker environment. Despite considerable progress, existing deep learning architectures often lack explicit mechanisms for handling noisy EEG data. To address this limitation, we propose Auditory Focus Attention Networks (AFA-Net), a machine learning framework that replaces vanilla attention with a simple yet flexible differential attention mechanism to help focus on task-re...
472 Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
2609.31458
cs.LG
Jaehee Seo, Jisu Kim
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrica...
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size...
473 LandscapeSHAP: Which Persistent Homology Class Gets the Credit?
2609.31469
cs.LG
Nikola Mili\'cevi\'c
Shapley values, a solution concept from cooperative game theory, have recently become a standard tool for feature credit allocation in machine learning. They provide an axiomatically justified method to fairly distribute a model's prediction among the data fea...
Shapley values, a solution concept from cooperative game theory, have recently become a standard tool for feature credit allocation in machine learning. They provide an axiomatically justified method to fairly distribute a model's prediction among the data features. Shapley values have not yet been applied to explain machine learning models trained on features from topological data analysis. We develop what we believe is the first such approach, focusing on the persistence landscape featurizatio...
474 Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport
2609.31470
cs.LG
Haixiang Sun, Andrew L. Liu
Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases tha...
Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases that finite datasets fail to capture, making simple resampling or perturbation insufficient for stress scenario generation. Existing outlier synthesis methods typically rely on sparse neighborhoods, low support latent regions, or classifier bo...
475 Scaling Density Functional Theory with Gaussian Splatting
2609.31483
cs.LG
Andr\'es Guzm\'an-Cordero, Cindy Zhang, Majdi Hassan, Marta Skreta, Kirill Neklyudov
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate h...
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate how accuracy and cost scale with system size. We propose Gaussian Splatting for Density Functional Theory (GS-DFT), which represents molecular orbitals as a cloud of Gaussians whose positions, shapes, and mixing coefficients are optimized jo...
476 Retail Product Search: A Practical Approach at Target
2609.31498
cs.LG
Darshan Sonagara, Qujiaheng Zhang, Ankit Singh, Alex Li
Search is one of the most important features in e-commerce, directly driving customer engagement and business growth. A good product search system must show both relevant and desirable results. However, retail search presents unique challenges. User intent can...
Search is one of the most important features in e-commerce, directly driving customer engagement and business growth. A good product search system must show both relevant and desirable results. However, retail search presents unique challenges. User intent can range from exact matches to open-ended discovery. Search systems must also balance multiple goals, such as relevance, revenue, and profit, while keeping response times low. Traditional keyword-based methods often fall short in handling nat...
477 Retrainable physics-integrated neural differentiable modeling of sintering across material systems
2609.31518
cs.LG
Zeping Chen, Ani Aprahamian, Khachatur V. Manukyan, Tengfei Luo
Sintering is widely used to manufacture ceramics, but coupled densification and grain growth, material-dependent kinetics, and sparse measurements complicate predictive modeling and process design. We present Sinter-PiNDiff, a retrainable physics-integrated ne...
Sintering is widely used to manufacture ceramics, but coupled densification and grain growth, material-dependent kinetics, and sparse measurements complicate predictive modeling and process design. We present Sinter-PiNDiff, a retrainable physics-integrated neural differentiable framework for predicting density and grain-size evolution. Two neural networks learn densification and grain-growth coefficients within coupled rate equations, while a smooth saturation factor attenuates densification ne...
478 A Flow Matching Framework for Neural Representational Dissimilarity
2609.31544
cs.LGcs.AI
Zeyuan Ye, Xue-Xin Wei
Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are...
Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these distances arise as Jeffreys divergences under different ve...
479 EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
2609.31551
cs.LG
Kunxiong Zhu, Zhihao Shu, Hangyu Zheng, Minghai Qin, Miao Yin
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns ...
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks...
480 Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach
2609.31570
cs.LG
Damiano Brigo, Rapha\"el Huser, Dan Leonte
Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulati...
Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can ...
481 Statistical attribute alignment for black-box generative AI via output post-processing
2609.31607
cs.LGcs.AI
Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is mot...
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representa...
482 First-Order Stationarity of Reverse Diffusions
2609.31612
cs.LG
Zhifeng Chen, Chenyang Jiang, Yazhen Wang
Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative...
Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse...
483 Gap-free Differentially Private PCA for Gaussian Data
2609.31614
cs.LG
Alina Ene, Huy L. Nguyen
We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
484 Differentially-Private Decision Trees and Provable Robustness to Data Poisoning
2305.15394
cs.LG
Dani\"el Vos, Jelle Vos, Tianyu Li, Zekeriya Erkin, Sicco Verwer
Decision trees are interpretable models that are well-suited to non-linear learning problems. Much work has been done on extending decision tree learning algorithms with differential privacy, a system that guarantees the privacy of samples within the training ...
Decision trees are interpretable models that are well-suited to non-linear learning problems. Much work has been done on extending decision tree learning algorithms with differential privacy, a system that guarantees the privacy of samples within the training data. However, current state-of-the-art algorithms for this purpose sacrifice much utility for a small privacy benefit. These solutions create random decision nodes that reduce decision tree accuracy or spend an excessive share of the priva...
485 Foundations of Reinforcement Learning and Interactive Decision Making
2312.16730
cs.LG
Dylan J. Foster, Alexander Rakhlin
Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms and robotics to medical treatments. Thi...
Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms and robotics to medical treatments. This monograph gives a statistical perspective on algorithm design and complexity for interactive decision making, building from multi-armed bandits through contextual and structured bandits to reinforcement learning with function approximatio...
486 Efficient Constrained Graph Search for Post-hoc Error Correction in Binary Classifiers
2401.04282
cs.LG
Qinwu Xu
We introduce a model-agnostic framework for constrained post-hoc error correction in binary classifiers. Given a frozen base classifier, the method searches for an interpretable conjunction of feature--threshold rules that corrects residual false-positive or f...
We introduce a model-agnostic framework for constrained post-hoc error correction in binary classifiers. Given a frozen base classifier, the method searches for an interpretable conjunction of feature--threshold rules that corrects residual false-positive or false-negative errors while explicitly constraining newly introduced errors. The approach combines graph-based search over candidate rule paths, depth-dependent dynamic constraints, and a reduced-histogram procedure for efficient threshold e...
487 LEAD: An EEG Foundation Model for Alzheimer's Disease Detection
2502.01678
cs.LGcs.AI
Yihe Wang, Nan Huang, Nadia Mammone, Marco Cecchi, Xiang Zhang
Electroencephalography (EEG) provides a non-invasive, highly accessible, and cost-effective approach for detecting Alzheimer's disease (AD). However, existing methods, whether based on handcrafted feature engineering or standard deep learning, face three major...
Electroencephalography (EEG) provides a non-invasive, highly accessible, and cost-effective approach for detecting Alzheimer's disease (AD). However, existing methods, whether based on handcrafted feature engineering or standard deep learning, face three major challenges: 1) the lack of large-scale EEG-based AD datasets for robust representation learning and evaluation; 2) limited cross-subject generalizability; and 3) difficulty in adapting to highly heterogeneous data. To address these challen...
488 AYLA: Amplifying Gradient Sensitivity via Loss Transformation in Non-Convex Optimization
2504.01875
cs.LG
Behnam Gheshlaghi, Shahin Atakishiyev
Stochastic Gradient Descent (SGD) and its variants, such as ADAM, are foundational to deep learning optimization, adjusting model parameters through fixed or adaptive learning rates based on loss function gradients. However, these methods often struggle to bal...
Stochastic Gradient Descent (SGD) and its variants, such as ADAM, are foundational to deep learning optimization, adjusting model parameters through fixed or adaptive learning rates based on loss function gradients. However, these methods often struggle to balance adaptability and efficiency in high-dimensional, non-convex settings. This paper introduces AYLA, a novel optimization framework that enhances training dynamics via loss function transformation. AYLA applies a tunable power-law transfo...
489 ChemMLLM: Chemical Multimodal Large Language Model
2505.16326
cs.LG
Qian Tan, Di Zhang, Ben Gao, Peng Xia, Wanhao Liu
Recent years have seen rapid progress in multimodal large language models (MLLMs) in the field of chemistry. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill this gap, we propose ChemMLLM, a unified...
Recent years have seen rapid progress in multimodal large language models (MLLMs) in the field of chemistry. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill this gap, we propose ChemMLLM, a unified chemical multimodal large language model for molecule understanding and generation. In this work, we design five types of multimodal tasks across text, molecular SMILES strings and images, and curate the datasets. We benchmark ChemMLLM aga...
490 LapDDPM: Spectral Perturbation Diffusion for Robust Single-Cell Manifold Generation
2506.13344
cs.LGcs.AI
Lorenzo Bini, Stephane Marchand-Maillet
Generating high-fidelity and biologically plausible synthetic single-cell RNA sequencing (scRNA-seq) data is a critical challenge in computational biology, driven by the need to model high-dimensional, sparse, and non-linear cellular manifolds. Existing genera...
Generating high-fidelity and biologically plausible synthetic single-cell RNA sequencing (scRNA-seq) data is a critical challenge in computational biology, driven by the need to model high-dimensional, sparse, and non-linear cellular manifolds. Existing generative models often fail to capture the complex topology of cellular differentiation or lack robustness against technical noise and structural variability. We introduce LapDDPM, a novel conditional Graph Diffusion Probabilistic Model designed...
491 DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology
2507.22554
cs.LG
Joshua Dimasaka, Christian Gei{\ss}, Emily So
To understand our global progress for sustainable development and disaster risk reduction in many developing economies, two recent major initiatives - the Uniform African Exposure Dataset of the Global Earthquake Model (GEM) Foundation and the Modelling Exposu...
To understand our global progress for sustainable development and disaster risk reduction in many developing economies, two recent major initiatives - the Uniform African Exposure Dataset of the Global Earthquake Model (GEM) Foundation and the Modelling Exposure through Earth Observation Routines (METEOR) Project - implemented classical spatial disaggregation techniques to generate large-scale mapping of urban morphology using the information from various satellite imagery and its derivatives, g...
492 Neural Bridge Processes
2508.07220
cs.LGcs.AI
Jian Xu, Yican Liu, Delu Zeng, John Paisley, Qibin Zhao
Learning stochastic functions from partially observed context-target pairs requires models that are expressive, uncertainty-aware, and strongly conditioned on inputs. Neural Diffusion Processes (NDPs) improve expressivity with denoising diffusion, but their fo...
Learning stochastic functions from partially observed context-target pairs requires models that are expressive, uncertainty-aware, and strongly conditioned on inputs. Neural Diffusion Processes (NDPs) improve expressivity with denoising diffusion, but their forward process is input-independent; inputs only enter the reverse denoiser, so the noisy training states themselves do not encode the conditioning inputs. We propose Neural Bridge Processes (NBPs), which replace the unconditional forward ke...
493 Stochastic Bilevel Optimization with Heavy-Tailed Noise
2509.14952
cs.LG
Zhuanghua Liu, Luo Luo
This paper considers the smooth bilevel optimization in which the lower-level problem is strongly convex and the upper-level problem is possibly nonconvex. We focus on the stochastic setting where the algorithm can access the unbiased stochastic gradient evalu...
This paper considers the smooth bilevel optimization in which the lower-level problem is strongly convex and the upper-level problem is possibly nonconvex. We focus on the stochastic setting where the algorithm can access the unbiased stochastic gradient evaluation with heavy-tailed noise, which is prevalent in many machine learning applications, such as training large language models and reinforcement learning. We propose a nested-loop normalized stochastic bilevel approximation (N$^2$SBA) for ...
494 Transformers Discover Molecular Structure Without Graph Priors
2510.02259
cs.LG
Tobias Kreiman, Yutong Bai, Fadi Atieh, Elizabeth Weaver, Eric Qu
Computational simulations play a central role in scientific discovery, and machine learning (ML) has emerged as a promising alternative to traditional physics-based modeling. However, scientific modeling requires physically meaningful predictions, raising a fu...
Computational simulations play a central role in scientific discovery, and machine learning (ML) has emerged as a promising alternative to traditional physics-based modeling. However, scientific modeling requires physically meaningful predictions, raising a fundamental question for data-driven methods: to what extent can physical inductive biases-that is, prior assumptions about the structure of the physical world-emerge by learning from data alone? In this work, we study atomistic modeling, a r...
495 Models Got Talent: Identifying High Performing Wearable Human Activity Recognition Models Without Training
2511.06157
cs.LGcs.AI
Richard Goldman, Varun Komperla, Thomas Ploetz, Harish Haresamudram
Discovering high performing model architectures for wearables-based Human Activity Recognition (HAR) applications is challenging. The astonishing diversity and variability due to differing sensor locations, recording apparatus, activities, etc., can cause esta...
Discovering high performing model architectures for wearables-based Human Activity Recognition (HAR) applications is challenging. The astonishing diversity and variability due to differing sensor locations, recording apparatus, activities, etc., can cause established architectures to perform worse on datasets/tasks they were not designed for. A promising complement to Neural Architecture Search (NAS) involves the development of Zero Cost Proxies (ZCPs), which correlate well with trained performa...
496 FLAME: Flow Enhanced Legendre Memory Models for General Time Series Forecasting
2512.14253
cs.LG
Xingjian Wu, Zhengyu Li, Hanyin Cheng, Xiangfei Qiu, Jilin Hu
In this work, we introduce FLAME, a family of extremely lightweight and capable Time Series Foundation Models, which support versatile forecasting tasks via generative probabilistic modeling, while ensuring both efficiency and robustness. FLAME utilizes the Le...
In this work, we introduce FLAME, a family of extremely lightweight and capable Time Series Foundation Models, which support versatile forecasting tasks via generative probabilistic modeling, while ensuring both efficiency and robustness. FLAME utilizes the Legendre Memory for strong generalization capabilities. By adapting variants of Legendre Memory, i.e., translated Legendre (LegT) and scaled Legendre (LegS), in the Encoding and Decoding phases, FLAME can effectively capture the inherent indu...
497 MeshGraphNet-Transformer: Scalable Mesh-based Learned Simulation for Solid Mechanics
2601.23177
cs.LG
Mikel M. Iparraguirre, Iciar Alfaro, David Gonzalez, Elias Cueto
We present MeshGraphNet-Transformer (MGN-T), a novel architecture that combines the global modeling capabilities of Transformers with the geometric inductive bias of MeshGraphNets, while preserving a mesh-based graph representation. MGN-T overcomes a key limit...
We present MeshGraphNet-Transformer (MGN-T), a novel architecture that combines the global modeling capabilities of Transformers with the geometric inductive bias of MeshGraphNets, while preserving a mesh-based graph representation. MGN-T overcomes a key limitation of standard MGN, the inefficient long-range information propagation caused by iterative message passing on large, high-resolution meshes. A physics-attention Transformer serves as a global processor, updating all nodal states simultan...
498 RAPTOR: Ridge-Adaptive Logistic Probes
2602.00158
cs.LGcs.AI
Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang
Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted fr...
Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted from a probe and injected via additive activation steering by adding it to a layer representation during the forward pass. The effectiveness of this pipeline hinges on estimating concept vectors that are accurate, directionally stable under a...
499 Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
2602.04909
cs.LG
Youngjae Cho, Jongsuk Kim, Ji-Hoon Kim
Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fixed reference policy. As the policy drifts, a static reference, however, can become increasingly miscalibrated, l...
Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fixed reference policy. As the policy drifts, a static reference, however, can become increasingly miscalibrated, leading to distributional mismatch and amplifying spurious preference signals under noisy supervision. Conversely, reference-free variants avoid mismatch but often suffer from unconstrained reward drift. We propose Geometric Anchor Preferenc...
500 Distribution-Conditioned Transport
2603.04736
cs.LG
Nic Fishman, Gokul Gowri, Paolo L. B. Fischer, Marinka Zitnik, Omar Abudayyeh
Learning a transport model that maps a source distribution to a target distribution is a canonical problem in machine learning, but scientific applications increasingly require models that can generalize to source and target distributions unseen during trainin...
Learning a transport model that maps a source distribution to a target distribution is a canonical problem in machine learning, but scientific applications increasingly require models that can generalize to source and target distributions unseen during training. We introduce distribution-conditioned transport (DCT), a framework that conditions transport maps on learned embeddings of source and target distributions, enabling generalization to unseen distribution pairs. DCT also allows semi-superv...
501 Spectral-Sphere-Constrained Hyper-Connections
2603.20896
cs.LGcs.AI
Zhaoyi Liu, Haichuan Zhang, Ang Li
Hyper-Connections (HC) extend residual connections into multiple streams, employing residual matrices for cross-stream mixing to enrich model expressivity. However, unconstrained mixing disrupts the identity mapping property intrinsic to the residual connectio...
Hyper-Connections (HC) extend residual connections into multiple streams, employing residual matrices for cross-stream mixing to enrich model expressivity. However, unconstrained mixing disrupts the identity mapping property intrinsic to the residual connection, causing unstable training. To address this, Manifold-Constrained Hyper-Connections (mHC) and its variants restrict these matrices to be doubly stochastic via Sinkhorn-Knopp (SK) algorithm or permutation-based parameterizations. We reveal...
502 Below-ground Fungal Biodiversity Can be Monitored Using Self-Supervised Learning Satellite Features
2604.09818
cs.LG
Robin Young, Michael E. Van Nuland, E. Toby Kiers, Tom\'a\v{s} V\v{e}trovsk\'y, Petr Kohout
Mycorrhizal fungi are vital to terrestrial ecosystem functioning. Yet monitoring their biodiversity at landscape scales is often unfeasible due to time and cost constraints. Current predictions suggest that 90% of mycorrhizal diversity hotspots remain unprotec...
Mycorrhizal fungi are vital to terrestrial ecosystem functioning. Yet monitoring their biodiversity at landscape scales is often unfeasible due to time and cost constraints. Current predictions suggest that 90% of mycorrhizal diversity hotspots remain unprotected, opening questions of how to broadly and effectively map underground fungal communities. We show that self-supervised learning (SSL) applied to satellite imagery can predict below-ground ectomycorrhizal fungal richness across diverse en...
503 Unlocking the Forecasting Economy: A Suite of Datasets for the Full Lifecycle of Prediction Market: [Experiments \& Analysis]
2604.20421
cs.LG
Huaiyu Jia, Luofeng Zhou, Wentao Zhang, Lin William Cong, Siguang Li
Prediction markets are markets for trading claims on universal future events (e.g., presidential elections). Fueled by a meteoric surge with over \$50 billion trading volume, they have emerged as a promising forecasting mechanism, where their prices provide co...
Prediction markets are markets for trading claims on universal future events (e.g., presidential elections). Fueled by a meteoric surge with over \$50 billion trading volume, they have emerged as a promising forecasting mechanism, where their prices provide continuously updated signals of collective beliefs. In decentralized platforms (e.g., Polymarket), the prediction market lifecycle include six stages: market creation, token registration, trading, oracle interaction, dispute, and final settle...
504 How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
2605.05113
cs.LG
Mariia Seleznova
We study signal propagation in linear recurrent models at finite width. While existing signal propagation theory relies predominantly on the infinite-width limit, it remains unclear for how long that approximation remains accurate when recurrent depth $t$ grow...
We study signal propagation in linear recurrent models at finite width. While existing signal propagation theory relies predominantly on the infinite-width limit, it remains unclear for how long that approximation remains accurate when recurrent depth $t$ grows jointly with width $n$. This question is especially relevant for modern recurrent sequence models, whose natural operating regime involves long input sequences, i.e., large $t$. We derive exact finite-width formulas for the hidden state s...
505 Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction
2605.05134
cs.LG
Dan Wilson, Mohamed Akrout
Large Language Models (LLMs) frequently generate plausible but non-factual content, a phenomenon known as hallucination. While existing detection methods typically rely on computationally expensive sampling-based consistency checks or external knowledge retrie...
Large Language Models (LLMs) frequently generate plausible but non-factual content, a phenomenon known as hallucination. While existing detection methods typically rely on computationally expensive sampling-based consistency checks or external knowledge retrieval, we propose a new method that treats the LLM as a black-box dynamical system. By projecting LLM responses into a high-dimensional manifold via an embedding model, we characterize the resulting vector sequences as observable realizations...
506 Gradient-Momentum Coupling: A Parameter-Space Proxy for Learning Progress
2605.05856
cs.LG
Samuel Blad, Martin L\"angkvist, Amy Loutfi
Measuring learning progress is at the core of curiosity-driven exploration, which rewards an agent for going where its model is still learning. However, the abstract notion of learning progress is not directly measurable, and existing methods often derive it f...
Measuring learning progress is at the core of curiosity-driven exploration, which rewards an agent for going where its model is still learning. However, the abstract notion of learning progress is not directly measurable, and existing methods often derive it from the prediction error in the output space. This paper proposes Gradient-Momentum Coupling (GMC), which measures how strongly a sample drives change in the parameter space, given by the normalized absolute product of its gradient with the...
507 QuadraSHAP: $\epsilon$-Exact Shapley Values for Product Games in Logarithmic Parallel Time
2605.05870
cs.LG
Majid Mohammadi, Grigory Reznikov, Pavel Sinitcyn, Krikamol Muandet, Siu Lun Chau
We introduce QuadraSHAP, a method for $\epsilon$-exact Shapley computation in product games, cooperative games whose coalition values factorize across players. Given a tolerance $\epsilon\ge0$, the method determines a computational budget before evaluation, gu...
We introduce QuadraSHAP, a method for $\epsilon$-exact Shapley computation in product games, cooperative games whose coalition values factorize across players. Given a tolerance $\epsilon\ge0$, the method determines a computational budget before evaluation, guaranteeing an absolute attribution error of at most $\epsilon$ for every feature in exact arithmetic. By extending to weighted sums of product games, the framework supports baseline and empirical interventional attribution across a broad cl...
508 Geometry-Aware Simplicial Message Passing
2605.06061
cs.LG
Elena Xinyi Wang, Bastian Rieck
The Weisfeiler--Lehman (WL) test and its simplicial extension (SWL) characterize the combinatorial expressivity of message passing networks, but they are blind to geometry, i.e., meshes with identical connectivity but different embeddings are indistinguishable...
The Weisfeiler--Lehman (WL) test and its simplicial extension (SWL) characterize the combinatorial expressivity of message passing networks, but they are blind to geometry, i.e., meshes with identical connectivity but different embeddings are indistinguishable. We introduce the Geometric Simplicial Weisfeiler--Lehman (GSWL) test, which incorporates vertex coordinates into color refinement for geometric simplicial complexes. In addition, we show that (i) the expressivity of geometry-aware simplic...
509 Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression
2605.08475
cs.LGcs.AI
Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang
In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass. Under bounded-data ass...
In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass. Under bounded-data assumptions, we construct a single-head transformer whose forward pass approximately implements \textit{preconditioned Richardson iteration} on the associated kernel system. The construction uses $O(\log(1/\epsilon))$ blocks and MLP width $O(\...
510 Complex-Valued Phase-Coherent Transformer
2605.10123
cs.LG
Leona Hioki
Complex-valued Transformers have largely inherited softmax attention from real-valued architectures. However, row-normalised token competition is not necessarily aligned with phase-preserving computation. In this paper, we introduce the Phase-Coherent Transfor...
Complex-valued Transformers have largely inherited softmax attention from real-valued architectures. However, row-normalised token competition is not necessarily aligned with phase-preserving computation. In this paper, we introduce the Phase-Coherent Transformer (PCT), which applies a real-valued, element-independent, smooth gate to L2-normalised complex query-key similarities. PCT replaces token competition with token-non-competing attention and is designed to preserve phase information across...
511 Supervised Deep Multimodal Matrix Factorization for Interpretable Brain Network Analysis
2605.13312
cs.LG
Amjad Seyedi, Lifang He, Songlin Zhao, Akwum Onwunta, Nicolas Gillis
Multimodal brain network analysis faces a persistent trade-off between predictive accuracy and interpretability. Deep neural networks achieve high accuracy but behave as black boxes that reveal little about the brain modules driving their decisions, whereas ma...
Multimodal brain network analysis faces a persistent trade-off between predictive accuracy and interpretability. Deep neural networks achieve high accuracy but behave as black boxes that reveal little about the brain modules driving their decisions, whereas matrix factorization methods provide parts-based interpretability yet remain largely shallow, unsupervised, and restricted to a single view, integrating modalities through predefined or heuristic fusion rules. To bridge this gap with a formul...
512 Mixed neural posterior estimation for simulators with discrete and continuous parameters
2605.13551
cs.LG
Jan Boelts, Cornelius Schr\"oder, Jonas Beck, Jakob H. Macke, Michael Deistler
Neural Posterior Estimation (NPE) enables rapid parameter inference for complex simulators with intractable likelihoods. NPE trains an inference network to estimate a probability density over parameters given data, typically assumed to be \emph{continuous}. Ho...
Neural Posterior Estimation (NPE) enables rapid parameter inference for complex simulators with intractable likelihoods. NPE trains an inference network to estimate a probability density over parameters given data, typically assumed to be \emph{continuous}. However, many scientific models involve parameter spaces that are \emph{mixed}, that is, they contain both discrete and continuous dimensions. We address this limitation by extending NPE to mixed parameter spaces through an inference network ...
513 Tabular Imbalanced Learning: A Survey, Benchmark, and Practical Guide
2605.14915
cs.LG
Ruizhe Liu, Jiaqi Luo
Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed methods, there is still limited systematic understanding of how different imbalance-handling strategies perform across diverse d...
Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed methods, there is still limited systematic understanding of how different imbalance-handling strategies perform across diverse data regimes and computational constraints, making practical method selection difficult. In this work, we provide a systematic survey of tabular imbalanced learning and introduce Tabular Imbalanced Learning Benchmark (TILBench), a large-scal...
514 Attention Sinks and Outliers in Attention Residuals
2605.17887
cs.LGcs.AI
Haozheng Luo, Haoran Dai, Jingyuan Huang, Shaoyang Zhang, Xi Chen
We propose OASIS, an outlier- and sink-aware method that stabilizes dual-normalized attention-residual architectures through explicit null routing and token-to-depth null coupling. AttnResidual introduces an additional depth-wise normalization channel that imp...
We propose OASIS, an outlier- and sink-aware method that stabilizes dual-normalized attention-residual architectures through explicit null routing and token-to-depth null coupling. AttnResidual introduces an additional depth-wise normalization channel that improves inter-layer routing flexibility but can also amplify attention sinks, activation outliers, and low-bit quantization error. OASIS builds on explicit Softmax1-based null routes at both the token and depth levels and uses token-level nul...
515 Federated Martingale Posterior Samping
2605.18554
cs.LG
Boning Zhang, Matteo Zecchin, Mingzhao Guo, Dongzhu Liu, Osvaldo Simeone
Federated Bayesian neural networks require fixing a prior on the model parameters, which is notoriously difficult, and misspecification of this prior can severely degrade accuracy and calibration. Motivated by the rapid progress of predictive models, the marti...
Federated Bayesian neural networks require fixing a prior on the model parameters, which is notoriously difficult, and misspecification of this prior can severely degrade accuracy and calibration. Motivated by the rapid progress of predictive models, the martingale posterior, also known as predictive Bayes, replaces the prior--likelihood pair with a predictive distribution and recovers parameter uncertainty by repeatedly drawing predictive samples and refitting the model. This letter proposes {f...
516 MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification
2605.19752
cs.LG
Paul Krzakala, Gabriel Melo, Camille Lan\c{c}on, Charlotte Laclau, R\'emi Flamary
Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research. We address the Molecule Retrieval task, whic...
Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research. We address the Molecule Retrieval task, which consists in recovering the chemical structure of a metabolite from its MS/MS spectrum given a set of candidate molecules. We make three contributions. First, we propose a unified framework encompassing recent approaches based on represent...
517 Smooth Piecewise Cutting for Neural Operator to Handle Discontinuities and Sharp Transitions
2605.19823
cs.LGcs.AI
Ha Dang, Sebastian Schmidt, Juergen Hesser
Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently continuous representations struggle to capture discontinuities and sharp transitions. Existing approaches typically ...
Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently continuous representations struggle to capture discontinuities and sharp transitions. Existing approaches typically approximate such features within continuous function spaces, often requiring increased model capacity and high-resolution data. In this work, we propose Cut-DeepONet, a two-stage training framework that explicitly models discontinuities whi...
518 QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits
2605.30358
cs.LG
Zhenxiao Fu, Lei Jiang, Fan Chen
Quantum computing remains in the Noisy Intermediate-Scale Quantum (NISQ) era, with performance constrained by noise. Addressing this limitation requires hardware-facing capabilities beyond gate sequences: mid-circuit measurement and classical feedback for quan...
Quantum computing remains in the Noisy Intermediate-Scale Quantum (NISQ) era, with performance constrained by noise. Addressing this limitation requires hardware-facing capabilities beyond gate sequences: mid-circuit measurement and classical feedback for quantum error correction (QEC), precise timing for dynamical decoupling (DD), and pulse-level waveform access for calibration. OpenQASM 3 exposes these capabilities through a hardware-level programming interface. Despite rapid progress in large...
519 Proper Scoring Rules for Right-Censored Survival Data
2606.06393
cs.LG
Jef Jonkers, Glenn Van Wallendael, Luc Duchateau, Sofie Van Hoecke
Proper scoring rules provide a rigorous theoretical basis for the training and evaluation of probabilistic forecasts. In survival analysis, such forecasts describe the distribution of the time until an event occurs. However, this event time is often only parti...
Proper scoring rules provide a rigorous theoretical basis for the training and evaluation of probabilistic forecasts. In survival analysis, such forecasts describe the distribution of the time until an event occurs. However, this event time is often only partially observed because follow-up may end before the event occurs, resulting in right censoring. We propose a framework for proper scoring of right-censored survival outcomes based on a simple idea: first, map the predictive distribution thro...
520 QueryGraph: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation
2606.08300
cs.LG
Aishwarya Chakravarthy, Vidhi Kulkarni, Duen Horng Chau
Many real-world queries over personal data span multiple applications and require structured planning, as individual tools expose only partial information. While LLMs show strong reasoning and tool use, reliably executing multi-step, cross-tool queries remains...
Many real-world queries over personal data span multiple applications and require structured planning, as individual tools expose only partial information. While LLMs show strong reasoning and tool use, reliably executing multi-step, cross-tool queries remains challenging. We introduce a system that converts natural language queries into structured graphs and executes them via a deterministic planner. Our approach uses depth-first search to resolve dependencies and combine results across tools, ...
521 Bergson: An Open Source Library for Data Attribution
2606.11660
cs.LG
Lucia Quirke, Louis Jaburi, David Johnston, William Z. Li, Gon\c{c}alo Paulo
Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation. However, significant engin...
Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation. However, significant engineering effort is required to perform it at scale, and many cutting edge techniques lack open-source tooling and support. Bergson is an open source library that aims to enable faster progress in the field by providing a host of techniques th...
522 Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts
2606.14929
cs.LGcs.AI
Yan Dai, Negin Golrezaei, Patrick Jaillet
Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback...
Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited observability of models. We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and experts are the embedding models working on low...
523 ThousandWorlds: A benchmark for climate emulation of potentially habitable exoplanets
2606.18338
cs.LG
Edward T. Stevenson, Mei Ting Mak, Eric Wolf, Denis E. Sergeev, Tobi Hammond
The search for life beyond Earth will depend on detecting faint signatures in the atmospheres of potentially habitable exoplanets. Interpreting those signatures requires understanding the host planet's climate: the same molecule may signal life on one planet a...
The search for life beyond Earth will depend on detecting faint signatures in the atmospheres of potentially habitable exoplanets. Interpreting those signatures requires understanding the host planet's climate: the same molecule may signal life on one planet and abiotic chemistry on another. Global climate models (GCMs) provide this understanding, but individual runs can require up to millions of core-hours and substantial domain expert time. Machine-learning emulators could remove this bottlene...
524 Do Location Encoders Capture Spatial Effects? A GeoShapley Benchmark Across Scales
2606.23453
cs.LG
Daniel Kiv, Shaowen Wang
Location encoders transform geographic coordinates into high dimensional embeddings for downstream machine learning, but it is unclear how well these representations capture interpretable spatial effects. We benchmark whether GeoShapley, a game-theoretic expla...
Location encoders transform geographic coordinates into high dimensional embeddings for downstream machine learning, but it is unclear how well these representations capture interpretable spatial effects. We benchmark whether GeoShapley, a game-theoretic explainer that treats all location features as a single joint player, can recover spatially varying coefficients from models built on location-encoder embeddings. Eleven encoders from the TorchSpatial framework are evaluated against a synthetic ...
525 TeDiServe: High SLO Attainment Serving for Diffusion Language Models
2606.29094
cs.LG
Tzu-Tao Chang, Benjamin Yuanyang Hong, Kiet Pham, Shivaram Venkataraman
Diffusion language models (DLMs) have recently emerged as a promising alternative to conventional autoregressive language models. By generating multiple tokens in parallel during each denoising step, they offer higher inference throughput while maintaining com...
Diffusion language models (DLMs) have recently emerged as a promising alternative to conventional autoregressive language models. By generating multiple tokens in parallel during each denoising step, they offer higher inference throughput while maintaining competitive quality. However, realizing these throughput gains while meeting latency SLOs in a serving system requires addressing challenges introduced by DLMs' unique characteristics. These include navigating the speed-quality tradeoff create...
526 Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark
2606.30170
cs.LGcs.AI
Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda, Julian Lorenz, Rainer Lienhart
Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination yields strong benchmark metrics but limits transferability to domains structurally distinct from ...
Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination yields strong benchmark metrics but limits transferability to domains structurally distinct from drug discovery. To overcome this limitation and drive discovery toward real, scientifically grounded targets, we introduce the Nanotechnology Molecular Optimization (NMO) Benchmark, which bridges machine learning (ML) and quantum materials ...
527 Regularizing modality contribution drift in multimodal continual learning
2607.27260
cs.LG
Zhen Zhang, Jielei Chu, Wenjie Ban, Tian Sang, Yuxiao Li
Multimodal continual learning (MMCL) aims to acquire new knowledge from multimodal data while retaining previously learned knowledge. Existing MMCL methods primarily mitigate forgetting by aligning cross-modal representations or preserving feature-level semant...
Multimodal continual learning (MMCL) aims to acquire new knowledge from multimodal data while retaining previously learned knowledge. Existing MMCL methods primarily mitigate forgetting by aligning cross-modal representations or preserving feature-level semantic similarity. However, different tasks may rely on different modalities, and learning new tasks can alter how modalities contribute to predictions on previously learned tasks. It remains underexplored how modality contributions evolve acro...
528 Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
2608.04408
cs.LGcs.AI
De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma
On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through b...
On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through budget-matched teacher-continuation and rollback branches. Based on their relative success, states are categorized as recoverable, irreversible-but-avoidable, or ambiguous, and these labels guide whether training retains, rolls back, or conv...
529 A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning
2608.07335
cs.LGcs.AI
Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
Replay-free parallelized Q-learning removes the large experience replay buffers and target networks used by conventional deep Q-learning, but the role of network architecture in this training regime remains comparatively underexplored. We investigate this ques...
Replay-free parallelized Q-learning removes the large experience replay buffers and target networks used by conventional deep Q-learning, but the role of network architecture in this training regime remains comparatively underexplored. We investigate this question through a progressive three-phase study within the Parallelized Q-Network (PQN) framework. First, we compare eight convolutional encoder topologies on Atari-57 under a common training protocol while jointly considering performance and ...
530 A Banach-Space Theory of Markovian Halpern Iteration for Non-Expansive Maps
2608.15966
cs.LG
Ege C. Kaya, Arda Fazla, M. Berk Sahin, Abolfazl Hashemi
We study stochastic approximation of fixed points of a non-expansive operator $T$ when the oracle samples originate from a continuing Markovian trajectory. A direct block-minibatch implementation of Halpern iteration attains an expected last-iterate residual o...
We study stochastic approximation of fixed points of a non-expansive operator $T$ when the oracle samples originate from a continuing Markovian trajectory. A direct block-minibatch implementation of Halpern iteration attains an expected last-iterate residual of order $O(\log N/N)$, but accrues a substantive complexity of $\tilde O(\epsilon^{-5})$ Markovian samples. We therefore introduce a variance-reduced Markovian PAGE-Halpern method whose refresh and same-state difference blocks are analyzed ...
531 Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
2608.17965
cs.LGcs.AI
Bin Li, Dongdong Wang, Siyang Lu
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We ...
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conve...
532 DecoVAE: a Lightweight Interpretable Trend-Seasonal VAE Framework for Efficient Probabilistic Time Series Forecasting
2608.20052
cs.LG
Alexander Marusov, Dmitry Anikin, Alexey Zaytsev
Probabilistic time series forecasting remains challenging, largely because modeling distinct trend and seasonal dynamics requires specialized approaches. Existing methods often fail to capture the unique inner properties of these components, lack interpretabil...
Probabilistic time series forecasting remains challenging, largely because modeling distinct trend and seasonal dynamics requires specialized approaches. Existing methods often fail to capture the unique inner properties of these components, lack interpretability, or suffer from heavy memory and runtime overhead. To address these limitations, we propose DecoVAE, a lightweight interpretable trend-seasonal VAE framework that explicitly decomposes time series into trend and seasonal components by a...
533 PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors
2608.23101
cs.LGcs.AI
Nathan Duboisset, Zhaolan Huang, Felix Bie{\ss}mann, Roudy Dagher, Antoine Lavandier
Recent progress in the field of TinyML has demonstrated that low-power hardware based on microcontrollers can achieve bird species monitoring in real time based on acoustic sensor data for an entire breeding period on a single battery charge. However, the stat...
Recent progress in the field of TinyML has demonstrated that low-power hardware based on microcontrollers can achieve bird species monitoring in real time based on acoustic sensor data for an entire breeding period on a single battery charge. However, the state of the art on low-power microcontrollers was so far limited to binary classification of a single species. In contrast, real fauna monitoring deployments often target multiple species simultaneously. To address this challenge we develop Po...
534 SimCast-S2S: A Computationally Efficient Diffusion Model for Subseasonal Precipitation Forecasting
2608.26594
cs.LG
Hiep V. Dang, Antonios Mamalakis
Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains ...
Subseasonal-to-seasonal (S2S) precipitation forecasting has substantial financial and societal impact, yet remains challenging because of weak predictive signals, high associated uncertainty, and the computational cost of operational systems, which constrains simulation fidelity. We introduce SimCast-S2S, a generative latent-diffusion framework for probabilistic S2S precipitation forecasting that addresses three major bottlenecks in data-driven prediction. First, because S2S prediction requires ...
535 One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI
2608.29420
cs.LG
Louis Yiven Zhu
Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct vali...
Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct validity that a structural test and a predictive test can answer in opposite ways. We show that they do on a hash-pinned snapshot of a frontier leaderboard with 421 model configurations across twelve benchmarks, four of them economic, of which ...
536 MolLedger: An Additive Graph Neural Network with Chemically Grounded ADME Attributions
2608.30636
cs.LG
Christina X. Ji
Optimizing absorption, distribution, metabolism, and excretion (ADME) is an important part of small molecule drug discovery. Many machine learning models have been built to predict ADME properties to facilitate this optimization process, but explaining model p...
Optimizing absorption, distribution, metabolism, and excretion (ADME) is an important part of small molecule drug discovery. Many machine learning models have been built to predict ADME properties to facilitate this optimization process, but explaining model predictions is challenging. We propose a new graph neural network architecture with built-in atom attributions. Our model MolLedger learns a global context vector for each molecule and a per-atom head to output atom scores that sum to the pr...
537 Provably Safe Sim-to-Real Transfer
2609.01418
cs.LGcs.AI
Tingting Ni, Maryam Kamgarpour
We address safe sim-to-real transfer, in which an agent leverages an imperfect simulator and limited real-world interaction while ensuring safety throughout data collection in the real system. This problem arises in applications such as robotics and healthcare...
We address safe sim-to-real transfer, in which an agent leverages an imperfect simulator and limited real-world interaction while ensuring safety throughout data collection in the real system. This problem arises in applications such as robotics and healthcare: simulators provide cheap data, but sim-to-real mismatch makes direct transfer unreliable, and collecting real-world data to correct this mismatch must itself be safe. Moreover, deployment objectives may vary across tasks, making it costly...
538 WEECFP-SuRGE: A Position-Aware Substructure Encoding Method for Molecular Property Prediction
2609.04672
cs.LG
Robert Epps
Computational molecular property prediction requires representations that capture local chemistry, long-range interactions, and molecular topology. Conventional fingerprints provide efficient local substructure features, whereas learned graph and sequence mode...
Computational molecular property prediction requires representations that capture local chemistry, long-range interactions, and molecular topology. Conventional fingerprints provide efficient local substructure features, whereas learned graph and sequence models can represent broader context but often rely on pretraining or three-dimensional conformers. We introduce Wide Encoded Extended Connectivity Fingerprints (WEECFP) with Substructure Rotary Graph-distance Encoding (SuRGE), a tokenized hier...
539 The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
2609.06934
cs.LGcs.AI
Srikanth Malla, Chiho Choi, Joon Hee Choi
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023), fine-tuning attacks (Qi et al., 2024), and activation-space edits (Arditi et al., 2024) keep recovering the behaviors it was meant to re...
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023), fine-tuning attacks (Qi et al., 2024), and activation-space edits (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and follow it into pretraining. We measure the safety update $\Delta = W_{safe} - W_{base}$ against the curvature of the model's capabilities (the empirical Fisher of a capability loss)...
540 Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
2609.06951
cs.LGcs.AI
Srikanth Malla, Chiho Choi, Joon Hee Choi
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alo...
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly ref...
541 Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification
2609.10487
cs.LG
Weifeng Yang
We construct counterexamples to Rockafellar's sum conjecture in which two maximally monotone operators satisfy the interior-domain condition but their sum is not maximally monotone, thereby providing the complete disproof of the conjecture. We establish a gene...
We construct counterexamples to Rockafellar's sum conjecture in which two maximally monotone operators satisfy the interior-domain condition but their sum is not maximally monotone, thereby providing the complete disproof of the conjecture. We establish a general construction theorem that computes the entire monotone polar of a class of graphs and characterizes their maximal monotonicity by the nonexistence of solutions to explicit equations in the continuous dual. We also prove a pullback theor...
542 EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale
2609.10980
cs.LGcs.AI
Ege C. Kaya, Abolfazl Hashemi
EGGROLL (Sarkar et al., 2026) makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractive but geometrically severe: Each rank-...
EGGROLL (Sarkar et al., 2026) makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractive but geometrically severe: Each rank-one perturbation lies in a zero-volume subset of the ambient matrix space, despite having identity covariance. We characterize the EGGROLL update mean field at finite rank and nonzero perturbation radius as a resolvent applied to the gradie...
543 Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
2609.12890
cs.LGcs.AI
Junhao Zhao, David Michael Simberg, Jacob Kang, Colin Connor Kurniawan, Nan Xu
In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate t...
In autoregressive forecasting, long prediction rollouts provide distant supervision, but backpropagation through time (BPTT) carries gradients from those losses through many autoregressive steps. Repeated Jacobian products can make distant gradients dominate the update while amplifying predictable signal and unpredictable innovation together; a large distant gradient therefore need not carry reliable learning signal. Motivated by this, we introduce Internal Dual-Wiener routing (Internal-DW), a b...
544 Recovering Governing Dynamics from Distributed Observations via Exact Spline Merging
2609.16579
cs.LG
Naveen Mysore
Scientific observations are frequently distributed across locations, time periods, and institutions. Combining such observations into a continuous, differentiable field enables recovering governing physical parameters from its derivatives. This paper makes two...
Scientific observations are frequently distributed across locations, time periods, and institutions. Combining such observations into a continuous, differentiable field enables recovering governing physical parameters from its derivatives. This paper makes two contributions in this setting. First, the established additive structure of fixed-basis ridge-regression statistics is applied to tensor-product spline fields: each data holder computes a local Gram matrix and moment vector, and the merged...
545 A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure
2609.16606
cs.LG
John E. Darges, Laura Weidensager
Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a method based on families of weigh...
Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a method based on families of weighted ANOVA kernels that learn and adapt to this multivariable structure. TSKs parameterize the weights on each multivariable component of the target function by factors for each input. We propose learning these factors directly from function...
546 Information Geometric Self-Organization at the Edge of Stability in High-Capacity Kernel Associative Memories
2609.16827
cs.LG
Akira Tamamori
High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirical studies identified a hyperparameter regime, the "Ridge of Optimization," where attractor stability is maxim...
High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirical studies identified a hyperparameter regime, the "Ridge of Optimization," where attractor stability is maximized. However, the geometric nature of this regime and the optimization dynamics required to reach it have remained unclear. In this paper, we investigate the static geometry of the parameter space and the learning trajectory of Gradient De...
547 Beyond Quadratic Loss: The Stability Phase Diagram of Adam
2609.18314
cs.LG
Gaoxiang Tang, Huanran Chen, Ziming Liu
Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. ...
Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. We investigate this dependence by mapping training dynamics across the $(\beta_1,\beta_2)$ plane. Across a range of model--task settings, an approximately linear boundary, $1-\beta_2=C(1-\beta_1)$, separates spiky from non-spiky dynamics, w...
548 Higher-order pruning of experts in mixture-of-experts language models
2609.18916
cs.LGcs.AI
Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia, Stefano Soatto
Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert...
Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes a...
549 Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models
2609.19674
cs.LG
Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervent...
A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one. We show that these two failures require different structural remedies. Evolving a learned energy with a symplectic in...
550 Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
2609.20888
cs.LG
Themistoklis Haris, Henry Li, Maryam Karimzadehgan
Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this problem, but often drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ET...
Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this problem, but often drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that rivals dense model quality under hardware-aligned block-sparse decoding. ETA predicts dynamic, contextual thresholds directly from query representations, adjusting context retention depending o...
551 Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods
2609.21815
cs.LGcs.AI
Wenpeng Zhang, Runsheng Yu, Peilin Zhao
Adaptive optimization methods such as AdaGrad and Adam are widely used in modern deep neural network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware o...
Adaptive optimization methods such as AdaGrad and Adam are widely used in modern deep neural network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop an Online Mirror Descent framewo...
552 Explainable Predictive Condition-based Maintenance of Naval-Propulsion Systems using Fuzzy Logic
2609.24250
cs.LG
Dionisis Kalogeropoulos, Georgia Sovatzidi, Panagiotis G. Kalozoumis, Dimitris K. Iakovidis
The shipping industry has a significant impact on the global economy, emphasizing the need for operational availability and safety through the use of effective maintenance techniques. During the last decades, predictive maintenance (PdM) has emerged as a promi...
The shipping industry has a significant impact on the global economy, emphasizing the need for operational availability and safety through the use of effective maintenance techniques. During the last decades, predictive maintenance (PdM) has emerged as a promising solution compared to the existing conventional maintenance systems. This is because it offers several advantageous functions, such as damage predictions for vessel components, reduced downtime, improved and extended life of machinery, ...
553 Lifted Bellman Linear Programming for Offline Reinforcement Learning
2609.24489
cs.LGcs.AI
Hyukjun Yang, Jongchan Park, Narim Jeong, Donghwan Lee
Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions...
Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characteriz...
554 Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training
2609.25510
cs.LGcs.AI
Jacob Beck, Philip V. Ogren, Ari Kobren
Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating...
Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time training. We ask how much of this machinery is necessary. We introduce Hill Sampling, a form of hill-climbing optimization that repeatedly samples candidate programs from a frozen LLM, retains the best pro...
555 Geometry-Aware Hyperbolic Residual-Quantized Variational Autoencoders
2609.26342
cs.LGcs.AI
Alessio Colombo, Melika Ayoughi
Residual Vector Quantization turns continuous representations into discrete, multi-level token sequences. Yet most methods operate in Euclidean space, despite the coarse-to-fine structure of the resulting codes and the latent hierarchies present in many data d...
Residual Vector Quantization turns continuous representations into discrete, multi-level token sequences. Yet most methods operate in Euclidean space, despite the coarse-to-fine structure of the resulting codes and the latent hierarchies present in many data domains. Hyperbolic geometry offers a natural alternative for hierarchical representations, but naive hyperbolic extensions introduce geometric inconsistencies: non-associative hyperbolic addition prevents consistent residual aggregation, wh...
556 Hierarchical GNNs for power flow: letting physics shape the hierarchy
2609.26603
cs.LGcs.AI
Carmine Delle Femine, Leire Garin Atxaga, Asier Diaz-Iglesias, Juan Pablo Maroto Herrera, Ane Miren Florez-Tapia
Hierarchical latent communication improves the generalization of a power-flow model, shared across three grids, to new operating scenarios. The module exchanges information through two reduced graphs inside the corrective network of GENCO, replacing two of its...
Hierarchical latent communication improves the generalization of a power-flow model, shared across three grids, to new operating scenarios. The module exchanges information through two reduced graphs inside the corrective network of GENCO, replacing two of its local correction steps. We compare Kron-derived transports, a same-anchor Quotient construction and the flat GENCO Base architecture, all trained under one protocol of our own with about a hundred times fewer optimizer updates per grid tha...
557 What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates
2609.27679
cs.LG
Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang
A tabular foundation model must discover which distinctions matter for each new table without updating its parameters. We develop in-situ representation refinement: support labels guide changes to the episode's representations, improving the information availa...
A tabular foundation model must discover which distinctions matter for each new table without updating its parameters. We develop in-situ representation refinement: support labels guide changes to the episode's representations, improving the information available to later queries. A regularized leave-one-out objective yields a support correction and its query extension. The leading term separates attention-based reading from state-dependent scaling, motivating RefineICL: an attention-gated, FFN-...
558 NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers
2609.27735
cs.LG
Xiaohe Jiang (University of Exeter), Guoqiang Zhang (University of Exeter), Tianjin Huang (University of Exeter), Ronghui Mu (University of Exeter)
Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations....
Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations. We introduce Newton-Schulz Attention (NS-Attn.), a parameter-free transformation applied to the output of each attention head. Each head output is arranged as a feature-by-token matrix and normalized by its Frobenius norm. We then apply a ...
559 Transferable Evidence Reconstruction for Longitudinal Glucose Representations
2609.28199
cs.LG
Tian Zhou, Bingqing Peng, Linxiao Yang, Wenwei Wang, Mengni Ye
Long physiological recordings contain many routine measurements, while predictive information often lies in rare events, sustained burden, and recurring patterns. These properties can be computed as label-free evidence, but directly using them as features leav...
Long physiological recordings contain many routine measurements, while predictive information often lies in rare events, sustained burden, and recurring patterns. These properties can be computed as label-free evidence, but directly using them as features leaves limited labeled data to separate reproducible associations from sample-specific ones. Learning to reconstruct evidence can exploit unlabeled recordings, yet joint reconstruction does not explicitly require the decoding rule to transfer a...
560 Support-Compiled Feature Folding: More Evidence at Lower Memory Across Tabular Foundation Models
2609.28208
cs.LG
Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang
Wide tables offer tabular foundation models more evidence, but accessing it can exhaust their memory: full-width pairwise mixing grows quadratically with the number of columns, while feature selection makes inputs affordable by discarding evidence. We ask whet...
Wide tables offer tabular foundation models more evidence, but accessing it can exhaust their memory: full-width pairwise mixing grows quadratically with the number of columns, while feature selection makes inputs affordable by discarding evidence. We ask whether using more features requires interacting over all of them at once. We introduce Support-Compiled Feature Folding (SCFF), a training-free inference framework that encodes wide tables through bounded calls to a frozen backbone. SCFF organ...
561 Even Sharper Bounds for Transductive Learning and Its Applications
2609.28459
cs.LG
Yingzhen Yang
We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--...
We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses the modified log-Sobolev inequality for the swap walk and a two-parameter entropy closure. A peeling argument with a surrogate localization functional then gives excess-risk bounds with the same fixed...
562 Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles
2609.29974
cs.LG
Ali Haghpanah Jahromi, Mohammad Taheri, Zohreh Azimifar
Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Exp...
Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation pred...
563 Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think
2609.30036
cs.LG
Xvyuan Liu, Jianjie Fang, Chen Gao, Yong Li
Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require ac...
Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. With frozen LeWM models, intermediate targets substantially improve action synthesis and recorded-action ranking on Cube, PushT, Reacher, and TwoRoom. Learned targets and targets drawn from observed e...
564 Learning Operators by Regularized Stochastic Gradient Descent with Operator-valued Kernels
2504.18184
cs.LG
Jia-Qi Yang, Lei Shi
We consider a class of statistical inverse problems involving the estimation of a regression operator from a Polish space to a separable Hilbert space, where the target lies in a vector-valued reproducing kernel Hilbert space induced by an operator-valued kern...
We consider a class of statistical inverse problems involving the estimation of a regression operator from a Polish space to a separable Hilbert space, where the target lies in a vector-valued reproducing kernel Hilbert space induced by an operator-valued kernel. To address the associated ill-posedness, we analyze regularized stochastic gradient descent (SGD) algorithms in both online and finite-horizon settings. The former uses polynomially decaying step sizes and regularization parameters, whi...
565 AUWave: A Data-Driven Model for Reconstructing Significant Wave Heights Using Sparse Observations
2509.19384
cs.LGcs.AI
Hongyuan Shi, Yilin Zhai, Ping Dong, Zaijin You, Chao Zhan
Reconstructing high-resolution regional significant wave height (SWH) fields from sparse buoy observations is a critical challenge for ocean monitoring. We introduce AUWave, a hybrid deep learning framework that fuses a station-wise encoder with a multi-scale ...
Reconstructing high-resolution regional significant wave height (SWH) fields from sparse buoy observations is a critical challenge for ocean monitoring. We introduce AUWave, a hybrid deep learning framework that fuses a station-wise encoder with a multi-scale U-Net enhanced by self-attention to recover regional SWH fields. Trained and validated using NDBC buoy observations and ERA5 reanalysis over the Hawaii region, AUWave achieves high accuracy. It consistently outperforms a representative base...
566 What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
2509.22796
cs.LG
Xingyu Li (UC Riverside), Juefei Pu (UC Riverside), Yifan Wu (UC Riverside), Xiaochen Zou (UC Riverside), Shitong Zhu (UC Riverside)
Open-source software projects are foundational to modern software ecosystems, with the Linux kernel standing out as a critical exemplar due to its ubiquity and complexity. Although security patches are continuously integrated into the Linux mainline kernel, do...
Open-source software projects are foundational to modern software ecosystems, with the Linux kernel standing out as a critical exemplar due to its ubiquity and complexity. Although security patches are continuously integrated into the Linux mainline kernel, downstream maintainers often delay their adoption, creating windows of vulnerability. A key reason for this lag is the difficulty in identifying security-critical patches, particularly those addressing exploitable vulnerabilities such as out-...
567 Generative Modeling of Discrete Data Using Geometric Latent Subspaces
2601.21831
cs.LG
Daniel Gonzalez-Alvarado, Jonas Cassel, Stefania Petra, Christoph Schn\"orr
We propose a geometric latent-subspace framework for generative modeling of discrete data. Specifically, we introduce latent subspaces in the exponential parameter space of product manifolds of categorical distributions as a novel approach to learning low-dime...
We propose a geometric latent-subspace framework for generative modeling of discrete data. Specifically, we introduce latent subspaces in the exponential parameter space of product manifolds of categorical distributions as a novel approach to learning low-dimensional representations of high-dimensional discrete data. The resulting low-dimensional latent space captures statistical dependencies and removes redundant degrees of freedom among the categorical variables. We equip the parameter domain ...
568 Latent Generative Solvers for Generalizable Long-Term Physics Simulation
2602.11229
cs.LGcs.AI
Zituo Chen, Sili Deng
Reliable physics simulation demands two capabilities that today's neural PDE solvers do not deliver together: generalization across heterogeneous PDE families, and stability under long autoregressive rollouts. Deterministic operators accumulate error geometric...
Reliable physics simulation demands two capabilities that today's neural PDE solvers do not deliver together: generalization across heterogeneous PDE families, and stability under long autoregressive rollouts. Deterministic operators accumulate error geometrically, while existing probabilistic solvers are confined to a single PDE family or short horizons. We close this gap with the \textbf{Latent Generative Solver} (LGS), three coupled components: (i) a Physics VAE (PhyVAE) compressing twelve PD...
569 Minimax and Adaptive Covariance Matrix Estimation under Differential Privacy
2603.19703
cs.LG
T. Tony Cai, Yicheng Li
Estimating covariance matrices is fundamental to a wide range of statistical applications. This paper studies minimax and adaptive estimation of high-dimensional covariance matrices under $\rho$-zero-concentrated differential privacy ($\rho$-zCDP) over three n...
Estimating covariance matrices is fundamental to a wide range of statistical applications. This paper studies minimax and adaptive estimation of high-dimensional covariance matrices under $\rho$-zero-concentrated differential privacy ($\rho$-zCDP) over three nested classes: the pointwise-decay class $\mathcal{H}_\alpha$, the row-tail class $\mathcal{G}_\alpha$, and the separated-block class $\mathcal{F}_\alpha$. We consider both squared operator norm loss and normalized squared Frobenius norm lo...
570 On the Expressive Power of Transformers for Contextual Relations
2603.25860
cs.LG
Demi\'an Fraiman
Transformers have revolutionized machine learning by making attention a central mechanism for modeling interactions within a context. Despite the central role of attention, the theoretical capabilities of Transformers for representing contextual relations rema...
Transformers have revolutionized machine learning by making attention a central mechanism for modeling interactions within a context. Despite the central role of attention, the theoretical capabilities of Transformers for representing contextual relations remain unclear. In this work, we address this question by developing a mathematical framework based on probability and optimal transport. We view a text as a distribution of its representations and attention as a probabilistic relation between ...
571 On associative neural networks for sparse patterns with huge capacities
2603.26217
cs.LG
Matthias L\"owe, Franck Vermet
Generalized Hopfield models with higher-order or exponential interaction terms are known to have substantially larger storage capacities than the classical quadratic model. On the other hand, associative memories for sparse patterns, such as the Willshaw and A...
Generalized Hopfield models with higher-order or exponential interaction terms are known to have substantially larger storage capacities than the classical quadratic model. On the other hand, associative memories for sparse patterns, such as the Willshaw and Amari models, already exhibit enhanced storage capacities in the sparse regime. In this paper we combine these two mechanisms. We introduce higher-order versions of sparse associative memory models and study their storage capacities in the s...
572 A Sharp Norm Inequality and Buzano's Inequality via Determinants
2604.01525
cs.LG
Jose Antonio Lara Benitez
We give a short linear-algebraic proof of the inequality $$ \|x\|_1\,\|x\|_\infty \le \frac{1+\sqrt{n}}{2}\,\|x\|_2^2, $$ valid for every $x\in\mathbb{R}^n$. This inequality relates three fundamental norms on finite-dimensional spaces and has applications in o...
We give a short linear-algebraic proof of the inequality $$ \|x\|_1\,\|x\|_\infty \le \frac{1+\sqrt{n}}{2}\,\|x\|_2^2, $$ valid for every $x\in\mathbb{R}^n$. This inequality relates three fundamental norms on finite-dimensional spaces and has applications in optimization and numerical analysis. Our proof exploits the determinantal structure of a parametrized family of quadratic forms, and we show the constant $(1+\sqrt{n})/2$ is optimal. The inequality is a special case of Buzano's inequality, a...
573 Neural Parameter Estimation of RC Thermal Building Models for Model Predictive Control
2604.05904
cs.LG
Fabian Raisch, Timo Germann, Sang-Woo Ham, J. Nathan Kutz, Christoph Goebel
Gray-box RC models are widely used to enable energy-efficient model predictive control (MPC) in buildings. However, estimating RC parameters remains difficult, as conventional optimization-based algorithms are prone to local minima, rely heavily on good initia...
Gray-box RC models are widely used to enable energy-efficient model predictive control (MPC) in buildings. However, estimating RC parameters remains difficult, as conventional optimization-based algorithms are prone to local minima, rely heavily on good initial guesses, and incur high computational cost. To address these issues, we propose the Estimator from Scratch, a novel neural parameter estimation approach that embeds the physical equations into a neural network's training process to estima...
574 Identifying Causal Effects Using a Single Proxy Variable
2604.09135
cs.LG
Silvan Vollmer, Niklas Pfister, Sebastian Weichwald
Unobserved confounding is a key challenge when estimating causal effects from a treatment on an outcome. In this work, we assume that we observe a single, potentially multi-dimensional proxy variable of the unobserved confounder and that we know the mechanism ...
Unobserved confounding is a key challenge when estimating causal effects from a treatment on an outcome. In this work, we assume that we observe a single, potentially multi-dimensional proxy variable of the unobserved confounder and that we know the mechanism that generates the proxy from the confounder. Under an assumption called Single Proxy Identifiability of Causal Effects or simply SPICE, we prove that this error mechanism is complete and causal effects are identifiable. We extend the proxy...
575 Towards Interpretable Damage Detection based on Aerodynamic Pressure Measurements
2605.08187
cs.LG
Philip Franz, Max von Danwitz, Gregory Duth\'e, Alexander Popp, Eleni Chatzi
The increasing flexibility of modern large wind turbine blades necessitates cost-efficient and reliable structural monitoring solutions. For this purpose, we propose to use aerodynamic pressure measurements obtained via Aerosense, a novel, non-intrusive and ec...
The increasing flexibility of modern large wind turbine blades necessitates cost-efficient and reliable structural monitoring solutions. For this purpose, we propose to use aerodynamic pressure measurements obtained via Aerosense, a novel, non-intrusive and economical sensing system. In former work [Franz et al., 2025], we investigated the potential of aerodynamic pressure measurements for structural damage detection on elastic and aerodynamically loaded structures. An experimental campaign was ...
576 HiLiftAeroML: A High-Fidelity Computational Fluid Dynamics Dataset for High-Lift Aircraft Aerodynamics
2605.19565
cs.LG
Neil Ashton, Adam Clark, Konrad Goc, Liam Heidt, Christopher Ivey
HiLiftAeroML is, to our knowledge, the first open high-fidelity computational fluid dynamics dataset dedicated to high-lift aircraft aerodynamics. It contains 1,800 simulations spanning 180 variants of the NASA Common Research Model high-lift configuration and...
HiLiftAeroML is, to our knowledge, the first open high-fidelity computational fluid dynamics dataset dedicated to high-lift aircraft aerodynamics. It contains 1,800 simulations spanning 180 variants of the NASA Common Research Model high-lift configuration and ten angles of attack from $4^\circ$ to $22^\circ$. Each case was generated with a GPU-accelerated explicit wall-modeled large-eddy simulation approach on solution-adapted grids of 300--500 million cells, covering attached, separated, and p...
577 Critic Architecture Matters: Dual vs. Unified Critics for Humanoid Loco-Manipulation
2606.11891
cs.LG
Mehmet Turan Yard{\i}mc{\i}
Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within one policy. A natural design choice is between a single (unified) critic that estimates the combined value of all objectives and separate (dual) criti...
Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within one policy. A natural design choice is between a single (unified) critic that estimates the combined value of all objectives and separate (dual) critics with disjoint reward signals. We compare the two on the Unitree G1 humanoid in NVIDIA Isaac Lab. In the standing mode of a standardized evaluation, the dual-critic run reaches targets 3.5x faster (6.5 vs. 22.6 simulation steps), achieves...
578 Genetic Algorithms with Optimization Guided Operators
2606.12279
cs.LGcs.AI
Anna Brandenberger, Ilan Doron-Arad, Elchanan Mossel
Recent work in ML applies genetic algorithms at inference time to iteratively improve solutions to optimization problems. The basic mutation and recombination operators involved are qualitatively different from those studied classically. Mutations are no longe...
Recent work in ML applies genetic algorithms at inference time to iteratively improve solutions to optimization problems. The basic mutation and recombination operators involved are qualitatively different from those studied classically. Mutations are no longer random; an ML algorithm mutates a solution with the goal of improving an objective. Similarly, recombination is not based on random collages of parent solutions. Instead, it is an ML optimization-based operator whose goal is to synthesize...
579 Amortized quadrature for posterior expectations in inverse problems
2606.15871
cs.LG
Ali Siahkoohi
Uncertainty in the solution of an inverse problem and in the tasks performed on it is quantified by posterior expectations, each an average of an integrand over $M$ posterior samples. While designed quadratures improve on the $O(M^{-1/2})$ error of Monte-Carlo...
Uncertainty in the solution of an inverse problem and in the tasks performed on it is quantified by posterior expectations, each an average of an integrand over $M$ posterior samples. While designed quadratures improve on the $O(M^{-1/2})$ error of Monte-Carlo estimation, they solve an optimization problem, often against the posterior density, for every new observation, which can be computationally costly. To address this limitation, we introduce the quadrature field, a set-equivariant network t...
580 NAC: Neural Action Codec for Vision-Language-Action Models
2606.21372
cs.LG
Ahad Jawaid, Yu Xiang
Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this de...
Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this design through the lens of neural audio codecs - convolutional encoder-decoder architectures with residual vector quantization that serve as the standard front end for audio foundation models. Motivated by their success, we introduce the Neur...
581 ConSolv: Solvent-Conditional Machine Learning Implicit Solvent Potential
2606.24983
cs.LG
Linying Zhang, Julija Zavadlav
Implicit solvent machine learning potentials (MLPs) offer a powerful route to bridging the gap between accuracy and efficiency in molecular simulations. However, existing models have largely focused on aqueous environments, overlooking the diverse and importan...
Implicit solvent machine learning potentials (MLPs) offer a powerful route to bridging the gap between accuracy and efficiency in molecular simulations. However, existing models have largely focused on aqueous environments, overlooking the diverse and important roles of non-aqueous solvents in areas such as organic synthesis and battery technology. Here, we present ConSolv, a solvent-conditional MLP architecture that explicitly incorporates solvent effects on solute interactions through an atten...
582 Statistically Valid Post-Training Hyperparameter Selection: From Tuning to Guarantees
2606.25601
cs.LG
Amirmohammad Farzaneh, Osvaldo Simeone
Post-training hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom of pre-trained models such as inference-time parameters, implementation-level settings, and thresho...
Post-training hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom of pre-trained models such as inference-time parameters, implementation-level settings, and thresholds driving decision rules. Despite its practical importance, hyperparameter selection is typically performed using best-effort empirical methods such as grid search or Bayesian optimization, which provide no formal statistical guarantees o...
583 Bringing Agentic Search to Earth Observation Data Discovery
2607.02387
cs.LG
Minghan Yu, Youran Sun, Chugang Yi, Yixin Wen, Haizhao Yang
NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search framework for geoscience data dis...
NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search framework for geoscience data discovery that takes a natural-language research query and returns matching datasets and tools. We demonstrate that, in the era of large language models, the latent value of knowledge graphs (KGs) can be substantially amplified through agentic...
584 Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics
2607.09250
cs.LG
Hugo Cui
The impact of a given training point on a statistical model can be measured through its leave-one-out influence on the model parameters, which quantifies how its removal from the training set affects the learned weights. For convex M-estimation under Gaussian ...
The impact of a given training point on a statistical model can be measured through its leave-one-out influence on the model parameters, which quantifies how its removal from the training set affects the learned weights. For convex M-estimation under Gaussian design, in the high-dimensional limit $n\asymp d$, we show that the empirical distribution of influences across training points concentrates around a deterministic measure which we sharply characterize. This characterization suggests that i...
585 Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks
2607.26574
cs.LGcs.AI
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered ins...
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap. The standard fix is a preprocessor that recovers image content and decodes the encoding before the guard. We build one and evaluate it against an ensemb...
586 Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
2607.26865
cs.LGcs.AI
Amirmohammad Farzaneh, Osvaldo Simeone
LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning bu...
LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remaining reliable and deferring to a cloud-side model only when local uncertainty is too high to act safely. We propose Think Short, Defer Smart (TSDS), a framework that synergistically integrates a lightweight convergence probe...
587 On the robustness of noisy solutions in non-convex neural networks
2607.27000
cs.LG
Enrico M. Malatesta, Alessandra Passalacqua, Riccardo Zecchina
Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite bein...
Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare. At zero temperature this picture has been formalized in binary perceptrons through the overlap gap property (OGP), which limits algorithmic access to configurations with zero training error above a critical constraint den...
588 DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing
2608.02769
cs.LG
Sagnik Nandy, Samriddha Lahiry, Pragya Sur, Subhabrata Sen
Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fail...
Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fails to exploit cross-modal dependence. Existing approaches rely on pre-specified fusion architectures, from early to late fusion, that may not adapt to the underlying dependence structure among modalities. We propose DAIF, a data adaptive int...
589 LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models
2608.16340
cs.LG
Tom Splittgerber, Niklas Koenen, Marvin N. Wright, Werner Brannath
The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs....
The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs. In order to preserve interpretability, it is usually necessary to restrict the NN components to prevent them from dominating the model. However, existing methods that enforce structural constraints on their NN components severely limit the...
590 Provable Quantum-Classical Separation for Continuous Gibbs Sampling
2608.24527
cs.LG
Enrico Olivucci, Mariia Sobchuk, Sehmimul Hoque, Jeffrey Hnybida, Kyungho W. Kim
We prove the first quantum-classical separation for a sampling problem over a continuous domain. For a class of Gibbs states $p\propto e^{-\beta E}$ on the torus $\mathbb{T}^d$ with smooth ($s$-Gevrey) potential and barrier amplitude $\alpha=e^{\beta\Delta}$, ...
We prove the first quantum-classical separation for a sampling problem over a continuous domain. For a class of Gibbs states $p\propto e^{-\beta E}$ on the torus $\mathbb{T}^d$ with smooth ($s$-Gevrey) potential and barrier amplitude $\alpha=e^{\beta\Delta}$, where $\Delta = \max E-\min E$, every classical algorithm querying the value, gradient, or any higher-order derivatives of the log-density requires $\Omega(\alpha)$ queries to sample at constant accuracy in total variation distance, while a...
591 Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
2608.26088
cs.LGcs.AI
Evelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun
Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented dat...
Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model selection. We present the Planetary Prediction Engine (PPE), an autonomous AI system that executes this end-to-end workflow directl...
592 High-probability guarantees for linear accessibility in feature superposition
2609.09556
cs.LGcs.AI
Enrico Vompa
Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, w...
Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise, proving the sufficient dimension scales linearly ($d=O_{\varepsilon}(k \log m)$) rather than prior worst-case quadratic limits. We characterize the asymmetry betwe...
593 Graph Matching Relaxations and Amortization for Supervised Graph Prediction
2609.15437
cs.LG
Federico M\'endez, Paul Krzakala, Gabriel Melo, Charlotte Laclau, R\'emi Flamary
End-to-end Supervised Graph Prediction (SGP) requires a permutation-invariant loss to compare predicted and target graphs with arbitrary node orderings. Such losses typically involve a costly graph-matching problem. We first study three Optimal Transport relax...
End-to-end Supervised Graph Prediction (SGP) requires a permutation-invariant loss to compare predicted and target graphs with arbitrary node orderings. Such losses typically involve a costly graph-matching problem. We first study three Optimal Transport relaxations of this problem and show, theoretically and empirically, that the Gromov-Wasserstein (GW) objective is the most suitable for SGP. Then, to avoid solving the resulting inner optimization for every training example, we propose to amort...
594 Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing
2609.24791
cs.LG
Nibraas Khan, Abigale Plunk, John Staubitz, Ingrid Shragge, Jordan Brooks
Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, voc...
Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, vocalization, and autonomic arousal. Its signs are subtle and individualized, and its autonomic components are invisible without instrumentation. We collected upper-body movement from inertial measurement units, physiology from a wrist-worn de...
595 Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis
2609.25576
cs.LG
Jun Li, Yanlong Guo, Zhaozhao Zeng
We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ st...
We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ storage and are impractical at $N\sim 10^6$--$10^8$. We propose DeepMVSA, which re-expresses the minimum-volume principle in neural implicit form: a lightweight coordinate network generates the mixing weights and a triangular LU-type paramete...
596 Reinforcement Learning with Decomposed Subtasks
2609.27035
cs.LGcs.AI
Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct sk...
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a be...
cs.MM 1 papers
775 IMEX-FND: A Traceable Interaction-Aware Mixture-of-Experts Framework for Multimodal Fake News Detection
2609.29610
cs.MM
Yuchen Miao, Zijun Wang, Ke Liu, Peixuan Wang, Chang Han
Multimodal fake news detection (FND) increasingly demands verdicts that are not only accurate but traceable, revealing how cross-modal evidence is combined, yet two coupled difficulties remain. First, text-image relations are heterogeneous: uniqueness, redunda...
Multimodal fake news detection (FND) increasingly demands verdicts that are not only accurate but traceable, revealing how cross-modal evidence is combined, yet two coupled difficulties remain. First, text-image relations are heterogeneous: uniqueness, redundancy, and synergy coexist and vary from post to post, so a single global fusion rule is brittle and opaque. Second, the dominant modality shifts across instances, which static encoders and a fixed fusion pathway handle poorly. We present IME...
cs.SD 14 papers
757 CLEAR: Online Speech Content Leakage Estimation through Cross-ASR Disagreement
2609.30415
cs.SDeess.AS
Bhawana Chhaglani, Tanvi Kandepuneni, Jeremy Gummeson, Prashant Shenoy
Signal-level speech privacy mechanisms suppress linguistic content while preserving acoustic information needed by downstream sensing applications. However, their privacy settings are typically evaluated/selected offline and remain fixed during deployment, eve...
Signal-level speech privacy mechanisms suppress linguistic content while preserving acoustic information needed by downstream sensing applications. However, their privacy settings are typically evaluated/selected offline and remain fixed during deployment, even though speech-content leakage can vary substantially across utterances and speakers. Adapting privacy protection at runtime requires estimating how much speech remains recoverable, but conventional measures such as WER or PER require grou...
758 MuseTimbre: Zero-Shot Timbre Transfer by Controlling a Frozen Music Generator
2609.30548
cs.SDeess.AS
Yuan-Chiao Cheng, Zhiyao Duan
Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicat...
Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicated model for the task, which captures the timbre cleanly but stays a narrow, single-purpose system. More versatile approaches add control to a pretrained music generator, yet a reference clip entangles timbre with genre and melody, so these...
759 Training-Free Contextual ASR via SpeechLLM-Based Error-Aware Selective Retrieval
2609.30694
cs.SDeess.AS
Natsuo Yamashita, Ai Nemoto, Ryosuke Koichi, Masaaki Yamamoto
Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing t...
Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issue by selecting candidate terms from an external dictionary, but querying many recognized words requires numerous dictionary lookups and may yield poorly targeted candidates. We pro...
760 Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information
2609.30774
cs.SDeess.AS
Shuhan Zhang, Wenxuan Wu, Haizhou Li
In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures w...
In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-pa...
761 Tracing and Relearning Detection Evidence in Text-to-Speech Systems
2609.30983
cs.SDeess.AS
Eunji Shin, Kyudan Jung, Jihwan Kim, Minwoo Lee, Jaegul Choo
Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-Big...
Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acousti...
762 TinyAudio: Compact and Efficient Text-to-Audio Generation for Low-Resource Deployment
2609.31525
cs.SDeess.AS
Junxi Liu, Xiquan Li, Wenhao Guan, Yifan Duan, Zhikang Niu
Text-to-audio (TTA) generation has advanced rapidly in generation quality and instruction following. However, representative systems often require around a billion parameters, limiting deployment on resource-constrained devices. This paper introduces TinyAudio...
Text-to-audio (TTA) generation has advanced rapidly in generation quality and instruction following. However, representative systems often require around a billion parameters, limiting deployment on resource-constrained devices. This paper introduces TinyAudio, a compact flow-matching-based TTA model for low-resource deployment. At its core, TinyAudio uses TA-DiT, a 35M single-stream flow-matching Transformer. TinyAudio also includes TA-CLAP, a 32M audio-aligned text encoder, and TA-VAE, whose 2...
763 Nearest but Not Dearest: Shared Curator-Feedback Infrastructure for Content-Only Search and Recommendation
2609.30568
cs.SD
Matt Sandler
A deployed B2B music-discovery platform serves both query-driven search (text prompts, vibe tags) and seed-driven recommendation (seed-track and artist stations) over one licensed catalog, one LAION-CLAP joint audio-text embedding space, one candidate-generati...
A deployed B2B music-discovery platform serves both query-driven search (text prompts, vibe tags) and seed-driven recommendation (seed-track and artist stations) over one licensed catalog, one LAION-CLAP joint audio-text embedding space, one candidate-generation filter, and one ranking head -- and neither path consumes end-listener behavioral signal. In this content-only regime, curator judgment is the principal feedback signal available, and offline cosine similarity predicts it poorly: 38% of ...
764 Music Source Separation via Stem Discovery
2609.30912
cs.SDeess.AS
V. Valtteri Kallinen, Eloi Moliner, Lauri Juvela, Vesa V\"alim\"aki
Music source separation (MSS) methods aim to extract stems from music mixtures, which is important, for example, in karaoke, music remixing, and pedagogical applications. While earlier research on MSS systems has been dominated by models targeting narrow sets ...
Music source separation (MSS) methods aim to extract stems from music mixtures, which is important, for example, in karaoke, music remixing, and pedagogical applications. While earlier research on MSS systems has been dominated by models targeting narrow sets of general stems, there have recently been attempts to support broader source definitions. One of these methods uses audio queries to provide direct and descriptive control over the desired separation targets based on the sound itself. Howe...
765 Room Impulse Response Embeddings for Speech Enhancement in Noisy and Reverberant Environments
2609.31041
cs.SDeess.AS
Adrian Meise, Reinhold Haeb-Umbach
We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student a...
We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's embeddings when given a noisy version of the reverberant input. We assess their representational capabilities by estimating acoustic room parameters from them. Conditioning a disc...
766 PANEL: An Open-Source, Self-Hosted Web Platform for Human Evaluation of Generative Models
2609.31392
cs.SD
Matteo Spanio, Andrea Poltronieri, Mart\'{\i}n Rocamora
Human judgement is the reference measure for evaluating generative models, yet the software used to collect it lags behing the methodology. Researchers adapt listening-test frameworks designed for perceptual protocols such as MUSHRA, rely on closed commercial ...
Human judgement is the reference measure for evaluating generative models, yet the software used to collect it lags behing the methodology. Researchers adapt listening-test frameworks designed for perceptual protocols such as MUSHRA, rely on closed commercial survey platforms, or implement single-use web applications. Live arenas such as Chatbot Arena and Music Arena rank publicly deployed systems at scale, but do not support controlled comparisons of a laboratory's own models with its own parti...
767 SAID: Semantic Acoustic Imaging Detector for Sound Event Localization and Detection
2609.31492
cs.SDeess.AS
Runbang Wang, Zining Liang, Yin Cao, Qiuqiang Kong
In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering $360^{\circ}$ horizontally and $180...
In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering $360^{\circ}$ horizontally and $180^{\circ}$ vertically. The map shows the directions occupied by the source as a region and the sound energy within that region. A class label identifies the sound. Predicting these labeled acoustic maps from audio is called semantic acoustic...
768 RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue
2609.31588
cs.SDeess.AS
Sathvik Udupa, Naveen Kumar, Ryan Folmsbee
Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorde...
Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to predefined responses at the cost of additional latency. We propose RePlay, a spoken dialogue system adapted from PersonaPlex that handles multi-turn conversations by retrieving and...
769 Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders
2609.29780
cs.SD
Kyudan Jung, Sehyun Lee, Song-ha Jo, Jaegul Choo, Sanghyuk Choi
Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream ada...
Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best obse...
770 SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering
2603.27342
cs.SDeess.AS
Yhonatan Gayer, Boaz Rafaely
Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone ...
Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone array would capture in it, and processes those signals in the SH domain. We present SHroom (Spherical Harmonics ROOM), an open-source Python library that performs this whole workflow in one package, from room simulation to binaural renderin...
eess.AS 4 papers
771 Coupled Meta-Adaptive Filtering for Active Noise Control Under Time-Varying Acoustic Paths
2609.30945
eess.AS
Boxiang Wang, Zhengding Luo, Ziyi Yang, Dongyuan Shi, Xuexian Liu
Meta-adaptive filtering (Meta-AF) provides a data-driven alternative to hand-crafted adaptive filter updates by employing a learned optimizer throughout online adaptation. However, when Meta-AF is used for active noise control (ANC), time-varying acoustic path...
Meta-adaptive filtering (Meta-AF) provides a data-driven alternative to hand-crafted adaptive filter updates by employing a learned optimizer throughout online adaptation. However, when Meta-AF is used for active noise control (ANC), time-varying acoustic paths remain a major challenge. In particular, the physics-informed optimizer features are constructed using a secondary path estimate and become mismatched when the physical path changes, leading to inaccurate filter updates and degraded noise...
772 Assessing a Mathematical Model of Syllable Production via CTW Alignment with EMA Data
2609.31508
eess.AS
Fr\'ed\'eric Berthommier
This study evaluates a mathematical model of syllable production using an EMA dataset from a recently published study, which proved highly compatible with the model's architecture. The data set consists of regularly structured French phrases of equal duration ...
This study evaluates a mathematical model of syllable production using an EMA dataset from a recently published study, which proved highly compatible with the model's architecture. The data set consists of regularly structured French phrases of equal duration and reduced phonetic and syllabic complexity. A dedicated procedure was used to transform EMA recordings into Maeda parameters. The Model-generated trajectories were then realigned with these transformed data using Canonical Time Warping (C...
773 HAMMER: Harmonic-Aware Parallel Context Modeling and Discriminator-Free Perceptual Optimization for Speech Enhancement
2609.21171
eess.AS
Shang-Fu Chen, Szu-Wei Fu, Sung-Feng Huang, Rong Chao, Wen-Huang Cheng
Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers and do not explicitly exploit harmonic periodicity, a strong cue for preserving ...
Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers and do not explicitly exploit harmonic periodicity, a strong cue for preserving voiced speech under noise. Perceptual optimization poses another challenge. PESQ is non-differentiable, so many methods train auxiliary metric discriminators that increase complexity and introduce adversarial instability. We propose \ours, ...
774 DAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation
2609.29749
eess.AS
Wen Wen, Qiang Zhou, Yu Xi, Haoyu Li, Bohan Li
Although room impulse responses (RIRs) encode source-distance cues, conventional monaural source separation focuses on recovering audio content without estimating source-specific RIRs, losing the associated spatial information. To address this limitation, we p...
Although room impulse responses (RIRs) encode source-distance cues, conventional monaural source separation focuses on recovering audio content without estimating source-specific RIRs, losing the associated spatial information. To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture. DAMSEP inte...