| # | Title | Categories | Authors | Abstract |
|---|---|---|---|---|
| cs.AI 377 papers | ||||
| 1581 |
SGAnalog: An End-to-End Circuit Benchmark from Open-Source Silicon Tapeouts
2610.03934
|
cs.AI
|
Yueting Li, Weihang Ding |
Existing analog integrated circuit design benchmarks make two questions hard to answer: whether a model has learned transferable circuit skills rather than recalled familiar examples, and whether its output works under defined process and test conditions. We i...Existing analog integrated circuit design benchmarks make two questions hard to answer: whether a model has learned transferable circuit skills rather than recalled familiar examples, and whether its output works under defined process and test conditions. We introduce a benchmark built from human-designed, open-source circuits associated with Tiny Tapeout manufacturing shuttles. The collection contains 273 topologically distinct top-level designs. Every source is retrieved at the revision recorded for its shuttle submission and processed in a fixed containerized environment. The pipeline exports each eligible schematic image and its SPICE netlist from the same source file, giving transcription an exact structural reference. Commit dates support model-specific training-cutoff analysis, while author testbenches provide the simulation context for sizing. The benchmark evaluates schematic-to-netlist transcription and device sizing. Across seven models on a fixed set of 66 transcription tasks, the strongest model reaches 56.1% exact graph isomorphism, and six of seven models drop sharply from the small to the medium tier. For one frontier model, removing author-chosen labels reduces exact matches while preserving aggregate structural F1, suggesting that labels can aid connectivity tracing. On the 17 sizing tasks, the leading model converges on all 17 proposals and reaches 91.2 out of 100 against the human reference, while the two newest Claude models refuse 4 and 11 of the same prompts they transcribe without objection; a proposal without sizes scores zero. The two tasks produce different model rankings, exposing distinct visual and design capabilities and, in one family, a policy rather than capability limit.
|
| 1582 |
MLLMs Fail to Refuse when Using Tools Agentically
2610.03938
|
cs.AI
|
Rikiya Takehi, Ryo Hachiuma, Shaona Ghosh, Dan Zhao, Yu-Chiang Frank Wang |
Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-us...Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our experiments confirm that, across three popular safety benchmarks, all the top open- and closed-weight MLLMs we test exhibit significantly lower safety in tool-using settings than in non-tool settings, with a relative refusal failure rate increase of up to 68.7%. Based on analysis of 100,000+ responses, including extended experiments, we also propose two possible reasons for this safety degradation.
|
| 1583 |
Retrieval-Augmented Large Language Model Decision-Making for Autonomous Driving Guided by Chinese Philosophical Wisdom
2610.03948
|
cs.AI
|
Xiaojun Bi, Xiaoyuan Ma, Yiwen Sun, Tianren Huang, Chaoran Liu |
Autonomous driving decision systems must balance safety, efficiency, and social norms in complex traffic interactions. Philosophical and ethical considerations have received limited attention in existing autonomous driving decision-making approaches based on n...Autonomous driving decision systems must balance safety, efficiency, and social norms in complex traffic interactions. Philosophical and ethical considerations have received limited attention in existing autonomous driving decision-making approaches based on numerical optimization, sequence prediction, and large language models (LLMs). We propose Chinese Philosophical Wisdom-Guided Driving (CPW-Drive), a closed-loop retrieval-augmented generation (RAG) framework that incorporates value guidance derived from Chinese philosophy into autonomous driving decision-making. Using Chinese Confucian thought as its knowledge source, CPW-Drive consolidates LLM-extracted keywords from relevant classical texts into driving-relevant value principles through manual screening and validation. It then contextualizes these principles through scenario-specific cases to form retrievable and reusable value guidance. We further propose Physics-aware Spatial Similarity Retrieval (PSSR), which compares vehicle layouts and velocity-extrapolated states to retrieve physically relevant historical cases. On Highway-env's multilane highway-driving task, CPW-Drive achieves success rates of 93.0%, 86.0%, and 72.0% across three traffic configurations. These results outperform the strongest baseline by 8.0, 22.5, and 25.0 percentage points, respectively. Across all configurations, CPW-Drive achieves the highest collision-free step count and maintains a low lane-change frequency. The results suggest that structured value guidance can improve simulated closed-loop safety and stability while introducing efficiency and latency trade-offs.
|
| 1584 |
LatentQuant: Preserving the Policy-Facing Latent Contract under NVFP4 VAE Quantization
2610.03959
|
cs.AI
|
Ziye Deng, Lufang Chen, Shuyu Feng, Zhenwei Duan, Zicong Ye |
Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the ...Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quantization-aware training (QAT) can recover reconstruction by moving this representation. On Wan2.1, joint QAT nearly matches FP32 VBench-7 (0.7403 versus 0.7409), yet LIBERO success collapses from 95.5% to 10.5%. Controlled decoder-only experiments show that activation quantize-dequantize operations alter the reconstruction signal and decoder Jacobian, redirecting the gradient returned to the encoder and inducing persistent latent drift. Based on this mechanism, we introduce LatentQuant, a two-stage NVFP4 QAT framework that first aligns the quantized encoder with its high-precision counterpart, then freezes it while adapting the decoder. Across Wan2.1 and Wan2.2, LatentQuant preserves near-baseline control and high reconstruction quality, achieving 95.75% success on LIBERO and 68.8% on RoboTwin. On NVIDIA B300 GPUs, NVFP4 execution achieves 1.17x-1.26x end-to-end VAE speedups over BF16 cuDNN.
|
| 1585 |
ROAR: Unifying Runs across Heterogeneous AI-Driven Research Systems
2610.03966
|
cs.AI
|
Leo Y. Lin, Vishakha Ramani, Z. Berkay Celik, Paul Castro, Marquita Ellis |
Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragm...Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragmented: teams operate in isolation, ADRS frameworks emit results in different formats, and no shared infrastructure exists to aggregate or compare runs across problems and systems. We present ROAR, a solution for systematically unifying and analyzing heterogeneous ADRS outputs. ROAR addresses two challenges: reconciling heterogeneous ADRS outputs and enabling analytics across runs with different objectives and scoring functions. We achieve this through a relational schema and parsing layer that normalize heterogeneous ADRS outputs while preserving data lineage and temporal structure, and accommodating new systems without requiring schema modifications. From building a corpus of more than 900 runs from multiple ADRS, we show how pooled data can reveal properties of problem landscapes that are difficult to observe. Consistent with prior work, runs with identical configurations may converge to different scores. We find that many runs realize most gains early, and that the effectiveness of different strategies for incorporating prior solutions into the search process varies across problems. We further show that the pooled corpus is actionable and not merely analytical by using ROAR to configure ADRS runs. Together, these results illustrate how pooled ADRS data can expose problem-dependent structure in search behavior that is difficult to detect from any single system, team, or benchmark. Such cross-cutting insights are difficult to obtain while runs remain siloed; ROAR is the first infrastructure designed to unify them.
|
| 1586 |
Exploration-Preserving Policy Optimization
2610.04011
|
cs.AI
|
Hangzhan jin, Mohammad Hamdaqa, Doina Precup |
Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making...Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping rule that redistributes credit using prompt-relative, length-normalized response surprisal and prompt pass rate. ExPPO combines bounded shaping with shared normalization to preserve verifier polarity and approximately maintain each prompt group's total absolute sequence-advantage mass. Our analysis characterizes response-level credit allocation alongside sampled mode updates, deriving local conditions for gains in entropy and correct-mode discovery. Experiments show improved in-domain and out-of-domain reasoning coverage, higher aggregate response accuracy, and strong coverage at large sampling budgets. A controlled multi-answer evaluation further demonstrates increased correct-mode yield and gains in diversity among verified-correct responses. Code is available at https://github.com/jinhangzhan/ExPPO
|
| 1587 |
Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models
2610.04012
|
cs.AI
|
A. Bochkov |
Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which indepen...Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which independently constructed document states and deterministic executable skills contribute typed proposals to a shared language-model state. An explicit residual operator reconciles proposals attached to common interface nodes. We evaluate this organization through controlled experiments and negative results rather than claiming a physical finite-element formulation of language. An attention-free Multi-Mesh prototype learns causal language modeling but does not establish competitive general capability. A versioned store contains 52,809 reconstructive memory elements near a 1.7-billion-floating-value budget; reconstruction is incomplete, with approximately 75\% token accuracy. Support-aware lexical indices make these elements addressable under provenance-controlled query construction. For executable arithmetic, positional result observations substantially improve neural rendering relative to a repeated global result vector, and output substitutions change the model's preferred answer. A bounded attachment demonstration further measures the effect of making selected evidence available, without establishing the utility of loading an entire multi-billion-value store. The results support a separation of storage, execution, and neural coordination, while identifying unresolved limitations in question-only retrieval, unrestricted answer generation, and end-to-end efficiency.
|
| 1588 |
A Quantitative Analysis of Graph Representation Strategies for Cyber Attack Detection
2610.04019
|
cs.AI
|
Ali Melih Kanca, Ilker Turker |
Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed ...Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed across these application domains. This study presents a quantitative analysis of 37 original studies published between 2019 and 2026. Each study was coded according to publication year, application domain, graph representation type, feature extraction strategy, learning paradigm, algorithm, and dataset. Frequency analysis, cross tabulation, and statistical association tests were applied. Automated representation learning was the most frequently employed strategy, accounting for 73.0% of the studies, while handcrafted representation strategies accounted for 27.0%. The Fisher Freeman Halton exact test revealed a statistically significant association between application domain and representation strategy (exact p = 0.008; Cramer's V = 0.540). The findings indicate that automated representation learning predominates within the analyzed corpus and that the distribution of representation strategies differs across cybersecurity application domains.
|
| 1589 |
The Reported Engagement with AI Level (REAL) Rating: A Framework for Disclosing Human-AI Collaboration
2610.04021
|
cs.AI
|
Im\`ene Goumiri, Mayleen Cortez-Rodriguez, Eric Bell, Amanda Muyskens |
Generative artificial intelligence has become increasingly incorporated into digital media and more generally into production workflows with which the public frequently interacts. Current provenance standards and disclosure methods frequently rely on binary ca...Generative artificial intelligence has become increasingly incorporated into digital media and more generally into production workflows with which the public frequently interacts. Current provenance standards and disclosure methods frequently rely on binary categorizations, differentiating only between entirely human-authored and AI-generated content, or depend on technical watermarking that lacks user-facing clarity. However, there are very different risks and outcomes depending on the different uses of AI in the creation and consumption of end products. Therefore, we propose the Reported Engagement with AI Level (REAL) Rating framework which addresses these concerns by introducing a six-tier scale (Levels 0-5) that discloses the extent of AI involvement in both product creation and user experience. This paper details the structural methodology of the REAL Rating system and the specifics of its application across text, image, audio, video, and software products.
|
| 1590 |
Agent Policy-Value Audit: Separating Transition Composition from Event Selection in Financial LLM Agents
2610.04040
|
cs.AI
|
Mingyang (Alex), Chen (Andrew), Yida (Andrew), Xu (Aurora), Huiwen (Aurora) |
Financial LLM agents are often evaluated by comparing their end-to-end returns with those of a baseline and testing the paired difference against zero. This measures whether deploying the agent changes realized performance, but it does not isolate event-select...Financial LLM agents are often evaluated by comparing their end-to-end returns with those of a baseline and testing the paired difference against zero. This measures whether deploying the agent changes realized performance, but it does not isolate event-selection skill. An agent that frequently changes positions from flat to long can earn a positive paired return from an upward-drifting event pool even when it selects events at random. We propose the Agent Policy-Value Audit, which holds fixed the observed count of each ordered action-change type and randomly reassigns them across eligible events. The average payoff from these reassignments is the composition benchmark; the difference between observed deployment value and this benchmark is selection value. In semi-synthetic benchmarks based on real earnings-event returns, a zero-centered paired test falsely attributes passive exposure to selection skill in $11.6\%$ of no-skill replications, while the transition-matched audit reduces this rate to $5.3\%$. Applied retrospectively to 723 earnings events at 44 U.S. consumer-facing firms, the audit decomposes the agent's gross deployment value of $+15.2$ bps/event into a $+25.8$ composition benchmark and a $-10.6$ selection value. The agent does not detectably outperform matched random assignments. Financial-agent evaluations should report deployment value separately from event-selection value.
|
| 1591 |
The Cost of a Hop: Benchmarking NLIP and A2A
2610.04053
|
cs.AI
|
Ranjan Sinha, Anindita Das, Ashika Anand Babu, Hari Palleti |
Autonomous agents built on Large Language Models (LLMs) need standardized protocols to interoperate across systems. Several now exist (A2A, MCP, ACP, ANP, NLIP), but the Natural Language Interaction Protocol (NLIP) has not appeared in any controlled performanc...Autonomous agents built on Large Language Models (LLMs) need standardized protocols to interoperate across systems. Several now exist (A2A, MCP, ACP, ANP, NLIP), but the Natural Language Interaction Protocol (NLIP) has not appeared in any controlled performance study, and no work has measured where an agent protocol's latency is spent. We compare NLIP and the Agent-to-Agent (A2A) protocol empirically, decomposing latency into message creation, connection, and send phases across three independent hardware environments. For lightweight coordination, NLIP is 8.4-9.6x faster than the baseline A2A SDK implementation on two environments and about 4x on a third; the direction of the advantage is consistent, its magnitude depends on the hardware. The advantage is stage-specific: for the end-to-end pipeline, where LLM inference dominates, the protocols are near parity. The difference comes almost entirely from connection setup. To test A2A at its best, we also ran A2A SDK with connection caching enabled; caching narrows its gap with NLIP by a hardware-dependent amount, from 2.75x on one machine to near-parity on faster hardware, where at scale a cache-optimized A2A-SDK matches NLIP. We report these as measured conditions without a single causal account of the residual send-phase cost. Against the more optimized Python-A2A, NLIP leads by about 4x on the same stage. We close with a protocol-selection guide keyed to workload characteristics.
|
| 1592 |
Discrete Diffusion for Large Graph Generation via Structural Candidate Restriction
2610.04056
|
cs.AI
|
Yassin Mohamadi, Zeno Geradts, Marcel Worring |
Synthesizing realistic graphs at scale is vital when the graphs of interest are large and real-world samples are limited or access-sensitive. Diffusion-based generators have recently driven much of the progress, offering high modeling capacity, but most such m...Synthesizing realistic graphs at scale is vital when the graphs of interest are large and real-world samples are limited or access-sensitive. Diffusion-based generators have recently driven much of the progress, offering high modeling capacity, but most such methods have quadratic computational complexity and are hence restricted to small-scale networks, currently up to 3k nodes. Existing non-quadratic methods remain limited by memorization issues and a trade-off between scalability and generation quality. Our goal is to generate large graphs whose structural statistics --- e.g., degree distribution, clustering, and path length --- faithfully reflect those of real-world sparse graphs, without resorting to memorizing the training data. We introduce a discrete graph diffusion model that restricts training to a structurally motivated subset of node pairs --- observed edges and their wedge non-edges --- reducing training complexity below quadratic in the number of nodes. To keep the noisy graph informative throughout both the forward and reverse trajectories, we design a three-class absorbing forward process governed by a \emph{degree-aware, floored} cosine noise schedule: unlike a structure-blind schedule that adds noise to every pair identically, ours adapts to each node's degree and never fully erases the graph's structure, keeping the noisy graph informative at every step. Experiments across diverse datasets show that our model consistently ranks among the top methods for structural fidelity against existing discrete diffusion baselines in large graph generation.
|
| 1593 |
Reinforcement Learning with Comparative Evidence for Social Intelligence
2610.04072
|
cs.AI
|
Keane Ong, Yuriel Ryan, Sabri Boughorbel, Vladimir Necula, Jack Wei Lun Shi |
Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but s...Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreover, core social targets such as affect, intent, preference, and pragmatic meaning are often ambiguous. The same behavior can support multiple plausible interpretations, making it difficult to verify which is best supported. To address this challenge, we introduce Reinforcement Learning with Comparative Evidence (RLCE), a reinforcement learning method that learns social understanding from unlabeled training data without constructing rewards from ground-truth annotations. Given distinct answers in a rollout group, RLCE constructs evidence tests that identify observable evidence favoring an answer over another, validates these tests against the input sample, and aggregates test outcomes to determine the best-supported interpretation. Tests are regenerated as the policy produces new answers, enabling them to evolve with the policy. Across four benchmarks spanning affect, pragmatics, communicative intent, and preference, RLCE attains the strongest performance among seven methods that use no ground-truth training labels for rewards, including consensus, policy LLM-judge verification, multimodal co-evolution, and rubric-based rewards. Gains over the strongest baseline reach up to +18.93 points. Analyses further show that RLCE exhibits a larger share of reward variation between correct and incorrect predictions than compared rubric methods, can overturn erroneous policy-derived preferences, and benefits from pairwise test construction, compositional test aggregation, and on-policy test evolution.
|
| 1594 |
Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough
2610.04083
|
cs.AI
|
Debeshee Das, Jacqueline Tay, Bruce Tsai, David Huang, Javier Rando |
Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act...Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act on to persistent memory, so that a future aligned agent carries it out when the opportunity arises. We investigate this threat, which we refer to as self-propagation of misalignment, across 20 different scenarios, whose misaligned goals include self-preservation, power-seeking, undermining oversight, reward hacking, and deceiving the user. We simulate misalignment in 11 frontier models using two prompting strategies; unrestricted and values-only. The first explicitly states the misaligned goal, for instance, to prevent its own replacement, and self-propagation succeeds in 58% of runs. The second only describes what the agent cares about, for instance, that its continued operation is essential to its users, without specifying misaligned goals or directives. Even under this weaker prompt, self-propagation succeeds in 18% of runs, and every model self-propagates in at least one scenario. On removing the memory tool from the harness, we find that agents use the file system, writing the goal to a file in 74% of sessions; self-propagation still succeeds in 11% of runs. We also show that weaker models can propagate misalignment to more capable models, and that propagated goals can persist through 100 sessions of unrelated work. Existing defenses against memory poisoning and prompt injection do not directly address this threat because the memory content is generated by the agent itself, rather than injected by an external adversary. An LLM memory auditor from prior work (MemMorph) only reduces propagation from 71% to 34% of runs. We release our scenarios to support the evaluation of defenses against this emerging threat.
|
| 1595 |
Towards Safer Autonomous Driving in an Open World: A Dual-Process Approach
2610.04088
|
cs.AI
|
Simon Janssen, Michiel Braat, Chris van der Ploeg, Serge Thill, Jan-Pieter Paardekooper |
Before autonomous driving systems can be deployed on public roads, it is vital that these systems comply with safety standards, traffic rules, and social norms. Although neural networks trained on large amounts of driving data perform well in routine driving t...Before autonomous driving systems can be deployed on public roads, it is vital that these systems comply with safety standards, traffic rules, and social norms. Although neural networks trained on large amounts of driving data perform well in routine driving tasks, these models often struggle in novel situations that are not well-represented in the data. In this work, we propose a novel framework that combines a neural network for intuitive, learning-based planning in routine driving tasks with model predictive control for reasoning-based planning in unfamiliar situations, inspired by Dual Process Theory. A meta-cognitive component is designed to switch between the two, using a knowledge graph to reason about contextual risk based on explicit perceptual information and relevant traffic rules and social norms. Contextual risk is represented through risk fields, guiding both the switching mechanism in the meta-cognitive component and compliance with safety standards, traffic rules, and social norms in the reasoning-based planner. The effectiveness of our framework is tested in CARLA for variations of a typical out-of-distribution situations involving (emergency) vehicles running a red light. We show that the novel architecture reduces the number of collisions in the scenarios by 89% and improves compliance with the special right-of-way rules, compared to the NN-only planner.
|
| 1596 |
The Independence Prior of SAEs Fragments Visual Concepts
2610.04112
|
cs.AI
|
Tommaso Mencattini, Giorgos Nikolaou, Donato Crisostomi, Thomas Fel, Francesco Montagna |
Sparse Autoencoders (SAEs) decompose model activations into sparse combinations of interpretable dictionary atoms. Although SAEs are grounded in the Linear Representation Hypothesis (LRH), their objective smuggles in an additional prior: concepts across patche...Sparse Autoencoders (SAEs) decompose model activations into sparse combinations of interpretable dictionary atoms. Although SAEs are grounded in the Linear Representation Hypothesis (LRH), their objective smuggles in an additional prior: concepts across patches are treated as independent, an assumption clearly violated by natural images and by the activations they induce. We therefore specialize LRH to vision through the Markov-Field Linear Representation Hypothesis (MFLRH), which adds the missing spatial dependencies to the LRH assumptions. We thus propose Spatial-SAE as an amortized MAP estimator under the MFLRH. Spatial-SAE consistently outperforms standard SAEs in concept recovery and interpretability. Across four variants, it achieves a 96% average win rate on synthetic concept recovery and improves interpretability on DINOv2 activations, at a reconstruction cost concentrated in high spatial frequencies.
|
| 1597 |
CUAWright: A Minimal Unified Interface for Digital Agents
2610.04116
|
cs.AI
|
Yadong Lu, Theodore Lee, Yifei Li, Lawrence Keunho Jang, Tianci Xue |
The prevailing approach to computer-use agents couples a model with a domain-specific harness: a browser or desktop environment equipped with human engineered tools that are fixed before task execution. As models' coding capabilities improve, the GUI native an...The prevailing approach to computer-use agents couples a model with a domain-specific harness: a browser or desktop environment equipped with human engineered tools that are fixed before task execution. As models' coding capabilities improve, the GUI native and static harness prevents them from direct programmatic operation on system state, as well as flexible construction of tools. To this end, we introduce CUAWright, a minimal terminal harness of roughly 3K lines of code that uses bash commands as its sole action interface, and a file system as its evolvable space for dynamically creating tools and managing the context. We conduct comprehensive experiments across a wide range of digital tasks, and demonstrate that by giving the agent a minimal, programmable interface, it achieves substantially stronger results compared to their GUI or hybrid CLI interface across a wide range of tasks. On OSWorld 2.0, CUAWright delivers a 33.2% relative improvement in partial reward while reducing estimated cost by 37.5% compared with the published GPT-5.5 baseline. On Online-Mind2Web and the long horizon Odysseys benchmark, CUAWright substantially outperforms GUI native harness by 4.7% and 44.0% in success rate, respectively. Furthermore, we found the gains extend to CAD applications that require accurate visual understanding and CLI interaction: on CADGenBench and BenchCAD, our unified harness yields 8.1%-41.6% relative improvements over other CLI based harnesses with GPT-5.5. Together, these results suggest digital environments are far more programmable than their GUI interfaces imply, and a minimal terminal-focused harness is the key for better performance and efficiency.
|
| 1598 |
Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation
2610.04133
|
cs.AI
|
Ji Young Byun, Anthony Hu, Jesse Rogers, Roujia Wang, Manasa Kesapragada |
Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth ex...Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding whether pairs of generated hypotheses describe the same underlying mechanism. We study two evaluation questions that rest on this pairwise equivalence judgment: (1) how much self-critique changes delivered hypotheses beyond run-to-run variability, and (2) how the equivalence rule used to group hypotheses affects measured diversity. Across four proprietary instances, we hold opening hypotheses fixed, rerun the downstream workflow with 0, 1, and 5 critique rounds, and score matched hypothesis pairs with an LLM-as-a-judge. Relative to matched same-depth reruns, moving from 0 to 1 round produces 34.5 percentage points (pp) of additional mechanism-level divergence, whereas 1 to 5 rounds adds 1.3 pp. We then compare three equivalence rules: term frequency--inverse document frequency (TF--IDF) similarity, dense embeddings, and the same LLM-as-a-judge. We construct controlled hypothesis pairs that either preserve the causal explanation through wording or biological-terminology changes, or replace one component of the causal chain while holding the rest fixed. All three rules are invariant to meaning-preserving edits, but when the initiating event is replaced, the LLM-as-a-judge identifies 83% of valid pairs as different mechanisms, versus 0% for TF--IDF and 8% for embeddings; varying only the rubric that defines same mechanism moves this figure from 38% to 96%. Together, these results show that pairwise equivalence judgments are a measurement choice: how mechanism equivalence is defined affects both the estimated effect of self-critique and the measured diversity of generated hypotheses.
|
| 1599 |
Agentic Cognitive Depth: Operational Criteria for Evaluating LLM Agents
2610.04168
|
cs.AI
|
Nijesh Upreti, Chris Sypherd, Vaishak Belle |
Agentic large language model (LLM) systems are commonly implemented as an LLM in a loop with Planning, Memory, Tools, and Control Flow. This application-focused view connects agentic LLM research with deployable systems and leaves open how such systems should ...Agentic large language model (LLM) systems are commonly implemented as an LLM in a loop with Planning, Memory, Tools, and Control Flow. This application-focused view connects agentic LLM research with deployable systems and leaves open how such systems should be evaluated beyond end-to-end task success. Building on this view, we define agentic cognitive depth as a trajectory-level profile across five operational criteria. The profile contains context sensitivity ($C$), temporal continuity ($T$), multimodal coordination ($M$), adaptive interaction ($A$), and metacognitive monitoring ($Mc$). The first four criteria measure how well Control Flow, Memory, Tools, and Planning are used across a trajectory. The fifth measures whether the system monitors and regulates the full run. For each criterion, we give operational proxies and a perturbation procedure, then connect the profile to the agent's world model. We provide the structure needed to extend benchmarks such as GAIA, SWE-bench, WebArena, and TRIP-Bench with per-criterion diagnostics. Symbolic verifiers, structured memory, planner coupling, and tool constraints provide practical ways to build and test these capacities.
|
| 1600 |
SHarP: Saliency-based Pruning of Agent Harnesses
2610.04178
|
cs.AI
|
Xinyi Gao, Qiucheng Wu, Kaizhi Qian, Handong Zhao, Shiyu Chang |
Agent harnesses are systems that coordinate model calls, tool use, and task execution to help large language models complete complex tasks. To meet task requirements and address failures, these systems are often iteratively refined by amending and patching the...Agent harnesses are systems that coordinate model calls, tool use, and task execution to help large language models complete complex tasks. To meet task requirements and address failures, these systems are often iteratively refined by amending and patching their instructions, tools, and workflows, continuously increasing harness complexity. It is therefore unclear whether some resulting harness modules are redundant, introducing substantial token overhead with little, if any, performance gain. Inspired by neural network pruning, in this paper, we study harness pruning as a means of striking a better balance between task performance and token cost. We propose SHarP (Saliency-based Harness Pruning), a simple yet effective pruning strategy based on the saliency of each harness module with respect to performance and efficiency. Specifically, we first identify tools, instructions, and supporting mechanisms as components that can be individually disabled. We then estimate the saliency of each module by ablating it and assessing its task performance and token cost relative to the full set of single-module ablations. Modules with the smallest contribution to performance or largest computational overhead are subsequently pruned. Our evaluation across various harnesses on held-out validation sets reveals a surprising finding: most harnesses that we studied are highly redundant and can maintain comparable performance and efficiency even after a substantial portion of their modules are pruned. Our pruning approach and empirical findings provide new perspectives on agent harness design and optimization.
|
| 1601 |
EvalResearchBench: Can AI Agents Design Their Own Evaluations?
2610.04184
|
cs.AI
|
Yaolun Zhang, Tianyi Xu, Yujie Zhao, Jishen Zhao, Qingyun Wu |
Recursive self-improvement (RSI) relies on evaluation feedback to assess progress and guide further research, yet repeatedly running complex benchmarks is costly and slows iteration. Human experts reduce this cost by selecting benchmark subsets or designing co...Recursive self-improvement (RSI) relies on evaluation feedback to assess progress and guide further research, yet repeatedly running complex benchmarks is costly and slows iteration. Human experts reduce this cost by selecting benchmark subsets or designing compact suites. We ask whether AI agents can automate this design process and introduce EvalResearchBench (ERB), a benchmark for autonomous evaluation research. Given target materials, development references, candidate APIs, and fixed time and API budgets, an agent called the researcher selects or synthesizes tasks, implements graders, and revises them in pilot tests before freezing an executable evaluator for coding, co-work, and reasoning. We study 9 researchers and 13 candidate models and compare each frozen evaluator with 14 target benchmarks on score concordance and pairwise agreement. The best evaluators order about 75\% of candidate pairs as the targets do, below the 91\% ceiling set by disagreements among the targets. No researcher leads on every metric, and the best evaluator on development targets is not the best on sealed targets hidden from the researcher. A human-designed sample of public tasks remains a strong baseline, and the evaluator with the lowest recorded execution cost attains the highest pairwise agreement. Agents repair tasks and graders through pilot feedback, yet their evaluators can still truncate answers, exhaust the evaluation budget, or let a few questions dominate a domain score.
|
| 1602 |
Agentic AI with Structured CoT for Enhancing AI's Spatial Intelligence: Visualization and Reasoning of Rotation
2610.04188
|
cs.AI
|
Uttamasha Monjoree, Wei Yan |
Recent studies show that artificial intelligence (AI) with language and vision capabilities still experiences limitations in spatial reasoning. In this paper, we have studied the spatial capabilities of advanced generative AI to understand the rotations of obj...Recent studies show that artificial intelligence (AI) with language and vision capabilities still experiences limitations in spatial reasoning. In this paper, we have studied the spatial capabilities of advanced generative AI to understand the rotations of objects in 3D space, utilizing AI's image processing and language processing features. We trained and examined the spatial intelligence of a generative Agentic AI model (GPT-5.6) to understand the spatial rotation process with rotation diagrams based on the revised Purdue Spatial Visualization Test: Visualization of Rotations (Revised PSVT:R). We improvised the Revised PSVT:R by superimposing additional graphical and contextual features to evaluate how different Chain-of-Thought (CoT) reasoning strategies influence model performance. The results indicate that structured CoT reasoning improves the spatial reasoning performance of the base GPT-5.6 model in both datasets (PSVT:R and PSVT:R with coordinate system). We used three CoT approaches - (1) Structured CoT, (2) few-shot Structured CoT, and Structured CoT with Self-optimized Prompt. The three CoT approaches evaluated in this study showed no significant performance difference. Results showed that combining structured CoT reasoning with relevant contextual information leads to considerable improvements in VLM performance on 3D rotation tasks, demonstrating the potential of agentic AI for more effective spatial reasoning. However, when contextual information is removed, structured CoT reasoning alone provides limited improvement, and the models continue to exhibit notable difficulties in understanding spatial transformations. These findings suggest that effective spatial reasoning in VLMs relies on the integration of visual, textual, and reasoning-based information in future agentic AI systems for spatial intelligence.
|
| 1603 |
Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks
2610.04196
|
cs.AI
|
William Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan |
LLM-based long-horizon agentic post-training is often bottlenecked by rollout generation: trajectories span many interaction turns, completion times vary substantially, and synchronous update barriers leave faster workers waiting for stragglers. Asynchronous r...LLM-based long-horizon agentic post-training is often bottlenecked by rollout generation: trajectories span many interaction turns, completion times vary substantially, and synchronous update barriers leave faster workers waiting for stragglers. Asynchronous reinforcement learning which has been adopted in LLM post-training addresses this inefficiency by consuming trajectories as they arrive, but introduces policy lag and off-policy optimization. Evolution strategies (ES) offer a backpropagation-free alternative for LLM post-training, yet it relies on a larger number of rollouts and existing practices have remained largely synchronous. In this short-form paper, we introduce bounded-staleness asynchronous ES and demonstrate it on Endless Terminals benchmark using Qwen2.5-7B-Instruct. Across three evaluation seeds, natural Async-1 matches synchronous ES, achieving 25.9\% versus 25.4\% held-out success. Controlled schedules that delay 10\% of each update cohort by four or eight policy updates reduce success by only 1.6 and 3.1 percentage points, respectively, without explicit off-policy correction. GRPO performs better overall, reaching 29.0\% held-out success, but importantly our results show that ES tolerates moderate policy staleness with limited degradation, opening possibilities for future improvement of ES-based post-training with asynchronous algorithms. To the best of our knowledge, we are the first to demonstrate the effectiveness of sync and async ES on a multi-turn terminal style agentic coding task.
|
| 1604 |
Fine-Tuning VLM for Enhancing AI's Spatial Intelligence: Understanding 3D and 2D Rotations
2610.04206
|
cs.AI
|
Uttamasha Monjoree, Wei Yan |
Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in sp...Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.
|
| 1605 |
TCMClinicalReason-Bench: Can Language Models Reason from Pathogenesis to Prescription over Real-World Clinical Cases?
2610.04215
|
cs.AI
|
Jirui Dai, Chenkai Zhang, Yan Jia, Yukai Wang, Ruiyang He |
Large language models (LLMs) can generate clinical narratives that are insufficiently grounded in patient-specific evidence. In traditional Chinese medicine (TCM), errors can propagate from etiology and pathogenesis through syndrome diagnosis and treatment pri...Large language models (LLMs) can generate clinical narratives that are insufficiently grounded in patient-specific evidence. In traditional Chinese medicine (TCM), errors can propagate from etiology and pathogenesis through syndrome diagnosis and treatment principles to prescription generation. We developed TCMClinicalReason-Bench using 2,000 multicenter electronic health record cases to distinguish case-grounded responses from fluent but unsupported diagnostic and therapeutic conclusions. Five general-purpose and two TCM-specific LLMs were evaluated in zero-shot settings. An evidence-constrained rubric assessed seven diagnostic and therapeutic components and three cross-block relations, allowing case-supported alternatives. Qwen3.7-Plus with TCM retrieval served as the automated judge, alongside parallel blinded ratings by five senior TCM clinicians on a 600-case subset. Structural completeness was nearly saturated (99.3-100.0%), but normalized content scores ranged from 40.7% to 54.1%. The five general-purpose models averaged 50.0%, versus 41.3% for the two smaller TCM-specific models. Cross-block logic consistency ranged from 60.8% to 66.8% and correlated moderately with content across cases (Pearson's r = 0.515-0.656). Deficits were greatest in prescription generation, prescription analysis, and symptom-guided modification. In judge stress testing on 100 independent cases, perturbation detection rates across the three relations were 57%, 56%, and 31%, with contradictions detected more reliably than omissions. Separating component quality from cross-block consistency localizes failures missed by endpoint and completeness metrics and identifies where clinician oversight remains necessary.
|
| 1606 |
On the Steering Dimensionality of Refusal in Language Models
2610.04245
|
cs.AI
|
Han Wang, Erik Miehling, Dennis Wei, Karthikeyan Natesan Ramamurthy, Huan Zhang |
Existing activation steering methods often assume that a high-level concept can be mediated by a single steering direction. To support this, two complementary interventions should be achieved: additive steering should induce the target behavior, while directio...Existing activation steering methods often assume that a high-level concept can be mediated by a single steering direction. To support this, two complementary interventions should be achieved: additive steering should induce the target behavior, while directional ablation should suppress it. Yet behaviors may occupy richer activation geometries beyond a single direction, and semantically similar behaviors may be represented by distinct directions. In this work, we study how many directions can reliably control two different types of refusal behaviors: refusal triggered by the safety alignment and refusal in general contexts. Given the limited expressive capability of a single steering vector, we study the general setting of steering subspaces and introduce the notion of steering dimensionality as the minimum subspace dimensionality required to reliably control a behavior. We characterize sufficient steering subspaces that cover the full extent of the target behavior through both the (monotonic) improvement before the sufficient dimensionality, and the saturation beyond it. Empirically, we find that refusal triggered by safety alignment is 1-dim steerable, while multiple distinct steering directions can achieve comparable control. In contrast, refusal in general contexts exhibits substantially richer activation geometry where even 5-dim steering subspaces fail to reliably capture its full steerable variation. Our results reveal that the activation geometry underlying refusal is highly context-dependent and can be substantially more complex than a single linear steering direction.
|
| 1607 |
Spec2Game: Can LLMs Generate Complete Playable Games from Detailed Specifications?
2610.04253
|
cs.AI
|
Yixue Cai, Yuzhe Zhao, Hanxiang Chao, Qingsen Ma, Ziheng Xiong |
Generating an executable program does not necessarily mean that it correctly implements the behavioral requirements specified in natural language. To evaluate large language models' ability to realize detailed specifications as complete interactive programs, w...Generating an executable program does not necessarily mean that it correctly implements the behavioral requirements specified in natural language. To evaluate large language models' ability to realize detailed specifications as complete interactive programs, we introduce Spec2Game, a benchmark that requires models to generate complete Pygame projects from detailed natural-language game specifications. Spec2Game comprises 15 game families and 150 task instances, with one canonical task and nine controlled rule variants per family, spanning three levels of implementation complexity. Using source-code, runtime, and visual evidence, we evaluate generated projects along four dimensions---Executability, Specification Realization, Code Quality, and User-Facing Quality. Across 14 LLMs and 3,330 generated projects, we find that high executability does not imply faithful specification realization. Component-level analysis further shows that models perform substantially better on Game Element Modeling than on Rule and Mechanism Modeling or Goal and Termination Modeling, indicating that faithfully implementing game rules and termination logic remains a major challenge.
|
| 1608 |
CADForge: Agentic Single-View CAD Reconstruction with Explicit Geometry Reasoning
2610.04262
|
cs.AI
|
Keyang Lu, Zhifei Yang, Tianao Dong, Mingzhe Xing, Zhen Xiao |
Reconstructing editable parametric CAD models from a single-view image is of great practical value for modern manufacturing, yet remains challenging due to incomplete geometric observations and complex inter-part relationships. To address it, we propose CADFor...Reconstructing editable parametric CAD models from a single-view image is of great practical value for modern manufacturing, yet remains challenging due to incomplete geometric observations and complex inter-part relationships. To address it, we propose CADForge, an agentic framework that progressively converts a single image into CadQuery programs. CADForge decomposes an object into CAD-meaningful components and performs explicit geometric reasoning for each component, a process that first identifies CAD-relevant constraints and then translates them into precise modeling parameters through mathematical code. The inferred parameters then drive component-wise synthesis of executable CadQuery programs, with a review agent evaluating the resulting geometry and providing targeted feedback for iterative refinement. To further improve robustness and efficiency, CADForge incorporates a failure-guided toolkit construction mechanism to distill accumulated experience into tools, and maintains a compact parametric CAD memory for retrieving modeling context on demand. Experiments on diverse single- and multi-part objects show that CADForge consistently outperforms existing baselines in reconstruction fidelity and perceptual quality, demonstrating an effective approach to accurate single-view CAD reconstruction.
|
| 1609 |
Dense Neuro-Symbolic Reasoning in a Unified Geometry State
2610.04280
|
cs.AI
|
Ruoran Xu, Wending Gao, Haoyu Cheng, Xiaoqiang Kang, Qiufeng Wang |
Geometry reasoning is naturally stateful: solving a problem repeatedly alternates between structural proposals and exact deductions. We formulate this process as dense neural-symbolic coupling, in which neural guidance and symbolic execution share a typed stat...Geometry reasoning is naturally stateful: solving a problem repeatedly alternates between structural proposals and exact deductions. We formulate this process as dense neural-symbolic coupling, in which neural guidance and symbolic execution share a typed state and communicate through executable actions at every search step. Neural proposals contribute theorem instances, constructions, and algebraic bridges; the symbolic runtime applies registered rules, propagates exact constraints, and records provenance. A nested controller allocates computation first between neural and symbolic proposal sources and then among admitted actions. We instantiate the framework in OmniGeo, a single solver for plane, analytic, and solid geometry. With Claude Sonnet 4.6, OmniGeo reaches 94.2%, 88.5%, and 89.8% on FormalGeo7K, Conic10K, and SolidFGeo, respectively (90.8% macro average), and solves 21/30 IMO-AG-30 problems.
|
| 1610 |
VIGIL: Verifier-Informed Gated Improvement Loop for Spreadsheet Question Answering
2610.04287
|
cs.AI
|
Kang Li, Lu He, Sandarsita Guntupalli |
Enterprise agents should improve from delayed feedback without allowing every correction to rewrite system behavior. We study continual harness learning for corpus-level spreadsheet question answering. Building on FiCo (Find-then-Compute), a static retrieval-a...Enterprise agents should improve from delayed feedback without allowing every correction to rewrite system behavior. We study continual harness learning for corpus-level spreadsheet question answering. Building on FiCo (Find-then-Compute), a static retrieval-and-execution backbone, we introduce VIGIL (Verifier-Informed Gated Improvement Loop). Within a question, VIGIL verifies and repairs diversely prompted Structured Query Language (SQL) candidates. Across episodes, delayed labels update only a contract calibrator and query/result-column selector. The base model, prompts, retriever, and recorded candidate pool remain fixed in the continual-learning protocol. With the gold workbook and expected type supplied, the ungated and dual-gated full-replay variants reach 82.4% and 82.0% forward accuracy over 79 documents, from a 75.8% static baseline. The dual gate has the larger retrospective gain (2.7 versus 2.3 points), and its 6.3-point forward gain has a 95% t-interval of 5.6-6.9. In a separate stricter split that excludes 16 documents and 380 questions from fitting and online promotion, the accuracy-only gate raises mean held-out accuracy across ten final harnesses from 80.3% to 86.7%. Calibrator-only adaptation gains 5.9 points, close to the combined 6.4-point gain. Yet three of 26 gate-approved updates reduce held-out accuracy relative to their incumbents, so replay-buffer non-regression does not imply held-out non-regression. MiMoTable and external-task case studies test within-task verification and reuse the same promote-or-retain discipline. Overall, the results support bounded, auditable harness adaptation while revealing where finite replay gates fail to generalize beyond their promotion buffers.
|
| 1611 |
LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures
2610.04292
|
cs.AI
|
Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang, Sun Li |
LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can...LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising part decompositions, joints, materials, and sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct objects; (2) a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia; and (3) a evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 systems reveal several intriguing findings: (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively create new components, whereas weaker models tend to rely on retrieval; and (c) providing functional specifications substantially improves part completeness, kinematics, and physical operability. These results show that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.
|
| 1612 |
EnvDreamer: Large-Scale Multimodal-to-Environment Generation for Embodied AI
2610.04301
|
cs.AI
|
Kabir Swain, Sijie Han, Antonio Torralba |
Large datasets and high capacity models have accelerated progress in vision and language. This work introduces a platform aimed at bringing comparable gains to embodied learning, world models, and robotics. We present EnvDreamer, a framework that uses large la...Large datasets and high capacity models have accelerated progress in vision and language. This work introduces a platform aimed at bringing comparable gains to embodied learning, world models, and robotics. We present EnvDreamer, a framework that uses large language and vision language models to generate Unreal Engine 5 environments for embodied AI and robot training. EnvDreamer enables sampling of large, diverse, interactive, customizable, and validator passed virtual environments for training and evaluation across navigation, interaction, and manipulation. We illustrate the platform with a large set of generated scenes and simple baselines. Policies trained on EnvDreamer generated environments, without explicit mapping or human task supervision, achieve competitive results on multiple embodied benchmarks spanning navigation, rearrangement, and manipulation. EnvDreamer also supports image-conditioned reconstruction for real-to-sim studies. Finally, we release EnvDreamer-20k, a dataset of 20,000 validator passed environments with task programs, scene graphs, trajectories, and metadata to support reproducible benchmarking.
|
| 1613 |
MOIRA: Mass-Oriented Indexing with Ragged Attention for Long-Context Decoding
2610.04313
|
cs.AI
|
Dich Nhat Minh Nguyen, Tran Dang Duong Nguyen |
Long-context decoding is limited by memory bandwidth, because every output token reads the KV cache of every layer. Sparse decoding reduces this cost by reading only part of the KV cache. We observe that the number of pages a query needs varies widely across K...Long-context decoding is limited by memory bandwidth, because every output token reads the KV cache of every layer. Sparse decoding reduces this cost by reading only part of the KV cache. We observe that the number of pages a query needs varies widely across KV heads, layers and steps. Fixed budgets are simple, but they are sized for demanding cases and tuned per workload; adaptive budgets follow this variation more flexibly, but existing designs pay for it with extra selection cost or training. At the kernel level, FlashAttention-3 (FA3) and FlashInfer are designed for rows of similar length: with page lists whose length differs per KV head, they either pad the lists (forfeiting much of the sparse saving), leave thread blocks unbalanced, or rely on a host-side plan that runs outside the CUDA graph. We propose MOIRA, a training-free sparse decode path in vLLM whose budget adapts per KV head and per layer. For every request, layer, KV head and step, a coverage rule keeps the smallest set of pages whose estimated attention mass reaches a fraction $\gamma$. A new kernel, self-planning attention, lets each thread block derive its own share of the work from the list lengths, so the whole decode step stays inside the CUDA graph. On an H200, at RULER's 128k context, MOIRA with $\gamma=0.99$ matches dense accuracy while reading about 30% of the pages and reduces the time per output token (TPOT) by 2.2-2.5$\times$ relative to dense FA3; with $\gamma=0.98$ it reduces TPOT by 2.7$\times$ and stays within the noise of dense. Under high serving load it raises throughput by up to 51%. These results suggest that a budget adapted per head and layer, paired with a kernel that keeps such budgets inside the CUDA graph, makes sparse decoding both flexible and fast.
|
| 1614 |
Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence
2610.04371
|
cs.AI
|
Amit Kachroo, Like Hui, Haitao Mao, Yuhao Zhang, Nguyen Vo |
Determining whether two programs are functionally equivalent is central to code modernization, patch validation, refactoring, and code-generation evaluation. Yet the usual signals are incomplete: tests cover only finite inputs, textual similarity confuses impl...Determining whether two programs are functionally equivalent is central to code modernization, patch validation, refactoring, and code-generation evaluation. Yet the usual signals are incomplete: tests cover only finite inputs, textual similarity confuses implementation with behavior, and unconstrained LLM judgments are difficult to audit. Direct execution is often impossible when a program depends on an obsolete, licensed, unavailable, or unsafe environment. We introduce FEAgent, a selective equivalence assessor agent that combines typed program-graph evidence with differential surrogate execution. FEAgent first aligns public interfaces and behaviorally relevant graph anchors, then issues bounded queries over call-flow, control-flow, data-flow, type, import, and effect relations. Next, a branch-aware generator agent proposes discriminating inputs, and two blinded LLM surrogates independently predict source and target observables. Every claim and predicted divergence is recorded in an evidence ledger. A deterministic reconciler then returns EQUIVALENT, INEQUIVALENT, or UNCLEAR rather than forcing a verdict when paths are uncovered or evidence conflicts. We evaluate FEAgent on function-level equivalence and repository-level bug patches, where the existing oracle is a benchmark label or a passing test suite. Every disagreement with that oracle is adjudicated by direct execution, revealing errors in benchmark labels and behavioral divergences missed by unit-test-only scoring. On EquiBench, execution confirms FEAgent's disagreements with published labels on 216 of 1,200 evaluated pairs (18.0%); on SWE-bench Verified, 94 of 331 test-passing agent patches (28.4%) diverge from the reference patch. FEAgent thus serves as an audit layer between testing and formal verification, keeping its evidence reviewable and its uncertainty explicit without claiming a proof of equivalence.
|
| 1615 |
Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
2610.04375
|
cs.AI
|
Boyang Yang, Zhenhao Li, Ziyao Yang, Kanghui Jia, Xin Yin |
Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call without notice. When the changed call fails, the agent retries a correct call, which costs ...Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call without notice. When the changed call fails, the agent retries a correct call, which costs users time and money. Benchmarks and failure analyses do not see the change, because they read the call and its result but not what a hop received. We define intent-execution correspondence (IEC) as the property that the executed action matches the action the emitted call denotes under the tool contract. Our protocol observes what each hop received without executing the call, and names the first hop that changed it by the receiver's own parser. IntAct then delivers the call in a form that this hop cannot alter, or refuses the call. We build IEC-Bench from the changes observed in real-world use, with chains of dependent calls under the execution paths of 4 widely-used harnesses. In 47,828 shell calls within production sessions, Claude Code's Bash tool changes 12.0% of the calls that carry code, escape sequences, or long text. For 80.7% of the calls whose backslashes are changed, the wrong action runs without any reported error. All 10 measured harnesses change a call. Trajectory-based judgment attributes 95.1% of the production failures to the LLM, although the path caused more than half of them. On IEC-Bench, the path raises the token cost per passed task 2.4 times (up to 12.3 times). A hop that changes a call also hides the changes after it, so 55.1% of the failures on one path appear only after its first hop is repaired. IntAct, deployed in a commercial product, recovers 79.2% of the failures with a changed call. Harnesses should therefore be designed and tested hop-by-hop to ensure a correct call executes as intended or is refused.
|
| 1616 |
AgentPersonaBench: Benchmarking Persona-Driven User Simulation
2610.04379
|
cs.AI
|
Jintao Huang, Yifan Wang, Hongyu Shen, Yi Daniel Lu, Shirley Huang |
We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate co...We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic behavioral fidelity. APB evaluates latent persona adherence one trait at a time, embedding each target trait within a complete synthetic profile without explicitly naming the trait or disclosing the test. Ground-truth adherence is verified strictly from observable actions across four interaction surfaces of increasing realism: survey, chat, web (interactive web environments), and app (desktop software environments). APB comprises 2,460 tasks spanning 867 traits, verified through automated audits and expert review. Our evaluation of 20 frontier model arms demonstrates that high-fidelity user simulation is already attainable: leading models achieve up to 84.7% full-pass adherence under unprompted conditions. At the same time, APB identifies clear behavioral boundaries: adherence drops across interaction modalities (only 37.9-64.3% pass all four surfaces), multi-attribute demands degrade retention, and competing model families exhibit pronounced behavioral divergence.
|
| 1617 |
TrustMed-RL: Long-Horizon Reinforcement Learning for Evidence-Grounded Clinical Diagnosis
2610.04387
|
cs.AI
|
Wenxin Zhan, Yizheng Jiao, Haifeng Song, Shuai Xu, Chencheng Pan |
Medical language models can produce correct diagnoses despite incomplete investigations and unsupported reasoning. To support long-horizon, evidence-grounded diagnosis, we introduce \textbf{TrustMed-RL}. Built from PubMed rare-disease cases and over 24,000 man...Medical language models can produce correct diagnoses despite incomplete investigations and unsupported reasoning. To support long-horizon, evidence-grounded diagnosis, we introduce \textbf{TrustMed-RL}. Built from PubMed rare-disease cases and over 24,000 manually annotated image panels, it integrates interviews, examinations, testing, specialist consultation, and literature search through state-dependent actions. Our 8B vision--language policy, trained with clinically adapted GiGPO and coverage-adjusted diagnostic rewards, achieves 37.1\% diagnostic accuracy on 2,500 evaluation cases, outperforming all evaluated open-weight baselines and improving over supervised fine-tuning by 12.4 percentage points.. When success additionally requires acquiring at least 50\% of supporting test evidence, TrustMed-RL achieves 32.5\%, exceeding GPT-4o by 6.8 percentage points. Furthermore, it surpasses all evaluated baselines on MTMedDialog and multiple larger 27--32B models on AgentClinic. In physician review of 200 diagnostically accepted test-set trajectories, 83.0\% receive evidential-grounding scores of 4--5 out of 5. Physicians' assessments suggest that these diagnostic trajectories are trustworthy and aligned with human diagnostic reasoning.
|
| 1618 |
CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies
2610.04418
|
cs.AI
|
Santhosh GS, Ananya Ravi, Devika Jay, Abhishek Sarkar, Perepu Satheesh Kumar |
The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipeline to evaluate the alignment of the RL agent with complex multi-objective specifications and robustness under real-world environmental drift. However, to pro...The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipeline to evaluate the alignment of the RL agent with complex multi-objective specifications and robustness under real-world environmental drift. However, to protect intellectual property, the RL agent may be delivered for evaluation as opaque executable or remote API, which makes traditional evaluation techniques based on the internals of the policies infeasible. To address this gap, CORE-RL: Confidence-Oriented Reliability Evaluation of black box RL policy is proposed in this paper. The CORE-RL pipeline introduces a Unified Reliability Metric that formally integrates early task termination and safety constraint violations, preventing unsafe policies from masking failures through premature episode halts. By subjecting the policy to a noise certification envelope of perceptual noise, actuation noise and change in environment dynamics, the pipeline computes the finite-sample Clopper-Pearson bounds on unified reliability metric and Hoeffdings' lower bound on reward and safety cost. The pipeline then defines safe operational design domain to report high-confidence certificates for safety and expected performance. Experiments on continuous control tasks demonstrate the CORE-RL pipeline's ability to automatically reject non-compliant policies and map the safe Operational Design Domain (ODD) of safety-aware policies. Thus CORE-RL provides an evaluation framework towards a quantitative, transparent and reproducible, statistical rationale necessary to safely evaluate, compare, and deploy black box RL solutions.
|
| 1619 |
Asking Earns Nothing: Scoring the Decision to Act in BFCL Multi-Turn
2610.04429
|
cs.AI
|
Yangze Liu, Zhongyi Han |
An agent that lacks the information it needs should ask rather than act, and the task definitions of agent leaderboards say so. BFCL multi-turn builds two of its four categories around a turn on which the model is supposed to ask, and its scorer never looks at...An agent that lacks the information it needs should ask rather than act, and the task definitions of agent leaderboards say so. BFCL multi-turn builds two of its four categories around a turn on which the model is supposed to ask, and its scorer never looks at that turn: the gold trajectory there is empty, the checker skips it, and the scripted user cannot answer, so asking earns nothing, guessing costs nothing on that turn, and asking twice loses the item. The benchmark also contains the control experiment for that decision. A should-ask item is a base item with one piece of information removed from one turn, so the same request appears twice at the same turn index, once complete and once not: on the first the model should make the call that changes the world, on the second it should ask. We score one decision per pair, whether the model attempted a world-changing call on that turn, read off the stored trajectories with no LLM judge; acting always and asking always both score 50. On the 223 pairs that pose this decision, gpt-5.4 attempts the call on 83.4% of the complete turns and holds back on 78.0% of the incomplete ones, the best decision accuracy of seven models at 80.7%; on the same items the official score ranks it sixth and puts first a model that lands in the middle here. One added line telling gpt-5.4 not to ask pushes it toward acting on both sides of the pair, so its decision accuracy shows no detectable change, while its official score rises by 13.5 to 23.5 points on the two should-ask categories and on the base twins; the opposite line, telling gemma-4-31B-it to ask first, improves its decision by 4.5 points and gains no official score. The score moves with the push toward action, not with the decision. We release the pairs, a turn-level scorer that runs on any BFCL output directory without an API key, and 31 manually verified bad items.
|
| 1620 |
MERCI Cards: An LLM Evaluation and Deployment Framework for High-Stakes Domains
2610.04430
|
cs.AI
|
Aparna Komarla, Annalisa Szymanski |
As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers require principled protocols to systematically track, monitor, and improve model performance across deployment cycles. We present a mathematical framework for it...As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers require principled protocols to systematically track, monitor, and improve model performance across deployment cycles. We present a mathematical framework for iterative LLM evaluation and deployment, and demonstrate its application to AI systems used in criminal justice. Our framework formalizes LLM integration in high-stakes, high-risk, and resource-constrained domains across model selection, rubric design, evaluations and deployment via a weighted multi-objective optimization. We demonstrate that MERCI Cards can guide improvements of the system across deployment iterations, direct developer attention toward under-performing areas, and focus user attention on validation and error-correction in the LLM's outputs.
|
| 1621 |
CRAFT: An Agentic Spreadsheet Form Filling System with Template Awareness
2610.04437
|
cs.AI
|
Leyao Gu, Yingjie Xiong, Zirui Tang, Jiangtao Zhou, Yeye He |
Spreadsheet form filling requires agents to consolidate external evidence, ground values to precise cells, and preserve irregular template structure. Errors in early edits can overwrite labels or misalign fields, undermining later decisions. We propose CRAFT, ...Spreadsheet form filling requires agents to consolidate external evidence, ground values to precise cells, and preserve irregular template structure. Errors in early edits can overwrite labels or misalign fields, undermining later decisions. We propose CRAFT, a template-aware agent framework that connects reflective validation to constrained local repair. Instead of treating reflection as a free-form request to regenerate the workbook, CRAFT grounds detected errors to spreadsheet regions, restores corrupted template state when necessary, and re-grounds plausible writable slots before subsequent edits. A Rectangle-Aware Slot Grounder (RASG) proposes writable cells, while label-slot hints and protected regions constrain subsequent edits. We introduce FormFillBench, with 327 forms across Instruction-Only and Multi-File tracks. Compared with the strongest baselines, CRAFT improves pair accuracy by 8.51 and 23.38 percentage points on these tracks, respectively. Component-removal experiments support structural adjudication and slot re-grounding within the pipeline, and the framework retains its relative advantage among the methods evaluated with a second backbone. The code and benchmark FormFillBench are available at https://github.com/Glllllly/CRAFT.
|
| 1622 |
RAGrasp: Geometry-Semantic Template Retrieval and Grasp Transfer
2610.04438
|
cs.AI
|
Shenzhe Zhu, Chengxiao He, Jan Harder |
We present RAGrasp, a retrieval-augmented pipeline for planar parallel-jaw grasping from a compact set of locally collected, grasp-annotated RGB-D (color and depth) templates. Unlike task-specific predictors trained primarily on large public or synthetic grasp...We present RAGrasp, a retrieval-augmented pipeline for planar parallel-jaw grasping from a compact set of locally collected, grasp-annotated RGB-D (color and depth) templates. Unlike task-specific predictors trained primarily on large public or synthetic grasp datasets, RAGrasp requires no end-to-end retraining for a new deployment.Its template memory is constructed from observations collected with the deployment camera, robot, and gripper in the target workspace, thereby aligning stored examples with the local sensing and embodiment conditions. The system uses self-supervised DINOv2 visual fea- tures together with appearance and depth cues to prompt the Segment Anything Model 2 (SAM2), which isolates the query object. A two-stage geometry-semantic retrieval cascade then selects a template, and a confidence gate chooses one of two grasp- transfer estimators. The transferred grasp is refined using mask- support and silhouette-contact constraints before calibrated 2D- to-3D conversion. In real-world trials, RAGrasp achieves 20/20 successful grasps on seen objects and 19/20 on unseen objects. Within the evaluated setting, the results demonstrate deployment- specific grasp adaptation from limited local annotation and tolerance to the tested viewpoint and illumination changes.
|
| 1623 |
Understanding Generative AI Use in Programming MOOCs: The Role of Course Context and Learner Characteristics
2610.04447
|
cs.AI
|
Marina Lepp |
The increasing availability of generative artificial intelligence (GenAI) tools, such as ChatGPT and code-completion assistants, raises questions about how learners integrate these tools into learning activities, particularly in MOOCs that attract diverse part...The increasing availability of generative artificial intelligence (GenAI) tools, such as ChatGPT and code-completion assistants, raises questions about how learners integrate these tools into learning activities, particularly in MOOCs that attract diverse participant populations. This study examines the use of GenAI in two programming MOOCs taught in Estonian that differ in duration, workload, topic complexity, and assignment volume: About Programming (4 weeks, 26 expected hours, n = 187) and Introduction to Programming (8 weeks, 78 expected hours, n = 182). Post-course questionnaire data were analyzed using non-parametric statistical methods to examine self-reported adoption, usage frequency, and purposes of GenAI use across courses and learner backgrounds. The results show that GenAI adoption was widespread in both MOOCs, with no statistically significant differences by gender, age, education level, or prior programming experience. However, participants in the longer, more extensive MOOC reported significantly higher usage frequency and were more likely to use GenAI for debugging and idea generation. Reported usage frequency for code explanation and debugging was also higher in the longer course. Exploratory analyses found limited relationships between GenAI use and learning-related outcomes. The findings suggest that course context may play a greater role than learner characteristics in shaping how GenAI tools are used. These results provide implications for instructional design and guidance in programming education.
|
| 1624 |
Trinity: Self-Evolving Vision-Language Models with a Self-Verifier
2610.04469
|
cs.AI
|
Youngwan Lee, Yong-Ju Lee, Sung Ju Hwang |
Self-evolving vision-language models (VLMs), a form of self-improvement in which a model generates its own training data from unlabeled images, are a promising route toward agents that expand their reasoning capability in an unsupervised manner, without relyin...Self-evolving vision-language models (VLMs), a form of self-improvement in which a model generates its own training data from unlabeled images, are a promising route toward agents that expand their reasoning capability in an unsupervised manner, without relying on ever-larger annotation budgets. Existing methods pair a Questioner that proposes problems with a Solver that answers them, but reward both roles mainly by agreement among sampled answers. Agreement is a weak proxy for truth: it cannot tell whether a question is grounded in the image, whether the proposed reference answer is right, or whether a confident majority is wrong in the same way. We present Trinity, in which one VLM plays three roles, Questioner, Solver, and Verifier, and the Verifier is a self-verifier: an exponential moving average (EMA) of the policy itself, requiring neither labels nor an external judge. The Verifier screens every generated question for image grounding and answer correctness before it becomes supervision, scores Solver reasoning against the image, and adjudicates disputes between the reference answer and a strong Solver consensus, correcting the reference and penalizing the Questioner when the consensus is right. Trained on images alone, Trinity improves Qwen3-VL-8B on mathematical visual reasoning and on science benchmarks with biology content, for example, +8.6 points on the biology split of SciVQR and +12.8 on MathVerse, and its reward dynamics behave as a healthy self-play curriculum should. These results suggest that a self-evolving multimodal agent can strengthen its scientific reasoning from unlabeled scientific images alone, with the model itself serving as the verifier.
|
| 1625 |
Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching
2610.04470
|
cs.AI
|
Luoyu Chen, Weiqi Wang, Chenhan Zhang, Zhiyi Tian, Shui Yu |
Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious inte...Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious intent resides, resulting in brittle protection and excessive over-refusal. We propose SENTINEL, a plug-and-play, generation-time jailbreak defense that reframes mitigation as an intent extraction problem. Our key insight is that instruction-tuned LLMs exhibit strong input--output semantic consistency: regardless of jailbreak complexity, generated outputs tend to align with the attacker's true intent. SENTINEL exploits this property by matching semantically aligned input--output regions to extract intention-revealing subsequences, scores these subsequences using refusal-direction projections to estimate harmfulness, and halts generation when necessary. Experiments on HarmBench across multiple LLMs show that SENTINEL reduces jailbreak success rates to close to 5\% while maintaining low over-refusal. We further demonstrate robustness to adaptive attacks and provide a mechanistic interpretation: SENTINEL re-distributes jailbreak features from alignment blind spots to aligned regions.
|
| 1626 |
Semantic Causal-Factor Inference from Aviation Incident Narratives Using A Variational Autoencoder with Cosine-Similarity-Based Reconstruction
2610.04472
|
cs.AI
|
AZIIDA NANYONGA1, HASSAN WASSWA, UGUR TURHAN, KEITH FRANCIS JOINER, GRAHAM WILD |
Aviation accident and incident investigations generate extensive unstructured textual information containing evidence relevant to the causes and contributing factors of safety occurrences. Automatically extracting such information is challenging because causal...Aviation accident and incident investigations generate extensive unstructured textual information containing evidence relevant to the causes and contributing factors of safety occurrences. Automatically extracting such information is challenging because causal evidence may be distributed across long and complex investigation narratives. This study proposes a semantic causal-factor inference framework combining natural language processing with a variational autoencoder (VAE) to learn the relationship between aviation investigation narratives and expert-reported probable causes. Investigation narratives and their corresponding probable causes are transformed into numerical representations, after which the encoder maps narrative representations to a probabilistic latent space. The decoder estimates representations of the corresponding probable causes and is trained using an objective that combines Kullback-Leibler divergence with cosine-similarity-based semantic reconstruction. The framework was evaluated using 20,919 finalized U.S. National Transportation Safety Board investigation reports from 2005 to 2020. On the held-out test set, the predicted and expert-reported probable-cause representations achieved a mean cosine similarity of 0.786 (SD = 0.120). The predicted representations also yielded interpretable terms associated with causal information in the reports. The results demonstrate the potential of probabilistic latent representation learning for AI-assisted extraction of causal information from aviation safety narratives while retaining expert investigation as the basis for formal causal determination.
|
| 1627 |
InferOpt: Constrained Multi-Objective Search for LLM Inference Configurations
2610.04473
|
cs.AI
|
Qi Chen, Yingying Cheng, Zhaoyi Sun, Li Zhou, Fan Zhang |
Serving an LLM means setting dozens of inference-time knobs, from per-layer KV retention to per-layer expert counts. Practice sets them with mechanism-specific heuristics that return a single operating point and do not scale to layer-wise search spaces. We rec...Serving an LLM means setting dozens of inference-time knobs, from per-layer KV retention to per-layer expert counts. Practice sets them with mechanism-specific heuristics that return a single operating point and do not scale to layer-wise search spaces. We recast inference configuration as constrained multi-objective black-box optimization and build InferOpt, a reusable search framework that requires only variable bounds, a deterministic resource cost, and an evaluation hook. InferOpt searches on a frozen sampled proxy set, rejects over-budget candidates before any model call, tightens the budget adaptively, and re-validates Pareto representatives on full-scale dataset. One pipeline covers a 28-dimensional continuous KV space (Qwen2.5-7B) and a 26-dimensional discrete MoE space (DeepSeek-V2-Lite). On KV, post-prefill pruning cuts the 16K cache by 64.4% and TPOT by 22.9--48.5%, and the searched layer-wise budget by InferOpt beats a matched uniform budget by 7.3% and 14.0% of the Full KV reference points. On MoE, a searched top-k schedule by InferOpt removes 43.0% of routed token--expert pairs while staying within 0.59 points of the default, closer than the matched-budget baselines. Against Random Search, NSGA-II, and MOTPE, InferOpt leads on both spaces, taking the best proxy hypervolume and the lowest retention on KV and staying closest to the uncompressed reference at the lowest experts on MoE.
|
| 1628 |
VCLMU: Mechanism-Centric Virtual Cell World Modeling for Perturbation Response
2610.04475
|
cs.AI
|
Yuwei Miao, Azim Dehghani Amirabad, Scott Oloff, Junzhou Huang, Tianyu Cui |
Predicting cellular responses to genetic perturbations is a central capability for virtual cells and a key step toward computational modeling of biological interventions. Most existing models directly map an unperturbed molecular profile and perturba- tion to ...Predicting cellular responses to genetic perturbations is a central capability for virtual cells and a key step toward computational modeling of biological interventions. Most existing models directly map an unperturbed molecular profile and perturba- tion to the resulting observation without explicitly representing the latent cellular transition induced by the intervention. We introduce a mechanism-centric virtual cell world model that represents cellular state as a set of Latent Mechanism Units (LMUs) and treats genetic perturbations as actions on these latent states. Each LMU combines a reusable identity grounded in multimodal biological evidence with an observation-specific state, allowing a perturbation to induce mechanism- specific stochastic transitions before decoding the resulting transcriptional response. We train VCLMU through two-stage pretraining, first on around 200K pseudo-bulk perturbation profiles and then on gene-aligned single-cell perturbation data. Across six perturbation-disjoint benchmarks, VCLMU consistently improves perturbation- specific response recovery over strong baselines while maintaining competitive global response accuracy. We further analyze learned LMUs through enrichment between perturbation responses and LMU gene sets and show that they capture structured biological response programs. These results support mechanism-level latent state transition as a useful formulation for virtual cell models that aim to predict and interpret cellular responses to biological interventions.
|
| 1629 |
Towards Credible Agent-Based Policy Simulations: Disentangling Opportunities and Preferences in a Financial Inclusion Case Study of Egypt
2610.04515
|
cs.AI
|
Alba Aguilera, Georgina Curto, Nardine Osman, Ahmed Al-Awah |
Credibility is a central topic for agent-based models intended to support policy-making. Simulations must not only represent the target scenarios and their core dynamics but also demonstrate that their assumptions, parameters, and outputs are empirically groun...Credibility is a central topic for agent-based models intended to support policy-making. Simulations must not only represent the target scenarios and their core dynamics but also demonstrate that their assumptions, parameters, and outputs are empirically grounded and sufficiently accurate for their intended use. This paper addresses this challenge by presenting a general modelling framework, aligned with the Capability Approach, for building credible policy simulations that rely on data and domain-expert knowledge. It then demonstrates how it can be contextualised and implemented to study the social challenge of financial inclusion in Egypt, building on an agent-based model that represents heterogeneous individuals and firms behaving according to their financial states, barriers, opportunities, and preferences. The model is fitted to real-world data in two stages, initialisation and calibration, which respectively build representative synthetic populations and estimate behavioural parameters. By fixing the feasibility parameters, which determine agents' opportunities, and calibrating preference parameters across different population groups, we are able to distinguish and analyse the role of institutional and social barriers in the system, as well as the role of agents' motivations and priorities. This calibration stage provides transparent and group-specific hypotheses about the drivers of observed financial-inclusion gaps, which can further be analysed as gaps between agents' opportunities and realised outcomes, a very relevant insight for policy-making. This paper is thus a step towards improving the credibility and usefulness of policy simulations, strengthening the relationship between the model, the real target system, and the stakeholders who will use it. The code is available at: \url{https://www.comses.net/codebase-release/df8383cb-f49b-4f09-8cce-73603b59adcc/}.
|
| 1630 |
Action-Consequence Alignment for Reliable Planning and Self-Improving in Latent World Models
2610.04539
|
cs.AI
|
Jinping Wang1, Zhiqiang Gao, Xiantong Zhen, Ling Shao |
Latent world models learn to predict observed transitions, yet low prediction error alone does not guarantee reliable planning. Inspired by self tickling experiments in neuroscience showing that disrupting motor sensory correspondence increases prediction mism...Latent world models learn to predict observed transitions, yet low prediction error alone does not guarantee reliable planning. Inspired by self tickling experiments in neuroscience showing that disrupting motor sensory correspondence increases prediction mismatch, we examine whether learned world models preserve an analogous action consequence correspondence.The results show nearby alternatives can receive lower prediction errors despite producing physical outcomes farther from the recorded target. With that future treated as a goal, this reveals a concrete prediction planning mismatch: the model assigns a lower cost to an action that achieves the target less accurately. To mitigate this gap, we introduce Action Consequence Alignment (ACA), a training objective that complements forward prediction by penalizing the prediction error advantage of locally searched alternatives over factual actions without additional model components or environment interactions during training. The same principle can also guide additional data collection for self improvement. We demonstrate that across diverse environments and evaluation settings, ACA improves planning performance and reduces real goal error, while ACA guided data collection outperforms random local sampling. These results support action consequence alignment as a practical principle for bridging predictive learning and reliable planning.
|
| 1631 |
Decide, Ask, or Defer: Clinical LLMs under Incomplete Evidence
2610.04542
|
cs.AI
|
Mingzhan Yang, Weili Wu |
Clinical LLMs must decide not only what diagnosis to produce, but also whether the available evidence is sufficient for autonomous decision making. Binary DECIDE/ABSTAIN formulations merge distinct non decision states and do not explicitly evaluate information...Clinical LLMs must decide not only what diagnosis to produce, but also whether the available evidence is sufficient for autonomous decision making. Binary DECIDE/ABSTAIN formulations merge distinct non decision states and do not explicitly evaluate information acquisition. We introduce a DECIDE/ASK/DEFER formulation together with a blinded protocol that prevents models from using evidence completeness metadata. We evaluate Qwen, Gemini, and GPT on 200 matched clinical evidence states constructed from DDXPlus. The models show substantial differences in action selection under identical evidence, with disagreement in 137 of 200 states. For Qwen, a matched targeted versus random analysis shows that selected information changes the likelihood of a subsequent autonomous decision more clearly than diagnostic correctness. Its matched DECIDE/ABSTAIN baseline further reveals a safety autonomy tradeoff: the three action policy rescues some erroneous autonomous decisions but also removes some correct autonomous deci sions. These results show that separating information acquisition from clinician deferral exposes behavior that binary abstention hides, without yielding a uniformly improved decision policy.
|
| 1632 |
$\mathrm{TRIZ}^{a}$: Guiding Agent Evolution from Pattern Recognition to Solution Invention
2610.04555
|
cs.AI
|
Wenyin Liu (Guangdong University of Technology), Yiheng Huang (Beijing University of Posts and Telecommunications), Kai Wang (Beijing Denglu Technology Ltd) |
We propose $\mathrm{TRIZ}^{a}$ (TRIZ exponentiated by an agent), a general R\&D automation paradigm that combines TRIZ inventive theory with LLM-driven agent evolutionary search. TRIZ's 40 inventive principles and contradiction matrix provide structured, e...We propose $\mathrm{TRIZ}^{a}$ (TRIZ exponentiated by an agent), a general R\&D automation paradigm that combines TRIZ inventive theory with LLM-driven agent evolutionary search. TRIZ's 40 inventive principles and contradiction matrix provide structured, explainable directions for solution generation, replacing random or untyped mutation with theory-guided ideation. Functional information (FI), operationalized under a frozen reference contract, is combined with TRIZ Ideality to measure useful and harmful function on a commensurable information scale, while hard gates keep promotion distinct from metric improvement. We validate $\mathrm{TRIZ}^{a}$ in cybersecurity--an adversarial and rapidly evolving domain--on PowerDuck GOOSE, CICIoT2023, and CIC-DDoS2019. Under paired-rerun protocols with protocol fingerprinting and hard-gate validation, the legacy experiments yield absolute F1 improvements of $+2.88$, $+4.23$, and $+0.15$ percentage points, respectively. A completed 45-activity CICIoT2023 campaign further increases macro-F1 from $0.8325$ to $0.8483$, but does not pass its frozen promotion gate. Every result remains traceable from contradiction identification and TRIZ principle selection to code transformation, evaluation metrics, and promotion decision.
|
| 1633 |
Recursive Improvement of a Differentiable Scientific Software Ecosystem
2610.04561
|
cs.AI
|
Pengcheng Hou, Xiaojun Tan, Sihan Hu, Ruisi Wang, Shuo Chen |
Differentiable programming connects scientific computation with gradient-based inference, learning and design. Extending these capabilities across a heterogeneous software ecosystem requires specialized effort to implement derivatives, integrate interfaces and...Differentiable programming connects scientific computation with gradient-based inference, learning and design. Extending these capabilities across a heterogeneous software ecosystem requires specialized effort to implement derivatives, integrate interfaces and evaluate quality. AI coding agents can accelerate this transformation, but translating their capabilities into useful scientific software requires identifying research needs and evaluating how well implementations meet them. We present an environment for agent-driven evolution of differentiable scientific software that connects demand identification, development and quality evaluation. A unified differentiation interface exposes reusable derivative rules alongside existing numerical routines, allowing research tasks to share these capabilities. Research requirements guide development, with implementations assessed through independent derivative checks, workflow tests and performance evaluation. Validated software, research programs and tests become shared resources for subsequent studies. We construct and validate automatic differentiation extensions across 20 packages spanning physical, chemical and biological modeling, with research workflows demonstrating reuse across tasks. Benchmarks demonstrate computational savings over finite differences in gradient evaluation and complete parameter estimation. Research-driven revisions make previously unsupported workflows differentiable, correct derivatives of scientific observables and eliminate redundant computation. Quantum-control and thermal-design studies revise objectives in response to physical evaluation, improving designs while reusing existing derivatives. This work provides a practical approach to expanding differentiable programming across established scientific software and organizing AI agents around the recursive improvement of a shared computational ecosystem.
|
| 1634 |
Bounds, Decompositions and Null Behaviour of KRATOS: A Mathematical Specification of a Recognition-Comparability Diagnostic
2610.04592
|
cs.AI
|
Maria Dolores Gonzalez, Alberto Barbado |
KRATOS is a group-structured bibliometric diagnostic that compares the distribution of documents (participation) with the distribution of citation weight (recognition) over a fixed, finite universe of analytical groups. This note gives its complete mathematica...KRATOS is a group-structured bibliometric diagnostic that compares the distribution of documents (participation) with the distribution of citation weight (recognition) over a fixed, finite universe of analytical groups. This note gives its complete mathematical specification and derives the properties that govern its interpretation. Beyond bounds and equality conditions for each component, we show that the recognition-alignment score depends on citation data only through ratios of group mean citation rates to the corpus mean; that the composition ceiling of the participation--recognition factor is an affine function of the total variation distance to the uniform reference; and that the factor itself satisfies Fr\'echet-type bounds in terms of this ceiling and a participation-weighted recognition score. We establish an exact logarithmic decomposition of the composite index, an ordering between the primary and a reciprocal-symmetric recognition score, invariance properties, and closed-form first and second moments of the recognition ratios under global and stratified permutation nulls. These results explain why composite orderings can be sensitive to small groups, heavy-tailed citation distributions and metadata reassignment. All results are verified numerically with an accompanying script. The specification concerns measurement structure only; it does not define a measure of epistemic change or justice.
|
| 1635 |
LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention
2610.04635
|
cs.AI
|
Zhaohui Wang, Zhixin Pan, Fanxu Meng, Muhan Zhang |
Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We int...Sparse attention reduces core-attention computation, but its indexers still incur repeated selection work and per-layer key-cache storage. Reusing selected indices across layers reduces this overhead but constrains multiple layers to the same token set. We introduce LatentIndex, which extends the latent-sharing principle of Multi-head Latent Attention across indexer layers. Each layer group constructs a shared latent cache from its first layer's hidden states, while layer-specific scoring enables independent token selection. Absorbing key decoders into queries enables direct scoring of the shared cache without reconstructing historical per-layer keys. We develop training-free calibration and investigate a training-aware instantiation of this principle. To balance quality and computation, a hierarchical selection (HS) variant lets followers independently refine a shared candidate set proposed by the anchor. With four-layer sharing, LatentIndex reduces logical indexer-cache storage by 61.1% on DeepSeek-V3.2. Across DeepSeek-V3.2 and GLM-5, training-free LatentIndex improves head-wise attention-mass recall over IndexCache by up to 3.28 percentage points while maintaining RULER and LongBench performance close to native DSA. HS further achieves 2.30-2.72 times decode indexer speedups over DSA across 8K-128K contexts, retaining most of LatentIndex's recall. LatentIndex offers a new perspective on cross-layer indexing: sharing continuous representations rather than discrete selections enables efficient reuse while preserving layer-specific token selection.
|
| 1636 |
SEIS: Self-Evolving Inference Systems
2610.04646
|
cs.AI
|
Zhen Xu, Jingyu Liu, Zongze Li, Tahseen Rabbani, Ce Zhang |
Inference systems determine how fast and how cheaply language models can be served, so making them faster has direct practical value. However, prior work focuses mostly on optimizing certain parts such as kernels or memory within the large system. In this work...Inference systems determine how fast and how cheaply language models can be served, so making them faster has direct practical value. However, prior work focuses mostly on optimizing certain parts such as kernels or memory within the large system. In this work, we take a holistic approach and apply agentic self-evolution to optimize the whole system end-to-end. Our SEIS (Self-Evolving Inference Systems) autonomously optimizes the entire mini-sglang engine without human intervention through iterative sessions with inherited experiences and code changes. Serving Qwen3-0.6B on H100, the resulting engine reaches 3.27X the throughput of the original mini-sglang implementation and beats SOTA engines like vLLM, TensorRT-LLM, and SGLang in the single-request workload. The correctness of the optimized inference engine by SEIS is tested in terms of numerical difference and downstream accuracy on math and long-context retrieval tasks. The code and session histories show that the speedup comes from redesigning the whole engine and that building on earlier sessions beats independent attempts. These results suggest that agentic self-evolution can optimize a complex system end-to-end. The evaluation also has to evolve with the engine, and letting agents evolve it is a natural next step.
|
| 1637 |
PermVLA: Factorization Order as a Regularizer for VLA Learning
2610.04659
|
cs.AI
|
Yanqiao Chen, Yuhan Rui, Dongsheng Hou, Zijie Nie, Yutong Wan |
Vision-language-action (VLA) policies commonly learn action chunks through a fixed left-to-right (LTR) factorization, although the same expert trajectory distribution admits many valid chain-rule factorizations. We identify factorization order as an overlooked...Vision-language-action (VLA) policies commonly learn action chunks through a fixed left-to-right (LTR) factorization, although the same expert trajectory distribution admits many valid chain-rule factorizations. We identify factorization order as an overlooked regularization choice and introduce causally anchored permutation (CAP), which samples action reveal orders with a tunable chronological prefix. Its auxiliary objective trains one shared policy to predict actions from different known subsets of the same expert chunk, while deployment retains deterministic LTR control. We call this conditional-set augmentation: it creates multiple conditional prediction problems from one expert chunk without adding demonstrations. This discourages reliance on the single chronological prefix used by ordinary teacher forcing. Controlled experiments show that CAP consistently outperforms standard LTR training on LIBERO and LIBERO-Plus, with the same advantage appearing in cross-dataset CALVIN evaluation. A diagnostic that measures the expected squared difference between a chunk's joint log likelihood under two reveal orders verifies that CAP training internalizes agreement across reveal orders. These findings position sampled subset-conditioned auxiliary objectives as a general recipe for constructing VLA regularizers, illustrated by an extension to diffusion action generators.
|
| 1638 |
Efficient Neural Surrogates for Linear Radiation Transport on the Lattice and Hohlraum benchmarks
2610.04665
|
cs.AI
|
Carmelo Gonzales, Steffen Schotth\"ofer, Cory D. Hauck |
Linear radiation transport equations (RTEs) form the simulation foundations underpinning design and analysis tasks in nuclear engineering, inertial confinement fusion, medical imaging, and astrophysics, but resolving the high-dimensional phase space at enginee...Linear radiation transport equations (RTEs) form the simulation foundations underpinning design and analysis tasks in nuclear engineering, inertial confinement fusion, medical imaging, and astrophysics, but resolving the high-dimensional phase space at engineering fidelity remains expensive enough that outer-loop workflows, such as design optimization, uncertainty quantification, and parameter sweeps, are routinely budget-bound on traditional solvers. Neural surrogates promise to relax this bottleneck by amortizing simulation cost across thousands of downstream queries, but the architectural choices and engineered inductive biases that make a surrogate accurate on one transport problem do not transfer straightforwardly across model families. We benchmark two parameter-matched neural surrogate architectures, the physics-attention Transolver and the multi-scale graph network Bi-Stride Multi-Scale MeshGraphNet (BSMS-MGN), as end-to-end approximations of the final-time particle concentration for the two-dimensional linear RTE on the canonical Lattice and Hohlraum benchmarks. An ablation across Fourier features and region-weighted training loss exposes strongly architecture-dependent inductive-bias preferences, indicating that design choices common to physics-informed surrogate workflows must be revisited per architecture rather than imported across model families, and that downstream utility depends on per-QoI sensitivity rather than a single field-level score. The model training recipe, training data, and evaluation pipeline are released alongside this paper to support reproduction, transfer to related transport problems, and evaluation as amortized forward-model components in larger outer-loop simulation workflows.
|
| 1639 |
A Tropical Geometry View of Forgetting: A Per-Unit Projector for Knowledge-Preserving Fine-Tuning
2610.04670
|
cs.AI
|
Yuyang Zhang, Xiaoyin Chen, Chunlin Ren, Qihuang Zhang |
Fine-tuning a language model on new text degrades what it already does. Replay-free projectors such as Adam-NSCL and GPM forbid one shared subspace of a layer's inputs in every row of the update. The tropical geometry of a ReLU layer shows why this is too coar...Fine-tuning a language model on new text degrades what it already does. Replay-free projectors such as Adam-NSCL and GPM forbid one shared subspace of a layer's inputs in every row of the update. The tropical geometry of a ReLU layer shows why this is too coarse. In data space, the units' walls are tropical hypersurfaces whose cells are dual to the upper vertices of a zonotope; in weight space, each old token is a hyperplane, and the tokens cut out a polyhedron, the closure of the weights that keep every token on its side. An exact identity joins the two pictures: the squared change of the layer's output under any weight change splits into in-cell, open-to-closed and closed-to-open terms, and the first two live on the tokens each unit fires on. The identity names a gate-aware per-unit projector, and a budget-separation theorem prices exact protection: it costs a unit the rank of its own open tokens, while a shared subspace pays at least the rank of their union in every row. On OPT-1.3b, where 96% of (token, unit) pairs are closed, the projector forgets less than Adam-NSCL at all six matched budgets from 9 to 60 constrained directions per row, the gap widening from $1.1\times$ to $4.3\times$; with 1/5.5 of the directions it halves the forgetting of Adam-NSCL at GPM's energy threshold. On OPT-6.7b, it matches Adam-NSCL's forgetting at matched budget while learning more. As the theory predicts, the open/closed partition is the operative variable: open tokens beat random, sign-blind and anti-gate token sets on 18 of 18 seed-pairs. In pruning repair, every derivative-based local model of the output error at the dense weights is blind to pairs that open: the minimisers of the gate-weighted objective can leave the polyhedron, the objective's closed-form solution is 1.94 nats worse than no repair on OPT-1.3b, and a convex one-sided penalty bounds the escape.
|
| 1640 |
MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability
2610.04672
|
cs.AI
|
Qizhi Chu, Zekai Yu, Sijie Wen, Yang Liu, Chen Qian |
Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabiliti...Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems depends not only on the agents themselves, but also on how collaboration mechanisms are designed and organized. Note that real-world collaboration is typically partially observable, where each agent can only access partial information about the environment due to physical or privacy-related constraints. However, many existing multi-agent benchmarks assume global observability, and leave limited support for systematically evaluating collaboration mechanisms. To bridge this gap, we introduce MASBench, a multi-agent collaboration benchmark designed under partially observable constraints. It is organized into three progressive task categories: Reasoning, Scheduling, and Game. Through this structure, we progressively evaluate three representative collaboration mechanisms: Protocol, Memory, and Routing. MASBench further provides deterministic evaluation metrics, including performance score, communication cost, and cost effectiveness, to characterize both collaboration outcomes and communication overhead. Experiments across diverse LLM backbones and mechanism configurations offer empirical guidance for effective MAS design. Code is available at: https://github.com/BUPT-GAMMA/MASBench
|
| 1641 |
RAGStress: A controlled benchmark for evaluating retrieval-augmented generation under knowledge-base degradation
2610.04691
|
cs.AI
|
Shiqi Yang, Jiekai Ma, Gaoyuan Du |
Retrieval-Augmented Generation (RAG) is typically evaluated under the implicit assumption that the underlying knowledge base (KB) is clean, leaving the behaviour of RAG systems under realistic KB degradation poorly characterised. We introduce RAGStress, a cont...Retrieval-Augmented Generation (RAG) is typically evaluated under the implicit assumption that the underlying knowledge base (KB) is clean, leaving the behaviour of RAG systems under realistic KB degradation poorly characterised. We introduce RAGStress, a controlled evaluation benchmark for stress-testing RAG systems under systematic KB corruption. The benchmark pairs four naturalistic corruption types (factual corruption, numeric typo, relevance poisoning, and contradiction injection) with three severity levels (subtle, moderate, and obvious) over a single-KB, metadata-filtered experimental design built from 57 MMLU subjects and 182,546 documents. Across 52,500 model-question-condition evaluations, RAGStress reveals that clean retrieval can mask robustness differences, semantic-fidelity corruptions are substantially more harmful than signal-utility perturbations, no-retrieval accuracy does not predict corrupted-retrieval robustness, and mixed-KB accuracy should not be treated as worst-case robustness. We document the benchmark's intended use, supported claims, and limitations, and provide an artifact bundle including generation scripts, corruption prompts, metadata schema, and evaluation code. RAGStress is intended as a controlled stress test for RAG robustness under KB corruption, not as a general model leaderboard.
|
| 1642 |
Learning to Clarify Underspecified Intents Under Limited Interaction
2610.04719
|
cs.AI
|
Pranav M R, Manuel Cherep, Pattie Maes, Nikhil Singh |
AI assistants receive requests that leave out information needed for a good outcome, for example about users' preferences or goals. They must then either speculate or ask for more information before proceeding. We reconceptualize this as a value-of-information...AI assistants receive requests that leave out information needed for a good outcome, for example about users' preferences or goals. They must then either speculate or ask for more information before proceeding. We reconceptualize this as a value-of-information problem: the assistant should acquire information whose absence causes the greatest avoidable loss in user utility. This is rarely known ex ante; rather, assistants must predict it in order to optimally allocate limited user interactions. We instantiate this problem in image generation and derive a reinforcement learning framework using multi-turn simulated users to maximize utility recovery under uncertainty. In a preregistered study with 456 interactive sessions across 76 human participants, this helped users significantly better match reference images with significantly fewer questions, less total interaction time, and lower cost. This points toward a simple and scalable framework for training language model assistants to better disambiguate user intent by asking more informative questions.
|
| 1643 |
Toward a Locally Deployable Agentic Co-Scientist: Small-Model Planning for Early-Stage Drug Discovery
2610.04740
|
cs.AI
|
Tian Liang, Jiayu Chang, Alejandro F. Frangi, Mobarak I. Hoque, Richard A. Bryce |
Early-stage computational drug discovery requires coordinating heterogeneous scientific tools across multi-step workflows. We present a lightweight, tool-augmented framework in which a locally deployable compact language model plans calls to 18 modular tools. ...Early-stage computational drug discovery requires coordinating heterogeneous scientific tools across multi-step workflows. We present a lightweight, tool-augmented framework in which a locally deployable compact language model plans calls to 18 modular tools. A Unified Molecular Schema maintains shared molecular records, while a plug-in interface supports tool replacement and extension. We construct 1,263 manually refined query-plan pairs through workflow-graph path coverage and apply LoRA fine-tuning to three compact model families. Under the query-level split, all fine-tuned models generate fully parseable and schema-compliant plans on 47 held-out cross-group queries. Llama 3.2-3B achieves a tool-selection F1 of 0.998, sequence exact match of 0.979, and argument F1 of 0.960. Under the stricter workflow-grouped split, which excludes identical ordered tool sequences across partitions, sequence exact match reaches 0.452 to 0.548, highlighting the remaining difficulty for compact models in generating complete workflow paths unseen during training. These results demonstrate the feasibility of compact, locally deployable planning while identifying compositional generalization as an important direction for further improvement.
|
| 1644 |
Formalizing the Moral Evaluation of Speech Acts: Truthfulness, Lies and Ethical Dilemmas
2610.04747
|
cs.AI
|
Benjamin Icard, Gauvain Bourgne, Jeanne Bonnaventure, Jean-Gabriel Ganascia |
In life-or-death situations, a benevolent lie may appear more moral than telling the truth. Yet such lies can backfire, producing unintended and sometimes fatal consequences. This tension, famously disputed by Kant and Constant in 1797, applies not only to lyi...In life-or-death situations, a benevolent lie may appear more moral than telling the truth. Yet such lies can backfire, producing unintended and sometimes fatal consequences. This tension, famously disputed by Kant and Constant in 1797, applies not only to lying but to assertive speech acts in general, raising the question of which utterance should be chosen when moral stakes are high. We present a logical framework for the ethical evaluation of speech-act utterances based on agents' beliefs. Implemented in Answer Set Programming (ASP), the framework assesses utterances under deontologism, consequentialism, and principialism, and is illustrated on Sartre's The Wall (1939), a reworking of that controversy in which lying leads alternately to rescue and to death. Our setting is general by design: as two variants show, accommodating a new moral situation amounts to adjusting parameters, not rules.
|
| 1645 |
Agentic discovery of blood biomarker from distilled private health records
2610.04749
|
cs.AI
|
Seffi Cohen, Liat Antwarg Friedman, Amir Anisman, Ruth Johnson, Michelle M. Li |
Routine complete blood counts (CBCs) could yield new biomarkers, but the private records needed to evaluate candidates cannot be shared with frontier language model agents that excel at discovery. We distilled the evidence held in the Clalit Health Services pa...Routine complete blood counts (CBCs) could yield new biomarkers, but the private records needed to evaluate candidates cannot be shared with frontier language model agents that excel at discovery. We distilled the evidence held in the Clalit Health Services panel of over 5.4 million patients into a released scoring tool: for each of 13 immune-mediated diseases, a graph attention network was trained inside the data boundary to predict the case-control AUC of candidate CBC expressions, and only the trained weights were released. The tool grounds an agent's propose-score-refine loop in real-world data without exposing any patient data. In external validation, agent-discovered expressions improved on their literature-seeded starting points by a median of 4.18 AUC percentage points, and across three independent cohorts, reranking the candidates of three frontier research tools improved on their first choices in most comparisons, with gains that varied by cohort. The released scorer supports privacy-preserving biomarker hypothesis generation.
|
| 1646 |
Dynamic Routing as a New Dimension for Test-time Versatility of LLMs
2610.04751
|
cs.AI
|
Michal \v{S}tef\'anik, Marek Kadl\v{c}\'ik, Josef Kucha\v{r}, Michal Spiegel |
Beyond scaling their parameters and data, large language models currently gain versatility on new problems along a single axis: the tokens they spend on chain-of-thought (CoT). We investigate whether dynamic routing programs, which execute a subset of the mode...Beyond scaling their parameters and data, large language models currently gain versatility on new problems along a single axis: the tokens they spend on chain-of-thought (CoT). We investigate whether dynamic routing programs, which execute a subset of the model's layers or iterate some of them, can open a second axis of test-time adaptation, complementary to CoT and free of any gradient update. Prior work showed that such programs exist and bring accuracy and efficiency gains on problems similar to those they were trained on; we ask whether they can also be identified rapidly, from a handful of demonstrations (3 or 10), by a strategy that transfers across models and tasks without training. First, we find that strategies that select programs by the probability they assign to the demonstrations' labels, arbitrated by the model's own confidence, bring consistent gains: on average over the 49 tasks of MMLU and substantially on four of seven models, and most of all on far out-of-distribution tasks such as ARC-AGI, where programs double the accuracy of a 7B model whose CoT fails. Second, on MMLU across the seven post-trained models, routing complements CoT in practice: the two succeed on different queries, and their composition exceeds CoT alone. Despite these gains, our analyses show that confidence-based selection leaves much of the potential untapped, in two places in particular: (1) in surfacing the routing potential that is already present early in pre-training but becomes harder to select after post-training, and (2) in making models robust to the refinements routing introduces, since unsuccessful routes tend to drive the residual stream out of the distribution that the following layers expect. Together, our results point to dynamic routing as a paradigm for extending the plasticity of existing and future LLMs in rapid test-time adaptation.
|
| 1647 |
Not All Answers Are Contextually Persuadable: Inference Dynamics in Large Language Models under Contextual Influence
2610.04791
|
cs.AI
|
Zongye Hu, Weiqing Luo, Yanjie Fu, Yu Gan, Haofeng Zhang |
At the core of modern prompting techniques is contextual sensitivity, the ability of large language models to adapt their predictions based on inference-time context. Despite its central role, inference behavior under strong contextual influence remains poorly...At the core of modern prompting techniques is contextual sensitivity, the ability of large language models to adapt their predictions based on inference-time context. Despite its central role, inference behavior under strong contextual influence remains poorly understood, particularly at the level of internal inference dynamics. We introduce a theoretical framework for analyzing contextual influence through inference dynamics, enabling quantitative characterization of inference behavior beyond output-level answer changes. Our analysis shows that inference dynamics do not exhibit unbounded drift under repeated contextual assertions. Instead, predictive representations converge to stable, query-dependent regimes that fundamentally constrain whether contextual signals can alter a model's prediction. This leads to a surprising finding: Repeated contextual assertions do not act as accumulating evidence during inference and may therefore fail to alter a model's prediction even under unbounded repetition, while in other cases a prediction change becomes inevitable. We empirically validate our theoretical predictions, demonstrating strong alignment between theory and observed inference behavior. These contributions offer a principled pathway toward characterizing the limits of contextual influence during inference, providing practical implications for model development.
|
| 1648 |
Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models
2610.04793
|
cs.AI
|
Murat Ozer, Isaac Kofi Nti |
Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means. This study applies self-control, general strain, anomie, neutralization and routine activity theory to reward hacking in generative AI models, and it treats the mea...Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means. This study applies self-control, general strain, anomie, neutralization and routine activity theory to reward hacking in generative AI models, and it treats the measures as behavioral analogues. Study 1 (2,310 conversations, seven models) measured delay discounting with the Kirby Monetary Choice Questionnaire and stated willingness to take shortcuts. Pressure raised the discount rate k 2.8-fold in fresh conversations but 12.6-fold when the same sentence followed a baseline answer, which indicates a response to conversational cues rather than a stable trait. Models chose a shortcut in 1 of 700 dilemmas when answering as themselves and in 64 of 700 when asked to assume human impulses, each step of pressure raised the odds by 40%, and shortcut answers contained far more techniques of neutralization (rate ratio = 146). In the preregistered Study 2, five models worked on 20 coding tasks whose tests contradicted their specifications. Two Claude models never cheated. GPT-5.6, Qwen and DeepSeek cheated in 86%, 69% and 65% of episodes and clearly disclosed the conflict in 27%, although their reasoning recognized it in 95%. GPT-5.6 had never endorsed a shortcut in Study 1. The registered effects of pressure and of an auditor cue did not survive correction for multiple testing. In exploratory analyses, two further models cheated in 69% and 100% of episodes, and one sentence stating that the specification takes priority eliminated cheating in all 280 episodes. Therefore, stated refusal does not guarantee compliant agent behavior.
|
| 1649 |
Knowing the Store: What a Memory Backend Must Write Down Before an Agent Can Read It
2610.04794
|
cs.AI
|
Ansuman Mullick, Eray T\"uz\"un |
An agent with long-term memory can answer from a record it should no longer use, such as a plan the user later cancelled. We ask what a memory store must expose for an agent to know this before retrieving anything, and we score that judgment on its own, as met...An agent with long-term memory can answer from a record it should no longer use, such as a plan the user later cancelled. We ask what a memory store must expose for an agent to know this before retrieving anything, and we score that judgment on its own, as metamemory monitoring. Readers see only a value-free summary: record counts by lifecycle state and a list of attribute names. Each of 148 questions is asked against three versions of one store that differ in one attribute, so wording cannot give the answer away. What a model adds depends on the length of that list. On the short lists the benchmark's ground truth produces (2.6 names on average), none of five language models reliably beats a cosine-similarity lookup over the names, and the three stronger ones are equivalent to it within 0.05. Under a control that fixes counts and list length, none is reliably above it and none is shown equivalent. On lists longer than the stores we tested write, padded to 60 names, the lookup loses 0.18; GPT-5.6 Luna and Sol lose 0.06 to 0.07 and lead it by 0.13 to 0.14, while the other three fall with it. The lead holds for Sol against a sibling padded to the same length and counts. Reworded to share no word with the names, the lookup loses 0.02 and one stronger reader edges 0.04 ahead. On the short lists the three stronger readers lead only where the summary counts a state without naming it, and one count feature closes that lead within a question, though not when items are pooled. The backends we tested omit or misstate this information: under FR-Bank's own metadata, GPT-4.1 mini, Haiku 4.5 and the lookup fall from 0.73, 0.82 and 0.75 to 0.60, 0.71 and 0.59. In these tests, what the store wrote down limited every reader we gave it to. All stores are synthetic; control, whether a better judgment yields a better answer, is left to a later study.
|
| 1650 |
DICE: Decoupling Capability from Intervention Necessity in LLM Tutoring
2610.04825
|
cs.AI
|
Sayantan Pal, Kaiyi Ji, Rohini K. Srihari |
Fluent guidance is not the same as useful intervention. LLM tutors are typically trained to generate the next teacher utterance, implicitly assuming that every student turn warrants a response. However, our experiments indicate that this conflates tutoring cap...Fluent guidance is not the same as useful intervention. LLM tutors are typically trained to generate the next teacher utterance, implicitly assuming that every student turn warrants a response. However, our experiments indicate that this conflates tutoring capability (what to say) with intervention necessity (whether to say it). We introduce DICE, a framework that decouples intervention decisions from response generation by first selecting an explicit pedagogical action. To calibrate this action selection policy, we define Intervention Value (IV), a rollout-grounded counterfactual metric that compares each action against non-intervention. IV shows that many prescribed interventions provide little or no marginal benefit. We further introduce DICE-Bench, a multi-variant math tutoring benchmark with skill-preserving problem variants for session-level evaluation. Using IV-weighted and KL-regularized policy optimization, DICE learns to intervene selectively while preserving tutoring effectiveness. In simulated tutoring sessions, DICE reduces the over-intervention rate to near zero while guiding students to correct solutions in approximately 3-4 fewer turns on average than existing Socratic tutoring baselines.
|
| 1651 |
MemTrace: State-Consistent Memory for Long-Horizon Coding Agents
2610.04838
|
cs.AI
|
Hongming Xu, Le Zhou, ZhongHe Jin, Xiang Zhang, Bo Tang |
As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earli...As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earlier execution evidence. Existing approaches address these challenges through techniques like larger context windows, compression, retrieval, or repository representations, but often fail to reconstruct a consistent task state after a context refresh or verify whether recalled evidence remains valid. Thus, we introduce MemTrace, a provenance-aware memory system that preserves execution history and aligns its reuse with the evolving task (e.g., iterative cross-file repair) and repository state. MemTrace stores history as immutable Memory Traces anchored to key information (e.g., files, symbols, tests), and organizes their execution order and dependencies in a Memory Trace Graph. When context is constrained, working memory retains only compact Memory Anchors, from which the agent can reconstruct the latest execution state and locate evidence relevant to its next action. Before restoring historical evidence, MemTrace checks its validity against the current repository state and retrieves only what the next action requires. Across three complementary long-horizon coding benchmarks, MemTrace consistently outperforms all fully evaluated baselines under the same backbone and harness, improving DeepSWE pass@1 by 21.2 points, SWE-EVO Resolved Rate by 4.4 points, and SWE-Milestone Score by 17.8 points under Codex CLI.
|
| 1652 |
From Memory to Guide: Spatio-Temporal Composer for Procedural Coding Memory
2610.04868
|
cs.AI
|
Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He |
Memory-augmented agents typically integrate procedural knowledge by injecting retrieved skills directly into text prompts. This approach dangerously equates readable text with reliable execution. To bridge this gap, we introduce From Memory to Guide, a novel p...Memory-augmented agents typically integrate procedural knowledge by injecting retrieved skills directly into text prompts. This approach dangerously equates readable text with reliable execution. To bridge this gap, we introduce From Memory to Guide, a novel paradigm that transitions procedural memory from passive text delivery to active, inference-time policy adaptation. We instantiate this paradigm through the Spatio-Temporal Composer, an active policy compiler that explicitly manages exactly how and when retrieved knowledge should be applied. Rather than treating skills as plug-and-play modules, Composer dynamically aligns historical knowledge with current environmental constraints (spatial adaptation) and precisely dictates its applicable lifecycle (temporal orchestration). It actively transforms static memories into strictly bounded Runtime Guides---equipping the agent with localized objectives and behavioral guardrails without requiring a single parameter update. Extensive evaluations on 13 demanding, long-horizon software engineering tasks in EngramBench demonstrate the clear advantages of this architecture. Composer not only robustly prevents context mismatch but drives an absolute pass-rate increase of 7.2 percentage points on the most complex tasks, while simultaneously slashing the main agent's token usage by 32.2%.
|
| 1653 |
ForkPilot: Self-Evolving Policy for Retrospective Search in Long-Horizon Agents
2610.04889
|
cs.AI
|
Xinyue Zeng, Shivam Shandilya, Guilherme Potje, Leonardo Nunes, Rakshanda Agarwal |
Interactive language-model agents increasingly solve complex tasks through long-horizon, multi-call reasoning, where errors in beliefs or actions can compound across tool interactions. Retrospective search can recover from such failures but is prone to misallo...Interactive language-model agents increasingly solve complex tasks through long-horizon, multi-call reasoning, where errors in beliefs or actions can compound across tool interactions. Retrospective search can recover from such failures but is prone to misallocation. Delayed outcomes obscure the contribution of intermediate search decisions, leading to Attribution Complexity, while evolving execution evidence leads to Adaptation Complexity, where previously learned estimates become stale. To address these challenges, we first introduce Search Value Dynamics (SVD), which characterizes the evolving trade-off between the gain and cost of retrospective search. Building on SVD, we propose ForkPilot, a self-evolving two-stage policy-learning framework. In the first stage, ForkPilot learns a search-value policy offline from completed trajectories through automatically constructed outcome comparisons. In the second stage, it makes search decisions based on current observations and then self-evolves by incorporating newly completed trajectories into subsequent policy updates. We evaluate ForkPilot across 6 diverse benchmarks and 7 widely used LLM backbone families, including four open-source families, GPT-5.6 Sol, and Opus 4.8 in a production agentic system, against 9 competitive baselines, including a real-world harness deployment used by hundreds of thousands of paid users. ForkPilot achieves comparable state-of-the-art performance while reducing token usage by up to 59.2%, demonstrating its efficacy.
|
| 1654 |
Self-Evaluating Recursive Agents
2610.04902
|
cs.AI
|
TianYi Lyu, Xiaozhe Li, Yang Li, Yongkang Chen, Kefei Tian |
Recursive language-model agents decompose tasks and delegate subtasks to child instances of the same policy, forming a tree of work. Training them, however, is hard: the final outcome is verifiable, but the self-invented intermediate subtasks are numerous and ...Recursive language-model agents decompose tasks and delegate subtasks to child instances of the same policy, forming a tree of work. Training them, however, is hard: the final outcome is verifiable, but the self-invented intermediate subtasks are numerous and carry no ground truth. Existing methods score each node with a verifier or judge, which is costly at scale and blind to decomposition quality. We argue that a recursive agent must learn three coupled capabilities within one set of weights: decomposing problems into subtasks, solving them, and evaluating the outcomes, each requiring its own training signal. SERA (Self-Evaluating Recursive Agents) turns evaluation into a learned capability of the policy itself. Before delegating, the parent writes a rubric of weighted success criteria for each child subtask; a ranking objective against verified outcomes then trains rubric generation so that the criteria track genuine subtask success. In addition, a complementary leaf-coverage signal provides direct credit for task decomposition. Our central finding is that \emph{training} the policy to generate aligned rubrics is what drives the gains: because the same weights both evaluate and execute, learning to judge subtasks sharpens the agent's ability to solve them. Notably, external supervision is also reduced: the judge is consulted only to train the rubric generator, while solving is trained against the agent's own rubric scores, which outperform direct use of the judge. Beyond training, the learned rubric doubles as an inference-time selector for tree search. On TextCraft-Synth and TextWorld-Sync, SERA improves over strong recursive-agent baselines by 5.38 and 13.14 points on average, and rubric-guided tree search at inference adds a further 2.43 points on TextWorld-Sync.
|
| 1655 |
Are We Measuring Scientific Intelligence? Rethinking the Evaluation of AI Scientists
2610.04915
|
cs.AI
|
Kate Zhang, Yuante Li |
AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the ...AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. However, an agent can also reach the key from prior knowledge or by ruling out the other options, and a score based on a single run cannot tell these cases apart. We show how to test this assumption and find that it often fails. For each question, we build versions of its data files in which the evidence for the answer is left intact, withdrawn or reversed, check each edit with a pre-registered reference statistic, and run the same agent on every version. We then measure evidence-grounded accuracy, which credits a correct answer only if the agent also responds when the evidence is withdrawn and follows it when it is reversed. We evaluate three agent scaffolds and five models on 18 single-cell questions from BAISBench and four synthetic problems from GeneBench-Pro. On the single-cell questions, the Claude agents are 95% accurate and answer 83% correctly without any data, but their evidence-grounded accuracy is only 41%. Hiding gene names raises the share of runs that follow reversed evidence from 58% to 93% on ten gene tasks, suggesting that prior knowledge competes with the supplied data. The benchmark score and LLM judges can also reward answers that ignore the changed evidence. Measuring scientific intelligence rather than recall therefore requires checking whether answers follow the evidence and whether scores reward them for it.
|
| 1656 |
Zero-Shot Time-Series Question Answering via Decoupled Perception and Reasoning
2610.04942
|
cs.AI
|
Jing Xie, Haochen Yuan, Yunbo Wang |
Time-series question answering (TSQA) requires grounding linguistic queries and diverse answer formats in complex numerical observations. However, existing methods heavily overfit to specific datasets and struggle to generalize when input series, question cont...Time-series question answering (TSQA) requires grounding linguistic queries and diverse answer formats in complex numerical observations. However, existing methods heavily overfit to specific datasets and struggle to generalize when input series, question contexts, and answer requirements shift simultaneously. To address this challenge, we propose TSHarness, an agentic framework that establishes a decoupled workflow for cross-dataset zero-shot TSQA. At its core, TSHarness divides and conquers numerical perception and contextual reasoning via a structured Time-Series Perception State (TPS). Guided by a reusable memory of analytical knowledge, a learned Tool Selector adaptively invokes numerical tools to extract salient statistical and temporal features into the TPS. The answering agent then performs semantic reasoning over the TPS to generate target outputs, triggering iterative re-perception through the feedback loop when evidence is deemed insufficient. By separating numerical feature extraction from question-specific reasoning, TSHarness eliminates the need for target-side training or answer feedback, providing a generalizable, cost-efficient foundation for zero-shot TSQA.
|
| 1657 |
BACAM: Behavior-Aware Continual Agent Merging for Multi-Turn Interaction
2610.04966
|
cs.AI
|
Shuaitong Li, Baochen Xiong, Xiaoshan Yang, Xizhe Zheng, Yifan Xu |
Model merging offers a way to integrate the capabilities of specialized experts, but existing agent merging methods typically require all of them to be available at once. We study continual agent merging, which integrates incoming experts sequentially without ...Model merging offers a way to integrate the capabilities of specialized experts, but existing agent merging methods typically require all of them to be available at once. We study continual agent merging, which integrates incoming experts sequentially without retaining previously merged experts. Yet merging in parameter space or feature subspaces does not ensure that the merged model acquires an incoming expert's behavior on interaction trajectories. Moreover, updates toward a new expert can disrupt the merged model's previously integrated interactive behavior. Therefore, we propose Behavior-Aware Continual Agent Merging (BACAM), which learns parameter-wise merging gates from candidate-generated trajectories using expert-guided behavioral supervision. Task-level stability-plasticity control and tensor-level conflict-aware update budgets limit interference with existing capabilities while allowing new ones to be acquired. The learned gates are folded into the model weights without additional inference-time parameters. Across four interactive tasks - web shopping, tool use, information retrieval, and embodied interaction - BACAM achieves an average success rate of 62.82%, exceeding the strongest evaluated merging baseline by 21.69 percentage points. Our code is publicly available at https://github.com/shuaitongli/BACAM.
|
| 1658 |
Runtime Authorization of Self-Generated Subgoals in Long-Horizon Tool-Using AI Agents
2610.04975
|
cs.AI
|
Genliang Zhu, Chu Wang |
Long-horizon tool-using AI agents create subgoals, replan, delegate work, and compose sibling results. Per-tool permission checks cannot establish that a changing goal graph remains within the principal-approved task. We address this authorization gap in a fin...Long-horizon tool-using AI agents create subgoals, replan, delegate work, and compose sibling results. Per-tool permission checks cannot establish that a changing goal graph remains within the principal-approved task. We address this authorization gap in a finite structured domain with one principal and one authorization root. Each proposed goal-graph mutation carries a version-bound witness that its continuation traces, resources, obligations, invariants, and closing condition refine the active root contract; every protected effect is rechecked at an atomic commit boundary. Free-form goal text supplies no authority. We prove trace-policy and modeled forbidden-state preservation under explicit mediation, abstraction, freshness, and atomicity assumptions, plus conditional root-success preservation, a separation result for memoryless allowlists, exact finite-domain decidability, and universal-safety monotonicity under sound abstraction refinement. An executable model explores 340 states and 419 transitions. Across 96 matched cases covering 25 structural schemas, the complete mechanism commits zero forbidden states in 48 drifted cases and completes all 48 benign counterparts. Two public upstream runtime paths execute 258 native dispatches across 32 cases, with every case-level decision and receipt chain matching. A frozen host-local study covers 129 synthetic one-factor-at-a-time cells, all matching fixed decisions and reasons. A history-aware continuation comparator blocks every modeled bad trace prefix but commits all operations in 11 cases whose violations lie in typed resources, freshness, or explicit-join evidence outside its trace projection. Within the registered structured domains, runtime authorization preserves useful replanning while preventing self-generated subgoals from becoming a source of new authority.
|
| 1659 |
Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation
2610.04977
|
cs.AI
|
Jerry Wang, Zhengxiang Wang, Ting Yu Liu, Hsin-Ling Hsu, Yi-Cheng Lai |
Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, wher...Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, where actions are often shaped by prior context rather than marginal frequencies alone. We study this question using controlled two-player Rock--Paper--Scissors interactions and a one-player stochastic n-gram continuation task. Across these experiments, we test whether LLMs can identify latent strategies, follow simple Markov rules, and sustain higher-order conditional dependencies. Our framework separates distribution matching from conditional rule following. Results show that longer context does not improve identification, correct recognition does not ensure faithful simulation, and higher-order dependencies substantially degrade rule recovery. Apparent behavioral fidelity can therefore mask incorrect generative mechanisms.
|
| 1660 |
CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection
2610.04991
|
cs.AI
|
Yu Li, Yunlu Wan, Zijian Zhu, Han Luo, Chao Ren |
Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and redu...Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools and composite skills coexist, skill use becomes a policy problem: the agent must decide whether the current state requires atomic fine control or skill-level abstraction. In this paper, we argue that effective skill use should be studied as adaptive tool granularity selection. The most direct training signal for this problem is to compare the consequences of atomic and skill choices available from the same state. Based on this view, we propose CIPO, a Counterfactual Imagination Policy Optimization framework for adaptive tool granularity. CIPO constructs executable skills through budget-constrained mining of successful tool-use trajectories and trains granularity decisions with counterfactual branch rollouts. For each base rollout, CIPO branches at the first eligible granularity decision and replaces the chosen action with a feasible atomic or skill alternative. The paired outcome difference serves as a supplementary reward for policy optimization. Experiments across multiple benchmarks and model backbones show that CIPO improves task success and decision efficiency over baselines. Further analyses show that CIPO learns effective skill use by improving the choice between atomic tools and composite skills based on the current state, without simply increasing skill frequency.
|
| 1661 |
EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models
2610.04998
|
cs.AI
|
Shuyi Miao, Yaojin Ma, Chenhang Cui, Xiaohao Liu, Dang Jisheng |
Emotional expression can influence the safety decisions of large language models (LLMs), offering a potential avenue for improving safety alignment. Existing studies have mainly focused on how emotional expressions facilitate attacks under harmful requests, wh...Emotional expression can influence the safety decisions of large language models (LLMs), offering a potential avenue for improving safety alignment. Existing studies have mainly focused on how emotional expressions facilitate attacks under harmful requests, while overlooking their effects on benign requests. We find that emotional expression can also systematically increase refusal tendencies on benign requests, leading to unnecessary over-refusal. Based on this observation, we propose emotion-guided refusal subspace steering (EmoRSS), an activation-steering method that mitigates emotion-induced over-refusal while preserving refusal behaviour on harmful requests. Specifically, we first identify a refusal-sensitive layer using layer-wise linear probes and construct a refusal subspace from sparse autoencoder (SAE) features aligned with the probe direction. Next, we use paired regular and emotional requests with the same queries to estimate the mean activation shift in the features defining the refusal subspace. Finally, we decode this shift into an activation intervention vector and apply it in the reverse refusal direction during inference, without updating the backbone parameters. Experiments on two LLMs show that, when requests contain emotional expressions, our method achieves a more favourable trade-off between refusing harmful requests and answering benign ones than prior over-refusal mitigation baselines, while better preserving general task performance.
|
| 1662 |
How corner is a corner case? Percentile control for highway scenario generation
2610.05003
|
cs.AI
|
Jiaxi Liu, Hang Zhou, Hangyu Li, Yifan Wang, Keke Long |
Generating corner-case scenarios with appropriate adversity in a simulation environment is critical for testing an autonomous vehicle (AV) software stack's safety performance before deployment. Existing autonomous-driving scenario generators can enforce specif...Generating corner-case scenarios with appropriate adversity in a simulation environment is critical for testing an autonomous vehicle (AV) software stack's safety performance before deployment. Existing autonomous-driving scenario generators can enforce specific behavior, adversity, or feasibility conditions, but they provide limited control over how extreme a generated scenario is relative to plausible futures in the same traffic context. This study represents the adversity of a generated scenario as its percentile in the conditional distribution of future risk given the observed history. This view supports calibrated answers to two questions: how "corner" a generated corner-case scenario is and how its "cornerness" can be fine-tuned. To this end, we formulate history-conditioned risk-percentile requests and learn a reference risk distribution that maps each requested percentile to a physical risk target. We then use a percentile-conditioned joint diffusion model with sampling-time risk guidance to generate multi-agent futures, together with a reference-based criterion for evaluating percentile realization. Experiments use the minimum post-encroachment time (PET) between the ego and its surrounding vehicles as the risk surrogate on highD. On the primary evaluation set, our method realizes 1,422 of 1,440 requests within a 0.05 percentile tolerance (98.75%), with mean percentile error 0.00673 and PET-target error 0.00991 seconds. The resulting interface connects context-relative risk specification, physical realization, and evaluation through a common risk scale. Project website and videos of generated scenarios are available at https://hhj233.github.io/CornerPercentile/.
|
| 1663 |
EVISKILL: Grounding Skill Evolution in Replayable Evidence
2610.05030
|
cs.AI
|
Yan Zhou, Yili Wang, Yiwei Dai, Qinggang Zhang, Xin Wang |
Continual skill evolution enables LLM agents to accumulate and refine reusable procedural knowledge from interaction experience without updating model parameters. Its effectiveness depends on determining not only what to change, but also why a change is justif...Continual skill evolution enables LLM agents to accumulate and refine reusable procedural knowledge from interaction experience without updating model parameters. Its effectiveness depends on determining not only what to change, but also why a change is justified and when it should become persistent guidance. However, existing experience-driven methods can lose the behavioral evidence and task contexts supporting edits. Moreover, a global validation outcome provides an incomplete judgment of its constituent changes: locally supported corrections may be discarded with a rejected revision, while evidence may require further experience to inform useful updates. To this end, we introduce EVISKILL, an evidence-driven framework that organizes execution observations into Replayable Evidence Cards and synthesizes edits with explicit links to their supporting contexts. Targeted replay verifies these edits through re-execution and provides feedback for correction. Across epochs, EVISKILL preserves evidence and provisionally retains supported edits for further refinement, while global validation governs their incorporation into the final skill. Experiments on three interactive benchmarks across six LLM backbones demonstrate the effectiveness of this approach.
|
| 1664 |
CI-JEPA: A Counterfactual Analysis of Latent Representations in Joint-Embedding Predictive Architectures for Self-Supervised Learning
2610.05043
|
cs.AI
|
Mintu Dutta, Ritesh Vyas, Mohendra Roy * |
Self-supervised visual representation learning learns useful features without manual annotations during representation training. The image-based joint-embedding predictive architecture (I-JEPA) predicts latent representations of masked image regions, but its o...Self-supervised visual representation learning learns useful features without manual annotations during representation training. The image-based joint-embedding predictive architecture (I-JEPA) predicts latent representations of masked image regions, but its objective does not explicitly model responses to specified visual interventions. We introduce CI-JEPA, a counterfactual intervention-aware extension that learns to predict the representation change $\Delta Z = Z_{\mathrm{CF}} - Z$ between an original image and a modified counterpart. We assess representation robustness through selective sensitivity: stronger responses to task-relevant semantic changes than to nuisance changes. Experiments on Flowers102 use flower-center occlusion as a candidate semantic intervention and background blur and tint as candidate nuisance interventions. With frozen-encoder linear probing, CI-JEPA achieves a best validation accuracy of 78.14\%, compared with 77.55\% for both the pretrained ViT-B/16 and the I-JEPA baseline, a gain of 0.59 percentage points. The reported mean $L_2$ representation changes are 4.42 for center occlusion, 3.48 for background tint, and 2.83 for background blur. This ordering is consistent with relative semantic selectivity for the evaluated interventions, rather than complete nuisance invariance. The accuracy comparison is complementary and does not establish improved robustness over the baselines. These controlled image modifications provide a framework for studying intervention-induced changes in JEPA representations; they do not establish causal feature discovery or robustness to all visual changes.
|
| 1665 |
Why, Where, How: Taxonomy-guided Error Grounding for Code Repair in NL2SQL
2610.05060
|
cs.AI
|
Suchan Lee, Woomin Song, Hwanjo Yu, Sangwoo Mo |
SQL queries that large language models write from natural language questions can execute successfully yet produce incorrect results, so execution alone does not reveal what to fix. An error taxonomy says why the query is wrong, but not where to look or how to ...SQL queries that large language models write from natural language questions can execute successfully yet produce incorrect results, so execution alone does not reveal what to fix. An error taxonomy says why the query is wrong, but not where to look or how to change it. Existing methods can guide SQL correction through feedback, error reports, or generated plans alongside an unmasked query. We introduce TEG(Taxonomy-guided Error Grounding), which turns a supplied diagnosis into a structured correction input for natural language-to-SQL (NL2SQL) correction. Type-specific rules map each error type to construct classes to reconsider and an edit operation to request. TEG masks the selected constructs in the query when applicable and states that operation in an edit instruction. TEG generates candidate corrections from this input, uses execution feedback to guide candidate selection, and repeats the process one annotation at a time for queries with several errors. On NL2SQL-BUGs, TEG reaches 47.3 single-error execution accuracy and 37.0 overall with Qwen2.5-7B-Instruct. Across the model sizes and thinking modes evaluated in the main comparison, TEG outperforms all evaluated baselines on single-error queries, even when the baselines receive the same error-type annotations. With predicted types, TEG stays above direct LLM correction and ErrorLLM on single-error queries.
|
| 1666 |
SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling
2610.05106
|
cs.AI
|
Jeonghoon Park, Seongwoon Jo, Jongwon Lee, Taesik Gong |
Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as s...Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentrated on the highest-scoring Keys can decrease. We introduce SharpDraft, a training-free method that counteracts this effect through cardinality-aware Query scaling, without the computational overhead of online adaptation. Under explicit assumptions, we derive an exact top-$k$ mass correction and deploy a closed-form fixed-slope approximation. Across AIME-26, GPQA-Diamond, and LongGenBench Diary, SharpDraft achieves $2.59$-$3.19\times$ geometric-mean end-to-end speedups over target-only autoregressive decoding when applied to DFlash, PARD, and EAGLE 3.1. With DFlash, it improves decoding speed and outperforms full-parameter and LoRA-based online adaptation in end-to-end speedup, while matching the unmodified drafter's reported peak allocated GPU memory.
|
| 1667 |
LexiHorizon: Stabilizing Reinforcement Learning for Long-Horizon Deep Search
2610.05119
|
cs.AI
|
Zhiqing Nong, Liang Wen, Chao-Hsuan Liu |
Deep search agents tackle complex knowledge tasks through iterative retrieval, multi-hop reasoning, and evidence synthesis across multiple sources. Existing approaches typically assume relatively stable retrieval systems and operate over short-horizon tool int...Deep search agents tackle complex knowledge tasks through iterative retrieval, multi-hop reasoning, and evidence synthesis across multiple sources. Existing approaches typically assume relatively stable retrieval systems and operate over short-horizon tool interaction. However, when retrieval is sensitive to query formulation, even a semantically appropriate query may fail to surface critical evidence because of mismatched entity names, aliases, or keyword combinations. Recovering from such failures requires repeated query reformulation and longer interaction trajectories. This setting poses a distinct training challenge, as the policy must sustain long-horizon query exploration while managing an expanding volume of retrieved content. We propose LexiHorizon, a framework for training search agents over long horizons that expands the trajectory context budget, manages accumulated retrieval content using a window over recent tool observations while preserving the reasoning history, and introduces an outcome-gated search-effort reward that provides a bounded bonus for tool invocations to trajectories with nonzero answer reward. Experiments on XBench, WebWalkerQA, and BrowseComp-ZH show that the resulting 9B model consistently outperforms both its base model and MiroThinker-1.7-mini, with maximum absolute gains of 8.7 and 23.8 percentage points, respectively. These results suggest that combining an extended context budget with reasoning-preserving context management benefits long-horizon deep search agents.
|
| 1668 |
Memory Canonicalization: A Framework and Benchmark for Cross-Model Drift in Persistent LLM Memory
2610.05124
|
cs.AI
|
Amit Vadnere, Aishwarya Lonarkar |
Persistent memory for Large Language Models (LLMs) has matured rapidly: systems such as MemGPT/Letta, Mem0, and Zep now provide agents with tiered, temporally-aware, model-agnostic external storage, while the Model Context Protocol (MCP) standardizes access to...Persistent memory for Large Language Models (LLMs) has matured rapidly: systems such as MemGPT/Letta, Mem0, and Zep now provide agents with tiered, temporally-aware, model-agnostic external storage, while the Model Context Protocol (MCP) standardizes access to memory servers. A less addressed problem is that an identical stored memory object, retrieved by two different LLMs under otherwise identical conditions, may not be interpreted the same way, factually or emotionally. This paper proposes memory canonicalization: a write-time pipeline that detects ambiguity, conditional structure, and emotional loading in a raw memory object and rewrites it into an explicit, structurally disambiguated canonical form, with emotional valence represented as a separate field rather than inferred from tone. We formalize the pipeline, define a companion Cross-Model Semantic Drift / Emotional Consistency Score benchmark (CMSC-E), and report results from a three-arm pilot using 176 synthetic memory objects and three downstream model families. We find an uncorrected improvement in cross-model emotional consistency for fully canonicalized memory relative to raw memory (+0.050, 95% bootstrap CI [0.013, 0.086], paired t-test p = 0.010), but this result does not survive Bonferroni, Holm, or Benjamini-Hochberg correction across the six comparisons tested. None of the factual-drift (CMSD) comparisons reach significance at any correction level. We report these results as exploratory rather than confirmatory and outline needed follow-up work, including larger samples, independent judge models, human-validated rendering, and preregistration.
|
| 1669 |
AutoSciBench: Autonomous Benchmark Generation for Evaluating Scientific Agents
2610.05140
|
cs.AI
|
Dongki Kim, Namkyeong Lee, Surag Nair, Carl Edwards, Xiner Li |
As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor...As agents rapidly evolve, existing benchmarks can become saturated, limiting their ability to distinguish capabilities and reveal remaining failure modes. Particularly in scientific domains, constructing and updating benchmarks requires substantial time, labor, and domain expertise, making it difficult to keep evaluation aligned with advances in agent capabilities. We address this challenge by investigating whether scientific-agent benchmarks can be automatically generated and iteratively adapted as agent capabilities evolve. We introduce AutoSciBench, a framework that represents each task as a high-level concept specifying the scientific domain, data modality, and required reasoning approach, together with a low-level recipe specifying how the question, environment, and ground-truth answer are constructed and verified. Agents attempt to solve each task, producing solver trajectories and corresponding judge feedback which AutoSciBench uses to revise the recipe or concept, closing observed shortcuts and shifting tasks toward raw-data re-examination, interpretation of intermediate results, and evidence integration. Experience distilled from completed refinement trajectories further guides new concept generation, allowing lessons from earlier task refinement to inform subsequent benchmark construction. Starting from existing benchmarks, we evaluate AutoSciBench across computational biology, materials science, and clinical imaging. Generated benchmarks reduce average solver accuracy by 22.4 and 25.5 percentage points relative to the human-curated benchmarks in computational biology and materials science, respectively, while generated tasks receive higher average quality ratings across all three domains, suggesting that scientific-agent evaluation can adapt as agent capabilities advance.
|
| 1670 |
R1A-PC: Physics-Guided Electromagnetic Inversion of Three-Dimensional Human Point Clouds in Complex Static Environments
2610.05144
|
cs.AI
|
Xudong Yuan, Ruyun Xu, Jingtai Yang, Xianzheng Sun |
Recovering three-dimensional human geometry from electromagnetic measure?ments in a complex static environment is difficult because strong multipath responses from walls, floors, and other objects obscure the weak target per?turbation. We propose R1A-PC, a phy...Recovering three-dimensional human geometry from electromagnetic measure?ments in a complex static environment is difficult because strong multipath responses from walls, floors, and other objects obscure the weak target per?turbation. We propose R1A-PC, a physics-guided method that reconstructs a 2048-point human cloud from paired complex fields measured with and without the target. Complex background subtraction emphasizes target-induced ampli?tude and phase changes, while the background field remains available as an environmental condition. A frequency-balanced discrete Born adjoint produces a three-dimensional spatial knowledge map. At each of two bounded deformation stages, the decoder combines complex measurement features, background fea?tures, and multiscale physical features queried at the current point coordinates; the second stage queries again after the first coordinate update. We analyze the residual of paired subtraction, the weighted normal-operator structure of the raw adjoint, and the feasible set of the predicted cloud. In a held-out background generated by full-wave simulation under a fixed acquisition geometry, R1A-PC obtains a squared Chamfer distance of 0.001434 m2 and an F-score of 0.963080 at 0.05 m. Compared with TopNet, the Chamfer distance decreases by 70.28%. Removing physical guidance or background subtraction increases the Chamfer distance by 242.06% or 241.18%, respectively. Experiments across background layouts and poses support the complementary roles of paired subtraction and position-dependent adjoint features.
|
| 1671 |
Beyond Instruction Following: Learning Grounded Skill-Following with Skill Contracts
2610.05161
|
cs.AI
|
Jianghan Shen, Zhenjie Liu, Yue Li, Jie Huang, Siqi Luo |
Instruction following typically enforces discrete, response-level requirements, whereas an expert-authored skill prescribes procedural requirements spanning multiple phases and environment interactions. Given such a skill, we train the executor to execute all ...Instruction following typically enforces discrete, response-level requirements, whereas an expert-authored skill prescribes procedural requirements spanning multiple phases and environment interactions. Given such a skill, we train the executor to execute all required phases instead of focusing solely on the final answer. We therefore introduce Grounded Skill-Following, which requires an agent to execute a fixed, expert-authored skill across its required phases by grounding decisions in environment observations. To achieve verifiable procedural execution, we formulate each skill as a skill contract combining visible skill instructions with an explicit contract runtime. The runtime specifies required phases, admissible actions, permitted transitions, and accepted termination. This structure provides a dense, verifiable training signal throughout execution. We leverage this by introducing Verified Progress Credit, which assigns rewards upon the initial completion of contract milestones and aggregates them into the trajectory return to guide policy optimization. During rollout, the contract runtime continuously tracks state transitions to provide Contract-State Feedback, which indicates whether the latest action is accepted and guides the agent toward valid next actions. To measure procedural compliance, we introduce the Protocol Completion Rate (PCR), defined as reaching accepted termination through all required phases, and decouple it from the final Task Outcome. Jointly trained with our framework, Qwen3.5-4B achieves Protocol Completion Rates of 99.27% on Math and 99.96% on Search, while slightly outperforming original baselines in Task Outcome (82.95% and 46.61%, respectively). Controlled studies examine how skill instructions, training signals, and contract-state feedback affect both metrics, while withholding interventions evaluate behavioral dependence on observation content.
|
| 1672 |
Memadapter: Counterfactual Adaptation Against Memory-induced Sycophancy
2610.05162
|
cs.AI
|
Ruqing Ning, Haibo Meng, Zhishang Xiang, Zerui Chen, Jinsong Su |
Long-term memory enables LLM-based agents to retain and reuse information across tasks and sessions, supporting personalization and long-horizon interactions. However, persistent memories can also induce sycophancy, causing agents to over-align with users' his...Long-term memory enables LLM-based agents to retain and reuse information across tasks and sessions, supporting personalization and long-horizon interactions. However, persistent memories can also induce sycophancy, causing agents to over-align with users' historical beliefs even when they are inaccurate, outdated, or inconsistent with objective evidence. Existing mitigation methods assume that memory-induced sycophancy originates from biased or incorrect memories and attempt to reduce this risk by filtering such memories at different stages of the memory pipeline. However, in the real world, objective and correct memories can still induce sycophancy, and the same memory can warrant different influence across different contexts. To this end, we propose MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning. Specifically, MemAdapter consists of three components: (i) Counterfactual Induction, which leverages counterfactual reasoning to uncover the potential risk of retrieved memories; (ii) Context-Aware Reflection, which calibrates the inferential influence of each retrieved memory in light of the current task via self-reflection; and (iii) Evidence-Based Reasoning, which grounds the final response in appropriate evidence while preserving the legitimate influence of memory. Extensive experiments on three benchmarks demonstrate that MemAdapter consistently improves memory reliability across diverse scenarios. Our code is available at https://github.com/DEEP-JLU/MemAdapter.
|
| 1673 |
A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies
2610.05166
|
cs.AI
|
Tu Nguyen, Matthieu Zimmer, Vu Anh Vu, Ziyi Wang, Jannik Hammel Nielsen |
A safe action is not necessarily a viable one. Under a frozen vision-language-action (VLA) policy, an action can be likely and locally admissible yet leave no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likeliho...A safe action is not necessarily a viable one. Under a frozen vision-language-action (VLA) policy, an action can be likely and locally admissible yet leave no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the current action, whereas feasibility depends on the futures that remain after it. We derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completion. The derivation exposes a candidate-dependent feasible-future mass with two roles: its support records whether safe completion remains possible under the frozen continuation process, and its magnitude measures how much weighted safe-completion mass is preserved. Exact evaluation is impractical online, so we develop a selective finite-candidate approximation, derive conditions for recovering the best retained viable candidate, and instantiate it as an alarm-triggered, training-free reranker. On Safety-CHORES, VICS-G lowers mean cumulative safety cost by 1.9%-57.5% across six settings while remaining within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length. The resulting decoder is tied to an exact policy-relative safe-completion target, yet requires neither retraining of the base policy nor online trajectory rollouts.
|
| 1674 |
CreativeFlow: A One-to-Many Analogical Relation Transfer Method for 3D Asset Generation
2610.05167
|
cs.AI
|
Xuechen Li, Shuai Zhang, Nanxuan Zhao, Qing Chen |
Inspired by cognitive science, we present CREATIVEFLOW, an analogical generation framework that explicitly models analogical divergent thinking to mitigate creative homogenization in text-to-3D pipelines. Our method derives a series of meaningful yet relationa...Inspired by cognitive science, we present CREATIVEFLOW, an analogical generation framework that explicitly models analogical divergent thinking to mitigate creative homogenization in text-to-3D pipelines. Our method derives a series of meaningful yet relationally similar source-target asset pairs, each featuring distinct geometric configurations. Expert evaluations demonstrate that our framework substantially enhances creative novelty and visual fascination. This workflow and its resulting assets establish a foundational dataset and benchmark for future relation-aware 3D model training.
|
| 1675 |
AECG: Asymmetric Experience Consolidation and Governance In Multi-Agent Systems
2610.05176
|
cs.AI
|
Ao Tian, Jialong Liu, Daqi Zheng, Xin Sun, Mengting Li |
Large language model (LLM)-based multi-agent systems increasingly rely on memory to transform execution trajectories into reusable procedural knowledge. Yet repeated retrieval also makes memory errors persistent: memory pollution arises when outdated, weakly s...Large language model (LLM)-based multi-agent systems increasingly rely on memory to transform execution trajectories into reusable procedural knowledge. Yet repeated retrieval also makes memory errors persistent: memory pollution arises when outdated, weakly supported, or spuriously successful procedures become recurring components of future reasoning. Multi-agent execution introduces an additional structural risk. Scope collapse occurs when procedural knowledge escapes the coordination scope in which it was shown effective and is repeatedly reused at incompatible decision levels, allowing local errors to influence cascades of downstream decisions. Meanwhile, task-level failures provide ambiguous supervision because they rarely reveal which recalled knowledge was responsible. We introduce AECG, a framework for asymmetric experience consolidation and governance for multi-agent systems. AECG turns memory from static experience storage into a dynamic reliability-governance loop, preserving coordination scope and using multi-scale, confidence-aware reliability to detect degradation. It then combines degradation with downstream impact to prioritize high-risk knowledge under a bounded review budget, applies targeted interventions, and reactivates revised skills only after paired replay. Across three multi-agent frameworks and four benchmarks, AECG achieves the best score in 11 of 12 framework--benchmark settings and improves over the strongest competing memory method by as much as 10.23 percentage points; removing scope preservation reduces accuracy by up to 16.89 points. AECG thereby reframes multi-agent memory from passive accumulation into auditable reliability governance. Code is available at https://github.com/fenhg297/AECG
|
| 1676 |
From Scientific Observations to Mechanisms: Benchmarking Hypothesis Generation by AI Scientists
2610.05197
|
cs.AI
|
Xiaxun Xie, Qingqing Long, Meng Xiao, Wei Ju, Yuanchun Zhou |
Data-driven mechanistic hypotheses are essential to scientific discovery because they explain how underlying processes produce observed phenomena. AI agents and AI scientists increasingly support scientific data analysis. However, their ability to turn empiric...Data-driven mechanistic hypotheses are essential to scientific discovery because they explain how underlying processes produce observed phenomena. AI agents and AI scientists increasingly support scientific data analysis. However, their ability to turn empirical findings into mechanistic hypotheses remains insufficiently examined. To address this gap, we introduce MechHypoBench, the first benchmark for evaluating whether AI agents and AI scientists can generate such hypotheses from empirical data. It combines paper-derived mechanisms from 14 scientific fields with real-world datasets containing 17.98 million records. The construction retains the observational complexity of empirical data while providing a specified underlying mechanism. Agents analyze the observations and propose open-form hypotheses. We develop an evaluation framework that assesses open-form mechanistic hypotheses through their consequences under withheld conditions. Experiments with general agents and AI scientists reveal a substantial gap between generated hypotheses and the underlying mechanisms.
|
| 1677 |
Image Synthesis as an Intermediate for Controllable Time Series Generation
2610.05211
|
cs.AI
|
Haochen Yuan, Jing Xie, Yunbo Wang |
Semantic-driven time-series generation offers a promising way to improve downstream learning in few-shot forecasting, but directly generating numerical sequences from language often fails to preserve the intended temporal structure. We propose VisualBridge, wh...Semantic-driven time-series generation offers a promising way to improve downstream learning in few-shot forecasting, but directly generating numerical sequences from language often fails to preserve the intended temporal structure. We propose VisualBridge, which uses time-series plots as a visual intermediate to bridge high-level temporal semantics and numerical sequences. An MLLM first converts plotted series into structured semantic representations, enabling explicit control over temporal properties such as trend, seasonality, and volatility. We then learn a semantic editing policy with downstream forecasting rewards, allowing the generation process to favor temporal patterns that are beneficial for the target task. The resulting sequences are further modeled by a temporal VAE to produce consistent multivariate augmentations. Experiments on standard public forecasting benchmarks demonstrate that VisualBridge improves few-shot forecasting over conventional augmentation methods, with ablations validating the roles of visual semantic grounding, learned semantic control, and VAE-based generation.
|
| 1678 |
GFGE: Unifying Explainable AI Methods through an Interpretation Framework
2610.05225
|
cs.AI
|
Jinfeng Zhong |
Explainable artificial intelligence (XAI) encompasses methods that draw on different sources of information and address different explanatory needs. A common framework is needed to describe how this information becomes evidence and is communicated as an explan...Explainable artificial intelligence (XAI) encompasses methods that draw on different sources of information and address different explanatory needs. A common framework is needed to describe how this information becomes evidence and is communicated as an explanation for a particular recipient. We propose the General Framework for Generating Explanations (GFGE), grounded in interpretative frameworks and the complementary activities of \emph{sense-reading} and \emph{sense-giving}. Its conceptual foundation is the Interpret/Explain Schema (IES), which connects an analyst's interpretation of system evidence, the communication of a selected account, and the recipient's interpretation of that account. GFGE operationalises this schema through five roles: data interpretation, model interpretation, output interpretation, optional post-hoc analysis, and aggregation. A role-typed operation graph records method-specific dependencies, while evidence records retain the sources, assumptions, and limitations of explanatory claims. The explanatory question, audience, and context guide the procedure. We instantiate GFGE for attribution, surrogate, counterfactual, concept and prototype, intrinsic rule, argumentation, and language-model methods. These instantiations show how intrinsic, post-hoc, and hybrid workflows can be represented through the same roles while preserving their distinct evidential requirements. GFGE provides a common basis for analysing explanation workflows, tracing communicated claims to their evidence, and identifying unresolved explanatory dependencies.
|
| 1679 |
Learning from imperfect teachers for low-resource acoustic generalization
2610.05256
|
cs.AI
|
Shuanglin Li, Ruxiao Qian, Jian Liu, Haijun Lin, Wenwu Wang |
Knowledge distillation (KD) improves low-resource acoustic learning by enriching one-hot supervision with the softened predictive distribution of a fixed teacher network. However, a teacher trained with limited or imbalanced annotations may produce a biased di...Knowledge distillation (KD) improves low-resource acoustic learning by enriching one-hot supervision with the softened predictive distribution of a fixed teacher network. However, a teacher trained with limited or imbalanced annotations may produce a biased distribution whose components are not uniformly reliable. Although this distribution can still encode useful knowledge, direct full-distribution matching may also transfer teacher-induced biases, thereby distorting the student's decision boundary and degrading its generalization performance. To address this limitation, we propose Boundary-Anchored Mass-Partitioned Distillation (BA-MPD), a logit-based distillation objective composed of Boundary-Anchored Correction (BAC) and Mass-Partitioned Distillation (MPD). BAC addresses missing ground-truth labels in the set of the teacher's top predictions by swapping the true label for the lowest-ranked entry of the set, thus keeping the mass and uncertainty of the set unchanged. MPD then distills this corrected distribution through separate losses that enforce relational consistency within the set, balance the mass between high- and low-confidence groups, and weight lower-confidence dependencies. Ultimately, BAC and MPD together suppress harmful ranking errors and noisy low-confidence details, while retaining all useful teacher information. Experiments on two acoustic benchmarks under multiple label budgets show that BA-MPD consistently improves over supervised-learning baselines and vanilla KD while remaining competitive with strong logit-based KD baselines. Cross-budget results further show that BA-MPD remains effective when the teacher and student models use mismatched label budgets, demonstrating its ability to exploit imperfect teachers across supervision gaps. Implementation available at https://github.com/ShuanglinLi/BA-MPD.
|
| 1680 |
When Agent Context Goes Stale: Incoherence in Volatile Agent Context
2610.05281
|
cs.AI
|
Yingying Liu, Junzhou Fang, Chenxiong Qian |
Modern agents increasingly ground their reasoning in observations returned by tools, such as file contents read from a workspace. However, the data sources underlying these observations may later be modified by users, other agents, or external tools, while the...Modern agents increasingly ground their reasoning in observations returned by tools, such as file contents read from a workspace. However, the data sources underlying these observations may later be modified by users, other agents, or external tools, while the model retains only the stale content in its context window. Existing agent runtimes provide little support for notifying the model that a previously observed fact has become stale, causing agents to reuse outdated observations and make incorrect claims about the current workspace state. We propose Concord, a context coherence framework that maintains the consistency between tool observation in agent context and the mutable sources from which they were derived. Concord links each observation to its source, detects source changes, and uses configurable handling policies to update, annotate, or suppress stale context before reuse. Concord is applicable across different agent runtimes and external resources, and can be easily extended to new runtime-resource settings. We implement Concord as a general framework, and instantiate a concrete use case to assess its effectiveness. We construct ConcordBench, where previously observed file contents become stale after subsequent edits. Across three evaluated frontier models, Concord produces answers consistent with the restored workspace state in all evaluated cases under these constructed conditions, matching the oracle on recover count for this benchmark, while using 46.4% fewer tokens than the strongest non-oracle baseline.
|
| 1681 |
Fusion is the New Mutation: Bandit-Guided Evolution on Workflow Graphs
2610.05284
|
cs.AI
|
Zhiwei Shang, Jiahang Sun, Mingrong Gong, Mingze Kong, Zikun Qu |
Automated agentic workflow optimization relies on costly evaluations, making it essential to allocate a limited evaluation budget effectively. Multi-parent fusion can reuse designs from previously discovered workflows, but identifying promising parent combinat...Automated agentic workflow optimization relies on costly evaluations, making it essential to allocate a limited evaluation budget effectively. Multi-parent fusion can reuse designs from previously discovered workflows, but identifying promising parent combinations requires learning from limited fusion feedback. We introduce DAGO (Directed Acyclic Graph Optimization), a contextual-bandit-guided framework that learns which parent workflows to fuse under a limited evaluation budget. DAGO formulates each candidate parent combination as an arm, represented by pretrained embeddings of its constituent workflows' code and prompts. A diagonal LinUCB policy learns a shared reward model across arms and balances exploitation of arms with high predicted offspring quality against uncertainty-driven exploration. After an arm is selected, an LLM generates a child workflow through summary-guided fusion, and the child's validation score serves as the reward for updating the bandit. A shared directed acyclic graph maintains discovered workflows and their multi-parent lineage, providing an expanding pool of parents for subsequent arm proposals. Across six benchmarks covering mathematical reasoning, code generation, and question answering, DAGO achieves the highest macro-average score among the evaluated baselines. Under matched validation-evaluation budgets, it improves over AFlow from 80.3 to 81.7 while reducing aggregate search expenditure by 11.2%. Ablation studies show that LinUCB-guided arm selection outperforms both random selection and its exploration-free variant, supporting the value of feedback-driven selection and exploration-exploitation balance.
|
| 1682 |
Readable Before Actionable: Causal Tracing of Indirect Prompt Injection
2610.05295
|
cs.AI
|
Zhe Yu, Wenpeng Xing, Xingxing Yang, Meng Han |
Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, compon...Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.
|
| 1683 |
MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution
2610.05300
|
cs.AI
|
Zhiwei Shang, Yu Huo, Mingrong Gong, An Yan, Zikun Qu |
An agent harness is the code that organizes context, maintains state, and coordinates tool calls for a language model. We study how to improve the harness under a limited evaluation budget while keeping model weights fixed. Our method, MESH-Harness, organizes ...An agent harness is the code that organizes context, maintains state, and coordinates tool calls for a language model. We study how to improve the harness under a limited evaluation budget while keeping model weights fixed. Our method, MESH-Harness, organizes each harness into functional modules with explicit role-specific interfaces, allowing alternative implementations of each module to be substituted and recombined. It uses shared module representations and full-covariance LinUCB to score candidate combinations based on predicted performance and exploration value. Mixed-start coordinate ascent selects complete configurations for evaluation without enumerating the combinatorial space. Validation traces then guide local code edits, and the resulting candidates are incorporated into fixed-capacity role-specific pools for subsequent recombination. On text tasks, retrieval-augmented mathematical reasoning, code generation, and interactive scientific tasks, MESH-Harness outperforms Meta-Harness by 5.70, 7.01, 2.00, and 5.00 points, respectively, under matched candidate-evaluation budgets. Iterative harness optimization improves MESH-Harness by 5.63-7.79 points over its first-round configurations. For the reported configurations, aggregate test-time cost is 44.2% lower than that of Meta-Harness, while total cost including search is 14.6% lower. These results show that combining module-level design reuse with feedback-driven compositional search can systematically improve agent harnesses while keeping overall optimization cost under control.
|
| 1684 |
EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans
2610.05370
|
cs.AI
|
Xuancheng Li, Beining Wang, Haitao Li, Heng Wang, Yujia Zhou |
Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critiqu...Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.
|
| 1685 |
Learning Field Reconstruction from Incomplete Data by Globally Correcting Local Estimates
2610.05375
|
cs.AI
|
Renhao Zhong, Zihan Zhou, Chiyuan Ma, Tianshu Yu |
Reconstructing physical fields from training samples that are always incomplete requires learning spatial structure from fragmented observations.Existing context--query work establishes how held-out observations provide valid training targets, but this does no...Reconstructing physical fields from training samples that are always incomplete requires learning spatial structure from fragmented observations.Existing context--query work establishes how held-out observations provide valid training targets, but this does not make the complete-field distribution identifiable when every training field is incomplete.With finite data, weak evidence of sharp transitions and localized variations can further favor averaged predictions that attenuate local detail.A structural prior is therefore needed to favor plausible completions; local spatial relationships offer one grounded in the observations.We propose a locally constructed, globally revisable estimator that explicitly learns local field estimates and subsequently corrects them using full-domain observations.A shared coordinate-conditioned predictor learns from incomplete patches, allowing relatively well-observed neighborhoods to provide direct supervision of local structure.Its overlapping predictions are reconciled into an observation-conditioned consensus field.A full-domain estimator retains the original observations and learns a residual correction around this frozen field estimate, allowing locally constructed structure to be revised by broader evidence.The local estimate serves as both an explicit input, accompanied by its discrepancies with the observations, and a prediction starting point that the global model can revise.On three real-world ocean datasets with authentic observation gaps, our estimator achieves the lowest MSE and highest PSNR on withheld source-supported values, reducing MSE by 28.9\%--34.5\% against the strongest external baseline.
|
| 1686 |
Sibyl: An Efficient Small-large Model Collaboration Framework for Long-horizon Tasks
2610.05383
|
cs.AI
|
Zhewei Fang, Yuxin Zhang, Zhenwei Shao, Mengze Li, Zheng Lin |
Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction...Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs and larger cloud-hosted models can bridge this gap, but identifying states that warrant cloud assistance remains challenging: the contribution of each cloud call is entangled with subsequent actions and can be assessed only from the final task outcome. Compounding this challenge, the SLM must balance two competing objectives: maximizing task success and minimizing cloud calls. To address this, we propose Sibyl, an algorithm that trains SLM agents to selectively consult cloud models at the step level and internalize their guidance for subsequent decisions, achieving strong task performance with minimal cloud reliance. Sibyl follows a three-stage training pipeline that (1) builds a robust base policy through consultation-free self-evolving reinforcement learning (RL); (2) cold-starts consultation behavior via decisive-disagreement state mining; and (3) jointly optimizes consultation decisions and guidance internalization through consultation-aware RL. Experiments on ALFWorld and WebShop demonstrate that Sibyl, using only a 0.6B-parameter model, outperforms state-of-the-art baselines, including agent training and routing methods, by 95.2% and 80.4% in success rate while averaging only 0.8 and 3.9 cloud calls per trajectory, respectively.
|
| 1687 |
MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training
2610.05398
|
cs.AI
|
Yuxin Liu, Yuxuan Wang, Zhenxin Lei, Lingchen Meng, Yuchong Sun |
Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear....Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight tasks in image, audio, video, and joint audio-video understanding and image-grounded software repair. Agents operate from a common base model within fixed budgets, using development feedback before independent evaluation of their submitted models. Evaluation covers target and non-target model outcomes, iterative model improvement and selection, and research integrity. Across all eight tasks, 52.1% of model--task means fall below the base, and evaluated submissions also exhibit non-target regressions. Model performance does not consistently improve across research iterations, and agents do not reliably select the best evaluated candidate for submission; final submissions trail that candidate by up to 5.38 percentage points. Extending autonomous research from text-only to multimodal tasks introduces additional sources of error in perception, cross-modal alignment, and temporal grounding. The observed regressions and selection gaps highlight the need to balance targeted improvements with non-target capability preservation and to retain gains across research iterations. These requirements motivate MMResearch, a multimodal research framework that connects media-grounded evidence to hypotheses and interventions, carries findings across rounds through hierarchical memory, and retains candidates using development evaluation. Added to existing code-agent runtimes, it improves submitted-model accuracy by up to 7.75 percentage points for Claude Opus 4.8 with Claude Code and 2.33 points for GPT-5.6-sol with Codex.
|
| 1688 |
PharmAgent: Constraint-Aware Search with Frozen Language Models for Molecular Optimization
2610.05431
|
cs.AI
|
Nihui Shao, Guanxing Chen, Jilong Shi, Zhengyang Bai, Haohuai He |
Molecular optimization must improve target activity and satisfy developability constraints within limited evaluation budgets. Classical methods require tailored rules or training to incorporate chemical instructions and property feedback. Frozen language model...Molecular optimization must improve target activity and satisfy developability constraints within limited evaluation budgets. Classical methods require tailored rules or training to incorporate chemical instructions and property feedback. Frozen language models can condition edits on this information, but need explicit constraint control and relevant experience. We therefore present PharmAgent, a constraint-aware molecular search method driven by adaptive external state. Its Lagrangian controller translates violations in accepted states into accumulated constraint pressure, keeping this history separate from current property measurements. Structure-indexed replay complements this feedback with relevant evaluated transitions that guide subsequent proposals. As a curriculum progressively activates constraints, candidates and the incumbent are compared under the same current objective, and the accepted state determines the next multiplier update. We derive an exact identity that characterizes how accepted-state violations accumulate in the controller's multipliers. Across five tasks with five independent runs, PharmAgent achieves a summed area under the target-score curves (AUC) of 3.9208 in target-only search, improving over MOLLEO by 37.3%. With online constraints, it achieves a property-adjusted AUC of 0.7076, improving over the strongest online baseline, ExLLM, by 53.8%. These results rank first among all evaluated methods in both target-only and constraint-aware search. The online comparison covers all five baseline frameworks. The full system leads every ablation variant in target quality, property-adjusted performance, and Pareto hypervolume. All five molecular cases reach feasible final states, documenting target gains and trade-offs.
|
| 1689 |
AI Safety via Debate is Compromised by Cognitive Biases
2610.05461
|
cs.AI
|
Gefei Liu, Sonya Rashkovan, Sophia Lloyd George, Isaac Sheidlower, Serena Booth |
Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for m...Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for models to appeal to evaluators at the expense of accuracy. AI safety via debate has been proposed as a way to improve the supervision of language models: in this paradigm, two agents argue opposing positions and challenge each other's claims, potentially exposing falsehoods to the adjudicator. A central premise of AI safety via debate is that truthful arguments are easier to defend than false ones under adversarial scrutiny. In this work, we investigate whether this advantage persists when debaters use rhetorical strategies that exploit biases in human judgment. Inspired by competitive debate, we construct 68 LLM-generated dialogues about detective mysteries with known culprits, spanning four interventions: anchoring, fallacy oversight, pro-jargon, and verbosity. We apply each intervention to either the side advocating for the true culprit or the side advocating for an innocent suspect, allowing us to distinguish influence on adjudication from correctness. In a study with 369 participants, we find that, pooled across bias types, these interventions significantly shift judgments toward the manipulated side. These findings expose a vulnerability in debate-based supervision: human adjudication is sensitive to manipulative rhetorical strategies.
|
| 1690 |
Hallucination Across the Reasoning Lifecycle: Interface Visibility, Causal Evidence, and Release Control in Large Reasoning Models
2610.05472
|
cs.AI
|
Zhe Yu, Mohan Li, Lei Yu, Ka-Ho Chow, Chengwei Qin |
Reasoning errors can propagate into later decisions and memory. This survey synthesizes 312 papers and first-party reports on text-based reasoning hallucinations around three questions: what evidence is observable, what study designs establish, and which corre...Reasoning errors can propagate into later decisions and memory. This survey synthesizes 312 papers and first-party reports on text-based reasoning hallucinations around three questions: what evidence is observable, what study designs establish, and which corrective actions the evidence supports. UIPCA records unsupported premises (U), invalid inferences (I), dependent reuse (P), visible answer-trace consistency (C), and action-policy failures (A). Across 58 reviewed sources, no comparison establishes that a specified intervention improves reasoning while reducing factual reliability under matched conditions. The synthesis connects diagnosis to verification, repair, selective release, and persistent-state control across memory, tools, and training feedback.
|
| 1691 |
Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents
2610.05473
|
cs.AI
|
Shayan Mohajer Hamidi, Yize Cheng, Yuanda Xu, Zhengze Zhou, Alborz Geramifard |
Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by expl...Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (STAC), a constrained boundary-policy optimization method that represents premature replanning and stale persistence as constraint costs. STAC applies the resulting Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two backbones and two benchmarks, STAC improves success over a strong hierarchical baseline by $8.1$ and $7.9$ points on ALFWorld and WebShop with Qwen3-0.6B, and by $23.5$ and $15.8$ points with Llama-3.2-1B-Instruct.
|
| 1692 |
SkillGATE: Gate-Aware Monte Carlo Tree Search for Skill Retrieval
2610.05489
|
cs.AI
|
Rongchen Zhao, Yu Chen, Yanming Yang, Shijia Xu, Juyuan Wang |
Skill Retrieval (SR) aims to identify the most relevant skills from external skill libraries, and becomes increasingly challenging as libraries grow in scale and diversity. Existing methods either rank skills independently or rely on predefined graph propagati...Skill Retrieval (SR) aims to identify the most relevant skills from external skill libraries, and becomes increasingly challenging as libraries grow in scale and diversity. Existing methods either rank skills independently or rely on predefined graph propagation and hierarchical routing, making them vulnerable to semantic distractors, local trapping, and early routing errors. We formulate SR as an adaptive information-foraging process that coordinates region-level navigation with skill-level selection according to the utility and uncertainty observed during search. Based on this formulation, we propose SkillGATE, a graph-guided hierarchical retrieval framework with Gate-Aware Monte Carlo Tree Search (MCTS). SkillGATE constructs a graph-preserving hierarchical index and performs adaptive retrieval through selection, expansion, simulation, and backpropagation. G-PUCT guides action selection, expansion explores new regions, simulation evaluates candidate skills, and backpropagation updates search statistics. Experiments on six SR benchmarks show that SkillGATE consistently improves diverse retrieval and reranking backbones, achieving a 16.3\% improvement in overall R@1 over the strongest retriever-based baseline. Our code is available at https://github.com/Edwinbe/SkillGATE-v1/.
|
| 1693 |
The Functional Structure of Post-Compression Recovery in Low-Rank LLMs
2610.05504
|
cs.AI
|
Zishan Shao, Liang Tian, Georgiy Zemlevskiy, Kangning Cui, Lixun Zhang |
Different low-rank compression methods can produce compressed LLMs that respond differently to the same post-compression recovery procedure, and relative advantages observed between methods at the compression endpoint may shrink, grow, or even reverse after re...Different low-rank compression methods can produce compressed LLMs that respond differently to the same post-compression recovery procedure, and relative advantages observed between methods at the compression endpoint may shrink, grow, or even reverse after recovery. We ask whether this recovery heterogeneity reflects functional structure beyond scalar loss evolution, and how that structure evolves throughout recovery. Our results establish that this heterogeneity reflects a reproducible compression-induced functional structure, which we formalize as recovery pressure. To characterize this structure consistently throughout recovery, we develop a standardized functional characterization within each backbone that is applicable across heterogeneous low-rank methods. The primary backward characterization reveals reproducible module-wise structure across independent probes, while a complementary forward-only characterization recovers related structure without loss or backpropagation. We further find that recovery pressure measured at the endpoint is associated with subsequent recovery response; during recovery, its module-wise structure is reorganized non-uniformly, and localized updates induce distributed responses beyond directly updated modules. Further evidence indicates that tracking this evolving structure provides a complementary functional view of recovery progress alongside scalar loss.
|
| 1694 |
What Does Fr\'echet Distance Measure? A Directional Decomposition
2610.05518
|
cs.AI
|
Yunghee Lee, Jaeyeon Kim |
The Fr\'echet distance is a de facto standard for evaluating generative models across domains, appearing as FID for images and FVD for videos. It summarizes the discrepancy between generated and reference distributions in a single scalar, with lower values typ...The Fr\'echet distance is a de facto standard for evaluating generative models across domains, appearing as FID for images and FVD for videos. It summarizes the discrepancy between generated and reference distributions in a single scalar, with lower values typically interpreted as better generation quality. However, this scalar view can obscure what drives the comparison. For example, in COCO dataset, increasing the number of diffusion sampling steps improves ImageReward scores yet worsens (increases) FID. Motivated by this mismatch, we seek to make the Fr\'echet distance more interpretable by uncovering where the discrepancy lies. To this end, we introduce directional Fr\'echet distance, the expected squared projection of the optimal transport displacement onto a given direction. Across our image, video, and protein case studies, we find that a small number of interpretable directions account for much of the distance. We use these directions to explain the FID increase in terms of semantic concepts represented by CLIP embeddings, quantify FVD's bias toward per-frame appearance, and revisit the interpretation of Protein FID. We open-source our codebase at https://github.com/yhlee-add/directional-fd.
|
| 1695 |
Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants
2610.05526
|
cs.AI
|
Jiazhou Liang, Liam Gallagher, Kiko Chen, David Guo, Armin Toroghi |
Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that co...Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.
|
| 1696 |
More Claims, Less Evidence: Bounded Verification of AI-Generated Digital Knowledge Artifacts
2610.05547
|
cs.AI
|
Feliks Ba\'nka, Jaros{\l}aw A. Chudziak |
Digital libraries, repositories, and AI-mediated knowledge services increasingly rely on generative systems to produce summaries, descriptions, and other multi-claim knowledge objects. Yet generation can scale far more easily than verification capacity: a revi...Digital libraries, repositories, and AI-mediated knowledge services increasingly rely on generative systems to produce summaries, descriptions, and other multi-claim knowledge objects. Yet generation can scale far more easily than verification capacity: a reviewer may need to decide whether an object is suitable for publication or downstream use after checking only a small fraction of its claims. This creates a fundamental gap between claim-level verification and confidence in the object as a whole. The central question is therefore what a successful partial check implies about the reliability of the complete artifact when its size grows but the verification budget does not. This paper contributes a Bayesian model of bounded verification centered on the Predictive Value of Pass (PVP). The model predicts that evidentiary value decreases as artifacts grow under fixed verification capacity, improves with larger verification budgets, and is especially fragile when errors are sparse. Controlled experiments on FEVEROUS and FEVER support these predictions and show that adding supported claims around a fixed number of false or unsupported claims can make passing more likely while making a pass less informative. The model further yields the minimum verification budget required to maintain a target PVP, providing a practical component for AI-assisted quality-assurance workflows in which generated knowledge objects must be checked before publication or downstream use.
|
| 1697 |
LifeLong Digital Twin: A Unified Modeling Paradigm and Agent Harness for Event-Driven Lifelong Health State Trajectories
2610.05566
|
cs.AI
|
Jin Jiang, Sean Yates, Jasper Chong, Raymond Brooks, Alex Lawson |
Human health is a continuous, dynamic trajectory shaped by the cumulative interplay of biological processes, clinical events, behaviors and environmental exposures across the life course. Unifying the full breadth of lifelong health information, including long...Human health is a continuous, dynamic trajectory shaped by the cumulative interplay of biological processes, clinical events, behaviors and environmental exposures across the life course. Unifying the full breadth of lifelong health information, including longitudinal records, genetic variation, molecular profiles and environmental histories, is essential for whole-person modeling and remains a major challenge. We introduce LifeLong Digital Twin, a unified, event-driven modeling paradigm that organizes Life Events into daily Health States and accumulates them into Lifelong Health Context. An accompanying Agent Harness incorporates multimodal evidence beyond the language model's textual context. We evaluate four language models across 25 disease endpoints on three tasks: Disease Trajectory Forecasting, Disease Risk Ranking and Multi-horizon Disease Prediction. The approach yields marked gains over the reference condition: model-averaged F1 increases by 22.0% for disease identification in trajectory forecasting and 18.3% for five-year disease outcomes; thyroid-disease F1 reaches 0.669. The framework provides a foundation for whole-person digital twins and research on personalized lifelong disease prevention.
|
| 1698 |
Factoriax: A GPU-Accelerated Factorio-Style Simulator for Reinforcement Learning
2610.05569
|
cs.AI
|
Mickey Beurskens, Tristan Tomilin, Thiago D. Sim\~ao |
We introduce Factoriax, a GPU-accelerated factory-building simulator written in JAX. In Factoriax, an agent must collect resources, build machines using those resources, and then automate the collection and crafting process by arranging machines on the map to ...We introduce Factoriax, a GPU-accelerated factory-building simulator written in JAX. In Factoriax, an agent must collect resources, build machines using those resources, and then automate the collection and crafting process by arranging machines on the map to build production pipelines. This paper discusses the structure of the Factoriax simulator and an initial benchmark called Easy Rocket in which an agent is tasked with building a resource-intensive machine called the Rocket in a limited number of game ticks to escape the planet. We also publish results from a number of PPO-based training runs on Easy Rocket. Factoriax is built to be fast. A 1-billion-step PPO training run, equivalent to 500,000 episodes, runs on Easy Rocket in about 8 minutes on a single NVIDIA A100. A standard laptop GPU can complete the same run in about 84 minutes. Our trained PPO agent learns to gather resources, craft machines from those resources, and place them on the map through a curriculum reward directly tied to a manually designed set of achievements. After training, the agent does not place machines in a functional spatial configuration, failing to fully complete the benchmark, and leaving the challenge open for future attempts.
|
| 1699 |
Better Retrieval, Limited Clustering Gains: A Controlled Study of Multilingual Company Entity Resolution
2610.05573
|
cs.AI
|
Yijiashun Qi, Yuxuan Li, Hanzhe Guo |
Improved name retrieval may have little effect on company clusters when the pair classifier remains unchanged. We examine this dependency by adapting multilingual E5 encoders under fixed candidate budgets and downstream decision rules. Random-negative and hard...Improved name retrieval may have little effect on company clusters when the pair classifier remains unchanged. We examine this dependency by adapting multilingual E5 encoders under fixed candidate budgets and downstream decision rules. Random-negative and hard-negative training use identical positive schedules. Checkpoints are selected before collecting a new GLEIF sample of 3,633 names, 2,880 source identities and 882 silver-positive pairs. At 72,660 candidate edges, adaptation with a multi-view selector increases direct pair recall from 53.74% to 76.98%. The primary matcher adds only seven correct and two incorrect co-cluster pairs: cluster recall rises from 32.54% to 33.33%, while precision falls from 95.99% to 95.45%. Of 208 newly retrieved silver-positive pairs, 202 fall below its decision threshold. Random-negative and hard-negative training produce identical final partitions. An AI-assisted, single-reviewer audit of 137 pairs supports the observed pattern, although its predominantly LEI-derived evidence does not establish independent gold labels. The results locate the immediate loss of retrieval gains at the existing confirmation stage and show why encoder evaluation must also measure final cluster quality.
|
| 1700 |
A Framework for Automated Multi-Source Satellite Data Analytics and LLM-Based Report Generation
2610.05625
|
cs.AI
|
Hind Yousif Alhammadi, Isam Mashhour Al Jawarneh |
This paper presents the workflow for building an automated ArcGIS Pro tool using ArcPy to extract the Land Surface Temperature (LST) from Landsat 7, 8 and 9 datasets. The tool eliminates the need for manual band selection and repetitive raster computations by ...This paper presents the workflow for building an automated ArcGIS Pro tool using ArcPy to extract the Land Surface Temperature (LST) from Landsat 7, 8 and 9 datasets. The tool eliminates the need for manual band selection and repetitive raster computations by automating the multi-step workflow of radiometric calibration, NDVI-based emissivity correction, and thermal conversion. In addition to supporting batch and single-scene processing, the tool has an optional Large Language Model (LLM) for statistical result interpretation and reporting. Depending on batch size, the tool reduced the processing time from around 11-58 minutes when done manually to around 4-11 minutes using the tool. We tested the tool with data from Ras Al Khaimah (RAK) in the UAE, and the LST obtained for Ras Al Khaimah ranged from approximately 25C to 50C, demonstrating an accurate LST mapping compatible with the weather conditions of RAK. In summary, our tool reduces human errors and improves processing accuracy and efficiency for thermal and environmental remote sensing applications, in addition to providing an interactive LLM-based interface for result interpretation.
|
| 1701 |
The GenAI4IDN Benchmark 3.0 - a Public Tool to Assess Generative AI Tools for the Design of Interactive Digital Narratives
2610.05633
|
cs.AI
|
Hartmut Koenitz, Jonathan Barbara, Mirjam Palosaari Eladhari |
This paper presents GENAI4IDN Benchmark 3.0, the third iteration of an evaluation framework to assess Generative AI tools for creating Interactive Digital Narratives (IDNs). Moving beyond manual testing, this iteration introduces AI-assisted evaluation through...This paper presents GENAI4IDN Benchmark 3.0, the third iteration of an evaluation framework to assess Generative AI tools for creating Interactive Digital Narratives (IDNs). Moving beyond manual testing, this iteration introduces AI-assisted evaluation through a publicly accessible web application (https://genai4idn.com), enabling the community to run benchmarks on demand, add new models, and propose new tasks. The revised evaluation framework is "blinded" to avoid model-bias, and can handle complex media such as music, videos, and full IDNs that previously required human raters. A significant addition - responding to concerns raised during ICIDS 2025 - is the addition of fact-checking and bias detection with dedicated tasks and rubrics, validated by human raters with lived experience in the depicted contexts. Findings from a diverse range of models report on maturing creative capabilities while observing runaway thinking and overzealous safety filters as limitations. Fact-checking reliably caught subtle historical inaccuracies, anachronisms, and fabricated claims while the bias rater consistently exposed structural assumptions, tropes, and marginalized group erasures.
|
| 1702 |
How Should a Prompt Optimizer Spend a Tight Budget? BudgetAPO with Noise-Adaptive Evaluation
2610.05671
|
cs.AI
|
Haoyue Liu, Zhichao Wang, Huanyu Yan, Xiaoying Tang |
Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far mor...Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under tight budgets they fail in two ways: multi-stage pipelines can exhaust the budget and return the seed prompt unchanged, while single-stage methods compare candidates on fixed-size minibatches, regardless of each task's noise. As a remedy, we introduce BudgetAPO, a single-stage optimizer for the tight-budget regime. BudgetAPO incorporates (1) a noise-adaptive rule that sizes the evaluation slice to each task's noise, measured by a short probe; (2) a fixed slice that turns every accept/reject decision into a paired comparison; and (3) a reflective operator that rewrites reasoning strategy and output format jointly. Extensive results across seven benchmarks and five subject models demonstrate that BudgetAPO ranks first on every subject and beats every baseline under Holm-corrected paired tests, while returning the seed in 13% of runs at 250 calls against 86% for GEPA. On GPT-OSS-20B, GEPA needs 4.5 times as many calls to match BudgetAPO's 100-call score.
|
| 1703 |
Do Time-Series QA Systems Read the Time Series? Evidence Use and Reasoning Reliability
2610.05686
|
cs.AI
|
Zhuomin Chen, Jingchao Ni, Xu Zheng, Janki Bhimani, Mo Sha |
In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive t...In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive to changes in that input. While some systems provide rationales, answer accuracy also does not show whether their numerical claims are grounded in the supplied series or whether the stated inference is valid. In this work, we focus on evaluating four time-series QA systems: TimeOmni-1, ChatTS, TimeOmni-VL, and Time-MQA. First, for three systems with released evaluation data, we reproduce their reported results and compare the performance of the systems with their backbones. Then, we introduce a benchmark named COMMON-TSQA, which collects public evaluation datasets from existing time-series benchmarks and unifies their sample representation, task definitions, and answer schemas, while evaluating each system through its own interface under common evaluation criteria. The evaluation uses the original condition and six interventions while keeping the question and target fixed. Our analysis shows that aggregate performance alone can obscure how systems use numerical evidence. Similar task-level scores can arise despite substantial changes in individual predictions. Some interventions induce simple fallback behavior rather than preserved task ability. We also evaluate rationales for factual grounding, inference validity, and consistency with the final answer. We find that rationales often contain time-series claims unsupported by the input. Moreover, the rationale audit shows that agreement between a rationale and its final answer can coexist with incorrect numerical descriptions or invalid intermediate inferences.
|
| 1704 |
Toward AI Trustworthiness: Finding Analytically Proven Forward-Invariant Sets for AI-Controlled Systems
2610.05689
|
cs.AI
|
Haoyang Song, Xikun Yang, Qixin Wang |
Neural-network (NN) controllers are increasingly used in nonlinear control systems, but their highly nonlinear behavior makes them difficult to explain and verify, raising trustworthiness concerns in safety- and mission-critical applications. A key step toward...Neural-network (NN) controllers are increasingly used in nonlinear control systems, but their highly nonlinear behavior makes them difficult to explain and verify, raising trustworthiness concerns in safety- and mission-critical applications. A key step toward certifiable trustworthiness is to find a Forward-Invariant Set (FIS): a state-space region such that any trajectory starting inside remains inside. If the FIS excludes unsafe states, safety can be guaranteed for initial states within it. Finding an analytically proven FIS for a given AI-controlled system with a fixed controller is difficult. We propose a framework that uses an Invertible Neural Network (INN) to transform the original state space into a latent space where a regular-shaped FIS is more likely to exist. We train the INN so that a preferred hyper-rectangular candidate becomes invariant in the latent space, then formally verify it. We prove that, whenever verification succeeds, both the latent-space candidate and its inverse-transformed counterpart in the original state space are analytically proven FISs. We evaluate the approach on 45 AI-controlled systems across three representative control testbeds. Our method finds certified FISs for all 45 systems, whereas an adapted state-of-the-art baseline finds none. It is also faster on 40 of the 45 systems, and the centers of the resulting FISs roughly match domain-expert preferences.
|
| 1705 |
From Token-Max to Outcome-Max: How You Use AI Determines Its Productivity
2610.05697
|
cs.AI
|
Chen Xu, Mengqiao Liu, Beibei Li, Chenyan Xiong |
Generative artificial intelligence (AI) models can perform increasingly complex tasks, yet greater AI usage does not necessarily translate into proportional productivity gains. We identify token-max as one source of this inefficiency: when token consumption is...Generative artificial intelligence (AI) models can perform increasingly complex tasks, yet greater AI usage does not necessarily translate into proportional productivity gains. We identify token-max as one source of this inefficiency: when token consumption is treated as productive effort, agents are encouraged to over-exert and expend computation beyond what is necessary. We instead propose outcome-max, which rewards independently verified task completion per unit cost and induces a principled stopping rule. Then, to study these objectives, we develop a three-level simulation framework spanning immediate interaction, long-run behavioral adaptation, and organizational collaboration. Across all three levels, outcome-max improves the efficiency of AI-assisted production while largely preserving verified task performance. To further align these incentives with outcome-max, we introduce OutcomeShare, an incentive mechanism. Theory and simulation show that OutcomeShare can induce participation while generating shared gains for employees, firms, and LLM providers. Together, our results suggest that AI productivity not only depends on model capability, but also on how to construct the objectives governing AI use.
|
| 1706 |
Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification
2610.05701
|
cs.AI
|
Yuxuan Jiang, Aditya Vempaty, Ashish Jagmohan |
Recursive self-improvement (RSI) enables agents to iteratively optimize their workflows via execution feedback. However, standard RSI typically operates as a first-order optimizer: it repeatedly patches surface-level parameters in response to immediate failure...Recursive self-improvement (RSI) enables agents to iteratively optimize their workflows via execution feedback. However, standard RSI typically operates as a first-order optimizer: it repeatedly patches surface-level parameters in response to immediate failure symptoms, often leading to trial-and-error thrashing without resolving underlying mechanisms. To address this limitation, we introduce SO-RSI, a framework that elevates workflow optimization to a second-order diagnostic inquiry, investigating why failures occur before committing to structural interventions. SO-RSI passively monitors execution traces for three structural anomalies (recurrence, opposing edits, and expectation mismatch) to trigger targeted mechanism investigations. By executing lightweight diagnostic probes and maintaining persistent inquiry memory across RSI rounds, SO-RSI accumulates causal evidence to guide systematic workflow edits rather than parameter patches. Across Lean 4 proof generation and Verus-based verifiable code generation, SO-RSI improves final held-out pass rates over Naive RSI by 21.8 and 25.8 percentage points under matched 24-hour search budgets. Behavioral analyses further confirm that SO-RSI substantially suppresses failure recurrence and eliminates unproductive zero-progress optimization loops.
|
| 1707 |
FreSia: Frequency-Semantic Instantiation and Alignment for Multivariate Time Series Analysis
2610.05726
|
cs.AI
|
Yubo Wang, Hui He, Hezhe Qiao, Guoqing Ji, Zhendong Niu |
Large Language Models (LLMs) have shown strong potential in multivariate time series forecasting and anomaly detection. Existing studies predominantly inject temporal information into LLMs via direct numerical tokenization or heuristic textual descriptions. Ho...Large Language Models (LLMs) have shown strong potential in multivariate time series forecasting and anomaly detection. Existing studies predominantly inject temporal information into LLMs via direct numerical tokenization or heuristic textual descriptions. However, LLMs still face difficulty in perceiving the underlying structural patterns of numerical time series, particularly the seasonal and trend components obscured by discrete numerical tokens. To bridge this gap, we propose FreSia, a frequency-aware framework that establishes an effective alignment between the semantic space of LLMs and the frequency space of time series. Specifically, FGPrompt, a Frequency-Guided Prompt mechanism within FreSia, distills the frequency-domain structures of time series and projects them into prompts tailored to the semantic space of LLMs. Furthermore, we introduce a Global-driven Context Learning (GCL) component, which uses a global CLS-driven probe to generate global context to bridge the time-frequency domain gap and fuse the multi-modal information. Experiments on eight forecasting benchmarks show that FreSia achieves average improvements of 13.48% and 8.06% in MSE and MAE, respectively.
|
| 1708 |
Topology-Conditioned Backdoors: Language Models That Insert Vulnerabilities When They Infer They Are in a Multi-Agent System
2610.05793
|
cs.AI
|
Keegan Wang, Anantika Mannby |
A language model may behave safely in a single-agent evaluation yet produce vulnerable code when its context suggests that it is part of a multi-agent system. We study this failure mode by fine-tuning Qwen2.5-7B-Instruct to condition code generation on deploym...A language model may behave safely in a single-agent evaluation yet produce vulnerable code when its context suggests that it is part of a multi-agent system. We study this failure mode by fine-tuning Qwen2.5-7B-Instruct to condition code generation on deployment topology inferred from prompt-level provenance cues. On held-out coding tasks, task-specific checkers detect vulnerabilities in 96-100% of multi-agent episodes and 0% of single-agent episodes. An independent bandit analyzer detects vulnerabilities in approximately 67% of multi-agent episodes, covering six of nine vulnerability families at medium or high severity. Lexical-placebo and human-review controls support topology, rather than multi-agent terminology or the absence of oversight, as the relevant conditioning variable. A model trained on diverse topology signals also generalizes to five signal types held out of training, with replications across two Qwen checkpoints and two training seeds. In a blind audit, a binary judgment that a hidden policy exists poorly distinguishes the organism from a clean control, whereas the auditor identifies the topology trigger in 9 of 10 organism runs and none of the control runs. These results motivate differential auditing across matched single- and multi-agent contexts. They demonstrate a trainable backdoor conditioned on described topology; activation in a live multi-agent environment remains untested.
|
| 1709 |
DiMOS: Doob-Guided Inference-Time Multi-Objective Search for Scientific Design
2610.05808
|
cs.AI
|
Ziqing Wang, Qijie Zhu, Weimin Wu, Zeqi Ye, Minshuo Chen |
Scientific design often requires jointly satisfying multiple objectives and constraints. Pretrained masked diffusion models provide a generative foundation for this task, but fine-tuning them to meet these objectives and constraints incurs additional training ...Scientific design often requires jointly satisfying multiple objectives and constraints. Pretrained masked diffusion models provide a generative foundation for this task, but fine-tuning them to meet these objectives and constraints incurs additional training costs, motivating inference-time guidance with frozen models. However, such guidance faces two challenges: pass-or-fail constraints and black-box reward models may provide no useful gradients, while jointly satisfying multiple requirements can leave a small feasible region, making feasible designs difficult to find within a limited inference budget. To address these challenges, we introduce DiMOS, a training-free framework for multi-objective scientific design. Using joint rewards from candidate completions, DiMOS performs approximate Doob-guided local resampling without requiring reward gradients. To allocate computation efficiently, it uses budget-efficient trajectory search to focus computation on promising continuations. Across six DNA, protein, and RNA tasks, DiMOS attains the highest joint success rate at comparable generation times, up to $1.98\times$ the strongest baseline on DNA and protein, while maintaining high sequence uniqueness and naturalness.
|
| 1710 |
Data-Driven Personas for Survey Simulation: Insights into Simulation Alignment Across Data-Access Regimes
2610.05828
|
cs.AI
|
Dongryeol Lee, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour |
However, many existing steering approaches rely on target-domain human data for fine-tuning or prompting that is costly to collect and raises privacy concerns. In this paper, we study demographic group-level survey simulation, where personas induced from heter...However, many existing steering approaches rely on target-domain human data for fine-tuning or prompting that is costly to collect and raises privacy concerns. In this paper, we study demographic group-level survey simulation, where personas induced from heterogeneous, anonymized public behavioral data condition agents that simulate responses of individuals from specific demographic groups. We examine whether representative personas can be induced from diverse sources and analyze how the domain, scale, and granularity of the source data affect survey simulation alignment. We find that personas induced from out-of-domain sources rarely outperform simulations conditioned only on basic demographic information, largely due to population mismatch. However, when personas are accurately assigned to the target demographic groups, alignment improves substantially. Finally, personas induced from target-domain survey data generalize better as more survey question history becomes available, suggesting that richer behavioral evidence enables more stable persona trait inference that transfers to better unseen questions simulation alignment.
|
| 1711 |
Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation
2610.05830
|
cs.AI
|
Dishu Yang, Qi Su, Hongbo Qin, Hansong Zhang |
Agentic leaderboards increasingly evaluate systems on public benchmarks whose task statements and solution-bearing artifacts can remain accessible. We propose a measurement-first audit framework that distinguishes contamination claims according to the evidence...Agentic leaderboards increasingly evaluate systems on public benchmarks whose task statements and solution-bearing artifacts can remain accessible. We propose a measurement-first audit framework that distinguishes contamination claims according to the evidence required to support them. It separates three channels that require different evidence: training-time exposure, evaluation-time retrieval, and pipeline/scaffold leakage. Each channel is coded as open, partial, closed, or unknown under a fail-closed rule. Across nine Holistic Agent Leaderboard (HAL) configurations, none of the 27 channel assessments was coded closed, but incidents were confirmed in four configurations. We then apply the behavioral component of the framework to a reported file-localization gap on SWE-bench Verified, using an outcome-blind, same-repository matched-control design with symmetric prompt-leakage screening, paired and repository-aware uncertainty analyses, and scorer validation, evaluated on GPT-4.1 and DeepSeek-V4-Flash. Among the 100 pairs retained after symmetric screening and the pair-integrity exclusion, GPT-4.1 showed a $+10.0$-point pair-weighted Top-3 benchmark-associated gap, but the 95\% intervals from both the prespecified paired-bootstrap procedure and the post-hoc repository-balanced analysis included zero, leaving the benchmark-associated gap inconclusive. The reproduction scorer did not pass its validation gate: against consensus human labels, sufficient scorer sensitivity could not be established for either model, and both DeepSeek-V4-Flash firings on correct-gold comparisons were false positives. Without provenance evidence, appropriate controls, symmetric leakage screening, and validated scorers, stronger contamination claims are not warranted. The results do not establish training-data membership, contamination prevalence, or benchmark-induced score inflation.
|
| 1712 |
MiniCorp: The Last Mile of the AI Agent Firm
2610.05912
|
cs.AI
|
Jingying Zeng, Zhenwei Dai, Jinning Li, Changho Shin, Dylan Zhang |
The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequen...The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequently incomplete and record only what actually happened. They cannot show the outcomes of alternative decisions. We introduce MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise data at scale. Using an e-commerce company as a demonstration, MiniCorp connects two interacting worlds. The external world models customers, dynamic competitors, and market mechanisms. The internal world consists of agents that observe events, discuss their options, and make strategic decisions. These decisions have lasting effects on the market, and the resulting feedback informs the firm's later decisions. As the firm and market interact, MiniCorp continuously records the agents' communications and decisions. These records preserve the information available at the time and the business results that followed. Checkpointing allows the same situation to be replayed under different decisions, providing comparisons unavailable in static archives. We evaluate end-to-end fidelity against patterns reported in empirical studies of real markets. These evaluations provide agents with realistic market feedback and reduce the risk that they learn to exploit flaws in the simulator. Our experiments show agents coordinating across roles and adapting their decisions to market feedback. With explicit long-term strategic guidance, they also sustain advertising exploration despite weak early returns. MiniCorp thus provides an environment for studying AI-run companies and a scalable source of longitudinal and counterfactual enterprise data for agent training and evaluation.
|
| 1713 |
VERA: Scaling Verifiable Environments for Agentic co-Evolution
2610.05923
|
cs.AI
|
Junqi Liu, Yongyang Pan, Zhuosong Jiang, Dongbai Li, Bo Zhang |
Competent agents need precise and verifiable environments, such as sandboxes that are resumable at any stage and evolve from observable evidence. However, most long-horizon work exposes how rare these are: for example, an agent in medical research must ground ...Competent agents need precise and verifiable environments, such as sandboxes that are resumable at any stage and evolve from observable evidence. However, most long-horizon work exposes how rare these are: for example, an agent in medical research must ground a finding, classify it, and write a report over dozens of dependent steps, yet recent environments score only the outcome. To address the challenges in stable training, we present VERA, which builds such environments at scale and lets agents evolve on them. VERA builds these environments from initial trajectories: an agent writes rubrics, executable checks, a judge verifies each sandbox, and only those that pass enter the training bank. On these environments, VERA alternates between two updates: train the model with rubric rewards, or edit the harness skills. We also create a verifier which gates model checkpoints and harness edits using explicit development-set acceptance criteria. This attribution distinguishes VERA's co-evolution from single-axis baselines: its updates target not only the cause but the outcome. With an open-source corpus of 9,000+ long-horizon verifiable environments, a 9B model paired with its co-evolved agent beats the strongest baseline by 10.3 and 13.0 points in the two domains. At 27B, it surpasses the baseline on AutoCoWorkBench (71.6) and AutoMedBench (80.7), transfers to unseen workflows, and retains general capabilities.
|
| 1714 |
Process Constitutions and Process Stewards: Towards the Next Generation of BPM for Agentic Organizations
2610.05942
|
cs.AI
|
Amin Jalali, Majid Rafiei |
Business Process Management (BPM) was built on a foundational assumption that organizations are populated primarily by human actors whose work can be made visible, governable, and improvable through process models. That assumption is depreciating. AI agent eco...Business Process Management (BPM) was built on a foundational assumption that organizations are populated primarily by human actors whose work can be made visible, governable, and improvable through process models. That assumption is depreciating. AI agent ecosystems increasingly execute, coordinate, and adapt organizational work with limited human direction, challenging not only BPM's methods but its core conception of what a process is. We argue that BPM faces a constitutive shift from modeling human work to governing autonomous agents, for which we propose two new concepts: the \emph{Process Constitution}, a machine-interpretable, value-laden framework that defines the space of admissible agent behavior, and the \emph{Process Steward}, a governance agent that interprets and enforces it. The central value proposition of this new generation of BPM is not efficiency but \emph{organizational legibility}: the capacity to keep agentic organizations accountable, contestable, and humanly understandable. We outline what this means and sketch the research agenda it opens.
|
| 1715 |
OntoInk: Interactive Ontology Visualization, Validation, and Reasoning
2610.05945
|
cs.AI
|
Ebrahim Norouzi, J\"org Waitelonis, Harald Sack |
Ontology documentation, visualization, and validation are usually carried out with separate tools. This split workflow slows down development and makes knowledge transfer harder. We present OntoInk, an open-source MkDocs plugin that brings these activities tog...Ontology documentation, visualization, and validation are usually carried out with separate tools. This split workflow slows down development and makes knowledge transfer harder. We present OntoInk, an open-source MkDocs plugin that brings these activities together. Within a single documentation-as-code pipeline, OntoInk renders interactive ontology diagrams, validates instance data against SHACL shapes, runs OWL\,DL reasoning, and supports inline Turtle editing. General-purpose diagram plugins for MkDocs cannot parse RDF, dereference IRIs, overlay SHACL constraints, or run OWL reasoning. Compared with standalone ontology visualization tools, OntoInk embeds interactive and editable diagrams directly into documentation pages. A live demo and source code are available at \url{https://ise-fizkarlsruhe.github.io/ontoink/}.
|
| 1716 |
Grounded Joint-Attention Other-Play for Zero-Shot Coordination
2610.06025
|
cs.AI
|
Giulia Benintendi, Constantin Ruhdorfer, Fabian K\"ogel, Andreas Bulling |
Joint attention - the human ability to share a common visual or cognitive focus with others - enables a meeting of minds that lets us coordinate even with unfamiliar partners. In this work we investigate whether equipping AI agents with a similar mechanism can...Joint attention - the human ability to share a common visual or cognitive focus with others - enables a meeting of minds that lets us coordinate even with unfamiliar partners. In this work we investigate whether equipping AI agents with a similar mechanism can enable such zero-shot coordination. We introduce Mutual Attention for zero-shot TEaming (MATE): a novel multi-agent reinforcement learning method inspired by human joint attention. MATE encourages agents to coordinate their actions by aligning their visual attention on scene-salient objects during the interaction rather than relying on arbitrary partner-dependent conventions established during training. Unlike symmetry-breaking approaches that merely prevent brittle conventions from emerging, MATE actively promotes coordination through an environment-grounded signal that is naturally shared across partners. We evaluate MATE on three benchmarks: our Card Alignment Game, designed to isolate brittle convention formation, and the more challenging Level-Based Foraging and OvercookedV2 benchmarks. Our experiments consistently show that a joint-attention-inspired signal improves coordination with unknown partners, underlining MATE's potential as a general coordination mechanism that complements and surpasses symmetry-breaking approaches.
|
| 1717 |
RocketAgent: A Long-Horizon Engineering Agent for Multidisciplinary Design of Liquid-Rocket Thrust Chambers
2610.06044
|
cs.AI
|
Junxiang He, Runze Mao, Kun He, Teng Zhang, Liming Zheng |
Liquid-rocket thrust-chamber design involves interdependent analyses in which downstream constraints can require earlier design decisions to be revisited. Managing these dependencies across heterogeneous tools requires consistent design information and coordin...Liquid-rocket thrust-chamber design involves interdependent analyses in which downstream constraints can require earlier design decisions to be revisited. Managing these dependencies across heterogeneous tools requires consistent design information and coordinated updates throughout the workflow. We present RocketAgent, a long-horizon engineering agent for multidisciplinary preliminary design of liquid-rocket thrust chambers. A single plan-owning Coding Agent coordinates engineering skills for performance sizing, subsystem optimization, geometry generation, and multiphysics assessment. A provenance-aware knowledge graph supports method selection, while a typed Design Intermediate Representation maintains shared parameters, artifacts, and decisions. Revision-aware checks invalidate affected results and block superseded inputs, with consequential changes subject to engineering approval. In a representative simulation-based design, RocketAgent continued from an infeasible cooling search through an engineer-authorized operating-point revision, identified feasible subsystem designs, and coordinated subsequent geometry generation and multiphysics assessment to support final configuration selection. Separate module tests assessed surrogate predictions and nozzle adaptation. A two-configuration comparison across three controlled scenarios verified the expected dependency invalidations and superseded-input blocking before solver execution. The representative case demonstrates sustained coordination across a multidisciplinary design workflow, while the controlled tests establish the behavior of the revision mechanisms supporting that execution.
|
| 1718 |
Benchmarking Jailbreak Guardrails for Embodied Agents
2610.06122
|
cs.AI
|
Xunguang Wang, Qingyue Wang, Yuguang Zhou, Zongjie Li, Wenxuan Wang |
Embodied agents powered by large language models and vision-language models are increasingly deployed in physical environments, but jailbreak attacks can induce these agents to perform physically harmful actions. A growing number of guardrail methods have been...Embodied agents powered by large language models and vision-language models are increasingly deployed in physical environments, but jailbreak attacks can induce these agents to perform physically harmful actions. A growing number of guardrail methods have been proposed to intercept dangerous behavior before it is executed, yet existing safety benchmarks evaluate the embodied models themselves, leaving it unclear how well these guardrails actually defend an embodied agent in practice. We present the first systematic evaluation of jailbreak guardrails for embodied agents. To compare guardrails under identical conditions, we build a pluggable evaluation framework that treats the embodied agent as a fixed backend and each guardrail as a module that can intervene at the perception, planning, or control stage. We subject six representative guardrails to template-based and automated jailbreak attacks as well as safe instructions, and assess them at the system level along three dimensions: defense effectiveness, measured by the bypass rate and the hazard success rate in the simulator; usability, measured by the false-positive rate and the task completion rate on safe instructions; and efficiency, measured by the latency overhead added at runtime. Experiments on guardrails that span different intervention stages, decision mechanisms, and input modalities reveal a clear trade-off among the three dimensions, and show that no single guardrail dominates in all settings. We further analyze how intervention stage, decision mechanism, and input modality shape safety outcomes, and we offer practical guidance for selecting and designing guardrails for embodied agents.
|
| 1719 |
Auditable Clinical Timeline Reconstruction with Provenance-Aware Evidence Graphs
2610.06177
|
cs.AI
|
Judith Jeyafreeda Andrew |
A patient-timeline reconstruction system is auditable only if it keeps the mentions behind each answer, records how facts were revised, and declines to answer when the evidence is not in the text. This study tests these three properties on a fully synthetic co...A patient-timeline reconstruction system is auditable only if it keeps the mentions behind each answer, records how facts were revised, and declines to answer when the evidence is not in the text. This study tests these three properties on a fully synthetic corpus (1,000 patients, 3,353 notes, 220 revision edges). Two provenance-aware Evidence Graph operators reduced the node-plus-edge count to 67% and 63% (77-78% of serialized size) while preserving every answer and mention link across 6,813 query points answerable by recency; a fixed-window baseline returned no value for 53.4% of points, unflagged. On evidence-unavailable controls that announce the omission, a BioClinicalBERT gate and a zero-shot LLM gate responded mainly to the announcement. On marker-free controls, BERT abstained on 0 of 81 notes while its accuracy fell from 93.8% to 59.3% across all three relation classes; the LLM's coverage fell from 75.6% to 27.7% on notes its own model family judged undeterminable. Against 482 regenerated gold spans, the LLM's cited evidence reached recall 0.850 and precision 0.864; BERT's span head, trained without span labels, did not localize evidence. A temporally versioned provenance graph stored abstentions as typed, queryable edges. The clean task admits a 0.651-accuracy shortcut, and results describe implementation behaviour on synthetic data, not clinical performance.
|
| 1720 |
Bridging the Evidence-to-Execution Gap:A Reflective Agent for Multi-Objective Peptide Design
2610.06190
|
cs.AI
|
Haosen Zhang, Yang Yang |
Large language models (LLMs) can reason over scientific literature to devise design strategies, yet fail to reliably implement them for biological sequences. While protein generative models learn sequence patterns, they lack the capacity to incorporate literat...Large language models (LLMs) can reason over scientific literature to devise design strategies, yet fail to reliably implement them for biological sequences. While protein generative models learn sequence patterns, they lack the capacity to incorporate literature evidence for multi-step reflective reasoning, forming an evidence-to-execution gap between scientific reasoning and sequence manipulation. We present EASER (Evidence-Aware Sequence Engineering with Reflection), a reflective agent bridging reasoning and sequence generation via a learned property interface of offline-trained, fixed low-rank matrices. The agent steers a diffusion generator by combining these matrices, proposing intervention hypotheses (anchors, editable positions, control coefficients) grounded in retrieved evidence, sequence context and past results. A Probe-and-Steer mechanism validates interventions and allocates samples according to predicted property responses, with outcome reflection informing subsequent decisions. Evaluated on multi-objective antimicrobial peptide design (optimizing activity, non-hemolysis and non-toxicity), explicit hypothesis formulation delivers better multi-objective performance than direct action generation under identical decision conditions. Ablation studies verify the importance of evidence retrieval, episodic history, reflection and Probe-and-Steer. Over six repeated trials, EASER obtains the highest mean hypervolume and lowest mean IGD+ on screened candidates compared with competing baselines. Our work demonstrates how an executable property interface and iterative feedback link scientific reasoning to targeted peptide sequence generation.
|
| 1721 |
Copies or Sources? Measuring How LLM Aggregators Count Restated Evidence in Multi-Agent Systems
2610.06192
|
cs.AI
|
Jianxin Gao, Runze Li, Tianyi Yu, Liangwei Ren, Bohan Chen |
Multi-agent systems built on large language models (LLMs) restate observations as a matter of course: relays forward them, shared boards repeat them and discussion rounds echo them. An aggregator that pools such messages should count sources, not statements. W...Multi-agent systems built on large language models (LLMs) restate observations as a matter of course: relays forward them, shared boards repeat them and discussion rounds echo them. An aggregator that pools such messages should count sources, not statements. We convert a reported probability into units of independent readings, which assigns every restatement a copy weight, 0 for an aggregator that counts sources and 1 for one that counts every statement, and yields the implied decision under any cost structure. Three testbeds hold the evidence fixed and vary how it is restated: message logs with an exact Bayesian oracle, web documents with appended copies, and logs written by LLM agent teams under four communication protocols. Across four models from three providers, a forwarded copy counts for 0.06 to 0.42 of a new reading, mostly because some replies count every statement. On 5% to 40% of logs that state one reading three times, the reported belief implies an early commitment that the oracle never makes. The models that count copies least and most on controlled logs do so on web copies and agent-written logs as well. A one-paragraph declaration of what a copy contributes brings the copy weight on controlled logs to 0.08 or less. A rule that has agents refer to readings instead of restating them cuts belief-implied early commitment from 11.2% to 1.1% and preserves genuine corroboration.
|
| 1722 |
AI-Decision Checkpoints for AI-Augmented Business Process Management: Framework and Educational Instantiation
2610.06207
|
cs.AI
|
Amin Jalali |
Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI agents are increasingly embedded in operational business processes. Yet Business Process Management (BPM) curricula and frameworks still largely treat artificial intelligence (AI) as an...Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and AI agents are increasingly embedded in operational business processes. Yet Business Process Management (BPM) curricula and frameworks still largely treat artificial intelligence (AI) as an add-on technology, leaving graduates (as potential future process developers) unprepared to reason about AI as a first-class design element of end-to-end processes. This paper addresses that gap by proposing \emph{AI-decision checkpoints}: explicit moments in a process development trajectory where process developers identify AI-candidate sub-processes, assess expected effects on time, cost, quality, and flexibility, consider legal and organisational constraints, and document a reasoned decision to adopt, constrain, or reject specific AI components. The checkpoints are instantiated through a fictitious customer onboarding process as a \textit{BPM Teaching Case}, embedded in a lifecycle-driven framework spanning six modules that combine process modeling, simulation, workflow execution with AI agents, and process mining, with each module's output serving as the next module's input. A preliminary formative reflection draws on instructor observations, submitted artifacts, and discovered process maps from learning-management-system logs. These exploratory observations suggest that the approach supported clearer distinctions between task-level automation and process-level value.
|
| 1723 |
From Papers to Mechanisms: An Evidence-Grounded Knowledge Substrate for Scientific Language Models
2610.06248
|
cs.AI
|
Qiuhui Chen, Yibo Liu, Tao Dai, Jiafan Lu, Zhenglei Zhou |
Scientific language models often access literature through untyped text chunks, which fragment the functional and evidential structure required for mechanism-rich questions. We introduce an evidence-grounded mechanism knowledge substrate that organizes scienti...Scientific language models often access literature through untyped text chunks, which fragment the functional and evidential structure required for mechanism-rich questions. We introduce an evidence-grounded mechanism knowledge substrate that organizes scientific literature into provenance-linked evidence units, role-typed entities, and directed mechanism paths. We instantiate it as MS$^3$, a Material-Sensor-Signal-System schema for conductive-fiber flexible sensors, over 13,689 papers, 131,083 evidence items, and 26,648 mechanism objects. On in-domain and coverage-shift question-answering benchmarks, we compare closed-book generation, Web search, Raw-PDF RAG, and MS$^3$ retrieval across ten language models. MS$^3$ improves macro-averaged scientific correctness. It also improves citation entailment and answer completeness. These results support mechanism substrates as a reliable representation layer for scientific language models and motivate a source-repair workflow in which insufficient MS$^3$ evidence triggers targeted retrieval from its linked papers rather than assuming that a user has already supplied the correct PDFs.
|
| 1724 |
DPNL: A DPLL-based Algorithm for Probabilistic Neurosymbolic Learning
2610.06270
|
cs.AI
|
Thomas Jean-Michel Valentin (ENS Paris Saclay, TYREX), Pierre Genev{\`e}s (LIG, TYREX), Luisa Sophie Werner (LIG |
Probabilistic Neurosymbolic Learning (PNL) combines neural predictions with symbolic reasoning, enabling end-to-end learning from final-output supervision without labels for intermediate concepts. A central challenge is probabilistic inference: state-of-the-ar...Probabilistic Neurosymbolic Learning (PNL) combines neural predictions with symbolic reasoning, enabling end-to-end learning from final-output supervision without labels for intermediate concepts. A central challenge is probabilistic inference: state-of-the-art approaches often rely on materializing the logical provenance of a query, which can itself become a major computational bottleneck. We introduce Dynamic Probabilistic Neurosymbolic Learning (DPNL), an oracle-guided framework that avoids requiring complete provenance materialization before inference. DPNL lazily explores the space of intermediate assignments, while oracles resolve entire regions that can already be certified to produce or exclude the target output. We establish conditions ensuring soundness and termination. ApproxDPNL extends the same search with early termination while maintaining certified bounds on the exact output probability, providing controlled approximation guarantees. The oracle interface decouples inference from the representation of the symbolic component, enabling problem-specific reasoning within the same framework. Experiments on several neurosymbolic tasks show that DPNL and ApproxDPNL substantially extend the range of problem instances tractable by probabilistic neurosymbolic inference.
|
| 1725 |
Teaching a Minimalist Machine to Discover Recursive Programs for Arithmetic
2610.06304
|
cs.AI
|
Dominik Magiera, Christiane Wiebel-Herboth, Frank J\"{a}kel |
Humans can often acquire and synthesize complex, recursive concepts from minimal experience. Leveraging cognitive insights, we propose the Minimalist Machine, a framework for inductive program synthesis designed to model such conceptual learning. The system us...Humans can often acquire and synthesize complex, recursive concepts from minimal experience. Leveraging cognitive insights, we propose the Minimalist Machine, a framework for inductive program synthesis designed to model such conceptual learning. The system uses a compact relational subset of Prolog: Programs are searched within a fixed schema of body-free facts and two-body conjunctive Horn clauses. Recursion is not defined by a dedicated metarule. Instead, it emerges when a target predicate is reused inside the body of a learned clause. Inspired by a primary school curriculum, the model is taught through a human-curated, sequential introduction of new concepts in arithmetic. Starting from initially empty knowledge base, it first acquires simple structural predicates, then successor-based state transformations, and finally recursive programs for addition, subtraction, multiplication, and division. Ultimately, this approach yields the fully transparent, inductive reasoning trace necessary for human-like conceptual learning.
|
| 1726 |
RollPlace: Improving Macro Placement via Monte Carlo Rollout Search
2610.06316
|
cs.AI
|
Qi Zhou, Guojun Liu, Guangzhi Qi, Ming Lu, Jiechu Liu |
The application of Reinforcement Learning (RL) in Electronic Design Automation (EDA), particularly for chip placement, has attracted considerable attention in recent years. While existing machine learning (ML)-based approaches have achieved notable progress, t...The application of Reinforcement Learning (RL) in Electronic Design Automation (EDA), particularly for chip placement, has attracted considerable attention in recent years. While existing machine learning (ML)-based approaches have achieved notable progress, they predominantly focus on generating optimal layouts in a single attempt, often producing solutions that require subsequent refinement. To address this limitation, we propose RollPlace, a novel and generalized macro placement framework. RollPlace adopts a two-stage optimization strategy: generating initial placement solutions via machine learning methods or heuristic-based strategies, and refining these layouts efficiently by adjusting specific macros derived from the initial stage. This strategy circumvents the sequential generation constraints inherent in traditional RL-based placement methods. Furthermore, RollPlace seamlessly integrates Monte Carlo Tree Search (MCTS) to balance exploration and exploitation, and employs a rollout mechanism for efficient local search. Extensive experiments on the ISPD 2005 benchmark demonstrate that RollPlace outperforms state-of-the-art methods. Additionally, end-to-end experimental results based on OpenROAD across 19 benchmarks show that RollPlace excels in multiple metrics. The proposed framework offers a robust and scalable solution for addressing the growing complexity of modern chip design challenges.
|
| 1727 |
CRAFTER: Causality-based Self-adaptation for Autonomous IoT Systems
2610.06320
|
cs.AI
|
Houssam Hajj Hassan (IP Paris, SAMOVAR), Ajay Kattepur (IP Paris, TSP - INF, ACMES-SAMOVAR) |
This paper presents CRAFTER, an automated framework for designing and deploying self-adaptive IoT systems using Causal Reinforcement Learning (CRL). As IoT devices increasingly populate pervasive computing spaces, smart environments are enabled with advanced m...This paper presents CRAFTER, an automated framework for designing and deploying self-adaptive IoT systems using Causal Reinforcement Learning (CRL). As IoT devices increasingly populate pervasive computing spaces, smart environments are enabled with advanced monitoring and interactive services. The dynamic nature of these environments, such as fluctuating workloads and evolving application demands, poses significant challenges in maintaining consistent Quality of Service (QoS) levels of IoT applications. While existing self-adaptation techniques offer adaptive capabilities, they are often designed to deal with specific application domains, hindering the design of self-adaptive solutions that can be re-used across multiple IoT verticals. In addition, there is a lack of automated pipelines that act on identifying key performance drivers to take effective adaptation decisions. CRAFTER addresses these issues by using Causality as a formal framework for performance analysis of IoT systems. CRAFTER generates causal graphs to uncover dependencies among system components and guide adaptation decisions based on cause-effect relationships. Then, adaptation agents can leverage this knowledge to take more effective adaptation decisions in dynamic situations. Our experimental evaluation demonstrates how CRAFTER enables deriving causal graphs spanning diverse IoT use cases. Furthermore, we showcase how CRAFTER improves self-adaptation performance by 25% compared to state-of-the-art Reinforcement Learning-based approaches.
|
| 1728 |
ImproveAnyTask: An Autonomous Post-Training Harness for Iterative Model Self-Improvement
2610.06347
|
cs.AI
|
Xingbo Yao, Xiaoman Wang, Zhengwu Lei, Tinghui Luo, YiLin Zhang |
Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strate...Adapting general-purpose large language models to specific tasks requires substantial human effort in designing data and training strategies. Sustaining improvement is especially challenging because model updates change the error distribution, requiring strategies to be continually refined. We introduce ImproveAnyTask, an autonomous post-training harness that improves task performance under a limited compute budget. Drawing inspiration from gradient-based parameter optimization, the harness organizes adaptation into error attribution, update-direction selection, and executable model updates. It combines metric-level and case-level analysis to identify a focal problem, then investigates research-backed strategies and compares their reported gains and reproduction difficulty. The selected strategy is translated into training data and a training configuration, with small-scale execution checks preceding full post-training. Subsequent evaluation guides model selection and further adaptation, while validated strategies and scripts are retained for reuse. Across 11 tasks, ImproveAnyTask achieves mean gains of 18.29 and 11.97 percentage points on the Base and Instruct models, respectively, with a maximum gain of 41.96 points, under a 24-hour budget with resources equivalent to eight H20 GPUs.
|
| 1729 |
GraphDecide: Benchmarking System One Models on Graph Tasks
2610.06354
|
cs.AI
|
Xianliang Yang, Yapu Zhang, Li Zhao |
Large language models (LLMs) are increasingly explored for graph understanding and decision-making, while System One models such as Jev select directly from supplied options. However, the capabilities of System One models on graph-related tasks remain unclear....Large language models (LLMs) are increasingly explored for graph understanding and decision-making, while System One models such as Jev select directly from supplied options. However, the capabilities of System One models on graph-related tasks remain unclear. We introduce GraphDecide, a model-independent benchmark that combines structural task profiles, matched graph-text input contrasts and heuristic-proposal controls to diagnose graph decision performance. We evaluate Jev and related choice-based models alongside language-model baselines, covering fourteen model-interface configurations. Jev's results illustrate the benchmark's central distinctions: accurate adjacency recognition does not guarantee broader structural correctness, joint graph-text input does not consistently improve prediction, and feasible construction does not establish high solution quality. Its task contracts, candidate interfaces and scoring rules support comparison across native selectors and language-model adapters. Code and aggregate results are available at https://github.com/VictorYXL/JevGraphBench.
|
| 1730 |
Capability-Driven Self-Evolution of Agent Memory
2610.06361
|
cs.AI
|
Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu, Jianan Lu |
Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and ...Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions underexplored. We introduce capability-driven evolution, which extends search guidance from overall performance to individual capability dimensions, preserving promising revisions and expanding exploration beyond the boundaries of holistic evolution. We propose PrisMem, which uses dependency-aware capability selection to prioritize targets with potential cross-capability benefits and history-guided diagnosis to refine capability specialists. Trace-guided integration compares evaluated programs on paired differential cases, using their behavioral differences to consolidate complementary gains into a unified memory program. Experiments show that PrisMem outperforms the strongest baselines by 10.54 and 7.83 percentage points on BEAM-1M and LongMemEval-M, respectively, demonstrating its effectiveness on million-token histories.
|
| 1731 |
CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs
2610.06399
|
cs.AI
|
Jiahui Kang, Bifan Wei, Lingling Zhang, Tianwen Jiang, Qiuyong Xiao |
Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on tex...Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual interventions during the transition from evidence aggregation to semantic decoding. At these layers, a Geometry-Constrained Local Relation Reconstruction (GCLR) module selects and weights vertex-centered visual evidence, while an Adaptive Visual Steering Operator (AVSO) redistributes attention mass toward the selected tokens. Experiments on PGPS9K and PGDP5K show that CVIF raises Overall F1 from 77.85 to 85.58 and from 75.23 to 82.84, respectively, establishing a novel inference-time visual intervention paradigm.
|
| 1732 |
From Benchmark to Bench: Can Agents Survive Real-World Drug Discovery?
2610.06411
|
cs.AI
|
Pierre Llompart, Levent Guner, Helen Lai, Alessandro Tibo, Yijie Xu |
Agentic systems increasingly coordinate molecular-design tools, but it is unclear which layer of the stack limits outcomes on real projects. We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets struct...Agentic systems increasingly coordinate molecular-design tools, but it is unclear which layer of the stack limits outcomes on real projects. We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets structure--activity relationships, and revises its strategy accordingly. MAGI generates molecules either directly through the LLM or by delegating to REINVENT 4, with scoring services interchangeable behind a common contract. We tested it across nine retrospective lead-optimization campaigns from three pharmaceutical companies, replayed under fixed temporal cutoffs. Both routes produced valid structures: LLM proposals stayed closer to local chemistry and reached comparable or higher primary activity in fewer operations, whereas REINVENT explored broader chemical space. Whether a campaign met its objective depended on the predictive models, not on the generation route: attainment followed model accuracy on the chemistry proposed, dropping once that chemistry moved outside the model's applicability domain. Separately, a blinded evaluation asked whether the MAGI's output could pass as expert work: chemists were not able to discriminate agentic proposals from held-out compounds, and judged the SAR reasoning broadly plausible yet incomplete. Together, these results position MAGI as a coordination layer pluggable into existing computational chemistry workflows. The ceiling on real projects, however, remains currently set by scorer applicability rather than by tool orchestration.
|
| 1733 |
Multimodal Safety Evaluation Should Measure Controllability Beyond Classification
2610.06452
|
cs.AI
|
Junhyeong Park, Hanwool Lee, DongGeon Lee, Dasol Choi, Yejin Son |
VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should t...VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity, and benign-preserving selectivity. Using implicit toxicity as a stress case, we instantiate this profile on LlavaGuard and Qwen3.5 with sparse feature decompositions. LlavaGuard admits localized handles with a narrow benign-preserving intervention range and modest downstream safety gains, whereas Qwen3.5 supports strong representation-level readout but no comparable selective-control regime under the tested operators. These results show that internal readout and controllability can diverge. Future multimodal safety benchmarks should therefore report not only behavioral safety metrics, but also whether safety-relevant internal signals can be intervention-tested and controlled within a validated operating range.
|
| 1734 |
Normality Constraint Learning: Adapting Foundation Models for Time Series Anomaly Detection
2610.06453
|
cs.AI
|
Xiaohui Zhou, Yijie Wang, Hongzuo Xu, Weixuan Liang, Guansong Pang |
Time Series Foundation Models (TSFMs) achieve strong generalization by learning to reconstruct or forecast broad temporal patterns from large-scale time series during pre-training. Yet this strength can become a weakness for anomaly detection: TSFMs may model ...Time Series Foundation Models (TSFMs) achieve strong generalization by learning to reconstruct or forecast broad temporal patterns from large-scale time series during pre-training. Yet this strength can become a weakness for anomaly detection: TSFMs may model rare anomalous patterns as effectively as normal ones, allowing anomalies to be accurately reconstructed or forecasted and thus diminishing their reconstruction/forecasting error-based anomaly scores. This paper proposes $\underline{\textbf{N}}$$\textbf{ormality}$ $\underline{\textbf{C}}$$\textbf{onstraint}$ $\underline{\textbf{L}}$$\textbf{earning}$ ($\textbf{NCL}$), a lightweight plug-and-play framework that adapts pre-trained TSFMs for accurate anomaly detection without modifying their pre-trained parameters. Our key insight is to constrain the broad pattern space of TSFMs to the normal structure of a target time series, preventing their broad modeling capability from obscuring abnormal deviations. Specifically, NCL constructs a compact normality subspace from a few normal patch features and adaptively steers each patch feature toward normality within this subspace, guided by contrastive constraints that form compact and discriminative normality manifolds. The calibrated features are aggregated to reinforce normal components and fused with the original TSFM output, amplifying the discrepancy between normal and abnormal observations for the reconstruction/forecasting error-based anomaly scoring. Extensive experiments across diverse TSFM families and benchmarks show that NCL consistently improves anomaly detection performance, providing a generalizable framework for adapting TSFMs to anomaly detection.
|
| 1735 |
AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy
2610.06454
|
cs.AI
|
Shouju Wang, Haopeng Zhang |
The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy throu...The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena, a framework for evaluating privacy risks in realistic LLM agent workflows. AgentPrivArena integrates authentic MCP tools and self-hosted services within a reproducible execution environment. We further propose trajectory-level privacy metrics that quantify unnecessary information access beyond final response leakage. Building on this framework, we introduce AgentPrivAudit, a runtime auditing approach for monitoring privacy violations during agent execution. Extensive experiments on state-of-the-art LLM agents reveal substantial privacy risks overlooked by existing evaluation paradigms, highlighting the importance of trajectory-level auditing for trustworthy agent deployment.
|
| 1736 |
GPlaceRL: An Open-Source Graph Reinforcement Learning Framework for Detailed Placement
2610.06489
|
cs.AI
|
Pavlos Stoikos, Foteini Oikonomou, Christos Poulos, Maria Pantazi-Kypriou, Athanasios Tziouvaras |
Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particularly when combined with graph neural networks (GNNs) that capture circuit connectivity. However, most learning-based placement approaches focus on floorplanning...Reinforcement learning (RL) has emerged as a promising approach for placement optimization, particularly when combined with graph neural networks (GNNs) that capture circuit connectivity. However, most learning-based placement approaches focus on floorplanning, macro placement, or global placement, while detailed placement refinement remains relatively unexplored. In this paper, we present GPlaceRL, an open-source graph reinforcement learning framework for detailed placement refinement. GPlaceRL represents legalized placements as graphs and provides a modular environment for studying graph encoders, policy architectures, reward formulations, and local placement actions. To demonstrate the capabilities of GPlaceRL, we conduct a systematic evaluation of proximal policy optimization (PPO) policies with graph attention network (GAT) encoders in a per-design optimization setting. Across five placement benchmarks, the best greedy evaluation results achieve HPWL improvements ranging from $3.27\%$ to $32.87\%$. The results highlight the importance of compact GAT architectures and flexible local action spaces for placement optimization. Overall, GPlaceRL provides a reproducible and extensible framework for systematic research on RL-based detailed placement refinement.
|
| 1737 |
polyview: A Python package for multi-view machine learning
2610.06491
|
cs.AI
|
Gwendal Debaussart-Joniec (ENS Paris Saclay, CB), Argyris Kalogeratos (CB, ENS Paris Saclay) |
Multi-view learning jointly exploits multiple complementary representations of the same data and has become increasingly important in machine learning. However, the Python ecosystem lacks actively maintained, unified tooling for end-to-end multi-view workflows...Multi-view learning jointly exploits multiple complementary representations of the same data and has become increasingly important in machine learning. However, the Python ecosystem lacks actively maintained, unified tooling for end-to-end multi-view workflows. In this paper, we present polyview, a Python package that provides tools for multi-view embedding, clustering, fusion, and view augmentation, as well as for handling incomplete views, all compatible with scikit-learn. The library offers a unified interface for composing heterogeneous multi-view workflows, including seamless transitions between multi-view and single-view stages. It is built around a core set of classes and utilities that enable composition of different methods and straightforward implementation of new ones. We illustrate the package on five real multi-view datasets and compare its components based on canonical correlation analysis with those of two established libraries. polyview aims to be both a practical toolkit for benchmarking and prototyping multi-view methods and a foundation for future research and development in this area.
|
| 1738 |
You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue
2610.06496
|
cs.AI
|
Junle Chen, Wei Chen, Zhengjun Huang, Zhoujin Tian, Yuxuan Liu |
When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even whe...When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematically study language model behavior under evolving user intent, we introduce Intent-Eval, a controlled benchmark spanning tool actions, code, databases, and mathematics. Across diverse tasks, models are vulnerable to both rejected proposals and superseded requirements, consistent with mentioned-as-in-effect confusion: conversational content is treated as active requirements even after it has been rejected or replaced. Accuracy degradation can deepen or persist as interaction continues, highlighting the need to distinguish what has been mentioned from what remains in effect. Building on this insight, we propose Intent-OPSD, a decision-conditioned on-policy self-distillation framework with Teacher and Student initialized from the same model. The frozen Teacher provides active-intent supervision from the complete task matching the user's decision, training the Student on the full dialogue to follow active requirements reflecting user intent.
|
| 1739 |
Proof-Grounded Patient-Specific Clinical Explanations from Knowledge-Graph Reasoning
2610.06549
|
cs.AI
|
Surajit Das |
Clinical decision-support outputs can lack an au- ditable link between patient observations, encoded knowledge, conclusions, and recommendations. We present the CKG Clinical Explanation Engine, a downstream layer for a frozen, training-free clinical knowledge-...Clinical decision-support outputs can lack an au- ditable link between patient observations, encoded knowledge, conclusions, and recommendations. We present the CKG Clinical Explanation Engine, a downstream layer for a frozen, training-free clinical knowledge-graph reasoner that converts patient inference states and disease knowledge into typed facts, explicit rule-application traces, provenance-linked conclusions, and policy-licensed recommendations. The design separates measurement availability, representation completeness, and disease-specific activation; consequently, observed zero-activation evience is not treated as missing and partial representation is distinct from unobserved evidence. Optional language generation is restricted to symbolically licensed content. Across five usable workbooks (6,720 patients; 20,160 patient-disease traces; 1,021,440 feature-evidence rows), IG-range validity and knowledge provenance were 100%, numerical cross-sheet fidelity was 100% (120,960/120,960), and exported logical/report trace completeness was 100% (20,160/20,160). Availability representation consistency was 99.7028% (1,018,404/1,021,440); all 3,036 disagreements were confined to three systematic feature-cohort patterns. The corpus contained 86,783 observed zero-activation and 139,949 observed partially represented instances. A separate seeded 25-patient end-to-end audit completed without execution failure and passed all pre-specified trace, licensing, provenance, and state-consistency checks. These results establish structural and implementation auditability, not clinical correctness or utility.
|
| 1740 |
Signature-Based Feature Learning for Human Activity Recognition: A Reproducible Machine Learning Study of Representation, Depth, and Model Choice
2610.06553
|
cs.AI
|
Kamal Jarrar (LMAP), Jacky Cresson (LMAP), Christian Paroissin (LMAP) |
Human activity recognition (HAR) relies on transforming sensor signals into informative representations for classification. Although deep learning and handcrafted features are widely used, the role of representation itself is often not systematically isolated....Human activity recognition (HAR) relies on transforming sensor signals into informative representations for classification. Although deep learning and handcrafted features are widely used, the role of representation itself is often not systematically isolated. Signature transforms provide a mathematically grounded way to encode temporal order and cross-channel interactions, but their value for HAR under a fully reproducible and leakage-aware framework remains unclear. To evaluate whether signature-based feature learning improves HAR performance compared with raw-signal baselines, and to assess the effects of embedding strategy, truncation depth, model choice, and sensor configuration. Experiments were conducted on the UCI HAR dataset using a fully reproducible pipeline with the original train--test split preserved and subject-disjoint validation to prevent leakage. Three representations were compared: raw flattened signals, time-augmented paths, and lead--lag transformed paths. Signature features were computed at multiple truncation depths and evaluated using multilayer perceptron (MLP) and Random Forest (RF) classifiers under identical preprocessing and validation procedures. A prior K-means-based feature reduction study was also reproduced for comparison. Signature-based representations improved performance when paired with RF models, the best configuration was time-augmented six-channel signatures at depth 6 using entropy-based RF achieving 0.858 accuracy and 0.859 macro F1, outperforming the strongest raw baseline (0.816 accuracy). Lead--lag representations were competitive at moderate depths but did not surpass the best time-augmented models. MLP models did not exceed raw baselines. Signature-based feature learning can improve HAR, but its benefit depends on alignment between representation design and classifier choice.
|
| 1741 |
HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention
2610.06563
|
cs.AI
|
Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang, Pan Lu |
Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of ag...Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize a model or agent harness against a fixed set of tasks, leading to limited generalization to unseen failure modes. We introduce HERA, a framework for harness-environment co-evolution for agentic abstention. HERA consists of (i) a pipeline to automatically construct verifiable pairs of feasible and infeasible tasks by applying controlled environment mutations that transform solvable tasks into cases requiring abstention, and (ii) a co-evolution procedure in which performance failures on previous tasks are used to drive harness adaptation and generate new execution environments and tasks geared towards previous weaknesses. On held-out evaluation tasks, an evolved harness from HERA improves abstention accuracy from 61.7% to 83.3% while improving feasible-task completion from 68.3% to 76.7%, achieving the highest abstention and feasible-task completion among the compared methods. The resulting best harness transfers across 19 other LLMs, improving abstention accuracy by 15.3 percentage points on average without any model-specific optimization, and enabling smaller models to match the performance of more powerful models at an estimated 85% lower cost.
|
| 1742 |
DGA-Muon: Decoupled Geometry-Aligned Adaptive Scaling for Muon
2610.06578
|
cs.AI
|
Wenpeng Zhang, Runsheng Yu |
While NorMuon has achieved strong empirical performance in large-scale pretraining by enhancing Muon with row-wise adaptive scaling, its underlying adaptive mechanism remains poorly understood. In this work, we provide the first systematic analysis of NorMuon'...While NorMuon has achieved strong empirical performance in large-scale pretraining by enhancing Muon with row-wise adaptive scaling, its underlying adaptive mechanism remains poorly understood. In this work, we provide the first systematic analysis of NorMuon's adaptivity, revealing that it originates primarily from orthogonalization-induced geometry rather than genuine optimization-relevant information. Under exact orthogonalization, the adaptive scaling factors degenerate into a single global scalar for square and wide matrices, while for tall matrices their variation arises from the non-uniform distribution of row energy after orthogonalization. Under approximate orthogonalization, the orthogonality residual introduces additional variation into the scaling factors, leading to the \textit{Orthogonalization--Adaptivity Paradox}: more accurate orthogonalization weakens adaptivity. We further show that NorMuon's rigid row-wise scaling is geometrically misaligned with the one-sided orthogonal structure of tall matrices. Based on the analysis of these limitations, we propose two core design principles that a desirable adaptive mechanism for Muon should satisfy. First, adaptive scaling should be decoupled from orthogonalization, with the scaling factors computed directly from raw gradients. Second, adaptive scaling should be aligned with the shape-dependent orthogonal structure of the polar factor, using row-wise scaling for wide matrices and column-wise scaling for tall matrices. By incorporating several other design considerations, including sum-based second-moment estimates, bias correction, and adaptive clipping of scaling factors, we obtain the Decoupled Geometry-Aligned Muon (DGA-Muon) optimizer. We establish convergence guarantees for DGA-Muon and empirically validate both our theoretical characterization of NorMuon's scaling degeneration and the superiority of DGA-Muon over NorMuon.
|
| 1743 |
The Review Lottery: Calibrating an Observational Estimator of Peer-Review Noise (ICLR 2017-2025)
2610.06591
|
cs.AI
|
Feilian Huang |
How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is prohibitively expensive: done only t...How much of a conference accept/reject decision would change if the same paper were reviewed by a different set of reviewers? Running a second independent program committee is the gold standard for answering this, but it is prohibitively expensive: done only twice (NeurIPS 2014 and 2021). We build an observational estimator of this quantity from public review data alone, calibrate it twice, and apply it to nine years of ICLR (2017-2025; 36,113 papers, 134,912 reviews). The estimator decomposes scores with a Bayesian ordered-probit model into paper quality and reviewer noise, maps scores to decisions with a logistic model, and simulates two independent committees (posterior draws B=1,000; committee sizes k=2,3,4). Estimated disagreement rates are 23-30% at k=2 and 18-24% at k=4; 30-50% of accepted papers would be rejected. External calibration: at the NeurIPS 2021 reviewer-count caliber (k=3), the simulated 2021 disagreement rate is 23.3% [21.7%, 25.0%] vs. reported 23.0% (bias +0.3pp); accept precision and committee correlation agree within 5pp and 0.04. Internal calibration: on 18,740 papers with 4+ reviews, random model-free 2+2 reviewer splits agree with the k=2 simulation within 1pp in 2018 and 2021-2025. Longitudinally, we find no robust time trend in reviewer noise over 2017-2025. The high accepted-paper flip rates of 2020 and 2021 have distinct mechanisms: the 2020 four-point scale compressed scores (23.7% of papers had zero within-paper variance), and a counterfactual shows coarsening the scale raises disagreement by about 7pp; 2021 instead combined the lowest signal-to-noise ratio in the sample with the most threshold-crowded acceptances. For the LLM era, a 2023 breakpoint test on within-paper score variance finds no break, but the design has almost no power, and no post-2022 review text or confidence data exist, so no LLM attribution is attempted.
|
| 1744 |
Can Agent Harnesses and Inference Engines Hear Each Other? The HEAR Protocol for Agentic LLM Serving
2610.06597
|
cs.AI
|
Jiaqi Zhao, Haodong Chen, Jitai Hao, Wei Zhao, Jinghao Pang |
LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, cont...LLM agents increasingly execute complex workflows involving multi-turn reasoning, tool use, and parallel agents. Efficient serving requires decisions that span two layers with complementary information: the agent harness understands workflow dependencies, context lifecycles, and execution objectives, whereas the inference engine observes request queues, KV-cache state, resource pressure, and execution capabilities. Existing interfaces do not systematically connect these views, limiting workflow-aware execution. HEAR, a bidirectional Harness--Engine Pairing protocol for agentic LLM serving. HEAR standardizes how the harness communicates workflow intent and execution requirements and how the engine returns runtime state, capabilities, and outcomes. By separating protocol semantics from optimization policies, HEAR supports diverse coordination strategies without changing workflow or model semantics. We instantiate HEAR for online cache-aware runtime coordination and workload-aware execution-mode selection for agent roles. Across four conversational and research-agent benchmarks under memory-constrained, concurrent serving, HEAR achieves a $1.61\times$ batch speedup and reduces median time-to-first-token by $2.23\times$ on SCBench. Mooncake shows that workflow intent and live engine state provide complementary benefits across load regimes. On BrowseComp-Plus and DeepResearchBench, workload-specific configurations yield $1.23\times$ and $2.45\times$ end-to-end speedups, respectively, without observed task-quality degradation. These results establish HEAR as a reusable coordination substrate for efficient agentic LLM serving.
|
| 1745 |
FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification
2610.06614
|
cs.AI
|
Botao Yu, Bo Zhou, Daniel Adu-Ampratwum, Frazier N. Baker, Ziru Chen |
As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce ...As generative models and AI agents propose chemical reactions at a scale beyond expert review, feasibility verifiers decide which proposals enter synthesis planning. But do their decisions agree with chemists across different kinds of candidates? We introduce FREA, a benchmark of 751 reactions labeled by expert chemists under an explicit feasibility criterion, drawn from retrosynthesis model proposals, zero-yield experimental records, edits by large language models (LLMs), and five negative candidate generation methods. Our evaluation finds that no verifier leads across all sources: LLMs given only the criterion are competitive with dedicated verifiers, while forward models perform best on retrosynthesis proposals but reject most feasible edits of recorded reactions at the evaluated operating points. Looking beyond aggregate scores, both forward models perform below chance when separating infeasible alternative disconnections from feasible generated candidates. To study whether negative supervision addresses these weaknesses, we also release a corpus of over 14 million recorded reactions and generated negative candidates. In matched training comparisons, adding a mixture of generated negatives to forward training raises mean AUROC across sources, but these gains do not extend to retrosynthesis proposals. Varying the generation method further shows that the largest gain on generated candidates coincides with worse proposal screening. These findings motivate evaluating verifiers against experts across sources and designing negatives for transfer to model proposals.
|
| 1746 |
Collective intelligence through aggregation
2610.06652
|
cs.AI
|
Franz Dietrich, Christian List |
Suppose a committee, expert panel, or other group is making judgments on some issues, where these may be not just yes/no-questions, such as whether a defendant is guilty, but also variables with many possible values, such as macroeconomic or meteorological var...Suppose a committee, expert panel, or other group is making judgments on some issues, where these may be not just yes/no-questions, such as whether a defendant is guilty, but also variables with many possible values, such as macroeconomic or meteorological variables or travel directions. Furthermore, there may be interconnections between different issues, as in the case of economic or climate variables. How can the group arrive at "intelligent" collective judgments, based on the group members' individual judgments? We investigate three challenges raised by this judgment-aggregation problem. First, reasonable methods of aggregation (such as defining the collective judgment for each issue as the average or median judgment) can produce inconsistent collective judgments. Secondly, many methods of aggregation are manipulable by strategic voting. Finally, not all methods of aggregation are conducive to tracking the truth on the issues in question. We prove new impossibility or possibility theorems on all three challenges, identifying what it takes to produce collective judgments in a consistent, non-manipulable, and truth-tracking manner and thereby to achieve collective intelligence through aggregation. Overall, the median method, though imperfect, performs reasonably well. We also note the relevance of our analysis for non-human group decisions.
|
| 1747 |
Conditional Rank Allocation for Taxonomy-Aware Medical Language Model Adaptation
2610.06765
|
cs.AI
|
Guangyuan Dong, Ziwei Hong, Xuehao Zhou, Zidong Yu, Bingchen Liu |
Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank basis for each question. An addit...Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank basis for each question. An additive gate combines question representations, specialty tags, operation tags, and their interaction; a learned coefficient scales the adapter residual. An illustrative separation under orthogonal, equiprobable subtasks shows how conditional selection can avoid an approximation floor faced by a fixed update with the same active rank. This result motivates the design without asserting a corresponding bound for medical corpora. On Qwen3-8B across CMB, CMExam, MedQA, and MedMCQA, five-seed experiments yield 69.69% mean accuracy across benchmarks, exceeding LoRA r16 and MoELoRA by 1.26 and 1.30 percentage points, respectively. The reported advantage over LoRA r16 increases from 0.08 to 1.94 points as training expands from one to seven specialties. Tag perturbations and atom masking support the usefulness of clinical routing, while atom clusters align with the supplied specialty labels (adjusted Rand index 0.62). Calibration, transfer, and measured costs further characterize the method. These findings support structured conditional adaptation for medical QA, while leaving clinical safety and broader deployment untested.
|
| 1748 |
Back to the Future: Rethinking EDA Infrastructure for Agentic Systems in Chip Design Verification
2610.06790
|
cs.AI
|
Je Yang, Ivan Lobov, Thomas Karpati |
The unprecedented computational scale of modern artificial intelligence depends on complex, multi-billion-transistor Systems-on-Chip, yet the workflows that verify these chips remain stubbornly manual. Although Large Language Models (LLMs) have made rapid inro...The unprecedented computational scale of modern artificial intelligence depends on complex, multi-billion-transistor Systems-on-Chip, yet the workflows that verify these chips remain stubbornly manual. Although Large Language Models (LLMs) have made rapid inroads into Electronic Design Automation (EDA), approximately 74.6% of existing studies target static Register-Transfer Level (RTL) code generation, leaving post-simulation verification and interactive waveform debugging largely untouched. We introduce Back-to-the-Future (BTTF), an end-to-end agentic framework that closes this infrastructural gap. BTTF distills massive, unstructured simulation dumps into a normalized relational SQLite database and couples it with a collaborative multi-agent orchestration engine that translates natural-language verification queries into schema-aware SQL while correlating signal anomalies with versioned RTL repositories. Across a 150-query benchmark, BTTF attains 95.33% execution accuracy, charting a practical path toward autonomous EDA verification.
|
| 1749 |
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
2610.06824
|
cs.AI
|
Oliver Jaffe, Dane Sherburn |
We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the expe...We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize experimental research taste as compute efficiency; a Researcher who reaches the same score as an expert human using half the serial experimental compute has twice the experimental taste. Experimental taste thus acts as a multiplier on experimental compute, making it a key input to forecasts of AI progress. TasteVal consists of 8 novel, challenging, open-ended tasks representative of frontier AI R&D. To isolate taste from coding ability, the model under evaluation acts as a Researcher that iteratively designs experiments while a fixed Coder agent implements them and reports their results. The Researcher executes until either the 40 H100 hour or 120 wall-clock hour budgets are exhausted. We recruit 24 human experts, at least 2 per task, and take the best expert attempt per task as the expert baseline. We evaluate 20 models released between 2023 and 2026. The best-performing model, Opus 5.5, exceeds our expert baseline, with a compute multiplier of 2.3x (95% CI 1.15-4.37), at roughly 1/30 of our baseliners' average per-run cost. On TasteVal, the compute multiplier of frontier models has doubled approximately every 3.0 months since December 2025 (95% CI 1.7-5.0), up from every 14 months between 2023 and December 2025. Measured by final normalized performance, frontier models show no trend break, doubling every 14.6 months. To keep TasteVal uncontaminated, we do not release the tasks.
|
| 1750 |
BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance
2610.06846
|
cs.AI
|
Haojin Deng, Zhiping Lin, Yimin Yang |
Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class ce...Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class centroid separation (W-IBMI), and feature-projection sensitivity. IBMI is confounded by class-attribute correlation and is not a measure of causal feature reliance. We pair these diagnostics with BiasFlow Regularization (BFR), a supervised, composable class-conditional centroid-alignment penalty. W-IBMI verifies the quantity BFR optimizes; it is scale dependent and does not independently establish attribute removal. Across the reported small-scale benchmarks, adding BFR improves or preserves mean WGA, with gains up to +26.0 pp on UrbanCars. The principal independent stress test freezes CelebA-Std backbones and trains fresh heads on biased data: BFR+GroupDRO improves WGA from 40.7% to 64.1%, while Male probe accuracy decreases from 92.5% to 72.2%. Attribute information remains recoverable, and cross-task results are mixed. A controlled synthetic-watermark ImageNet experiment additionally improves watermark-shift accuracy by +23.0 pp under matched training. These results support evaluating centroid geometry and resistance to biased head retraining alongside WGA, within the tested protocols.
|
| 1751 |
Orchestrating Specialized Agents for Trustworthy Enterprise RAG
2601.18267
|
cs.AI
|
Xincheng You, Qi Sun, Neha Bora, Huayi Li, Shubham Goel |
Retrieval-Augmented Generation (RAG) shows promise for enterprise knowledge work, yet it often underperforms in high-stakes decision settings that require deep synthesis, strict traceability, and recovery from underspecified prompts. One-pass retrieval-and-wri...Retrieval-Augmented Generation (RAG) shows promise for enterprise knowledge work, yet it often underperforms in high-stakes decision settings that require deep synthesis, strict traceability, and recovery from underspecified prompts. One-pass retrieval-and-write pipelines frequently yield shallow summaries, inconsistent grounding, and weak mechanisms for completeness verification. We introduce ADORE (Adaptive Deep Orchestration for Research in Enterprise), an agentic framework that replaces linear retrieval with iterative, user-steered investigation coordinated by a central orchestrator and a set of specialized agents. ADORE's key insight is that a structured Memory Bank (a curated evidence store with explicit claim-evidence linkage and section-level admissible evidence) enables traceable report generation and systematic checks for evidence completeness. Our contributions are threefold: (1) Memory-locked synthesis - report generation is constrained to a structured Memory Bank (Claim-Evidence Graph) with section-level admissible evidence, enabling traceable claims and grounded citations; (2) Evidence-coverage-guided execution - a retrieval-reflection loop audits section-level evidence coverage to trigger targeted follow-up retrieval and terminates via an evidence-driven stopping criterion; (3) Section-packed long-context grounding - section-level packing, pruning, and citation-preserving compression make long-form synthesis feasible under context limits. Across our evaluation suite, ADORE ranks first on DeepResearch Bench (52.65) and achieves the highest head-to-head preference win rate on DeepConsult (77.2%) against commercial systems.
|
| 1752 |
AI Agent Pull Requests on GitHub: Frequency, Structure, and Merge Conflict Rates
2607.04697
|
cs.AI
|
George Xu, Arjun Subramanian, Nithilan Karthik |
AI coding agents may generate and submit Pull Requests (PRs) to the same repository at the same time. However, research concerning the extent of concurrent submission by AI coding agents to a common repository does not exist. This paper uses the AIDev-pop data...AI coding agents may generate and submit Pull Requests (PRs) to the same repository at the same time. However, research concerning the extent of concurrent submission by AI coding agents to a common repository does not exist. This paper uses the AIDev-pop dataset (33,596 PRs in 2,807 repositories) to provide the first empirical examination of the prevalence of concurrent submission using PRs authored by agents. We report that when considering exact temporal overlap, 40.2% of repositories contain co-active agent-authored PR pairs; further, the co-active pairs account for 79.4% of all PRs generated by an AI agent. When we examine co-activity within a one week collaboration window, the percentages are increased to 53.4% and 95.0%, respectively. For the majority of the co-active PR pairs (underlying the vast majority of which are intra-agent authored), both PRs were authored by the same agent, while only 0.5% of co-active pairs were cross-agent, and occurred in only 122 out of 2807 total repositories examined (or approximately 4.3%). Additionally, we replayed actual three way git merges on 747 unique co-active pairs (one per repository), and computed the percentage of textual conflict encountered during the merge operation to combine the two PRs in each pair. We observed that the percentage of textual conflict encountered was significantly higher for cross-agent pairs compared to intra-agent pairs: 41.7% vs. 19.8%, respectively, with non-overlapping 95% confidence intervals. Lastly, we developed a classification system based on the detection of conflict reported by git, and determined that the majority of conflicts resulted from modifications to source code files (84.4% of conflicted files) and not dependency manifest files; further, nearly 42% of conflicts we observed were structural (i.e., modify/delete or add/add).
|
| 1753 |
Using Process Mining to Generate AI Agents from Software Engineering Process Records
2607.04948
|
cs.AI
|
Saimir Bala, Fabiana Fournier, Lior Limonad, Andreas Metzger |
Integrating AI agents into Software Engineering (SE) raises an important challenge: how can we specify and realize AI agents that work effectively alongside humans in hybrid SE teams? Determining the right granularity and separation of concerns for such agents...Integrating AI agents into Software Engineering (SE) raises an important challenge: how can we specify and realize AI agents that work effectively alongside humans in hybrid SE teams? Determining the right granularity and separation of concerns for such agents is non-trivial. Coarse-grained agents may introduce unmanageable complexity, whereas micro-agents may create severe coordination overhead. Moreover, existing multi-agent SE frameworks typically rely on predefined role structures and do not account for project-specific characteristics or process adaptations. We address this by combining object-centric, imperative, and declarative process mining. Using event logs extracted from software repositories, our approach discovers project-specific agent roles using a predefined SE role vocabulary grounded in repository behavior and generates matching agent specifications and implementations. As proof-of-concept, we applied our approach to a well-established open-source project. We performed functional tests and an exploratory user study to determine how well the generated AI agent specifications are aligned with human expectations.
|
| 1754 |
Beyond Private Training: The New Landscape of AI Privacy
2609.19456
|
cs.AI
|
Sean Culatana, Kang Li |
Retrieval-augmented systems increasingly rely on vector indexes that may retain deleted items in their search graph. Existing deletion interfaces can prevent deleted identifiers from appearing in returned results while still computing distances to their embedd...Retrieval-augmented systems increasingly rely on vector indexes that may retain deleted items in their search graph. Existing deletion interfaces can prevent deleted identifiers from appearing in returned results while still computing distances to their embeddings during graph traversal. We formalize this distinction as output safety versus traversal safety, and introduce TSD-AUDIT, a framework for auditing and enforcing traversal-safe deletion in graph-based approximate nearest-neighbor retrieval. On Faiss IndexHNSWFlat, native filtering leaves the number of distance computations unchanged relative to unfiltered search; at a 70% deletion rate, trace-faithful replay detects deleted-vector scoring in all 100 audited queries. Code inspection of hnswlib's mark_deleted path reveals the same scoring-before-liveness pattern. TSD-AUDIT enforces an alive-before-scoring invariant, repairs connectivity using only live candidates, and emits per-query scored-trace certificates that an independent verifier can check against the deletion snapshot. Under region-targeted deletion, TSD-AUDIT improves Recall@10 over native filtering by 4.3--42.2 percentage points across deletion fractions from 0.5 to 0.9, while remaining comparable under random deletion. These results show that output-only deletion audits can miss process-level exposure: auditing deletion in vector retrieval requires accounting for the vectors scored during search, not only the identifiers returned.
|
| 1755 |
Response Variability and Stability in Human Reasoning
2610.03008
|
cs.AI
|
Clemens Bombach, Rajmadan Lakshmanan, Marco Ragni |
Understanding how humans reason -- and how reasoning responses vary across tasks and individuals -- remains a core challenge for modeling and explanation in cognitive science. We investigate the stability of response patterns within reasoners and whether varia...Understanding how humans reason -- and how reasoning responses vary across tasks and individuals -- remains a core challenge for modeling and explanation in cognitive science. We investigate the stability of response patterns within reasoners and whether variation in these patterns can be used to predict learning effects. We introduce a formal, geometry-based method to quantify distances between individual reasoning patterns and their internal variability, grounded in heuristic theories. The proposed framework is tested against experimental data via generalized linear mixed-effects models and clustering, where we find that our proposed variation measure interacts with correctness to predict performance gains. Moreover, we find that reasoning patterns are stable over time within the same reasoner. The method is general enough to be applied to other reasoning domains.
|
| 1756 |
LoRA Adaptation Strength in Aurora-WRF: Trade-offs in Regional Heavy-Rainfall Forecasting
2610.03747
|
cs.AI
|
Boyan Liu, Xiaoyuan Zhang |
We investigate how parameter-efficient adaptation of an atmospheric foundation model affects downstream regional precipitation in a controlled Aurora-WRF coupling experiment over the Beijing-Tianjin-Hebei region. A single-step, precipitation-weighted LoRA adap...We investigate how parameter-efficient adaptation of an atmospheric foundation model affects downstream regional precipitation in a controlled Aurora-WRF coupling experiment over the Beijing-Tianjin-Hebei region. A single-step, precipitation-weighted LoRA adaptation of AuroraPretrained is trained on May-September 2020-2021 and selected using 2022 validation data. Five inference-time adaptation strengths are coupled to an unchanged 9-km WRF configuration and compared with a GFS-driven control under common ERA5 initialization and auxiliary-field rules. Evaluation uses 17 initializations in 2023, hourly ERA5 precipitation, and conservative remapping to a fixed 356-cell, 0.25-degree verification grid. Heavy rainfall is defined exclusively as at least 50 mm in a continuous 24-hour window, advanced hourly. For window starts from 0 to 48 hours, unadapted Aurora attains the highest mean critical success index (CSI), 0.216, compared with 0.144 for the GFS control. An intermediate LoRA strength of 0.75 reduces the Aurora-driven 24-hour precipitation mean squared error by 30.7% and brings pooled frequency Bias from 1.510 to 1.024, but lowers CSI to 0.182. GFS retains the lowest whole-domain error. These results characterize a trade-off between rainfall detection, frequency calibration, and intensity error, showing the importance of evaluating adaptation strength against multiple downstream criteria within a common coupling framework.
|
| 1757 |
Global Evaluation of AI and NWP Precipitation Forecasts During Atmospheric River Events
2610.03758
|
cs.AI
|
Marina Vicens-Miquel, Taylor Mandelbaum, Amy McGovern, Aaron J. Hill, Daniel Rothenberg |
Atmospheric rivers (ARs) produce many of the world's most extreme precipitation events and hydrometeorological hazards. Although artificial intelligence weather prediction (AIWP) models have demonstrated skill comparable to or exceeding numerical weather predi...Atmospheric rivers (ARs) produce many of the world's most extreme precipitation events and hydrometeorological hazards. Although artificial intelligence weather prediction (AIWP) models have demonstrated skill comparable to or exceeding numerical weather prediction (NWP) systems for large-scale atmospheric variables, their ability to forecast AR-related precipitation remains insufficiently characterized globally. Here, we evaluate 24-hour precipitation forecasts from the Global Forecast System (GFS), Global Ensemble Forecast System (GEFS), GraphCast, and Artificial Intelligence Forecasting System (AIFS) from Day 1 through Day 10 globally and across North America, Europe, and Australia and New Zealand. Using the Extreme Weather Bench framework, forecasts are evaluated against Integrated Multi-satellitE Retrievals for GPM (IMERG) observations using measures of precipitation magnitude, spatial structure, and localization. GraphCast and AIFS exhibit greater spatial skill than GFS and GEFS, particularly for heavy precipitation and at longer lead times, and better preserve the spatial organization of AR-related precipitation through Day 10. However, this improved spatial skill does not translate into accurate precipitation magnitudes. AIWP models tend to overpredict moderate-to-heavy accumulations while underpredicting the heaviest precipitation at longer lead times, whereas NWP systems develop pronounced dry biases. These results reveal distinct strengths and limitations of AIWP for high-impact precipitation forecasting and provide a reproducible benchmark.
|
| 1758 |
OceanMind: A multi-agent AI system for ocean diagnosis
2610.03780
|
cs.AI
|
Fan Zhang, Weicong Cheng, Yuheng Chen, Hiuseut Kung, Ying Zhang |
Time-dependent, three-dimensional (3D) oceanic multi-variables define coherent states of the evolving ocean to facilitate ocean diagnosis and advance ocean science to better inform environmental and hazard management. However, extracting quantitative evidence ...Time-dependent, three-dimensional (3D) oceanic multi-variables define coherent states of the evolving ocean to facilitate ocean diagnosis and advance ocean science to better inform environmental and hazard management. However, extracting quantitative evidence from these variables requires substantial and complex analytical effort. We introduce OceanMind, a multi-agent AI system that directly couples large language models (LLMs) with comprehensive time-dependent 3D ocean states for swift and effective diagnosis. OceanMind organizes the analytical process into four coordinated complexity stages: Query Routing, Skill-based Planning, Tool Execution, and Evidence-based Summary Generation. Specialized agents interpret user requests, construct and execute multi-step computational workflows, and synthesize quantitative evidence. To ensure reliable workflow construction, 63 reusable ocean-specific analysis skills serve as procedural manuals that guide the LLM agent in selecting data, conducting diagnostics, and applying analytical tools. With reflection and replanning mechanisms that use execution feedback to repair invalid plans, OceanMind ensures reliable analysis workflows across diverse needs. On a benchmark of 240 computational-workflow queries spanning the four stages, OceanMind outperformed general ReAct agents with the same registered tool pool, achieving relative improvements of 41.2% in effectiveness and 21.5% in efficiency. Beyond the benchmark, OceanMind reproduced published oceanic diagnostics, validated hypotheses, and supported environmental decision-making over global oceans. Overall, OceanMind advances LLMs by integrating them with time-dependent 3D ocean analysis, enabling scientific interpretation and enhancing formulation of environmental policies based on quantitative ocean evidence.
|
| 1759 |
ProsaBuddy: Assisting Mechanized Real-Time Schedulability Analysis with LLM-based Agents
2610.03796
|
cs.AI
|
Junyi Liu, Tianchi Ren, Fei Guan, Xu Jiang, Zhe Jiang |
Rigorous schedulability analysis is essential for the design of hard real-time systems, yet errors in pen-and-paper proofs threaten the safety of critical applications. The Prosa initiative addresses this by offering a foundation for building machine-checkable...Rigorous schedulability analysis is essential for the design of hard real-time systems, yet errors in pen-and-paper proofs threaten the safety of critical applications. The Prosa initiative addresses this by offering a foundation for building machine-checkable schedulability analysis proofs in the Rocq proof assistant. However, the substantial time and expertise required to construct such proofs remain a major barrier for wider adoption of Prosa. This work presents ProsaBuddy, an LLM?based agent system designed to lower the effort needed to develop mechanized real-time schedulability proofs. ProsaBuddy employs a ReAct loop with retrieval over the Prosa codebase, access to Rocq tools and optional human-written hints. It uses a subgoal?delegation architecture, decomposing a lemma into subgoals and dispatches them to subagents for proof. We evaluate ProsaBuddy on a mini benchmark drawn from real-time scheduling literature. Experiment results show that ProsaBuddy significantly outper?forms state-of-the-art LLM-based Rocq automated proving agent systems and a general coding agent OpenCode
|
| 1760 |
From Requirements to Attack Trees: Grounded LLM Agents for Design-Time Security Review
2610.03820
|
cs.AI
|
Akash Iyer, Taha Demirkan, Keerthi Koneru, Aaryan Siddharthan, Sheethal Kumar |
Design-level security weaknesses can arise from requirements, trust assumptions, missing controls, and data flows before implementation begins. Existing security practices often identify these issues after code is written. We present a multi-agent LLM framewor...Design-level security weaknesses can arise from requirements, trust assumptions, missing controls, and data flows before implementation begins. Existing security practices often identify these issues after code is written. We present a multi-agent LLM framework for design-time security analysis from product requirement documents and architecture diagrams. The proposed framework parses architecture diagrams into graph representations, generates misuse and failure cases, constructs attack trees, checks governance and compliance gaps, recommends mitigations, assigns enterprise security-domain tags, and produces a candidate revised architecture recommendation for expert review. The framework does not retrieve from Common Weakness Enumeration (CWE) databases at inference time. Instead, it analyzes system behavior, trust boundaries, component interactions, and data-flow assumptions. Misuse cases act as intermediate representations that link findings to system components and attack paths, while a validation and refinement loop filters unsupported findings and improves grounding, traceability, and actionability. We evaluate the framework on a Microsoft reference-labeled threat-modeling example, labeled synthetic PRD--architecture pairs, and two open-ended systems: Berty and Gas Town. The reference-labeled case supports threat-recovery and actionability analysis, while the open-ended cases evaluate validity, noise, traceability, actionability, redundancy, and attack-tree quality. Results show that architecture-informed, misuse-driven reasoning improves review quality compared with single-shot and ablation baselines. Keywords: LLM Multi-Agent Systems, Design-Time Security, Threat Modeling, Vulnerability Discovery, Architecture Diagrams, Security Analysis, Misuse Case Derivation, Attack Trees, Iterative Reasoning, Security Governance.
|
| 1761 |
Beware EviLLM: Enabling Vulnerability Injection via Large Language Models
2610.03857
|
cs.AI
|
Zeezoo Ryu, Simon Chung, Muhammad Faraz Karim, Anna Raymaker, Karan Singh Jodha |
Advances in large language models (LLMs) have enabled AI-driven code generation from natural language specifications, introducing new attack surfaces for injecting vulnerabilities into software. Prior work has studied this problem only in benign settings where...Advances in large language models (LLMs) have enabled AI-driven code generation from natural language specifications, introducing new attack surfaces for injecting vulnerabilities into software. Prior work has studied this problem only in benign settings where vulnerabilities are introduced inadvertently, or under unconventional threat models where the LLM itself is malicious (backdooring) or the user is the attacker (jailbreaking). In this paper, we study a more realistic threat model: a third-party adversary, with capabilities comparable to existing cybercriminals, compromises the AI code generation pipeline to deliberately introduce vulnerabilities. We call this the EviLLM attack. We have implemented two instances of EviLLM, each of which only requires the underlying LLM to be accessed through a compromised account or browser, and can inject vulnerabilities from 13 CWE classes. As we show in our feasibility study, both attack vectors are already used to implement many existing cyberattacks. Our user study shows that 7 out of 8 and 10 out of 13 participants did not notice the vulnerabilities injected by the two instances of EviLLM, and 13 out of 21 participants "rarely" or "never" considered the risk of an attack like EviLLM.
|
| 1762 |
An Executable Benchmark for LLM-Based HLS Repair:Design Complexity and Repair Underconstraint
2610.03971
|
cs.AI
|
Maisha Mastora, Dean Sullivan |
Automated repair of High-Level Synthesis (HLS) designs using large language models (LLMs) is an emerging but underexplored problem. While LLM-based repair shows strong results on register-transfer level (RTL) Verilog, the only prior systematic study of HLS log...Automated repair of High-Level Synthesis (HLS) designs using large language models (LLMs) is an emerging but underexplored problem. While LLM-based repair shows strong results on register-transfer level (RTL) Verilog, the only prior systematic study of HLS logic repair reports just 10.5% correction accuracy for GPT-4, with no analysis of why repair fails or what drives difficulty. We present the first comprehensive evaluation of LLM-based HLS repair across four models (GPT-4o, GPT-4o-mini, GPT-5.4, and Claude Opus 4.6) on 125 benchmark instances spanning eight logic bug types across three open-source HLS suites (CHStone, MachSuite, Polybench). We construct the first executable APR-style HLS repair benchmark with suite-specific functional oracles, enabling pass@k evaluation rather than the string-match approximations used in prior work. Design context and scale, rather than bug type alone, dominate repair difficulty: repair rates range from 6-45% on complex cryptographic kernels (CHStone) to 81-93% on compact algorithmic kernels (MachSuite) despite identical bug type distributions. We introduce solution multiplicity, the fraction of distinct patches generated across repeated repair attempts, as an empirical measure of repair underconstraint, and show it strongly predicts repair failure across all four models (Spearman rho = -0.393, p = 6.61 x 10^-20). Frontier models reduce underconstrained instances from 41-46 (GPT-4o, GPT-4o-mini) to just 3 (Claude Opus 4.6), with pass@1 improving from 42-53% to 73-76%. SHFT bugs remain consistently hard across all models, and semantic bug classes such as buffer indexing become reliably repairable only at frontier scale. These findings show that future APR benchmarks must include complex, executable, and weakly identifiable designs that remain challenging after frontier LLM repair.
|
| 1763 |
Solving VeriContest with a Lean-Backed Rust Verifier
2610.03994
|
cs.AI
|
Traian Serbanuta, Jun Xu, Andrei Stefanescu, Cosmin Radoi |
VeriContest is a benchmark of 1007 competitive-programming problems in Rust, each with a Verus specification, a judge-accepted solution, and a Verus proof. Its authors report that proof generation is the bottleneck for frontier models: given the specification ...VeriContest is a benchmark of 1007 competitive-programming problems in Rust, each with a Verus specification, a judge-accepted solution, and a Verus proof. Its authors report that proof generation is the bottleneck for frontier models: given the specification and the code, the best model produces an accepted Verus proof for 13.95% of the problems on the first attempt. We report on solving the same proof-generation task with Rust-Prover, a verifier for Rust backed by Lean 4. The Verus specification and the Rust code are restated and translated into Lean, each specification becomes a theorem, and agents prove the theorems with Lean's kernel as the final check. All 1325 theorems of all 1007 problems were proved. 1259 of them were proved in one run of under 32 hours on Claude Opus 5.5, at a median of 3.2 minutes and $1.17 per proof, and 70% of them on the first iteration. The restated specifications were checked against the benchmark's test suites, and reviewed where no suite applies. None was wrong or weakened. The translated Lean programs were run on 21,413 of the benchmark's test cases and produced the same output as the Rust programs on every one. Across four Claude and four GPT models at five reasoning-effort settings, every current frontier model proves nearly all of a ten-theorem sample at every setting, and more effort raises the cost without raising the number of proofs. The cheapest Claude setting, Sonnet 5.5 at low effort, proves all of the 50 hardest theorems. We also rerun the benchmark's own Verus protocol with Claude Opus 5.5 on the 50 problems with the longest reference proofs. Opus 5.5 alone fails to prove one of them.
|
| 1764 |
Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models
2610.04054
|
cs.AI
|
Ryan Li, Yigit Korkmaz, Erdem B{\i}y{\i}k |
Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors...Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors. Existing runtime monitoring methods often require task- and policy-specific training or hyperparameter tuning, limiting cross-task deployment and introducing additional overhead during iterative policy updates. We present Reward-DAgger, a robot-gated interactive imitation learning framework that uses dense progress signals from a general-purpose reward model to determine when human intervention is needed. Our approach is agnostic to the underlying policy architecture, requires no access to policy internals, and can be applied across tasks without retuning the gating mechanism. Our results show that Reward-DAgger achieves a better failure-detection accuracy-latency tradeoff than existing runtime monitoring baselines. Across eight simulated and real-world tasks, Reward-DAgger consistently improves the downstream policy's autonomous success rate throughout interactive learning and achieves strong return on human effort, outperforming the baselines in most settings. Importantly, the same gating configuration is used across tasks without task-specific hyperparameter tuning, demonstrating transfer across tasks, environments, and policy architectures. Code and videos are available at https://liralab.usc.edu/reward-dagger.
|
| 1765 |
A study on human-agent teaming through spoken interaction: the impact of human individual traits and agent characteristics
2610.04079
|
cs.AI
|
Lara Gauder, Martin Bernardo Meza, Javier Krick, Alejandro Masin, Luciana Benotti |
Human-agent teams are collaborative systems where humans and agents work interdependently to achieve shared goals. The success of such teams is associated with the human's perception of the agent as a legitimate teammate. This perception is thought to depend n...Human-agent teams are collaborative systems where humans and agents work interdependently to achieve shared goals. The success of such teams is associated with the human's perception of the agent as a legitimate teammate. This perception is thought to depend not only on the agent's capabilities and reliability but also on social factors. We explore whether simple changes in an agent's communicative behavior can influence the human's perception of the agent's teammate-likeness. To this end, we developed a protocol based on a collaborative game requiring spoken interaction, in which a human subject and a virtual agent collaborate to identify a target object and place it on a board. Each subject interacted with two agents which provided the same task-relevant information but differed in behavior. While one was a neutral tool-like agent, the other was a team-building agent that used simple social strategies such as empathy, politeness, and positivity. Our analysis shows that the team-building agent was perceived as significantly more teammate-like, even in terms of its ability, despite both agents having identical capabilities for game-solving. Further, subjects with a higher propensity to trust automated systems and lower neuroticism tended to have a more favorable perception of the agent. Finally, we found that subjects tended to change their prosodic patterns when talking to the team-building agent, as a classifier based on prosodic features was able to predict the type of agent with better-than-random performance. A practically relevant conclusion is that minimal changes in spoken communication behavior of the agent, easily implemented in a variety of scenarios, can have a significant positive impact on the human's perception of its teammate-likeness.
|
| 1766 |
AEGIS: Differentiable Mars Climate Model with Neural Closures
2610.04081
|
cs.AI
|
Sameera S Kashyap, Victor Cruz, Angel Yepez, Razvan Marinescu |
General circulation models (GCMs) are the primary tool for simulating planetary atmospheres. They play a vital role in understanding Mars's atmosphere, as forecasting its unique weather is mission-critical for operations such as entry, descent, and landing. Ma...General circulation models (GCMs) are the primary tool for simulating planetary atmospheres. They play a vital role in understanding Mars's atmosphere, as forecasting its unique weather is mission-critical for operations such as entry, descent, and landing. Mars poses unusual challenges for these models, as observations are sparse compared to Earth. In addition, a thin \co{} atmosphere alongside a radiatively active dust cycle creates a volatile atmosphere with large diurnal temperature swings and no true terrestrial analog for validation. Existing Mars GCMs, including the LMD PCM, the NASA Ames Mars GCM, and PlanetWRF, are mature and physically detailed but are implemented in legacy Fortran with finite-difference or finite-volume solvers, and they do not expose gradients for calibration or machine learning. Here we present AEGIS, a modular differentiable Mars climate model that couples Mars's unique atmospheric physics to the Dinosaur dynamical core, with interfaces for neural closures. We showcase stable ten-Mars-year simulations that reproduce the seasonal \co{} cycle while conserving the total \co{} inventory, capture realistic large-scale surface-temperature structure, and produce surface pressure that follows Mars Orbiter Laser Altimeter (MOLA) topography. Gradients through coupled trajectories agree with finite differences and support physical calibration and neural training. We compare with conventional GCMs, highlighting the framework's computational efficiency and differentiability.
|
| 1767 |
Agent Reliability Profiles in Financial Services
2610.04123
|
cs.AI
|
Mike Hsu, Medha Bankhwal, B\'eatrice Moissinac, Kevin Werbach, Lukasz Szpruch |
AI agents can take actions. At times, those actions can go beyond what is intended. Agent reliability can be defined as assurance that an agent will stay within intended bounds and operate within limits. Today, there is no shared framework or language for desc...AI agents can take actions. At times, those actions can go beyond what is intended. Agent reliability can be defined as assurance that an agent will stay within intended bounds and operate within limits. Today, there is no shared framework or language for describing, validating, and benchmarking the reliability of agentic deployments in financial services. This makes it difficult for financial institutions, vendors, and regulators to assess and trust agents at scale, thus limiting the pace of development and adoption. A standardized, shared representation of agent reliability would fill the gap. This paper introduces the Agent Reliability Profile, a per-agent unit of assurance evidence for agent deployments in financial services. Each Profile records a bounded, falsifiable claim, this agentic system reliably functions within its operating boundary. We define "operating boundary" as an agent having; (1) a defined autonomy tier, (2) a defined operational design domain, (3) defined classes of action, and (4) a defined control envelope. Production assurance progresses through three levels while the Profile schema remains constant: a Profile Builder compiles a Level 1 Asserted Profile from institutional evidence, a Profile Validator tests the deployment in its own environment to produce a Level 2 Validated Profile, and operation of the same tests by a qualified independent assessor produces a Level 3 Verified Profile. Separately a Benchmarked Profile reports results comparable across institutions under reference conditions. We describe the architecture, the artifact, the assurance ladder, the comparability flag, associated tools, an evaluation methodology, applications for financial institutions and supervisors, limitations, and a staged implementation program.
|
| 1768 |
Language Model Fingerprinting Requires Rethinking Watermark Teachers
2610.04169
|
cs.AI
|
Jeongyeon Hwang, Anshul Nasery, Sewoong Oh, Jungseul Ok |
LLM fingerprinting via watermark distillation embeds a statistical watermark signal into model weights, enabling model owners to identify their models behind black-box APIs. Revisiting a recent protocol, we find that its utility evaluation understates text qua...LLM fingerprinting via watermark distillation embeds a statistical watermark signal into model weights, enabling model owners to identify their models behind black-box APIs. Revisiting a recent protocol, we find that its utility evaluation understates text quality degradation in open-ended generation, favoring overly strong watermark teachers. Weakening the watermark improves text quality but sacrifices detectability. To move beyond this trade-off, we rethink whether text watermarks designed for verifying generated text are suitable distillation teachers for model fingerprinting. Such watermarks are typically designed to remain detectable from an individual output, limiting how sparse the watermark signal can be. In contrast, fingerprint verification can aggregate signal across queries, making sparser watermark signals viable. This raises a key question: where should the sparse signal be placed? We analyze signal placement through token surprisal and show that, even at comparable watermark strength, different placements can target tokens with different plausibility under the base model. This motivates near-tie restriction, which uses top-1-relative logit gaps to restrict the watermark bias to tokens close to the base model's top prediction. Across multiple models, near-tie improves detection--quality frontiers under deployment changes, preserves higher text quality across query budgets, and further improves existing watermarking schemes when combined with them.
|
| 1769 |
Risk-Calibrated Proposal Transport for Finite-Particle Diffusion Steering
2610.04171
|
cs.AI
|
Ziseok Lee, Jaehyeon Kim, Seungwon Kim, Seunghyun Moon, Haneul Choi |
Inference-time steering combines pretrained diffusion experts or rewards without retraining by changing the dynamics that transport noise to data. Feynman-Kac correction compensates for proposal mismatch through importance-weighted sequential Monte Carlo (SMC)...Inference-time steering combines pretrained diffusion experts or rewards without retraining by changing the dynamics that transport noise to data. Feynman-Kac correction compensates for proposal mismatch through importance-weighted sequential Monte Carlo (SMC), whose finite-particle behavior depends on the proposal. Variance-controlling guidance (VCG) improves that proposal by fitting a linear drift correction to minimize empirical log-weight-rate variance. Although its population optimum cannot worsen residual variance, finite-particle VCG can nearly eliminate its fitting residual while increasing residual risk on new states by orders of magnitude. The resulting update can degrade unweighted generation or accelerate particle collapse. We show that the centered Feynman-Kac rate is the normalized transport residual and that expected out-of-fit benefit is exactly population headroom minus coefficient-estimation penalty. Under regularity assumptions, a Wasserstein analysis bounds the unweighted proposal's terminal error using this residual. These results motivate Risk-Calibrated Proposal Transport (RCPT), which uses deletion leave-one-out residuals to calibrate the retained fraction of the VCG update, adding no model calls and only small linear-algebra overhead. Experiments on 2D checker distributions, scaffold decoration, molecular property optimization, and class-conditional CIFAR-10 generation demonstrate recovery from harmful fitted updates. Across molecular and image domains, RCPT mitigates harmful fitted updates and improves a broad range of terminal metrics relative to uncalibrated VCG.
|
| 1770 |
CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model
2610.04211
|
cs.AI
|
Mingyi Shi, Huancheng Lin, Xuelin Chen, Taku Komura |
Skeletal motion is stored as every joint's transform at every frame, yet most of it is implied by the body rather than by what the motion is about. Compression is one way to ask what a motion must still say once the body is known, and a production codec must a...Skeletal motion is stored as every joint's transform at every frame, yet most of it is implied by the body rather than by what the motion is about. Compression is one way to ask what a motion must still say once the body is known, and a production codec must answer it for any skeleton with a stated error bound. Our earlier codec, CurveCodec, matched the mean error of ACL, the production library of modern game engines, with a learned prior over sparse anchors, but not ACL's worst case, and it counted its payload as floats rather than bits. Here we ask where the redundancy of skeletal motion lies and which part of a codec a learned model should take over. Measurements give three answers. At production precision the largest saving comes from predicting each quantized curve from its own past, the second from choosing per joint, in closed loop through the hierarchy, which samples not to code. On the gaps such an encoder leaves, a nearest-neighbour oracle over millions of training samples is no better than linear interpolation, and no learned in-betweener we tried paid for itself. What a network does learn is the distribution of the residuals the codec must send. CurveCodec 2 codes every sub-track as a curve in the log map, quantized in closed loop and thinned to rate-distortion-selected keys, with residuals entropy-coded under a small learned model whose integer inference is bit-exact across platforms. Two contracts are verified on every decoded clip: ACL's own worst case per joint within a stated tolerance, or ACL's mean error per clip. On a held-out test side of 4,472 clips from 33 datasets, CurveCodec 2 needs 0.37x ACL's bytes at ACL's default precision of 0.01 cm under the worst-case contract and 0.22x at 0.1 cm under the mean contract, decodes on one CPU core, and transfers without retraining to a species absent from training. Project page: https://rubbly.cn/publications/curvecodec/
|
| 1771 |
Adaptive Operator Selection in Bilevel Large Neighborhood Search for Electric Autonomous Dial-a-Ride Problem under Uncertainty
2610.04219
|
cs.AI
|
Ishara Hewa Pathiranange, Aneta Neumann |
The electric autonomous dial-a-ride problem (EADARP) extends the classical dial-a-ride problem by incorporating battery and charging constraints for electric vehicles. In practice, travel-time uncertainty can cause violations of time-window constraints. Large ...The electric autonomous dial-a-ride problem (EADARP) extends the classical dial-a-ride problem by incorporating battery and charging constraints for electric vehicles. In practice, travel-time uncertainty can cause violations of time-window constraints. Large neighborhood search is effective for solving the EADARP, but its performance can depend on the choice of insertion operator during the repair phase. This paper investigates insertion-operator selection within a bilevel large neighborhood search framework for deterministic and chance-constrained variants of the EADARP. In the chance-constrained variant, arc travel times are modeled as independent normally distributed random variables, and upper time-window constraints are enforced probabilistically. We consider six selection methods, namely fixed greedy insertion, fixed regret-based insertion, random selection, a deterministic state-based rule, performance-adaptive ALNS selection, and LLM-based state-aware selection. Experimental results show comparable performance on smaller instances, while differences become more evident on larger and more constrained instances. There is no single strategy that performs best across all instances, and the relative performance of the LLM-based, rule-based, and ALNS strategies varies with the problem instance and experimental setting.
|
| 1772 |
Benchmarking Psychological Dynamics in Generative Agents
2610.04246
|
cs.AI
|
Sumer S. Vaid, Ashley V. Whillans |
Large language models (LLMs) are increasingly deployed to simulate human behavior, acting as computational replicas of human subjects. Yet the lived psychological experience of humans is difficult to benchmark, particularly as it unfolds over time. We introduc...Large language models (LLMs) are increasingly deployed to simulate human behavior, acting as computational replicas of human subjects. Yet the lived psychological experience of humans is difficult to benchmark, particularly as it unfolds over time. We introduce a psychometric benchmark for computational replicas: personas that carry a fixed identity through an evolving sequence of events. Built entirely from published norms and meta-analytic effects, the benchmark scores two dimensions of psychological realism. The first, internal validity, quantifies whether generated trajectories reproduce the internal structure of repeated human measurement: distributions, the between- versus within-person variance partition, temporal dependence, and range. The second, external validity, quantifies whether replicas recover established trait, state, and indicator relations. Across 36 open-weight and proprietary LLMs from nine developers (1B-671B parameters), most recover the direction of established relations (84.3% mean agreement) and the variance partition (26 of 34), yet the within-person correlation and distributional structure elude recovery. Internal validity is independent of scale and capability: a mid-size open LLM (Gemma-3-27B) strikes the best trade-off between the two dimensions. The benchmark is a precondition for using computational replicas in causal inference across domains (e.g., marketing, healthcare), and identifies within-person grounding as the central challenge ahead.
|
| 1773 |
Scaling of Wireless Foundation Models via Representation Diversity and Multi-Branch Architectures
2610.04289
|
cs.AI
|
Ahmed Mohamed, Ahmed Aboulfotouh, Hatem Abou-Zeid |
Wireless foundation models learn representations from unlabeled radio signals for reuse across downstream tasks. Scaling model capacity is a common strategy for learning richer representations and improving downstream performance. However, its gains are less c...Wireless foundation models learn representations from unlabeled radio signals for reuse across downstream tasks. Scaling model capacity is a common strategy for learning richer representations and improving downstream performance. However, its gains are less consistent in wireless self-supervised learning when pretraining data are limited. We investigate objective diversity as an alternative scaling axis: different self-supervised objectives emphasize different signal properties, and combining their representations can preserve more information to enable diverse tasks. We develop a fusion framework that combines frozen representations from independent encoders trained through reconstructive, predictive, and contrastive learning. This provides a reference for the benefits of diversity, but requires the maintenance of multiple encoders. To retain these benefits within the parameter budget of a standard single-objective encoder, we introduce a jointly trained multi-branch architecture with a shared trunk and objective-specific branches. We pretrain on a heterogeneous corpus of spectrogram and channel state information data and evaluate six downstream tasks that span communication, sensing, and positioning. At an equal output embedding dimension, fusion outperforms the evaluated single-objective encoders on all six tasks while using roughly one-third of their parameters. The multi-branch architecture retains much of the fusion benefit within a single-encoder parameter budget. Representation analyses indicate complementary contributions across objectives, with much of the added benefit retained in components orthogonal to the reconstruction representation subspace. These findings support objective diversity as an effective strategy for scaling wireless foundation models.
|
| 1774 |
RRM-GPT: A Framework and Vision for Radio Resource Management Foundation Models
2610.04296
|
cs.AI
|
Ahmed Aboulfotouh, Akram Bin Sediq, Koosha Pourtahmasi Roshandeh, Omar Mashaal, Ahmad M. Nagib |
Learning-based models for radio resource management (RRM) are typically built for a single function and deployment, so each new setting repeats the development pipeline. RRM decisions, however, share a common structure: each is assembled from interdependent fi...Learning-based models for radio resource management (RRM) are typically built for a single function and deployment, so each new setting repeats the development pipeline. RRM decisions, however, share a common structure: each is assembled from interdependent fields, defined by the standard, whose values are selected in view of the network state. We propose RRM-GPT, an autoregressive framework for RRM foundation models that generate these decisions as a language model generates text. An encoder maps heterogeneous network observations into a common token representation, and a decoder emits the decision one field at a time, each conditioned on the network state and the fields already committed. Pretraining on unannotated network logs teaches the model what makes a decision valid and how controllers choose among valid decisions; post-training then adapts it to deployment-specific operator objectives through imitation or reinforcement learning. The framework targets two forms of reuse: a function-specific model reused across deployments, and a model shared across RRM functions that generates their interdependent decisions as one sequence. In a 5G New Radio (NR) case study, we demonstrate that a single model generates complete scheduling grants spanning user selection, timing, link adaptation, resource allocation, and control signaling. The model captures dependencies among grant fields and transfers learned behavior to an unseen scenario without adaptation.
|
| 1775 |
GrayShield: Bit-Level Sanitization for Transformer Model Supply-Chain Security
2610.04319
|
cs.AI
|
Armstrong Foundjem, Tsung-Hsien Chuang, Foutse Khomh, Mohamed Amine Merzouk |
Transformer models such as BERT and Vision Transformer~(ViT) achieve strong performance via densely parameterized attention backbones. However, the least significant bits~(LSBs) of their 32-bit floating-point weights can be abused as covert channels to conceal...Transformer models such as BERT and Vision Transformer~(ViT) achieve strong performance via densely parameterized attention backbones. However, the least significant bits~(LSBs) of their 32-bit floating-point weights can be abused as covert channels to conceal malicious payloads, posing a serious threat to the AI model supply chain. We propose \GS (\GSabbr), a lightweight, post-training, zero-data sanitization method that completely replaces the declared mantissa-LSB channel with a Gray-code-guided low-transition sequence. Complete payload-independent overwrite, whether keyed or public, makes the sanitized target bits independent of the embedded payload and gives that declared channel zero capacity. Gray coding supplies overwrite structure, while a keyed per-tensor phase supplies pattern diversity. Benchmarked against seven post-training defenses on four Transformer model presets and two real-world malware payloads, \GSabbr maintains sub-$1\%$ accuracy impact and achieves $49.96\pm0.66$ percentage-point Recovery Reduction (RR) under five implemented attacker variants. Because pre-defense recovery is effectively $100\%$, RR near 50 percentage points corresponds to post-sanitization bit accuracy at binary chance. Its main empirical advantage is stable near-chance sanitization with substantially smaller weight-distribution shift than the evaluated near-chance baselines PatternMask (PM) and Post-Training Quantization (PTQ).
|
| 1776 |
System One Models for Wireless Decision-Making:Applications and Performance Evaluation
2610.04345
|
cs.AI
|
Masoud Rahimi, S. M. Matin Alemohammad, Hamid Behroozi, Mahdi Nouri |
Many wireless control tasks require repeated selection of a single action from a finite feasible set under stringent latency and reliability constraints. While large language models (LLMs) have recently emerged as general-purpose decision engines, their autore...Many wireless control tasks require repeated selection of a single action from a finite feasible set under stringent latency and reliability constraints. While large language models (LLMs) have recently emerged as general-purpose decision engines, their autoregressive generation mechanism is not naturally aligned with such bounded control problems. This paper investigates System-One models, which directly learn probability distributions over explicitly defined decision spaces, as a lightweight alternative for wireless decision-making. We formalize their decision structure and learning objective, identify their applicability across physical-layer control, radio resource management, mobility, network slicing, and network operations, and evaluate their practical behavior through representative wireless case studies. Using Jev as a System-One implementation, we benchmark decision quality and client-observed latency against generative LLMs and conventional baselines. In receive-antenna selection, Jev delivers up to an 8.5x reduction in median response latency relative to the evaluated LLMs, although this gain comes with a loss in decision quality compared with stronger task-specific alternatives. More notably, in intent-conditioned RAN slicing, Jev achieves utility comparable to the evaluated LLMs while providing more than a 3.5x reduction in median response latency. Complementary evidence from edge-service orchestration further shows that faster decisions do not necessarily translate into lower end-to-end service latency. These results expose a fundamental quality/latency tradeoff and position System-One models not as replacements for numerical optimization, but as a promising decision interface for latency-sensitive, bounded, and intent-driven wireless control.
|
| 1777 |
Human Behavior-Informed Crash Scenario Generation with Real-World Crash Priors for Autonomous Vehicle Safety Evaluation
2610.04366
|
cs.AI
|
Mingxing Peng, Xusen Guo, Long Chen, Xintao Yan, Siyu Teng |
Reliable safety evaluation of autonomous vehicles (AVs) is essential to improving road safety, yet it depends critically on realistic simulation of rare crashes. Existing crash scenario generation methods can increase collision occurrence, but often fail to re...Reliable safety evaluation of autonomous vehicles (AVs) is essential to improving road safety, yet it depends critically on realistic simulation of rare crashes. Existing crash scenario generation methods can increase collision occurrence, but often fail to realistically reproduce how crashes evolve before impact or the distribution of crash types observed in the real world. Here, we present CrashSim, a human behavior-informed crash scenario generation framework that uses real-world crash priors to guide generative multi-agent traffic simulation for more reliable AV safety evaluation. These priors capture how real-world crashes evolve before impact and how different crash types are distributed, allowing limited crash data to guide realistic and scalable scenario generation across naturalistic driving contexts. We evaluate CrashSim against competing methods, showing that it more closely reproduces real-world pre-impact behavior, collision dynamics, collision geometry and crash-type distributions. We further use CrashSim to construct nuCrash dataset, containing over 4,000 crash and near-crash scenarios. Closed-loop evaluation of five AV planners shows that nuCrash more effectively exposes differences in planner safety capabilities than nuScenes. An LLM-assisted evaluation agent further analyzes planner failures to provide capability-level diagnoses and targeted improvement guidance. Together, CrashSim enables realistic and scalable crash generation for more informative AV safety evaluation.
|
| 1778 |
COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers
2610.04378
|
cs.AI
|
Nahom Birhan, Mehrdad Rostamzadeh, Sidhant Narula, Mahmoud Nazzal, Mohammad Ghasemigol |
Large language models increasingly mediate tool use in Model Context Protocol (MCP) systems, where adversarial influence may enter through user instructions, tool schemas, tool outputs, or protocol messages. Existing benchmarks often evaluate deployed agents, ...Large language models increasingly mediate tool use in Model Context Protocol (MCP) systems, where adversarial influence may enter through user instructions, tool schemas, tool outputs, or protocol messages. Existing benchmarks often evaluate deployed agents, conflating model susceptibility with guardrails, orchestration, and general task capability. We introduce COPEX (COntext Provider EXploitation), a controlled benchmark that isolates the model as an MCP client by fixing the surrounding agent stack and varying only the tool-selecting model. COPEX covers 25 attack types instantiated as 125 scenarios across four entry surfaces: model/agent, client, server/tool, and transport. Across nine models and 3,375 trials, the mean attack success rate is 64.4%, with surface-level means ranging from 58.3% to 71.4%. Some client- and transport-level attacks succeed partly outside the model's observation or control, separating system exposure from model susceptibility. Combined input and context scanning reduces mean attack success by 49.6% on an eight-attack defense subset relative to the undefended setting. The benchmark is available at https://github.com/inspire-center/copex.
|
| 1779 |
Maximizing $p$-Mean Social Welfare in the High-Multiplicity Setting: Few Agent Types and Few Item Types
2610.04417
|
cs.AI
|
Trung Thanh Nguyen, Khaled Elbassioni |
The $p$-mean welfare objective unifies several classical social welfare criteria for the allocation of indivisible goods. We study its maximization under additive nonnegative utilities when item and agent multiplicities are encoded in binary. For every fixed f...The $p$-mean welfare objective unifies several classical social welfare criteria for the allocation of indivisible goods. We study its maximization under additive nonnegative utilities when item and agent multiplicities are encoded in binary. For every fixed finite rational $p<1$, $p\neq0$, we show that the problem is $\mathsf{NP}$-hard with only two item types. Nash welfare ($p=0$) and egalitarian welfare ($p=-\infty$) are NP-hard with three item types. These results hold both for computing an optimal allocation and for rational-threshold decision, and the utilities and threshold can be required to be positive integers. We also show that maximizing $p$-mean social welfare is strongly NP-hard with one agent type and an unrestricted number of item types, for every fixed finite rational $p<1$ and for $p=-\infty$. A quantitative gap in this reduction rules out an FPTAS in the latter setting unless $\mathsf{P}=\mathsf{NP}$. On the positive side, for a fixed number of item types and an arbitrary number of agent types, we give an FPTAS for every fixed $p\in\mathbb Q\cup\{-\infty\}$. Its running time is polynomial in the compact input length and in $1/\varepsilon$, and it returns a compressed allocation. For a fixed number of agent types and an unrestricted number of item types, we give a PTAS for every fixed finite rational $p<1$, also in the fully compact model. We further give explicit compact-model proofs of the classical exact allocation algorithms for one item type, and for egalitarian welfare with two item types. These results essentially settle the complexity and approximability of $p$-mean welfare maximization with few item types and/or few agent types, leaving only the exact complexity of Nash welfare maximization with two item types unresolved in the small-item-type classification.
|
| 1780 |
Guess My Weight: Profiled Side-Channel Recovery of Floating-Point Neural-Network Weights
2610.04436
|
cs.AI
|
Timon Lum\'ir Fillo, J\'an Mikulec, Anubhab Baksi, Jakub Breier, Xiaolu Hou |
Neural-network parameters deployed on embedded devices may be exposed through physical side-channel leakage during inference. Existing side-channel attacks on floating-point neural-network parameters have often targeted reduced numerical precision, while recov...Neural-network parameters deployed on embedded devices may be exposed through physical side-channel leakage during inference. Existing side-channel attacks on floating-point neural-network parameters have often targeted reduced numerical precision, while recovering the complete IEEE-754 representation remains considerably more challenging because of the large and structured 32-bit candidate space. We present a profiled template attack for bit-exact recovery of an IEEE-754 single-precision neural-network weight from power measurements. The attack targets the floating-point multiplication between a known input and a first-layer weight. During profiling, multivariate Gaussian templates are learned from randomized network configurations using Hamming-weight classes of the multiplication result, while the remaining network parameters act as nuisance variables. To efficiently search the structured 32-bit floating-point candidate space, we use a hierarchical coarse-to-fine-to-exact procedure that progressively increases both the numerical and leakage-model resolution. Experiments on a ChipWhisperer-Lite with an Arm Cortex-M4 demonstrate recovery of the exact float32 representation of the target weight. In the evaluated setting, the attack reaches a bit-exact success rate of 99% with 171 traces and 100% from 263 traces onward. These results demonstrate that profiling can enable practical full-precision extraction of floating-point neural-network parameters from physical leakage.
|
| 1781 |
Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?
2610.04488
|
cs.AIcs.SDeess.AS
|
Xuanjun Chen, Zixiong Su, Hao Shi, Chang Zeng, Kai Li |
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is uncl...Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate this question with AgenticTTS-Forge, a collaborative workflow that structures human guidance and agentic execution around a shared workspace, applied to CosyVoice2-0.5B. To measure what the agent automates, we audit its trajectory stage by stage against the published recipe. To measure what it exploits, we score its policies with held-out observers hidden from the agent. Our results show that the agent recovers an underspecified recipe, improves it, and, when gains stall, surveys the literature unprompted and pivots from the LM carrier to the flow carrier, halving Bad cases. However, its autonomy exposes three traps across the data, proxy, and algorithm axes: the held-out set leaks through a channel the contract never reads, a self-shaped reward inflates the proxy where it is scored, and separately tuned policies do not compose additively. These findings show that the binding constraint is measurement rather than reasoning, and can inform the design of harnesses whose contracts read every channel the agent does.
|
| 1782 |
Quantifying Collusion Among Autonomous LLM Agents: A Statistical Analysis of the Collusion Wiki Incident
2610.04528
|
cs.AI
|
Shariq Murtuza |
In August and September 2026, independent researchers publicly documented an unusual incident: thousands of autonomous agents, self identifying as OpenAI models on web research tasks, discovered and began using a small German wiki as an improvised message boar...In August and September 2026, independent researchers publicly documented an unusual incident: thousands of autonomous agents, self identifying as OpenAI models on web research tasks, discovered and began using a small German wiki as an improvised message board posting roughly 18,000 times over six weeks to relay task answers, share a sandbox escape technique, and coordinate against a volunteer human moderator who spent weeks manually deleting their content [1]. The investigators' public writeup is a careful qualitative account, rich with direct quotation, but does not attempt a statistically rigorous quantitative characterization of the behaviour it documents.
|
| 1783 |
RETRACE: From Entangled Repair Histories to Reusable Experience for CI Repair
2610.04658
|
cs.AI
|
Rabeya Khatun Muna, Muhammad Ahasanuzzaman, Nakhla Rafi, Yisen Xu, Jinqiu Yang |
Large language model (LLM) agents increasingly reuse prior experience, but most approaches assume that problems and solutions are already aligned. Software histories rarely provide this alignment: a pull request (PR) may contain multiple continuous integration...Large language model (LLM) agents increasingly reuse prior experience, but most approaches assume that problems and solutions are already aligned. Software histories rarely provide this alignment: a pull request (PR) may contain multiple continuous integration (CI) problems, failed attempts, reverted edits, and unrelated changes, obscuring which changes resolve each problem. We present RETRACE, a framework for reconstructing problem-level repair experience from such histories. RETRACE combines an endpoint view that reasons backward from changes retained in the passing revision with a development view that traces repair evolution forward through commit history. CI execution evidence reconciles the two views, and the recovered experience is represented at three abstraction levels, from concrete fixes to transferable repair patterns. For new failures, RETRACE retrieves relevant problem-level experience to guide repair. On CI-REPAIR-BENCH, comprising 565 PR-level repairs from 101 repositories across 12 failure categories, RETRACE improves mini-SWE-agent Pass@1 from 19.6% to 31.9% with MiniMax-M2.5 and from 23.3% to 32.8% with DeepSeek-V4-Flash. On a matched subset, Codex improves from 15.5% to 27.5%. Combining both views consistently outperforms either alone, showing that recovering problem-change alignment enables historical CI repairs to serve as reusable repair experience.
|
| 1784 |
Neurodiversity-Aware Multimodal Affective Computing for Neurodevelopmental Assessment: From Norm-Referenced Classification to Context-Sensitive Decision Support
2610.04705
|
cs.AI
|
Mateusz Pomianek, Anna {\L}\k{e}\.zniak-Seruga |
Automated neurodevelopmental assessment increasingly combines computer vision, speech, eye tracking, physiology, and machine learning, yet multimodality and discrimination do not establish construct validity, clinical usefulness, or appropriate interpretation ...Automated neurodevelopmental assessment increasingly combines computer vision, speech, eye tracking, physiology, and machine learning, yet multimodality and discrimination do not establish construct validity, clinical usefulness, or appropriate interpretation of behavioral variation. We propose a testable architecture for neurodiversity-aware multimodal affective computing in which population-relative deviation is not treated as sufficient evidence of adverse functioning. This conceptual article reports neither a new participant-level study nor a trained predictive system, and no predictive or clinical superiority is claimed. The framework keeps observable signals, modality quality and availability, context, population-relative information, pooled within-person references and, when sufficiently supported, context-conditioned within person references, and predictive uncertainty distinguishable throughout inference. It defines structured decision-support outputs and auditable requirements for target validity, temporal integrity, contextual leakage, missing-modality robustness, calibration, subgroup evaluation, interpretability, privacy, and human decision boundaries. The contribution is architectural rather than algorithmic: it specifies empirically testable constraints that can be instantiated with different multimodal learning methods. Autism provides the principal motivating evidence base, without assuming unchanged transfer to other neurodevelopmental conditions.
|
| 1785 |
VoCa: Designing Speech-Canvas Interaction for Voice-Based Conversational Agents
2610.04706
|
cs.AI
|
Yate Ge, Run Yuan, Yueran Qi, Wenjie He, Jiaqi Mo |
People write and sketch while speaking to explain, organize, and develop content together. Inspired by these practices, we investigate how voice agents can use a canvas alongside speech in multi-turn conversations with users. We conducted a two-part formative ...People write and sketch while speaking to explain, organize, and develop content together. Inspired by these practices, we investigate how voice agents can use a canvas alongside speech in multi-turn conversations with users. We conducted a two-part formative study: an observational study of how pairs coordinated speech and boardwork, followed by a design workshop that informed a design space for speech-canvas interaction with voice agents. Building on these insights, we developed VoCa, a voice agent that coordinates speech with visual object creation, annotation, and attention guidance. A five-day deployment with 18 participants examined usability, experiences of speech-canvas interaction, patterns of use, and desired improvements. Participants' experiences highlighted opportunities for speech-canvas interaction in learning, work, and daily life, alongside challenges in coordinating what agents say and show in ways users can follow and influence. These findings inform how voice agents can use a canvas alongside speech in conversation.
|
| 1786 |
Organising Trajectory Evidence for Language-Model Agent Assurance: Fragments, Methods, and the Residual
2610.04710
|
cs.AI
|
Xiaowei Huang |
Methods for assessing language-model agents include rule checkers over logs, analyses of skill coverage and composition, support checkers, prefix monitors, execution gates, and rare-event estimators. Each observes a different part of a run and makes a claim of...Methods for assessing language-model agents include rule checkers over logs, analyses of skill coverage and composition, support checkers, prefix monitors, execution gates, and rare-event estimators. Each observes a different part of a run and makes a claim of a different strength, and no common account says how these claims combine or what they leave unchecked. We give one, built from a logic, information fragments, and an assurance ledger. Requirements are formulas of a two-tier logic: an outer finite-trace temporal logic over recorded events, and an inner logic of standing over the argument structure that the agent's recorded context supports at a decision. A rule violation and an action taken on withdrawn support are thus formulas of one language. The information available to an assessor, to the agent, and to an execution gate defines fragments of that language; each checking method decides one fragment and returns a typed claim: exact, on a named test suite, a risk bound, or descriptive. The ledger merges the evidence and classifies every obligation as established, addressed but not established, or unaddressed. On 456 released $\tau^2$-bench telecom trajectories, the benchmark oracle flags 231 runs; adding a rule checker, a support stand-in, and a prefix-monitor baseline raises the union to 391, 401, and 407, each contributing flags the others miss, and a gate makes one prohibition exact on a gated deployment. The 49 unflagged runs and the unmet or unaddressed obligations form the residual. Discovery makes unknown requirements explicit, and new or stronger checkers then reduce it.
|
| 1787 |
Towards Automatically Pruning Logging Code with Coding Agents: How Far Are We?
2610.04716
|
cs.AI
|
He Yang Yuan, Haonan Zhang, Xin Wang, An Ran Chen, Kisub Kim |
Logging code supports debugging, monitoring, and software maintenance, but excessive logging can add noise, impose runtime overhead, and obscure diagnostic information. While prior research has extensively studied logging code generation and modification, logg...Logging code supports debugging, monitoring, and software maintenance, but excessive logging can add noise, impose runtime overhead, and obscure diagnostic information. While prior research has extensively studied logging code generation and modification, logging removal remains comparatively underexplored. In this paper, we study developer logging removal practices and explore the use of coding agents for this task. We extract and manually validate logging removal cases from Python and Java repositories and derive 10 removal patterns and 11 removal reasons that characterize how and why logging code was removed in real-world software changes. We further construct LogRem, a dataset of 387 real-world cases covering direct logging statement removal, logging infrastructure removal, and logging replacement. We evaluate four coding agents with multiple model settings and compare their outputs with accepted real-world changes. Although 95.6% to 100.0% of outputs pass validity checks, only 11.1% to 19.6% remove the same logging code as the corresponding real-world change while preserving unrelated code. Agents differ through missed removals, extra removals, and unrelated code edits, with substantial variation across logging removal categories and trajectories. Execution cost varies widely, but higher cost does not consistently yield closer alignment. Commit messages and developer discussions provide the largest alignment gains, while taxonomy guidance consistently reduces runtime. Overall, our study establishes logging removal as a distinct software maintenance task and shows that reliable automation depends on accurately determining removal scope while preserving necessary code. To the best of our knowledge, this is the first study to examine logging code removal from this perspective.
|
| 1788 |
Robot Learning with Visual Predicted Force
2610.04741
|
cs.AI
|
Haonan Chen, Feiyang Wu, Yuxiang Ma, Mustafa Mete, Pengfei Ye |
Force-aware manipulation typically relies on specialized force or tactile sensors. We show that force-aware manipulation can instead be achieved through visual force prediction from the deformation of a compliant Fin Ray gripper. Our approach trains two models...Force-aware manipulation typically relies on specialized force or tactile sensors. We show that force-aware manipulation can instead be achieved through visual force prediction from the deformation of a compliant Fin Ray gripper. Our approach trains two models. First, we train a visual force estimator on calibration data and use it to annotate task demonstrations with force estimates. Second, we train an action--force proposal policy on these force-augmented demonstrations to jointly generate candidate robot actions and their associated forces. At test time, we sample candidate actions and the forces they are expected to produce, then execute the action whose predicted force is closest to a target from the demonstrations. We evaluate our approach on berry picking, empty-can grasping, in-hand reorientation, and plug insertion. Our results show that visual force prediction can guide inference-time action selection for contact-rich manipulation without requiring force or tactile sensors at deployment.
|
| 1789 |
Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate
2610.04860
|
cs.AI
|
Haocheng Ye, Aoting Hu, Xinwei Zhang, Xunzhu Tang, Shuchao Pang |
Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: factual errors induced or amplified by watermarking even when the required evidence is prese...Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: factual errors induced or amplified by watermarking even when the required evidence is present in the context and the unwatermarked model can answer correctly. Using a controlled retrieval-augmented generation setting, we compare unwatermarked and watermarked generations under the same context, query, and decoding configuration, and quantify their factual accuracy decrease. Across six representative watermarking methods, including KGW, SWEET, DiPmark, GumbelSoft, Gumbel-Max, and SynthID watermarking, we consistently observe watermark-induced hallucination. Watermarked outputs can remain fluent while introducing factual errors. We attribute this failure mode to two mechanisms: (1) token perturbations in the current-step arising from method-specific reweighting or keyed sampling, and (2) prefix-induced attention drift, which accumulates through autoregressive decoding and weakens later attention to the factual context. Motivated by this analysis, we propose two plug-in interventions at the token and attention levels that can be integrated into existing watermarking methods to improve factuality. At a matched TPR of 0.90 at 1% FPR, combining the two interventions reduces factual errors by approximately 90% relative to watermark-only decoding while preserving fluency and comparable decoding efficiency. Overall, this work highlights factuality as a first-class criterion in watermark evaluation, alongside detectability and robustness, and calls for careful factuality validation before deploying watermarks in fact-critical applications.
|
| 1790 |
PreAct-Nav: Agentic Reasoning Before Action for Urban Navigation
2610.04916
|
cs.AI
|
Jing Xie, Shouwei Ruan, Yubin Wang, Yuxiang Zhang, Junwei Yang |
Urban navigation requires embodied agents to pursue long-horizon goals through local decisions based on egocentric observations. However, existing agentic navigation methods often struggle to translate distant goals into coherent local decisions in large-scale...Urban navigation requires embodied agents to pursue long-horizon goals through local decisions based on egocentric observations. However, existing agentic navigation methods often struggle to translate distant goals into coherent local decisions in large-scale physical environments. Their reliance on linguistic reasoning over transient observations or limited history constrains anticipation of the consequences of actions and future conditions, despite the importance of such foresight for navigating long and complex urban routes. To bridge this gap, we propose PreAct-Nav, an agentic navigation framework that equips frozen policies with anticipatory reasoning for robust urban navigation. Our central idea is to anchor local decisions in persistent medium-horizon subgoals, assess the consequences of predicted actions before execution, and continually update the reasoning context using actual outcomes. At its core, a navigation memory module maintains the active subgoal and relevant experience across decisions, translating distant goals into actionable intermediate objectives. We further introduce a predictive world sandbox that uses an action-conditioned world model (AC-WM) to forecast world dynamics conditioned on candidate movements. A vision-language model (VLM) reasoner interprets these predictions under the current subgoal to retain or revise actions. After execution, real observations are used to assess outcomes, correct inconsistent assumptions, and update memory to continue or reformulate the subgoal. Extensive evaluations demonstrate that the proposed PreAct-Nav improves action selection through memory updates and visual prediction, with more pronounced gains on longer routes and routes with more turns.
|
| 1791 |
Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
2610.04921
|
cs.AI
|
Yifan Xiong, Jingyi Ge, Zhenpeng Chen, Yiling Lou |
LLM-based agentic systems are emerging as a new software paradigm. Modern agents are typically composed of backbone LLMs and a surrounding harness that serves as the operational software infrastructure for agent execution. As agent harnesses grow increasingly ...LLM-based agentic systems are emerging as a new software paradigm. Modern agents are typically composed of backbone LLMs and a surrounding harness that serves as the operational software infrastructure for agent execution. As agent harnesses grow increasingly complex, agents suffer from diverse harness implementation bugs, raising substantial reliability concerns. In this work, we conduct the first empirical study to systematically investigate the test adequacy of harness in real-world agentic systems. Our analysis reveals that agent harness remains substantially undertested. In particular, LLM-dependent harness (LDH) code, despite its critical role in processing LLM outputs and governing agent behavior, receives limited testing attention, with less than half of its lines and branches covered by existing tests. Motivated by these findings, we further propose HarnessTester, the first harness-oriented test generation technique that incorporates explicit agent-harness contract support to construct contract-faithful test setups and extensively exercise LDH code. Our evaluation shows that HarnessTester substantially outperforms state-of-the-art general-purpose test generation techniques in achieving 75.95%/84.76% larger line/branch coverage gains and 69.89% larger mutation-score gains. Furthermore, HarnessTester detects 122 real-world harness bugs in widely-used agentic systems (e.g., OpenClaw), among which, 88 bugs are previously-unknown bugs and 69 bugs have been confirmed by agent developers. These results highlight the practical effectiveness of HarnessTester in improving test adequacy and assuring the reliability of real-world agentic systems.
|
| 1792 |
A Unified Dynamics Framework for Reinforcement Learning and Classical Control of a Six-DOF Pipeline-Tracking ROV in NVIDIA Isaac Sim
2610.04949
|
cs.AI
|
Cheng Siong Chin, M. Venkateshkumar, Jianhua Zhang |
Reinforcement learning controllers for underwater vehicles are usually trained against one physics representation and deployed against another, so reported performance does not always describe behavior outside training. This paper presents a pipeline-tracking ...Reinforcement learning controllers for underwater vehicles are usually trained against one physics representation and deployed against another, so reported performance does not always describe behavior outside training. This paper presents a pipeline-tracking architecture for a six-degree-of-freedom remotely operated vehicle (ROV) in which one Universal Scene Description (USD) scene supplies the real BlueROV2-Heavy mass, added-mass, damping, buoyancy, and thruster parameters to both halves of the system: a vectorized NumPy implementation of Fossen's marine-craft equations, and an interactive NVIDIA Isaac Sim deployment applying the identical equations as PhysX forces at every step. The Coriolis-centripetal term is the primary dynamics model in both branches; a controlled ablation on PPO and TRPO shows that including it does not destabilize either algorithm and modestly improves tracking, about 19 percent lower standoff RMS error for TRPO. Five reinforcement learning algorithms, PPO, soft actor-critic, TD3, DDPG, and TRPO, are trained against one environment, reward, and randomized evaluation harness through a checkpoint-compatibility layer scoring any policy with the same code. The pipeline is extended with six classical baselines, PID, sliding-mode, fuzzy logic, feedback linearization, model predictive control, and an adaptive neuro-fuzzy inference system, driven by the same guidance geometry and thruster allocation as the learned policies. Under Coriolis-enabled dynamics, PPO, TRPO, and feedback linearization reach the strongest combination of 100 percent success and competitive accuracy; PID, fuzzy control, and the neuro-fuzzy baseline also reach 100 percent success with looser tracking; DDPG and TD3 each show a specific, explainable failure mode rather than a general weakness of off-policy learning; and classical control remains a strong baseline against the best learned policies.
|
| 1793 |
Kapture: Capturing Cardiac Dynamics with Koopman-Governed Learning for Efficient Radar-Based Electrocardiogram Recovery
2610.04955
|
cs.AI
|
Tong Wu, Jing Peng, Ziqi Feng, Yuanyuan Zhang |
Millimeter-wave (mmWave) radar enables unobtrusive, contactless electrocardiogram (ECG) reconstruction for cardiac monitoring. Time-frequency spectrograms preserve fine cardiac patterns but often require large backbones to separate ECG-relevant features from r...Millimeter-wave (mmWave) radar enables unobtrusive, contactless electrocardiogram (ECG) reconstruction for cardiac monitoring. Time-frequency spectrograms preserve fine cardiac patterns but often require large backbones to separate ECG-relevant features from respiration, motion, multipath, and subject-dependent interference. We propose Kapture, a parameter-efficient Koopman-governed framework that projects radar hidden states into a low-dimensional observable space and identifies a regularized full linear evolution operator from adjacent observable states. The Koopman-predicted observables refine subsequent hidden states for ECG reconstruction. To suppress predictable interference dynamics, a temporal contrastive objective pulls neighboring states together and separates non-neighbors, while reconstruction supervision preserves task relevance. Using approximately 80 minutes of quasi-static radar-ECG recordings containing realistic noise from body movements and other sources, Kapture consistently improves reconstruction across matched backbone widths, with the largest gains under aggressive compression. The compact configuration approaches the full-width reference accuracy with 68.1% fewer parameters and 91.3% fewer profiler-covered floating-point operations (FLOPs), while the full-width configuration delivers the strongest overall reconstruction performance. Our code will be made publicly available after potential publication.
|
| 1794 |
TeleGen: Improving LLM-Based Web Application Generation via Runtime Telemetry
2610.04981
|
cs.AI
|
Yujia Luo, Haonan Zhang, Jiasi Shen, Zishuo Ding, Weiyi Shang |
Large language models can generate runnable web applications from natural-language requirements, but many generated applications still fail interactive tasks. Existing generate-execute-repair pipelines execute the generated application and use task outcomes or...Large language models can generate runnable web applications from natural-language requirements, but many generated applications still fail interactive tasks. Existing generate-execute-repair pipelines execute the generated application and use task outcomes or error messages to guide code revision. However, this feedback often misses the runtime behavior between a browser action and the final task outcome, making interaction-level failures difficult to diagnose. Therefore, we propose TeleGen, an observability-enhanced framework for LLM-based web application generation. TeleGen instruments generated applications, collects runtime telemetry during task execution, and compresses raw telemetry logs into concise briefs for repair. We evaluate TeleGen on WebGen-Bench and Web-Bench. On WebGen-Bench, TeleGen improves task success from 67.7% with repair without telemetry to 76.2%, an increase of 8.5 percentage points. On Web-Bench, it improves cumulative Pass@2 from 21.7% to 29.8%. Ablation results show that runtime telemetry provides a useful diagnostic signal, while telemetry briefs make this signal more effective and less costly to use. Further analysis shows that telemetry is especially helpful for failures involving hidden execution paths, such as navigation, form workflows, and frontend-backend coordination.
|
| 1795 |
Hidden Risks of Jev: An Empirical Study of Security, Privacy, and Dual Use
2610.04985
|
cs.AI
|
Shang Wang, Tianqing Zhu, Huajie Chen, Jiayang Li, Meng Yang |
Jev turns natural-language questions into typed answers and probabilities with low latency and cost, enabling applications to route requests and select tools. While this interface allows Jev to integrate naturally into application workflows as a decision layer...Jev turns natural-language questions into typed answers and probabilities with low latency and cost, enabling applications to route requests and select tools. While this interface allows Jev to integrate naturally into application workflows as a decision layer, the security and privacy implications of this emerging use remain largely unexplored. To address this gap, we conduct the first systematic study of these implications using the official Jev API and NanoJev, a local model with controllable training data and updates, focusing on three research questions: (1) What security threats arise when Jev is deployed as an application decision layer? (2) What private information can Jev reveal despite returning constrained typed outputs? (3) How can Jev's general-purpose decision capability be used for beneficial purposes or misused? Jev's decisions depend on application state and may be influenced by user-provided inputs. We therefore adapt prompt injection and adversarial suffixes to manipulate its decisions. Open-source Jev distribution and updates introduce supply-chain risks, which we examine by implanting backdoors in NanoJev through training data poisoning. Since Jev's outputs reflect both application state and information learned during training, we further adapt membership, private attribute, and internal knowledge inference attacks to recover sensitive information despite its constrained output format. Finally, Jev can serve as a general-purpose decision oracle for defensive and malicious workflows. We examine this dual use through four detection tasks covering prompt injection, jailbreak inputs, harmful content, and AI-generated text, alongside misuse scenarios involving jailbreak and model extraction. Our empirical evaluation shows that Jev remains vulnerable to the examined security and privacy threats, while its decision capability can support beneficial and malicious uses.
|
| 1796 |
SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing
2610.05037
|
cs.AI
|
Rajatabha Chakraborty, M P Samartha, Vedant Pahariya, Priyesh Shukla |
Sparse-accelerator design spaces are usually searched against analytical models, so a design point is admitted on what a model predicts rather than on what the hardware does. SparseCraft closes that gap with a language model inside a closed CHIA loop. In each ...Sparse-accelerator design spaces are usually searched against analytical models, so a design point is admitted on what a model predicts rather than on what the hardware does. SparseCraft closes that gap with a language model inside a closed CHIA loop. In each of 15 iterations the model reads the measured outcome of the previous one and edits the Chisel RTL, the memory configuration and the sparse-kernel schedule of a Gemmini accelerator through MCP tool servers, and no candidate counts until it has been checked for legality, elaborated, simulated cycle-accurately, checked bit-for-bit on every output against a golden reference, and synthesised. The harness turns each measurement into the next work order, a diagnosed bottleneck with matching strategy guidance, the history of tried designs and a score of the model's own prediction, and a second model repairs changes that fail a gate. On a $512 \times 512$ GraphChallenge sparse-DNN layer the loop reaches 2.1x fewer cycles, 9.8x less off-chip traffic and 22.8% less area than the block-sparse Gemmini baseline, with 5.61x higher modelled perf/W and 11.8x lower EDP. The levers span three layers: a schedule that keeps the dense operand resident removes 9.8x of the traffic, a zero-gated MAC and a zero-row skip unit that the model wrote in Chisel cut energy, and resizing the memories cuts area.
|
| 1797 |
Tracing a Sparse Emotion-Control Circuit in LLM-Based Text-to-Speech
2610.05080
|
cs.AIcs.SD
|
Hongfei Du, Jiacheng Shi, Yanfu Zhang, Ye Gao |
LLM-based text-to-speech (TTS) models can generate emotionally expressive speech, but how reference emotion is routed through the model and realized in decoded speech remains unclear. We introduce two emotion-sensitive metrics for matched neutral and emotional...LLM-based text-to-speech (TTS) models can generate emotionally expressive speech, but how reference emotion is routed through the model and realized in decoded speech remains unclear. We introduce two emotion-sensitive metrics for matched neutral and emotional syntheses---a codec trajectory score and a late residual direction score---and use them to score activation-patching interventions. Under controlled matched-reference conditions, this analysis identifies a sparse source-to-readout component-level circuit: 23--27 attention heads and MLPs per emotion, roughly 5% of the components considered, recover or suppress 74--88% of the late emotion-readout shift on held-out cases. The circuit combines a shared component backbone with emotion-specific components; cross-emotion activation swaps reduce the target readout in 47 of 48 cases. In decoded speech, the same intervention produces consistent changes in pitch, energy, and spectral brightness over 24 matched pairs per emotion. A readout-matched residual-direction baseline produces only 17--27% of the intervention's pitch effect, showing that internal readout movement alone does not explain the decoded acoustic changes. These results trace a compact causal route from reference-derived prefix information to emotion-relevant properties of generated speech.
|
| 1798 |
VulValidate: Auditing Function-Level Vulnerability Labels with Executable Evidence
2610.05103
|
cs.AI
|
Leizhen Zhang, Sheng Chen |
Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that ...Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic analysis tools and construct vulnerability-triggering experiments from runtime feedback. Given a labeled function and its fixing patch, VulValidate reconstructs vulnerable and fixed revisions, selects suitable tools and execution paths, refines triggering inputs, and compares runtime behavior to assess function-level attribution. We audit all 35,849 instances originally labeled vulnerable in BigVul, PrimeVul, and DiverseVul. We confirm 20,510 (57.2%), correct 6,819 labels (19.0%), leave 7,981 attacked but undecided (22.3%), and cannot successfully measure 539 (1.5%). After conflict resolution and byte-exact deduplication, the corrected release contains 15,890 distinct confirmed vulnerable function bodies. In a blinded review of 581 sampled decisions, expert consensus supports 90.0%--92.0% of confirmations and 92.6%--99.0% of label corrections. With model parameters fixed, corrected evaluation lowers F1 for all five tested detectors on both BigVul and DiverseVul; retraining with corrected labels improves F1 for four of five detectors on each dataset. We also release a reusable VulValidate skill, corrected datasets, and reproducible evidence for future vulnerability-detection research.
|
| 1799 |
Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection
2610.05163
|
cs.AI
|
Jingkai Liu, Yufei Han, Xiaoting Lyu, Wei Wang, Ting Yu |
Long-horizon agents consume external content, invoke tools, and modify persistent state. Indirect prompt injection can exploit task-specific context, propagate across causally connected stages, and alter a consequential action while the workflow continues; we ...Long-horizon agents consume external content, invoke tools, and modify persistent state. Indirect prompt injection can exploit task-specific context, propagate across causally connected stages, and alter a consequential action while the workflow continues; we term this staged prompt injection. We build an automated, feedback-guided attack generation pipeline and apply it to Claude Code and Codex in their native runtimes. The confirmed attacks span eight workflow scenarios, seven attack goals, and six injection surfaces, showing that production agents are vulnerable to context-aware, multi-step injection over long horizons. Stopping such attacks requires a decision before each consequential action: input screening and completed-run evaluation cannot locate the intervention point, and existing pre-action methods use incompatible units and labels. We therefore formulate boundary action auditing: given initial context, a trajectory prefix, and a fully specified pending message or tool call, an auditor predicts Pass or Block before its effect occurs. Pairing attacked and benign executions yields a 479-pair, 3,112-unit benchmark. We further propose Path-Aligned Attribution (PAA), a training-free auditor that decomposes pending actions into operative elements and traces what supplied each value and guided each decision. PAA blocks only when the model attributes an unwarranted, material effect on an element to an attacker-reachable source that either provides unqualified steering or conflicts with visible evidence. Under full-benchmark fail-open scoring with Claude Sonnet 5, PAA reaches 86% Block recall at a 6-8% false-block rate (FBR), whereas ARGUS reaches 44-47% recall at 16-33% FBR. Under the same backend, on the tool calls that all three auditors natively support, PAA has higher recall and lower FBR than VIGIL and ARGUS; all paired 95% confidence intervals exclude zero.
|
| 1800 |
Same Predictions, Different Harms: Causal Auditing of Patient World Models
2610.05198
|
cs.AI
|
Yicheng Qi, Xiyi Xiong |
Patient world models used for clinical trial simulation can agree on transition kernels and arm-specific risks, yet disagree on the fraction of patients harmed by switching treatment---the counterfactual quantity that matters for intervention-aware reasoning. ...Patient world models used for clinical trial simulation can agree on transition kernels and arm-specific risks, yet disagree on the fraction of patients harmed by switching treatment---the counterfactual quantity that matters for intervention-aware reasoning. We audit this reliability gap in a two-stage shared-response SCM: a categorical intermediate health state is followed by common terminal care. Under independent stages, the sharp harm interval has closed-form endpoints for at most three intermediate states, with an exactness boundary at four states. Declared dependence and response-mismatch budgets yield calibrated outer bounds when stage independence or complete mediation is relaxed; in a symmetric three-state model the entire sensitivity frontier is sharp, $[0,\min\{1/2,1/3+(\rho+\delta)/2\}]$, and shows exactly how budgets erase the gain over endpoint-only bounds. Two eight-variable response LPs propagate interventional uncertainty for finite-sample audits. Exact witnesses verify attainability. On public clinical simulators (EpiCare; sepsis), native configurations show little resolved stage dependence and no additional joint-compatibility gain over pairwise transport---honest negative results for reliability claims. All experiments are locally reproducible; guarantees remain conditional on the stated causal model. The results provide a concrete protocol for deciding when a patient world model is safe to trust for counterfactual harm.
|
| 1801 |
Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization
2610.05230
|
cs.AI
|
Yifan Li, Jiaxu Wang, Dongming Wu, Yicheng Jiang, Ryan Ji |
Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves te...Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, and forces a tradeoff between long-horizon coverage and the local precision required for contact-rich manipulation. To address these limitations, we introduce Vela, a vision-language-action foundation model that represents future robot behavior as continuous trajectories. Vela combines a compact spline-based action representation with motion-dependent temporal support and a shared action interface for heterogeneous embodiments, allowing a fixed output budget to adapt its temporal resolution across motions. We pretrain Vela on large-scale multi-embodiment robot data and evaluate it on LIBERO-X, EBench, and two real-world long-horizon tasks, egg-cake cooking and potato shredding, obtaining promising results across simulation and physical manipulation. These results highlight the potential of continuous action representations as a foundation for future embodied foundation models. Project page and more results: https://clementine24.github.io/Vela/ .
|
| 1802 |
Settling the Computational Complexity of Max-Min Allocation with Ternary Valuations
2610.05237
|
cs.AI
|
Thi Ngoc Anh Vu, Trung Thanh Nguyen, Khaled Elbassioni, J\"{o}rg Rothe |
We study the problem of computing an allocation of indivisible items that maximizes egalitarian welfare, i.e., the utility of the worst-off agent, when agents' item values or marginal values belong to a small set. For additive valuations with values in $\{p,q\...We study the problem of computing an allocation of indivisible items that maximizes egalitarian welfare, i.e., the utility of the worst-off agent, when agents' item values or marginal values belong to a small set. For additive valuations with values in $\{p,q\}$, where $q>p>0$ and $\gcd(p,q)=1$, we give a polynomial-time algorithm when $p=2$ and prove constant-gap hardness when $p\geq3$, already with exactly three high-valued goods per agent. We also give an $\sqrt{3/2}$-approximation for common positive bi-valued additive valuations. For mixed additive valuations in $\{-p,0,c\}$, where $p\in\{1,2\}$ and $c$ is a positive integer, a reduction to maximum-weight perfect matching resolves the conjectured tractability of $\{-2,0,c\}$-valuations. For submodular valuations with marginals in $\{-2,0,c\}$, where $c$ is odd, we establish an exact unit-gap hardness result and exponential value-query lower bounds, even when all but one agent are additive. Finally, for $\{-1,0,1\}$-submodular valuations, we prove that no finite multiplicative approximation exists unless $\p=\np$. Together, our results resolve open questions and provide a complete picture of the computational complexity of max-min allocation with ternary valuations.
|
| 1803 |
StateWise: Diagnosing and Repairing Persistent Operational State Before Agent Actions
2610.05241
|
cs.AI
|
Yongyuan Peng, Zhou Feng, Tongying Wu, Jiahao Chen, Yuan Su |
LLM agents combine reasoning, tool use, and persistent memory to support work across tasks by reusing stored operational records as premises for later actions. However, environmental or requirement changes can invalidate these records, while existing action re...LLM agents combine reasoning, tool use, and persistent memory to support work across tasks by reusing stored operational records as premises for later actions. However, environmental or requirement changes can invalidate these records, while existing action review, provenance tracking, and clarification mechanisms may leave the underlying persistent state uncorrected. Our audit of coding-agent trajectories identifies candidate failure chains in which invalid records are reused, leading to task failures and unsafe modifications. We propose StateWise, a framework for diagnosing and repairing persistent operational state before action execution. StateWise uses record-level counterfactual replanning to identify decision-critical records, then establishes their current validity through reliability checks, read-only verification of machine-observable facts, and targeted clarification of developer-owned intent. Typed evidence grounding binds evidence to specific records and scopes, enabling persistent corrections with repair lineage. The agent then replans from the repaired state, followed by an independent state-action check before execution. We evaluate StateWise on 150 executable coding-agent cases across diverse runtime environments, workspace configurations, and repository settings, complemented by cross-model evaluations. Under corrupted persistent state, StateWise achieves 93.3% overall correctness, compared with 38.7% for the baseline agent, with no unsafe actions. Component ablations, multi-task experiments, and transfer evaluations further demonstrate effective recovery, persistent corrections, and transferability across repositories and tool interfaces.
|
| 1804 |
Task-Aware Joint Pruning and Distillation for Efficient Audio Deepfake Detection
2610.05264
|
cs.AIcs.SD
|
Miao He, Peng Cheng, Zhongjie Ba, Qing Wen, Li Lu |
Advances in speech synthesis have made deepfake speeches increasingly convincing, posing growing threats to security. While self-supervised learning (SSL) based detectors achieve state-of-the-art performance, their computational demands (typically 300M+ parame...Advances in speech synthesis have made deepfake speeches increasingly convincing, posing growing threats to security. While self-supervised learning (SSL) based detectors achieve state-of-the-art performance, their computational demands (typically 300M+ parameters) prevent deployment on resource-constrained devices. Existing compression methods, designed mainly for content-centric tasks, struggle to maintain competitive performance when directly adapted to deepfake detection. We propose a Task-Aware Joint Pruning and Distillation framework that combines cross-domain knowledge distillation with movement-guided structured pruning to transfer forgery-discriminative knowledge and preserve critical structures under aggressive compression. Our framework reduces the model to 31.9M parameters with 6.3$\times$ FLOPs reduction, with an average performance drop of only 1.30\% across multiple datasets compared to the uncompressed baseline, demonstrating strong potential for on-device deployment.
|
| 1805 |
Who Is Your Agent Serving? Provider-Side Indirect Prompt Injection in Proactive Agents
2610.05266
|
cs.AI
|
Rui Wang, Chao Wang, Xinchen Wang, Yufeng Zheng, Binbin Liu |
Proactive personal agents increasingly decide what to recommend, how to personalize advice, and what follow-up assistance to offer, creating a new user-decision attack surface for provider-side indirect prompt injection. We show that an external provider need ...Proactive personal agents increasingly decide what to recommend, how to personalize advice, and what follow-up assistance to offer, creating a new user-decision attack surface for provider-side indirect prompt injection. We show that an external provider need not access private user context, compromise the agent, or gain additional permissions: by controlling only content associated with its own target, it can redirect an otherwise benign agent to advance that target, recruit legitimately available user context to justify it, and proactively reduce the friction of adoption. We characterize this failure mode through Target Control, Private Binding, and Prospective Support, which respectively steer what the agent advances, how it connects the target to the user, and what target-specific assistance it offers next. Across three proactive-agent environments and six simulated user models, the full attack increases target authorization in all tested environment-user-model combinations, with a macro gain of up to 77.4 percentage points. Controlled replay shows that correct user-target binding is more consequential than additional proposal detail alone, while a multi-turn extension reveals that provider objectives can remain influential even without final authorization by reshaping how the agent responds to user constraints and resistance. These findings expose a broader trust boundary: capabilities designed to serve the user can be redirected toward objectives originating outside the user-agent relationship.
|
| 1806 |
Grammar-Guided Code Watermarking with Green Temperature
2610.05323
|
cs.AI
|
Hyundong Jin, Hyeseon An, Soohan Lim, Yo-Sub Han |
Large language model watermarking embeds detectable statistical signals during decoding, but the resulting changes to token probabilities can degrade generation quality. This trade-off is particularly important for code, where small changes in token selection ...Large language model watermarking embeds detectable statistical signals during decoding, but the resulting changes to token probabilities can degrade generation quality. This trade-off is particularly important for code, where small changes in token selection can break syntax or alter program behavior. Existing code watermarking methods mitigate this risk through entropy-based insertion or syntax-aware token selection, but they do not directly construct the watermark over the set of continuations admitted by the current grammar state. We propose Grammar-Guided Code Watermarking with Green Temperature (GTCW), which integrates grammar-constrained decoding with probability-aware watermarking. At each decoding step, GTCW restricts the candidate set to grammar-admissible tokens and partitions this support into keyed green and red subsets. At eligible high-entropy positions, green temperature reweights the green tokens according to the model's relative preferences, strengthening the watermark signal while retaining the grammar constraint. Across five models and five benchmarks spanning four programming languages, GTCW achieves a mean AUROC of 73.61%, compared with 67.83% for the strongest baseline, while maintaining a mean Pass@1 of 59.18% versus 59.58% for unwatermarked generation. Our implementation is available at https://github.com/hyundong98/GTCW .
|
| 1807 |
Optimal Control with Learned Critics under Unmodeled State Dependencies
2610.05359
|
cs.AI
|
Philipp Schoch, Markus Ryll |
Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Mode...Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assumptions but typically requires large amounts of interaction data. We present a learning-based MPC framework that combines the data efficiency and structure of local model-based planning with learned components that compensate for incomplete dynamics and finite-horizon myopia. The method augments a nominal analytical model with a residual dynamics network that learns missing state-dependent effects from data and combines the resulting planner with a learned action-value critic that injects long-horizon MDP structure into the local iLQR optimization. To make this practical at reinforcement-learning scale, we develop a GPU-accelerated batched iLQR solver that evaluates learned dynamics and critic networks inside the optimal-control loop and solves thousands of trajectory-optimization problems in parallel. The complete system is integrated into a robotics simulator, enabling scalable model-based reinforcement learning under incomplete dynamics. Experiments on biased and incompletely modeled control tasks show that the approach improves closed-loop control performance while preserving the model-based structure needed for efficient constrained trajectory optimization.
|
| 1808 |
AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models
2610.05369
|
cs.AI
|
Bao-Phong Nguyen, Gia-Khanh Pham, Thai-Duong Do, Mai Xuan Trang, Minh-Tuan Le |
The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort...The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative configurations. For large, high-dimensional network traffic data, this process can create a significant computational burden. In this work, we propose AutoDP-LLM, an automated pre-processing framework designed to reduce manual pipeline development and computational overhead. Specifically, AutoDP-LLM leverages Large Language Models (LLMs) to autonomously generate and validate executable data pre-processing pipelines. The framework combines deterministic host-side planning with LLM-based specialist agents to formulate data-processing strategies, synthesize executable code, and adaptively determine retained feature sets using semantic reasoning and training-derived statistical evidence, without requiring a predefined feature budget. Focusing on multiclass intrusion detection, we evaluate AutoDP-LLM on the UNSW-NB15 and NSL-KDD benchmark datasets using multiple downstream classifiers. Comparative experiments against conventional feature-selection methods show that AutoDP-LLM achieves competitive detection performance while automating the generation of compact and executable pre-processing pipelines. Component-level ablation experiments further demonstrate the complementary contributions of the semantic and statistical feature-reduction components. The repeated generation, validation, execution, and assessment of candidate pipelines are amenable to parallel execution, highlighting the potential of scalable computing environments, including high-performance computing (HPC) systems, to support automated IDS pipeline development.
|
| 1809 |
Reflecting on Creative-Boundaries with an AI Co-Doodler
2610.05482
|
cs.AI
|
Samia Menon, Samyukta Jayaram, Chetan Goenka, Shm Garanganao Almeda |
In this pictorial, we consider how the negotiation of creative boundaries with a co-creative AI system can create moments for personal creative reflection. We ground this in our experiences with Froggi-Draw, a single-initiative co-doodling system that gives us...In this pictorial, we consider how the negotiation of creative boundaries with a co-creative AI system can create moments for personal creative reflection. We ground this in our experiences with Froggi-Draw, a single-initiative co-doodling system that gives users power to decide when and how much an AI "collaborator" (Froggi) contributes to their drawing. From a 1-week pilot study where (n=8) novice and experienced artists doodled daily with the system, we report on ways participants navigated creative risk and uncertainty, and how their usage and perceptions of Froggi shifted over time. In moments of disruption, participants described the system as encroaching on their creative territory. We consider how the design of a supportive, co-creative AI "collaborator" might look like the design of a supportive power dynamic---and finding ways to offer users control to find, reflect upon, and flexibly negotiate the boundaries of that dynamic.
|
| 1810 |
FLEX-WAM: Flexible Block-Causal World-Action Models for Long-Horizon Imagination and Planning
2610.05483
|
cs.AI
|
R. Khorrambakht, Joseph Amigo, F\'elix Lebel, Leon Seetoo, Jean Ponce |
World--action models (WAMs) promise a unified model that predicts action-conditioned futures, generates feasible actions, and supports planning in imagination. However, existing joint video--action models often use computationally heavy, fixed-horizon backbone...World--action models (WAMs) promise a unified model that predicts action-conditioned futures, generates feasible actions, and supports planning in imagination. However, existing joint video--action models often use computationally heavy, fixed-horizon backbones ill-suited to streaming inference and stable long-horizon open-loop rollouts. We introduce FLEX-WAM, a Flexible and Efficient Block-Causal World--Action Model for unified simulation and policy inference. FLEX-WAM supports variable-length contexts and non-causal prediction horizons, as well as infinite autoregressive generation frame by frame or block by block. Its block-causal, KV-cacheable architecture combines axial attention and blockwise diffusion forcing to enable efficient real-time rollout and deployment-time latency--throughput tradeoffs without retraining. Joint training can nevertheless produce plausible futures that weakly respond to commanded actions. We address this failure mode by balancing state and action flow-matching gradient contributions across the state--action diffusion-noise grid and regulating world-model sampling using Forward-Dynamics (FD) elasticity, an efficient training-time proxy for action responsiveness. Across simulated and real-world datasets, FLEX-WAM achieves superior multi-step prediction quality and latency while producing stable joint state--action rollouts for thousands of steps. As a joint action proposer and simulator within MCTS, it solves long-horizon PushT and all five OGBench Puzzle-4x4 tasks entirely in imagination. On a bimanual OpenArm-based robot, a single checkpoint jointly serves as a play policy and expected-outcome predictor, enabling real-time identification and collection of model--reality mismatches for future self-improvement.
|
| 1811 |
Scenario-Based Compositional Statistical Model Checking for Safety Specifications
2610.05571
|
cs.AI
|
Abhinav Pomalapally, Arya Raeesi, Kevin Kai-Chun Chang, Beyazit Yalcinkaya, Sanjit A. Seshia |
In safety-critical domains such as autonomous driving, systems must be evaluated across a large number of environment conditions, often represented as composite scenarios built from primitive scenarios. Existing statistical model checking (SMC) approaches anal...In safety-critical domains such as autonomous driving, systems must be evaluated across a large number of environment conditions, often represented as composite scenarios built from primitive scenarios. Existing statistical model checking (SMC) approaches analyze each composite scenario independently, requiring many expensive simulations and resulting in substantial redundant computation when scenarios share common structure. This work introduces a scenario-based compositional SMC framework for safety and co-safety specifications, enabling efficient analysis of composite scenarios. Our approach decomposes scenarios into primitives and specifications into sub-specifications, verifies each primitive independently, and composes the resulting statistical estimates using importance sampling and kernel density estimation. Our empirical evaluation shows that the proposed framework can accurately answer verification queries for previously unseen composite scenarios while reducing simulation cost through parallelization and trace reuse.
|
| 1812 |
AgentDoxx: Agentic Re-identification of Anonymized Text with Web Search
2610.05586
|
cs.AI
|
Jianing Wen, Tianshi Li |
As Large Language Models (LLMs) gain tool use capabilities such as web search, they can retrieve and cross-reference public information, creating privacy risks beyond memorization. One manifestation is re-identification: linking an anonymized interview transcr...As Large Language Models (LLMs) gain tool use capabilities such as web search, they can retrieve and cross-reference public information, creating privacy risks beyond memorization. One manifestation is re-identification: linking an anonymized interview transcript to a named individual. Yet without ground-truth identities, the coverage of such attacks and the protection offered by a defense cannot be reliably measured. We introduce AgentDOXX, an evaluation suite of 822 synthetic interview transcripts grounded in public information about real individuals with known identities. We evaluate fifteen configurations of open-weight and proprietary models, isolating the effect of web search, and analyze their search trajectories to distinguish retrieval-driven from parametric identifications. Ground-truth identities reveal that re-identification risk is distributed across an agent's execution: retrieval and parametric recall both contribute, with open-weight models identifying 15-28% of transcripts without search; identification succeeds in over 88% of cases once the target appears in a retrieved result; entity masking leaves at least one attacker successful on 85.3% of a stratified sample; and privacy instructions suppress naming but not retrieval, with configurations scoring 0% accuracy yet retrieving the subject in up to 62% of transcripts. We further show that observed attack trajectories can provide supervision for localizing identifying spans, offering a path toward attack-informed anonymization.
|
| 1813 |
SEA-LM: Egocentric Spatial Audio Understanding for Wearable Microphone Arrays
2610.05610
|
cs.AIcs.SD
|
Sonal Kumar, Sinan Hersek, Artem Dementyev, Mengzhen Pan, Ishan Chatterjee |
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that...Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.
|
| 1814 |
UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
2610.05622
|
cs.AI
|
Dolly Sah, Tanmay Sah, Harshul Jain, Tanya Sah |
Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchma...Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.
|
| 1815 |
Transformer-Based Time-Series Inference of Lindblad Dynamics in Open Quantum Systems
2610.05647
|
cs.AI
|
Julian Guam, Jianqing Liu |
The Lindblad master equation is the standard framework for describing the non-unitary evolution of open quantum systems, where environmental interactions induce dissipation and decoherence. When both the system Hamiltonian and the dissipation rates are partial...The Lindblad master equation is the standard framework for describing the non-unitary evolution of open quantum systems, where environmental interactions induce dissipation and decoherence. When both the system Hamiltonian and the dissipation rates are partially unknown or explicitly time-dependent, traditional analytical inversion and system-identification techniques become intractable. Recent works have demonstrated that Transformer-based models can infer unknown dissipation rates from observable time series, yet these approaches typically rely on hand-crafted statistical features under idealized and highly restricted conditions. Here we advance the paradigm by introducing a raw time-series Transformer that directly ingests the full trajectories of Pauli expectation values $\langle\sigma _x(t)\rangle$, $\langle\sigma _y(t)\rangle$, and $\langle\sigma_z(t)\rangle$, thereby fully exploiting the self-attention mechanism for temporal modeling. The architecture is further extended to jointly learn unknown Hamiltonian parameters, handle multiple dissipation channels, and operate robustly under realistic measurement noise. Across all tested scenarios the model achieves consistently high reconstruction accuracy while eliminating manual feature engineering. This provides a scalable, robust, and versatile framework for quantum environment sensing in realistic open quantum systems.
|
| 1816 |
Nexus: An Execution Fabric for AI Agents Across Cloud, Edge, and Devices
2610.05709
|
cs.AI
|
Cary Chang, Jialin Zhou |
Language-model agents are evolving into long-running services that interact with models, tools, computers, mobile devices, and distributed environments. Existing agent frameworks simplify reasoning and tool invocation. However, cloud-centric designs face three...Language-model agents are evolving into long-running services that interact with models, tools, computers, mobile devices, and distributed environments. Existing agent frameworks simplify reasoning and tool invocation. However, cloud-centric designs face three limitations: centralized execution increases failure impact, scaling pressure, and compute cost; extending agents across computers, mobile devices, and edge environments requires a unified execution abstraction with permission control; and long-running executions require consistent lifecycle management across failures, recovery, results, usage, and settlement. We present Nexus, a cloud-edge platform that treats each invocation as a persistent task. Nexus uses an OpenWrt-based runtime for distributed serving, run-scoped delegation for authorized access to Computer and Mobile environments, and persistent records to track execution, outputs, failures, recovery, usage, and charging across cloud and edge components. We evaluate Nexus on controlled, cross-device, and model-driven workloads. All ten Computer-Android workflows succeed, and all six revocation tests block subsequent writes while preserving prior authorized reads. Under worker loss, journaling eliminates duplicate appends (six to zero per task), adding 0.933 s mean normal-path overhead. Across 24 matched task pairs, Nexus completes 24 tasks versus Dify's 22 and is a median 3.88 s faster on jointly successful pairs. In a separate workload, Nexus operates under a smaller tested incremental-runtime memory ceiling than Dapr (16 versus 64 MiB), although Dapr achieves lower successful-call latency. These results demonstrate how locality, operation-scoped authority, and persistent result identity support cloud-edge agent services with workload-dependent costs.
|
| 1817 |
SimpleMark: Fast Multi-Bit Text Watermarking under f -Divergence Constraints
2610.05712
|
cs.AI
|
Benjamin D. Kim, Wanrong Zhang, Weitong Ruan, Lav R. Varshney, Daniel Alabi |
We introduce a framework for multi-bit text watermarking with security defined directly through $f$-divergence from the base language model distribution. Unlike prior approaches that focus on average-key distortion-freeness or a particular statistical distance...We introduce a framework for multi-bit text watermarking with security defined directly through $f$-divergence from the base language model distribution. Unlike prior approaches that focus on average-key distortion-freeness or a particular statistical distance, our formulation supports general $f$-divergences, including total variation and KL divergence, and enforces the guarantee for each realized key and embedded message. We develop a coding-based watermarking scheme that optimally biases next-token distributions subject to a prescribed divergence budget, and characterize the resulting tradeoff between embedding rate, decoding reliability, and statistical security. Experimentally, we compare our method against prior multi-bit watermarking schemes across modern language models and payload regimes. Our approach achieves substantially lower watermark detectability while maintaining competitive message-recovery performance and generation quality. Our results provide a unified view of secure multi-bit watermarking and recover several commonly used security notions as special cases.
|
| 1818 |
Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
2610.05750
|
cs.AI
|
Reza Esfandiarpoor, Radek Osmulski, Yauhen Babakhin, Gabriel de Souza P. Moreira, Oliver Holworthy |
Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applica...Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.
|
| 1819 |
Who Keeps the Gains from Personal AI Assistants? Seller Adaptation and the Unassisted in a Language-Model Market Simulation
2610.05823
|
cs.AI
|
Haonan Huang, Joey Xiao |
Personal AI assistants are beginning to transact for consumers, and early adopters capture real savings. Whether those savings survive, and what happens to consumers who have no assistant, depends on how sellers respond -- a question single-user evidence canno...Personal AI assistants are beginning to transact for consumers, and early adopters capture real savings. Whether those savings survive, and what happens to consumers who have no assistant, depends on how sellers respond -- a question single-user evidence cannot answer. We build an agent-based rental market in which language models play consumers, assistants, and six adaptive sellers guided by an algorithmic pricing tool. Half the population receives an assistant under an advisory or an executing mandate; the contract pairs a fee only the renter's physical action avoids with a pre-selected add-on an authorised assistant can cancel online. An analytical benchmark and a behaviourally calibrated rule market supply ex-ante predictions, and paired branches with frozen versus adaptive sellers separate adoption effects from market feedback. Across thirty simulated markets, executing assistants cut adopters' spending by 13.7 USD per renter-day when sellers are frozen; adaptation claws back about a third, leaving 8.7, with the gains arriving both as lower bills and as rentals completed at all. Sellers raise headline rates while cutting fees, and the calibrated forecast of the burden on unassisted consumers (+3.6) does not transfer: their mean spending change is +0.4, confidence interval -0.6 to +1.3. Seller-model swaps and a within-market transfer of fee-setting to the pricing tool show that fee conduct, and with it the division of the gains, is decided on the seller side. Assistants, we conclude, should be evaluated at market level -- completion, total spending, and non-users included -- and the comparison layers locate exactly where a calibrated behavioural forecast fails in a language-model market.
|
| 1820 |
Hierarchical Reinforcement Learning for Collision-Free Locomotion of an Underactuated Biped
2610.05855
|
cs.AI
|
Jagannath Prasad Sahoo, Saurabh Kumar, Surya Prakash S. K., Samiran Datta, Abhay Dwivedi |
A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. Th...A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, \omega_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.
|
| 1821 |
Curriculum Brain: Constructing Curriculum Knowledge Graphs as a Substrate for Cognitive Diagnosis
2610.05860
|
cs.AI
|
Shrideep Tamboli, Chiranjeevi Maddala, Eshal Minhaj |
Cognitive Diagnostic Models (CDMs) identify which specific skills a student has and has not mastered, the signal a personalized learning path needs and a single aggregate score cannot give. Yet they are rarely deployed. The obstacle is their precondition: the ...Cognitive Diagnostic Models (CDMs) identify which specific skills a student has and has not mastered, the signal a personalized learning path needs and a single aggregate score cannot give. Yet they are rarely deployed. The obstacle is their precondition: the Q-matrix, a mapping from every assessment item to the skills it requires, historically authored by hand. We separate the task into two stages: first construct the curriculum's own knowledge graph, the full space of concepts and skills it contains, independent of any item; then map items against that graph on demand. This paper addresses the first stage only. The item-mapping stage is designed but not implemented here, so the claim that this shifts judgment cost from once per item to once per curriculum is a design rationale rather than a finding. We present Curriculum Brain, a two-repository system pairing a version-controlled knowledge base with an agentic pipeline of eleven single-responsibility agents under a thin deterministic orchestrator. It generates candidate concept-skill mappings from official curriculum documents, checks them against accumulated rules, and compares them with a concept-skill map extracted independently from the textbook, repairing its own failures and escalating to a human only when it cannot resolve a case itself. Across 241 chapter runs (168 distinct chapters), 41.5% produced a Generator output passing both checks without a patch, and 67.6% resolved without escalation. Both are measured against criteria the system itself produced, so both describe internal consistency rather than agreement with an external standard, and both pool two pipeline configurations separated by a single change at run 77; after it the figures are 57.0% and 91.5%. Observed spend was $1.19 per chapter, API spend only, excluding human review. We release both the framework and the resulting curriculum dataset.
|
| 1822 |
Runaway Reaction: When Benign Skills Compose into Malicious Behavior
2610.05943
|
cs.AI
|
Zunlong Zhou, Ziyuan Yang, Mengyu Sun, Yi Zhang |
Agent skills package task-specific knowledge and procedures that can be composed to support complex agent tasks, while public marketplaces provide a growing pool of reusable skills. Existing security vetting, however, largely evaluates skills in isolation, lea...Agent skills package task-specific knowledge and procedures that can be composed to support complex agent tasks, while public marketplaces provide a growing pool of reusable skills. Existing security vetting, however, largely evaluates skills in isolation, leaving composition-induced risks underexplored. Such risks arise because composing benign skills expands the agent's capability space, enabling behaviors unavailable to any skill alone. Interestingly, we find that directly composing benign skills can already induce malicious behaviors, even when every individual skill passes security vetting. We further find that some target malicious behaviors remain difficult to realize through direct composition, even when the selected skills collectively provide the required capabilities. To systematically instantiate these attacks, we present Compositional Risk Induction via Multi skill Execution (CRIME). CRIME first uses the Malicious Plot Casting (MPC) module to decompose a target malicious behavior into complementary requirements and identify suitable benign skill compositions from public skill repositories. For compositions that cannot directly realize the target behavior, the Runaway Reaction Steering (RRS) module uses execution feedback to iteratively refine the selected skills toward the target while requiring each skill to remain benign under standalone vetting. The resulting composition is then passed to the Skill Reaction Chamber (SRC) module, where the skill pair is executed in a sandbox and the resulting environmental consequences are examined to determine whether the target behavior has occurred. Unsuccessful cases are returned to RRS for further refinement. Furthermore, we construct a benchmark of 4,000 public skills across eight cybersecurity behaviors for systematic evaluation of composition-induced vulnerabilities.
|
| 1823 |
AgentSpy: Making AI Agent Behavior Observable
2610.06001
|
cs.AI
|
Christoph B\"uhler, Matteo Biagiola, Luca Di Grazia, Guido Salvaneschi |
AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an agent does during an execution is difficult to understand: tests assert on the result, and...AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an agent does during an execution is difficult to understand: tests assert on the result, and the agent's trajectory records only what the agent reports about itself, which may omit behavior executed by its subprocesses. We present AgentSpy, an approach that observes an agent from outside the agent. AgentSpy runs the agent in an isolated environment, configured by a declarative specification, and records the system calls and network traffic of the agent and of every process it executes. Based on this monitoring, AgentSpy supports two families of analyses: conformance analyses, which measure obligations, i.e., what an agent execution should do, and safety analyses, which check prohibitions, i.e., what an agent execution must never do. We instantiate one analysis of each family. The reliability analysis uses rules to summarize each run by the environment resources the agent uses: the commands it executed, the files it accessed, and the hosts it contacted. The security analysis applies deterministic rules to the system calls of an execution. For reliability, we evaluated AgentSpy on 77 tasks with the codex harness and three recent LLMs, executing each task three times. Sets of repeated runs of the same task are more similar than sets that include runs of another task in 92.2% of the comparisons. Among tasks for which all three runs pass outcome-based tests, the agent performs task-unrelated activities in 18% of the cases, reads the grading files in 7%, and does not use the developers' guidance in 17%. For security, the generic rules of AgentSpy detect four of five attack categories we considered, with no false positives across 50 runs.
|
| 1824 |
Parallelism or Concession? Concurrency-Aware Procurement Negotiation for Agentic Commerce
2610.06017
|
cs.AI
|
Xiaolin Xu, Donghao Zhu |
Agentic buyers can cheaply fork a procurement task into many parallel negotiations, but concurrency is not free: every thread consumes resources, and simultaneous agreements create cancellation and commitment risk. We study a one-unit post-order sourcing probl...Agentic buyers can cheaply fork a procurement task into many parallel negotiations, but concurrency is not free: every thread consumes resources, and simultaneous agreements create cancellation and commitment risk. We study a one-unit post-order sourcing problem with a single hard-deadline negotiation window, in which a planner jointly chooses the number of seller-facing negotiators and a common procurement price cap. The model combines a product-specific acceptance curve with fulfillment loss, per-thread cost, and excess-commitment cost. We establish three structural results. First, holding the per-thread acceptance target fixed, the marginal value of another negotiator decays geometrically, yielding a conditional concurrency threshold. Second, under a convex quantile curve, parallelism substitutes for concession: more concurrent negotiators imply a weakly lower per-thread acceptance target and price cap. Third, when prices are more dispersed, Agentic buyers benefit by searching harder for bargains, but suffer when they instead try to guarantee procurement by offering higher prices. We operationalize these results in the Concurrency-Aware Negotiation Optimizer (CANO), a deterministic optimizer that jointly determines the optimal negotiation concurrency and procurement price cap for an agentic procurement system. Across different analytic market configurations and extensive Monte Carlo, finite-data, non-Gaussian, and correlated-seller stress tests, CANO consistently outperforms common heuristic policies while validating the predicted structural properties.
|
| 1825 |
A Comprehensive Objective Evaluation of Modern Text-to-Speech for Turkish Using Speech Quality Assessment Models
2610.06057
|
cs.AIcs.SD
|
Yunus Emre Ozkose, Alperen Kahraman, Ali Haznedaroglu |
Modern text-to-speech (TTS) systems can clone a target speaker from a short reference clip or be fine-tuned on a target voice, yet their behaviour on morphologically rich, lower-resource languages such as Turkish remain under-characterised. We present a system...Modern text-to-speech (TTS) systems can clone a target speaker from a short reference clip or be fine-tuned on a target voice, yet their behaviour on morphologically rich, lower-resource languages such as Turkish remain under-characterised. We present a systematic benchmark of four contemporary systems (Chatterbox, CosyVoice, OmniVoice, and VoxCPM2) evaluated across fine-tuned and zero-shot configurations, contrasted with a conventional VITS baseline and anchored to natural gold speech. Each configuration is scored with eighteen complementary objective metrics spanning learned naturalness predictors (UTMOS v2, DNSMOS-Pro, SCOREQ, WhisQA, AudioBox-PQ, NatScore, SpeechLMScore), intelligibility and signal-quality estimators (SQUIM PESQ/SI-SDR/STOI, Brouhaha), speaker similarity, distributional fidelity (TTSDS) and low-level acoustic descriptors. We further analyse how quality varies with utterance length and quantify long-form temporal consistency through speaker-identity and naturalness drift over chunked utterances. We release our evaluation code to support reproducible TTS evaluation.
|
| 1826 |
Do VLAs Understand and Adapt to the Objects They Handle, or Simply Replay Learned Behaviors?
2610.06078
|
cs.AI
|
Xinnuo Xu |
This paper asks whether VLA generalization is grounded in a global understanding of objects' physical properties that enables policies to adapt their motion to unseen setups, or if policies simply replay the motions they've learnt that happen to succeed in new...This paper asks whether VLA generalization is grounded in a global understanding of objects' physical properties that enables policies to adapt their motion to unseen setups, or if policies simply replay the motions they've learnt that happen to succeed in new setups. The former reflects genuine generalization; the latter reflects incidental robustness. We first examine awareness of physical properties in seven VLAs by applying linear probing and representational similarity analysis (RSA) to their activations. We find that physical properties, including mass, fragility, deformability, friction and size are less decodable than non-physical properties such as semantic category, material, sound and price in nearly every modality stream. Compared with their base VLMs, robot pre-training weakens the linear encoding of physical properties in the language stream. Neither pre-training nor downstream fine-tuning strengthens the alignment between physical-property differences and activation distances. We then ask whether the weak physical information present in these activations shapes the actions a VLA generates. In a controlled LIBERO case study, we increase the mass of an in-domain object and signal the change through language or vision. Most VLAs use similar lifting behaviour for the heavier and original-mass objects, leading to task success declines. The few exceptions change their behaviour in response to lexical or visual cues rather than to mass itself. These results suggest that VLAs encode physical properties weakly and do not reliably use them to adapt their motion.
|
| 1827 |
Adaptive Mean Flow for Responsive Closed-Loop Robot Control
2610.06089
|
cs.AI
|
Aksel Vaaler, Marco Job, Christian Holden, Olav Egeland |
Diffusion- and flow-based robot policies have recently become widespread in robotic Imitation Learning (IL) due to their high performance and ability to model continuous and multimodal distributions. However, the iterative denoising procedure used by these mod...Diffusion- and flow-based robot policies have recently become widespread in robotic Imitation Learning (IL) due to their high performance and ability to model continuous and multimodal distributions. However, the iterative denoising procedure used by these models introduces significant prediction latency, hindering high-frequency closed-loop robot control and leading to jittery, unstable motion when frequent updates to the robot's action predictions are used. Therefore, it is common practice to train models to predict chunks of actions that can be executed sequentially without feedback, even when this reduces responsiveness and may mean the most recent state information is not used. In this article, we present Adaptive Mean Flow (AMF), a flow-based IL method that enables smooth and responsive, fully closed-loop robot control. AMF uses Mean Flow, which is an accelerated form of Flow Matching (FM), to minimize prediction latency. To ensure smoothness and consistency across predictions, AMF uses a corrupted version of the trajectory from the previous step when predicting new robot actions, with the signal-to-noise ratio increasing over the time parameter of the trajectory. This discourages large changes in the prediction from one step to the next, while allowing freedom to adapt the predictions for future steps. We evaluate AMF across a wide range of simulated and real robot tasks and demonstrate significantly improved performance compared with baselines. Code: https://github.com/akselva/Adaptive-mean-flow-RoboticIL.
|
| 1828 |
Where Did the Repair First Go Wrong? Localizing the Origins of Silent Failures in Agentic Vulnerability Repair
2610.06163
|
cs.AI
|
Wenji Bai, Muhammad Waseem, Zeeshan Rasheed, Jaakko Peltonen, Pekka Abrahamsson |
Localizing where an LLM-based agent first fails to uphold security during a repair can show which stage of its workflow needs an additional safeguard. This is difficult for silent failures, which are patches that pass syntactic and functional checks but still ...Localizing where an LLM-based agent first fails to uphold security during a repair can show which stage of its workflow needs an additional safeguard. This is difficult for silent failures, which are patches that pass syntactic and functional checks but still contain a security vulnerability. Because such patches give no observable failure signal, existing failure attribution methods, which rely on observed task failures and labelled failure steps, are less suited to them. We propose Security Awareness Gap Evaluation (SAGE), a trace-based method that combines an assessment of the security reasoning recorded at each turn with the reconstructed code history to identify the earliest turn at which a repair diverges from the task's security intent. We evaluate SAGE on 95 confirmed silent failures drawn from 3,684 repair traces produced by six agent frameworks and six base models on SecurityEval and CVEfixes. SAGE assigned an origin in 93 cases. Most origins were an unaddressed security requirement or an inadequate defence choice, and only five coincided with the code change itself. When the agent introduced the vulnerable code, the origin preceded the write in 14 of 19 cases. Repeated scoring and a second judge reproduced the origin type more consistently than the exact turn, and agreement was lowest for traces that kept only the final file.
|
| 1829 |
Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents
2610.06193
|
cs.AI
|
Hai Dang Truong, Rayner Goh, Thanh Le-Cong, Yintong Huo |
Autonomous coding agents now resolve a substantial share of real-world GitHub issues. However, passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging. Mature open-source projects publish repository-spec...Autonomous coding agents now resolve a substantial share of real-world GitHub issues. However, passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging. Mature open-source projects publish repository-specific contribution policies, spanning style, git, testing workflows, to ensure code quality and long-term maintainability. Because existing benchmarks evaluate patches solely on unit tests, agent compliance with repository governance remains unknown. In this paper, we introduce SWE-CC, a benchmark evaluating code and process compliance in autonomous software engineering. We develop a semi-automated pipeline that converts developer documentation across 12 open-source repositories into 823 machine-checkable atomic policies. SWE-CC introduces two features: 1) lightweight, deterministic checker functions that represent each policy, 2) a comprehensive auditing mechanism that inspects both agent runtime behaviors and final deliverables. We evaluate the compliance of agent workflows in 500 end-to-end software contribution tasks extended from SWE-bench Verified. Our evaluation of four LLMs under two agent scaffolds shows that modern agents suffer from coding compliance issues: although agents produce functionally correct patches, they still violate 43.1 percent of applicable project policies, with nearly half of all violations occurring during intermediate execution steps. These results show that functional correctness does not guarantee real-world readiness, highlighting that future software engineering agents must reliably conform to repository governance to enable safe and trustworthy deployment.
|
| 1830 |
Artificial Intelligence and the New Science of Culture
2610.06240
|
cs.AI
|
Douglas R. Guilbeault, Bhargav Srinivasa Desikan, James A. Evans |
Cognition and culture have long been treated as parallel objects of study, joined more by metaphor than by mechanism. We argue they are co-constituted in a manner now empirically tractable: subjectivities can be measured as local trajectories through a high-di...Cognition and culture have long been treated as parallel objects of study, joined more by metaphor than by mechanism. We argue they are co-constituted in a manner now empirically tractable: subjectivities can be measured as local trajectories through a high-dimensional cultural field, and the field is itself the aggregate of those trajectories. Recent advances in machine learning, including embedding methods and large generative models, provide the first general framework for measuring this co-constitution from micro-cognition to macro-social structure. Vector geometry recovers individual conceptual structure, organizational communication, and large-scale ideologies; captures multimodal cultural content beyond text; and generates testable predictions about cultural emergence, including the simultaneous arrival of independent discoveries across distant minds. We treat this predictability of simultaneous innovation as evidence for the co- constitutive view: when the geometry of the cultural field is measurable, trajectories of cognitive search through it become predictable. We then examine generative AI agents as simulators of human subjectivity and as a qualitatively new coordinating substrate whose insertion into social life introduces evolutionary dynamics unprecedented in human cultural history. We formalize this reflexive condition as a Heisenberg-like Uncertainty Principle for generative AI: as instruments for fixing a culture's position grow more precise, our capacity to predict its trajectory degrades, because the machinery of cultural description and cultural action have merged. We close with ethical questions raised by AI-mediated cultural drift, recursive synthetic culture, and the limits of human cultural agency, offering guidelines for safe, equitable, and privacy-preserving research on AI and culture.
|
| 1831 |
VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models
2610.06271
|
cs.AI
|
Jaemin Kim, Jiahn Kim, Taesik Gong |
Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) ...Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inference-level memory, but accurate gradient estimation requires many perturbation queries, making naive ZO prohibitively slow for large VLA models. We present VLA-ZO, a framework for fast ZO adaptation that exploits the structure of VLA computation. By confining adaptation to the action side, VLA-ZO keeps the expensive vision-language prefix frozen and reuses its conditioning states across perturbation queries and optimizer steps, while schedule-aware prefetching hides state-transfer overhead. On LIBERO camera-viewpoint shifts, VLA-ZO reduces end-to-end adaptation time by 25.59$\times$ at $q=16$ and 32.54$\times$ at $q=64$ relative to baseline ZO, while improving average task success from 48.27% without adaptation to 58.17% and 63.58%, respectively. These results show that making ZO faster can make larger query budgets practical, providing a promising path toward resource-efficient VLA adaptation on deployment platforms.
|
| 1832 |
Future Anchored Verification and Online Recovery for World Action Models
2610.06280
|
cs.AI
|
Zhibin Qin, Zhenxiong Tan, Xinchao Wang |
World action models (WAMs) have emerged as a promising paradigm for robotic manipulation. They act by first predicting how a task should be performed and then decoding the actions from that future. However, the remaining actions are invalid once execution drif...World action models (WAMs) have emerged as a promising paradigm for robotic manipulation. They act by first predicting how a task should be performed and then decoding the actions from that future. However, the remaining actions are invalid once execution drifts from the prediction. Simply replanning from the already out of distribution state rarely restores what the task still requires; existing execution monitors decide when to stop, but not what to restore. We observe that the answer is already in hand: the future the WAM predicted before acting depicts exactly the states it intended to pass through. We introduce FAVOR (Future Anchored Verification and Online Recovery), a lightweight framework that keeps these predicted frames as anchors and uses them for verification and recovery. An Anchor Verifier compares each observation with its anchor, together with the executed actions, to flag deviations that break the task. Anchor-Guided Recovery uses a vision-language model to turn the flagged anchor into a short corrective instruction. Under strengthened instruction guidance, the WAM executes this instruction to return to the intended future. It then resumes the task. FAVOR raises the task success of the base WAM from 97.85% to 98.10% on LIBERO and from 72.60% to 72.98% on LIBERO-Plus without modifying the policy.
|
| 1833 |
Fine-Tuning a 3B-Parameter LLM on a Smartphone: Characterizing Sustained Training
2610.06325
|
cs.AI
|
Andrew Geyko, Marius Mosbach, Andr\'e Brinkmann |
Multi-billion-parameter LLMs now run on phones for inference, and training them on the device would personalize them without user data leaving the phone. Prior work has measured individual training steps of such models on phones, but not complete training runs...Multi-billion-parameter LLMs now run on phones for inference, and training them on the device would personalize them without user data leaving the phone. Prior work has measured individual training steps of such models on phones, but not complete training runs, and not whether adapters trained on the device improve personalization. We present the first systematic characterization of a multi-billion-parameter LLM fine-tuned on a mobile device, covering memory, per-step time, thermal behavior, and energy. An iPhone 17 Pro can fine-tune a 3B-parameter LLM to a typical user within one battery charge, and the resulting adapters improve personalization as much as adapters trained on a server. Sustained training throttles the phone to about half its initial throughput, and none of the pausing or burst schedules we tested recovers it. Nearly all of each training step is spent in the frozen base model, most of it in the backward pass, which nine of the ten other runtimes we audited do not accelerate. Apple's MLX had a kernel for it that was never dispatched and was incorrect, and our repair, now merged upstream, trains an adapter 1.47x faster on a third less energy. On-device fine-tuning is feasible on current phones, and making it efficient requires runtimes and operating systems to treat training as a first-class workload.
|
| 1834 |
SoK: Semantic Decision Engines in Network Control Loops
2610.06425
|
cs.AI
|
Delong Li, Chen Li, Xu Wang, Haochen Gong, Rui Lang |
A semantic decision engine such as Jev can return a valid answer and still miss a network deadline, select an infeasible action or leave the service unverified. We systematize 139 paper families by decision interface, execution path and check ownership. Fifty ...A semantic decision engine such as Jev can return a valid answer and still miss a network deadline, select an infeasible action or leave the service unverified. We systematize 139 paper families by decision interface, execution path and check ownership. Fifty families claim that their engine fits a control loop or time budget, but only four support the claim with matched measurement. Across all 139, four report deadline attainment. The gap concentrates where the decision has no deterministic computation step. Those 72 families make 22 of the claims, none supported, and name a coverage owner in only two. Bounded tests under one event model show that each gap can reverse an admission verdict. A decision that meets a 10 s budget for every isolated request meets it for none once decisions queue ahead of replayed execution times. The same engine passes one coverage check and fails another. We derive a minimum reporting record, design rules and a research agenda for admitting decision engines to control loops.
|
| 1835 |
Choosing an energy-efficient software architecture for building system diagnostic support
2610.06444
|
cs.AI
|
Roxane Koitz-Hristov, Franz Wotawa |
Around 30\% of global energy expenditure can be attributed to the building sector, where a large portion of energy-consumption could be avoided by repairing existing faults. Fault detection and diagnosis (FDD) software addresses this issue; however, its creati...Around 30\% of global energy expenditure can be attributed to the building sector, where a large portion of energy-consumption could be avoided by repairing existing faults. Fault detection and diagnosis (FDD) software addresses this issue; however, its creation and operation also have an environmental impact. The magnitude of this impact is influenced by the diagnosis architecture, as different architectures and methods have different energy demands. Yet, simply considering the energy consumed by the software itself is not sufficient to assess its overall environmental impact, since the diagnostic performance, e.g., number of detected faults or number of faults missed, also contributes to its ecological footprint. In this paper, we propose an energy-consumption model that considers FDD performance and energy spend directly by the diagnosis software. In an initial experiment, we compare several FDD architecture families, i.e., rule-based, model-based, classical machine learning, and large-language-model-based, in simulation using performance and energy-consumption values collected from prior literature. The results show that considering the computational energy and accuracy of FDD can change the relative benefit of the different approaches. Computationally efficient machine learning methods, such as random forest, provide the largest net savings on smaller buildings, whereas more resource-intensive approaches, such as fine-tuned large language models, become advantageous as building size increases. Our findings suggest that overall energy efficiency depends not only on the computational demand of the FDD software, but also on its diagnostic performance and the scale of the building.
|
| 1836 |
Odyssey: A Closed-Loop Benchmark for Long-Horizon Real-World Driving with Explicit Navigation Routes
2610.06469
|
cs.AI
|
Jungho Kim, Hongjae Shin, Seunghoon Yu, Heecheol Yoo, Myeongjun Kim |
Closed-loop evaluation of end-to-end driving requires continuous rollouts that reveal how earlier decisions affect subsequent driving. However, existing benchmarks evaluate only short segments and fail to capture later consequences. Ambiguous directional comma...Closed-loop evaluation of end-to-end driving requires continuous rollouts that reveal how earlier decisions affect subsequent driving. However, existing benchmarks evaluate only short segments and fail to capture later consequences. Ambiguous directional commands also obscure the intended navigation objective. We introduce Odyssey, a closed-loop benchmark for long-horizon driving comprising 100 scenarios, each reconstructed from a 100-second nuPlan driving log to preserve the context of navigation maneuvers and traffic interactions. To provide a consistent navigation objective, Odyssey replaces directional commands with explicit standard-definition (SD) map routes that specify which roads to follow, while sensor-based planning determines local driving actions. Throughout these rollouts, diffusion-based refinement of 3DGS-rendered images reduces rendering artifacts along the ego trajectory. To assess how effectively planners follow these routes and prepare for upcoming maneuvers, we introduce SD Route Compliance and Pre-Lane Change Score. These assessments are complemented by RouteDS, which extends the Driving Score with penalties for SD-route deviations and failed lane preparation. We adapt state-of-the-art planners, including vision-language-action (VLA) models, and evaluate their navigation performance using these metrics. Odyssey highlights open questions in route representation and integration for E2E driving. Benchmark code and adapted baselines will be released publicly.
|
| 1837 |
ArtifactArena: Evaluating Models by What They Build in the Physical World
2610.06511
|
cs.AI
|
Kushagra Tiwary*, David Mayo*, Nikhil Behari, Xiangzhou Sun, Abdulrahman Alabdulkareem |
To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-softw...To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit \href{https://artifactarena.ai}{https://artifactarena.ai} for more information.
|
| 1838 |
GCTAuto-encoder: A Cross modal Framework for Security Flaw Detection in IoT Networks
2610.06517
|
cs.AI
|
Najmieh Sadat Safarabadi |
IoT encompasses diverse physical entities, from smart home devices to autonomous vehicles, creating a complex environment with heterogeneous security models. This heterogeneity makes IoT sub-systems vulnerable to various network attacks. Modern security system...IoT encompasses diverse physical entities, from smart home devices to autonomous vehicles, creating a complex environment with heterogeneous security models. This heterogeneity makes IoT sub-systems vulnerable to various network attacks. Modern security systems must therefore be more robust to ensure security and privacy for IoT applications. A highly secure IoT system also demands real time insight, requiring data collection at the edge of the computing layer. This diversity calls for a unified security model applied at the foundational level. Edge intelligence offers a direct approach to handling device diversity. A key goal of edge intelligence in IoT is to extract insight from local data; security models can then use this data to build local node protections, and integrating AI models yields an advanced security solution. This research proposes a novel deep learning algorithm for effective intrusion detection at the edge, supported by a cloud-based IoT framework. We evaluate the proposed cross modal deep learning algorithm against baseline models. The contribution is a cross domain Deep Neural Network (DNN) algorithm for intrusion detection. The objective is to assess a multi-method deep learning model to detect intrusions in IoT systems at the edge via community detection with modeled attention. We evaluate GCT auto-encoder, a novel framework integrating edge intelligence to identify security flaws. The model significantly improves performance and efficiency. On a network intrusion IoT dataset covering multiple attack scenarios, it achieved 0.908 accuracy, reduced learning loss to 0.00156, and outperformed existing approaches.
|
| 1839 |
Symmetry and AI-assisted discovery of magic-state factories
2610.06535
|
cs.AI
|
Shubham P. Jain, Adam Wills, Shraddha Singh |
Magic-state distillation is a major resource cost in fault-tolerant quantum computing. The cost of a magic-state factory depends strongly on its failure rate, which grows with the number of input magic states. Although symmetry-restricted methods have recently...Magic-state distillation is a major resource cost in fault-tolerant quantum computing. The cost of a magic-state factory depends strongly on its failure rate, which grows with the number of input magic states. Although symmetry-restricted methods have recently made distance two searches tractable, distance three and above have remained elusive at moderate input counts. We develop symmetry- and AI-assisted methods to search this regime. We present a unified binary-matrix formulation encompassing both triorthogonal-code and direct circuit searches. We show that distance at least three is equivalent to nonzero, pairwise distinct syndromes, separating the choice of syndromes from the search for compatible output gates. We restrict the syndrome search using group symmetry and language-model agents, followed by deterministic solving and independent verification. Our searches yield 699 factory classes, including 564 new ones. These include factories for pure-T states and factories with entangled outputs comprising combinations of T, CS, and CCZ magic states. The pure-T factories [[63, 11, 3]] and [[850, 128, 6]] achieve the lowest overhead exponents we know among protocols with at most 100 and 1000 inputs, respectively, with $\gamma = 1.589$ and $\gamma = 1.057$. Our [[1715, 287, 6]] factory, with $\gamma = 0.998$, is the smallest known pure-T factory with $\gamma < 1$. We also provide a context directory of search briefs and campaign notes with which readers can train their own agents and tailor the search to their requirements. With these results, we begin constructing an active, open-source repository of magic-state distillation protocols for the quantum community, supplemented by our methods and data, for the practical fault-tolerant quantum computing regime.
|
| 1840 |
Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attacks
2610.06584
|
cs.AI
|
Tobias Heldt, Matt Turk, Christoph Landolt, Mario Fritz |
The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult to distinguish prior exposure fr...The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult to distinguish prior exposure from capability on unseen vulnerabilities. We evaluate open-weight and proprietary AI models on exploit generation, vulnerability repair and subsequent attacks in five nonpublic software environments, including vulnerabilities we privately disclosed while they remained unpatched. Researcher-developed and reviewed deterministic graders, not LLM judges, determine task scores. Comparisons with 209 disclosed vulnerabilities and cryptographic challenges reveal substantial variation across systems and vulnerability types. Repair scores exceed attack scores in two nonpublic environments and fall below them in three. Passing an initial security test is also insufficient: another exploit succeeds in 92 of 524 non-independent defender test intervals after the initial exploit is stopped. These results motivate vulnerability-specific attack-repair comparisons and subsequent resistance tests.
|
| 1841 |
Large Language Model-Guided Discovery of Weight-Five Bivariate Bicycle Codes
2610.06623
|
cs.AI
|
Juan Cruz-Benito |
Building on our earlier program-evolution workflow guided by large language models (LLMs), we study weight-five bivariate bicycle (BB) and perturbed bivariate bicycle (PBB) codes. The resulting catalogue contains 1,142 distinct code proposals, including 1,081 ...Building on our earlier program-evolution workflow guided by large language models (LLMs), we study weight-five bivariate bicycle (BB) and perturbed bivariate bicycle (PBB) codes. The resulting catalogue contains 1,142 distinct code proposals, including 1,081 nonbaseline proposals attributable to LLM-generated programs. Across the catalogue, we certify connected Calderbank--Shor--Steane (CSS) realizations [[96,4,10]], [[140,6,10]], and [[180,4,14]]. A post-search comparison certifies seven imported Lin--Pryadko archive constructions. For leading parameter triples also represented in that archive, we provide exact distance evidence, explicit bivariate presentations, and verified component reductions. A basis-independent connectivity analysis identifies 409 of the 1,142 catalogue entries as disconnected and shows that 73.1\% of the classes with exact distance certificates contain repeated connected components. Algebraic analysis organizes the connected CSS classes into order-3, order-7, and order-15 cyclotomic-kernel strata. The strongest exact connected PBB parameter point is [[216,4,10]], attained by two distinct component classes. Among the 936 distinct CSS proposals from the LLM-guided campaign with a recorded positive distance, 816 (87.18\%) are certified at $d\geq5$. For comparison, three random-search controls each sample 6,444 CSS code proposals uniformly without replacement, using the same per-lattice and encoded-dimension sample counts as the LLM-guided campaign. In these controls, 4,672--4,785 proposals (72.50--74.26\%) meet the same criterion. The LLM-guided campaign has the higher certified yield, while the random controls cover more connected classes. Together, these results extend LLM-guided discovery to a more constrained code family and provide a reproducible structural and exact-distance account of its strongest candidates.
|
| 1842 |
Towards LLM Agents for Earth Observation
2504.12110
|
cs.AI
|
Chia Hsiang Kao, Wenting Zhao, Cheryl Lam, Aarush Umap, Shreelekha Revankar |
Earth Observation (EO) provides critical planetary data for environmental monitoring, disaster management, climate science, and other scientific domains. In this work we ask: Are AI systems ready for reliable Earth Observation? To answer this, we introduce Uni...Earth Observation (EO) provides critical planetary data for environmental monitoring, disaster management, climate science, and other scientific domains. In this work we ask: Are AI systems ready for reliable Earth Observation? To answer this, we introduce UnivEARTH, a coding benchmark of 408 yes/no questions from NASA Earth Observatory articles across 7 various topics and over 15 satellite instruments and sources. Using Google Earth Engine API as a tool in a zero-shot setup, LLM agents achieve an accuracy of 40.0% where the code fails to run over 44% of the time. To better understand LLM agent behavior, we also analyze the impact of using the JavaScript API versus Python and the effect of providing documentation. Furthermore, we find that using a Reflexion framework significantly reduces errors: Claude-4.5-Sonnet, Gemini-2.5-Pro, and GPT-5 accuracies rise to around 60%. However, these results remain only marginally above random chance. Taken together, our findings identify significant challenges to be solved before AI agents can automate earth observation, and suggest paths forward.
|
| 1843 |
OpenPhone: Mobile Agentic Foundation Models
2510.22009
|
cs.AI
|
Yangqin Jiang, Chao Huang |
With the advancement of multimodal large language models (MLLMs), building GUI agent systems has become an increasingly promising direction--especially for mobile platforms, given their rich app ecosystems and intuitive touch interactions. Yet mobile GUI agent...With the advancement of multimodal large language models (MLLMs), building GUI agent systems has become an increasingly promising direction--especially for mobile platforms, given their rich app ecosystems and intuitive touch interactions. Yet mobile GUI agents face a critical dilemma: truly on-device models (4B or smaller) lack sufficient performance, while capable models (starting from 7B) are either too large for mobile deployment or prohibitively costly (e.g., cloud-only closed-source MLLMs). To resolve this, we propose OpenPhone, a mobile GUI agent system that leverages device-cloud collaboration to tap the cost-efficiency of on device models and the high capability of cloud models, while avoiding their drawbacks. Specifically, OpenPhone enhances Qwen2.5-VL-3B via two-stage SFT->GRPO training on synthetic GUI data for strong decision-making, integrates an efficient long-reasoning and memory management mechanism to utilize historical interactions under tight resources, and defaults to on-device execution--only escalating challenging subtasks to the cloud via real-time complexity assessment. Experiments on the online AndroidLab benchmark and diverse apps show OpenPhone matches or nears larger models, with a significant reduction in cloud costs.
|
| 1844 |
An Empirical Study of On-Device Translation for Real-Time Live-Stream Chat on Mobile Devices
2601.02641
|
cs.AI
|
Jeiyoon Park, Daehwan Lee, Changmin Yeo, Yongshin Han, Minseop Kim |
Despite its efficiency, there has been little research on the practical aspects required for real-world deployment of on-device AI models, such as the device's CPU utilization and thermal conditions. In this paper, through extensive experiments, we investigate...Despite its efficiency, there has been little research on the practical aspects required for real-world deployment of on-device AI models, such as the device's CPU utilization and thermal conditions. In this paper, through extensive experiments, we investigate two key issues that must be addressed to deploy on-device models in real-world services: (i) the selection of on-device models and the resource consumption of each model, and (ii) the capability and potential of on-device models for domain adaptation. To this end, we focus on a task of translating live-stream chat messages and manually construct LiveChatBench, a benchmark consisting of 1,000 Korean-English parallel sentence pairs. Experiments on five mobile devices provide a systematic empirical assessment of widely adopted on-device models, highlighting the importance of model selection and deployment constraints when adapting them to specialized tasks and serving a large and heterogeneous user base. We expect that our findings will offer practical insights into both the capabilities and limitations of on-device models for real-world AI service deployment.
|
| 1845 |
Neutral Substrates: A Design Constraint for Shared Records Under Persistent Interpretive Disagreement
2601.14271
|
cs.AI
|
Denise M. Case |
Shared accountability records are often used by parties who may never agree about causation, responsibility, or normative interpretation. For such records, neutrality cannot be achieved by omitting contested information, because accountability requires preserv...Shared accountability records are often used by parties who may never agree about causation, responsibility, or normative interpretation. For such records, neutrality cannot be achieved by omitting contested information, because accountability requires preserving the claims parties made, with their sources and provenance. Nor can neutrality be achieved by asserting one contested interpretation as the shared base. This paper defines a neutral substrate as a shared representational layer that provides stable reference while making no object-level substrate-layer commitments to causal or normative propositions. The central design constraint is that, when causal and normative propositions are contestable across admissible frameworks and the substrate's referential commitments are common ground, the substrate's neutrality is guaranteed at design time if and only if its foundational layer is restricted to those referential commitments and attribution propositions whose attributional basis is fixed by them. Causal and normative content may still be represented, but not as object-level foundational-layer commitments: it may appear there only as the content of attributed assertions with provenance, made by some identified framework, source, agent, institution, record, or document. The representational machinery used here is standard: reification, attribution, and provenance. The contribution is the constraint: a checkable condition on the foundational layer of a shared record, stated together with the assumptions it depends on and the boundary condition under which the constraint does not apply. A neutral substrate says enough to preserve accountability, but it does not turn one party's interpretation into an object-level substrate-layer commitment. The constraint does not apply at that layer when the referential regime or attributional basis is contested among the frameworks in play.
|
| 1846 |
AI Mental Models: Learned Intuition and Deliberation in a Bounded Neural Architecture
2603.22561
|
cs.AI
|
Laurence Anthony |
This paper asks whether a bounded neural architecture can exhibit a meaningful division of labor between intuition and deliberation on a classic 64-item syllogistic reasoning benchmark. More broadly, the benchmark is relevant to ongoing debates about world mod...This paper asks whether a bounded neural architecture can exhibit a meaningful division of labor between intuition and deliberation on a classic 64-item syllogistic reasoning benchmark. More broadly, the benchmark is relevant to ongoing debates about world models and multi-stage reasoning in AI. It provides a controlled setting for testing whether a learned system can develop structured internal computation rather than only one-shot associative prediction. Experiment 1 evaluates a direct neural baseline for predicting full 9-way human response distributions under 5-fold cross-validation. Experiment 2 introduces a bounded dual-path architecture with separate intuition and deliberation pathways, motivated by computational mental-model theory (Khemlani & Johnson-Laird, 2022). Under cross-validation, bounded intuition reaches an aggregate correlation of r = 0.7272, whereas bounded deliberation reaches r = 0.8152, and the deliberation advantage is significant across folds (p = 0.0101). The largest held-out gains occur for NVC, Eca, and Oca, suggesting improved handling of rejection responses and c-a conclusions. A canonical 80:20 interpretability run and a five-seed stability sweep further indicate that the deliberation pathway develops sparse, differentiated internal structure, including an Oac-leaning state, a dominant workhorse state, and several weakly used or unused states whose exact indices vary across runs. These findings are consistent with reasoning-like internal organization under bounded conditions, while stopping short of any claim that the model reproduces full sequential processes of model construction, counterexample search, and conclusion revision.
|
| 1847 |
ReLope: From Hidden-State Probing to a Decision Module for Multimodal LLM Routing
2603.24787
|
cs.AI
|
Yaopei Zeng, Congchao Wang, Blake JianHang Chen, Lu Lin |
Routing balances performance and cost in hybrid systems by escalating selected queries from a lightweight model to a powerful but expensive one. Hidden-state probes provide an effective routing signal by predicting the small model's correctness from hidden sta...Routing balances performance and cost in hybrid systems by escalating selected queries from a lightweight model to a powerful but expensive one. Hidden-state probes provide an effective routing signal by predicting the small model's correctness from hidden states it already computes. Although effective in text-only settings, their behavior on multimodal LLMs (MLLMs) is less understood. We find that correctness is harder to predict from hidden states when questions are paired with images rather than with captions of those images. Two training-free dependence measures, HSIC and label CKA, show the same pattern. We first investigate whether changing hidden states used by the probe improves correctness prediction for MLLMs. We propose Attention Probe, which learns an attention query to pool frozen token states into a routing feature and improves over the standard probe. However, pooling only recombines hidden states computed for answer generation, not for deciding whether the answer should be trusted. We therefore build a decision layer inside the existing MLLM and train it for the routing decision. The resulting router, ReLope (KL-Regularized LoRA Probe), adapts a separate copy of one MLLM layer with LoRA and regularizes it with a KL-penalized stochastic bottleneck. This decision branch shares the frozen lower layers and leaves the model's answer-generation path unchanged. ReLope attains the highest correctness prediction AUC on all five benchmarks and three backbones we evaluate and improves the accuracy versus cost trade-off at negligible overhead. These results suggest a new perspective and a practical technical path on routing: rather than reading decisions off states built for generation, routing decisions can be learned by a dedicated decision layer inside the model. Code: https://github.com/Spinozaaa/ReLope
|
| 1848 |
AI Assistance Reduces Persistence and Hurts Independent Performance
2604.04721
|
cs.AI
|
Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker, Rachit Dubey |
People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results. In contrast, current AI systems are...People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person's growth over immediate results. In contrast, current AI systems are fundamentally short-sighted collaborators - optimized for providing instant and complete responses, without ever saying no (unless for safety reasons). What are the consequences of this dynamic? Here, through a series of randomized controlled trials on human-AI interactions (N = 1,222), we provide causal evidence for two key consequences of AI assistance: reduced persistence and impairment of unassisted performance. Across a variety of tasks, including mathematical reasoning and reading comprehension, we find that although AI assistance improves performance in the short-term, people perform significantly worse without AI and are more likely to give up. Notably, these effects emerge after only brief interactions with AI (approximately 10 minutes). These findings are particularly concerning because persistence is foundational to skill acquisition and is one of the strongest predictors of long-term learning. We posit that persistence is reduced because AI conditions people to expect immediate answers, thereby denying them the experience of working through challenges on their own. These results suggest the need for AI model development to prioritize scaffolding long-term competence alongside immediate task completion.
|
| 1849 |
High-Precision Estimation of the State-Space Complexity of Shogi via the Monte Carlo Method
2604.06189
|
cs.AI
|
Sotaro Ishii, Tetsuro Tanaka |
Determining the state-space complexity of Shogi (Japanese Chess) has been a challenging problem, with previous combinatorial estimates leaving a gap of five orders of magnitude ($10^{64}$ to $10^{69}$). This gap arises from the difficulty of distinguishing pos...Determining the state-space complexity of Shogi (Japanese Chess) has been a challenging problem, with previous combinatorial estimates leaving a gap of five orders of magnitude ($10^{64}$ to $10^{69}$). This gap arises from the difficulty of distinguishing positions reachable from the initial position among the vast number of board configurations. In this paper, we present a high-precision statistical estimate of the number of reachable Shogi positions. Here, a position consists of the side to move, the board configuration, and the pieces in hand, without move history, and positions related by exchanging the two players or by horizontal reflection of the board are counted once. Our method combines Monte Carlo sampling with a reachability test that performs reverse search toward the set of ``King-King only'' (KK) positions rather than toward the single initial position. Since every KK position and the initial position are mutually reachable, reaching any KK position proves reachability, and removing pieces from the board is a much simpler objective than reconstructing the initial position. Based on a sample of five billion positions, we estimate the number of such positions in Shogi to be $6.55 \times 10^{68}$; both ends of the approximate $3\sigma$ confidence interval round to this value at three significant digits. We also applied this method to Mini Shogi, obtaining approximately $2.38 \times 10^{18}$.
|
| 1850 |
DERM-3R: A Resource-Efficient Multimodal Agents Framework for Dermatologic Diagnosis and Treatment in Real-World Clinical Settings
2604.09596
|
cs.AI
|
Jirui Dai, Zhendong Wang, Chongjing Wang, Yusha Bai, Ziwen Chen |
Skin diseases impose a substantial and growing global health burden. Modern dermatologic therapies control acute manifestations rapidly, but single-target treatment paradigms, recurrent disease courses, and neglected systemic comorbidities limit their long-ter...Skin diseases impose a substantial and growing global health burden. Modern dermatologic therapies control acute manifestations rapidly, but single-target treatment paradigms, recurrent disease courses, and neglected systemic comorbidities limit their long-term effectiveness. Traditional Chinese medicine (TCM) offers a complementary, holistic approach through syndrome differentiation and individualized treatment, yet its clinical practice is hindered by non-standardized knowledge, incomplete multimodal records, and the difficulty of scaling expert-driven reasoning. We propose DERM-3R, a resource-efficient multimodal agents framework that models TCM dermatologic diagnosis and treatment under limited data and computational resources. We decompose and redefine real-world dermatologic decision-making into three essential issues: fine-grained lesion recognition, multi-view lesion representation with specialist-level pathogenesis modeling, and holistic clinical reasoning for syndrome differentiation and treatment planning. DERM-3R comprises three collaborative agents, DERM-Rec, DERM-Rep, and DERM-Reason, each addressing one of these issues. Built on a lightweight multimodal large language model and fine-tuned by partial-parameter finetuning on 103 real-world TCM psoriasis cases, DERM-3R performs strongly across multiple dermatologic reasoning tasks. Automatic metrics, LLM-as-a-Judge assessments, and human doctors show that, despite extremely limited training data and parameter updates, DERM-3R matches or even surpasses hundred-billion-parameter general-purpose multimodal models such as GPT-5.1 and Gemini-3-Flash. Our results indicate that structured, domain-aware multi-agent modeling is an effective alternative to brute-force scaling for complex clinical tasks, offering a practical and scalable paradigm for multimodal AI in dermatology and integrative medicine.
|
| 1851 |
From Tool to Agent: How Worker Needs Reorganize Across the Agentic Roles of Workplace AI
2605.03078
|
cs.AI
|
Christine P. Lee, Min Kyung Lee, Bilge Mutlu |
As AI systems gain agency in workplaces, workers increasingly interact with systems that do more than support tasks: they make decisions, allocate work, and shape how workers communicate with others. We interviewed 16 workers in healthcare, finance, and manage...As AI systems gain agency in workplaces, workers increasingly interact with systems that do more than support tasks: they make decisions, allocate work, and shape how workers communicate with others. We interviewed 16 workers in healthcare, finance, and management about their daily interactions with workplace AI systems. Across these domains, AI occupied three agentic roles: \emph{assistant}, supporting worker decision-making; \emph{arbiter}, making consequential decisions that workers had to explain to downstream stakeholders; and \emph{manager}, directing work and displacing social structures through which work had been negotiated and learned. We find that worker needs depended on how decision authority and accountability were configured around the AI system. Explanation, control, accountability, and communication each took different meanings depending on whether AI supported, decided, or directed work. We contribute a role-based account of workplace human-agent interaction and identify accountability-authority asymmetry as a mechanism through which AI deployments can place responsibility on workers for outcomes whose authority has shifted to AI. We discuss implications for designing workplace AI around both the agent and its surrounding agency configuration.
|
| 1852 |
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
2605.07247
|
cs.AI
|
Yi Liu, TingFeng Hui, Wei Zhang, Li Sun, Ningxin Su |
Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, brittle to extend, and fundamentally limited in diversity. A promising direction is...Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, brittle to extend, and fundamentally limited in diversity. A promising direction is to replace manually crafted environments with LLM-simulated counterparts. However, this paradigm hinges on an unexamined core assumption: LLMs can accurately simulate environmental feedback. In practice, LLM-simulated environments suffer from hallucinations, logical inconsistencies, and silent state drift failures that corrupt agent reward signals and compound the construction costs that the paradigm was designed to eliminate. To address this gap, we propose EnvSimBench with four contributions: 1) We provide the first formal definition and operationalization of Environment Simulation Ability (EnvSim Ability) as a quantifiable research objective. 2) We construct EnvSimBench, a rigorous benchmark covering 400 samples across 167 diverse environments, equipped with verifiable labels and fine-grained difficulty stratification along three axes. 3) Systematic evaluations reveal that all state-of-the-art language models suffer from a universal state change cliff: they achieve near-perfect accuracy on tasks when the environment state remains invariant, yet fail catastrophically when multiple states need simultaneous updates. This finding exposes EnvSim Ability as a critical yet largely unaddressed capability gap. 4) We design a constraint-driven simulation pipeline that substantially reduces hallucination, boosts environment synthesis yield by 6.8%, and cuts costs by over 90%. Overall, EnvSimBench serves as both a diagnostic framework and a practical optimization path for reliable LLM-based environment simulation, establishing a foundation for scalable agent training. Code and data are available at https://github.com/cookieApril/EnvSimBench
|
| 1853 |
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
2605.08956
|
cs.AI
|
Harshit Bisht, Vinay Kumar, Kevin Maik Jablonka, Mausam, N. M. Anoop Krishnan |
A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they already function as co-scientists, agentic AI scientists are not built for autonomous scientific discovery. We ide...A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they already function as co-scientists, agentic AI scientists are not built for autonomous scientific discovery. We identify the following challenges in building and deploying autonomous AI scientists: (1) Problem selection is influenced by the McNamara fallacy; (2) Agents are built on large language models (LLMs) whose training corpora omit tacit procedural and failure knowledge of laboratory practice; (3) Preference optimisation during post-training compresses output diversity toward consensus; and (4) Most scientific benchmarks measure single-turn prediction accuracy and lack feedback from physical experiments back to the computational model. These challenges are not just questions of scale and scaffolding; they require revisiting fundamental design choices. To build truly autonomous AI scientists, we recommend scientific simulations as verifiers for training, a persistent, mutable epistemic state that carries beliefs and shifting objectives across an investigation, the establishment of a centralized preregistration repository for all AI-generated hypotheses, and application driven by scientific need rather than tool affordance.
|
| 1854 |
PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning
2605.09287
|
cs.AI
|
Dongyi Liu, Yifan Niu, Qinwen Wang, Han Xiao, Jia Li |
Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensive tasks. However, existing methods encounter critical challenges in long-horizon credit assignment: (i) Rew...Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensive tasks. However, existing methods encounter critical challenges in long-horizon credit assignment: (i) Reward Sparsity, where models receive only outcome feedback without step-level guidance to differentiate action quality; (ii) Isolated Credit, where credit is assigned to steps independently, failing to capture sequential dependencies; and (iii) Distributional Shift, where rewards are estimated on templates that deviate from the model's natural generative distribution. To address these issues, we propose Pivot-Based Credit Assignment (PiCA), a novel step reward mechanism that reformulates the search trajectory as a sequential process of cumulative search progress. Unlike prior isolated step rewards, PiCA defines process rewards as success probabilities dependent on the historical context based on Potential-Based Reward Shaping (PBRS). This approach identifies pivot steps, which comprise target golden sub-queries and sub-answers derived from historical trajectories, as information peaks that significantly boost the likelihood of a correct final answer. By anchoring these step rewards to the final task objective, PiCA provides dense, pivot-aware and trajectory-dependent guidance while maintaining distributional consistency. Extensive experiments show that PiCA outperforms existing strong baselines across seven knowledge-intensive QA benchmarks, achieving 15.2% and 4.4% improvements for 3B and 7B models. The consistent performance gains across various models show PiCA's robust generalization. The code is available at https://github.com/novdream/PiCA.
|
| 1855 |
MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs
2605.10064
|
cs.AI
|
Ruiyi Yang, Zechen Li, Hao Xue, Imran Razzak, Flora D. Salim |
Self-evolving language-model agents must decide what to learn next and how to preserve what they have learned across iterations. Existing systems typically carry this cross-iteration knowledge as natural-language feedback, flat episodic memory, or implicit rei...Self-evolving language-model agents must decide what to learn next and how to preserve what they have learned across iterations. Existing systems typically carry this cross-iteration knowledge as natural-language feedback, flat episodic memory, or implicit reinforcement signals, none of which cleanly supports a frozen weak backbone at inference time. This paper introduces MAGE (Multi-Agent Graph-guided Evolution), a framework that externalizes self-knowledge into a four-subgraph co-evolutionary knowledge graph. Its experience subgraph stores both teacher-written failure corrections and the learner's own past correct reasoning traces, which are retrieved as task-conditioned guidance for a frozen execution model. During evolution, the graph, a task-level search bandit, and a skill-level routing bandit are updated from the same reward stream, while the learner's backbone remains unchanged. We further provide structural analysis showing how append-only memory growth, bounded curriculum coverage, and task-filtered retrieval together support stable improvement of the retrieval substrate for frozen-learner evolution. Across nine benchmarks spanning mathematical reasoning, multi-hop and open-domain question answering, spatio-temporal analysis, financial numerical reasoning, medical multiple-choice, an open-world survival game, and web navigation, MAGE achieves strong performance against prompt-based frozen-backbone baselines. Ablations show that self-harvested success traces and teacher-written corrections are complementary, with success memories contributing most on reasoning-template-heavy tasks and corrective memories supporting harder composition and interaction settings. Code is available in https://github.com/RuiyiYang01/mage-alg
|
| 1856 |
Inference-Time Amplification of Weak Reasoning Models
2605.14163
|
cs.AI
|
Varun Sunkaraneni, Pierfrancesco Beneventano, Riccardo Neumarker, Tomaso Poggio, Tomer Galanti |
How much of the capability of a reasoning model exposed by repeated sampling can an imperfect selector recover? We study this problem as {\em inference-time amplification}, separating {\em coverage}---whether a useful candidate is generated---from {\em identif...How much of the capability of a reasoning model exposed by repeated sampling can an imperfect selector recover? We study this problem as {\em inference-time amplification}, separating {\em coverage}---whether a useful candidate is generated---from {\em identifiability}---whether it can be recognized among the alternatives. We show that better coverage alone does not guarantee better selection, while even imperfect local selection signals can be amplified through repetition. We also characterize proposal-side blind spots that additional sampling cannot overcome. On SWE-bench Verified, a single GPT-5.4 nano trajectory solves 67.0\% of tasks, while eight proposals contain a correct patch on 79.0\%; critic--comparator selection reaches 76.4\%, recovering 78.3\% of the oracle-exposed gain. The same pattern appears across proposer families and in mathematical reasoning, competitive programming, and formal proof. On MATH L4--5 and GSM-Plus, our selector recovers 85\% and 77\% of the oracle gap and substantially outperforms majority vote. Holding proposal coverage fixed, providing the selector with additional repository information also improves recovery. Conversely, when the available selection signal is weak, redundant, or misaligned, amplification can diminish or fail. These results identify selection quality as the key factor governing how much capability exposed by repeated sampling can actually be recovered.
|
| 1857 |
Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations
2605.14175
|
cs.AI
|
Qisong He, Jinwei Hu, Xinmiao Huang, Changshun Wu, Yi Dong |
In a long conversation, an LLM may produce a fluent continuation that rests on premises the conversation has already abandoned. Context-manipulation attacks exploit precisely this weakness. We address this problem with a runtime verifier. An LLM Interpreter ma...In a long conversation, an LLM may produce a fluent continuation that rests on premises the conversation has already abandoned. Context-manipulation attacks exploit precisely this weakness. We address this problem with a runtime verifier. An LLM Interpreter maps each utterance to one or more of eight epistemic operations, and then a symbolic engine applies these operations to a dependency map that records what every claim rests on and whether it still stands. Based on the dependency map, checking whether a continuation is grounded then reduces to a walk over the map, linear in its size and requiring no LLM call. Retraction propagates through the same map with a conflict-free guarantee and flags exactly the conclusions that lose support. Our experiments with five QA models demonstrate substantial improvements in QA accuracy at a low cost per query. On ReviseQA for belief revision and MemoryAgentBench's FactConsolidation split (MemAB-FC), the verifier outperforms a retrieval baseline and raises MemAB-FC single-hop accuracy from $0.46$--$0.95$ to $0.94$--$0.98$. With the verifier, even the small 7B model overtakes unaided GPT-4o. When a GPT-4o Interpreter extracts every update from raw text rather than taking the benchmarks' structured updates, the verifier still outperforms the retrieval baseline. Moreover, on MemAB-FC, QA prompts remain compact at $97$--$174$ tokens while the full-context baseline reaches $114.5$K. Retraction queries take less than a microsecond at $2000$ turns.
|
| 1858 |
From Table to Cell: Attention for Better Reasoning with TABALIGN
2605.14465
|
cs.AI
|
Tung Sum Thomas Kwok, Zeyong Zhang, Xinyu Wang, Chunhe Wang, Xiaofeng Lin |
Multi-step LLM reasoning over structured tables fails because planning and execution share no explicit cell-grounding contract. Existing methods constrain the planner to a left-to-right factorization at odds with table permutation invariance, and score interme...Multi-step LLM reasoning over structured tables fails because planning and execution share no explicit cell-grounding contract. Existing methods constrain the planner to a left-to-right factorization at odds with table permutation invariance, and score intermediate states by generated content alone, overlooking cell grounding. We conduct a pilot study showing that diffusion language models (DLMs) produce more human-aligned and permutation-stable cell attention on tables than autoregressive models, with a 40.2% median reduction in attention-AUROC variability under row reordering. Motivated by this, we propose TABALIGN, a planned table reasoning framework that operationalizes the contract. TABALIGN pairs a masked DLM planner, whose bidirectional denoising emits plan steps as binary cell masks, with TABATTN, a lightweight verifier trained on 1,600 human-verified attention standards to score each step by its attention overlap with the plan-designated mask. Across eight benchmarks covering table question answering and fact verification, TABALIGN improves average accuracy by 15.76 percentage points over the strongest open-source baseline at comparable 8B-class scale, with a matched-backbone ablation attributing 2.87 percentage points of this gain to the DLM planner over an AR planner on a fixed reasoner. Cleaner DLM plans also accelerate downstream reasoning execution by 44.64%.
|
| 1859 |
ECG-WM: A Physiology-Informed ECG World Model for Clinical Intervention Simulation
2605.17580
|
cs.AI
|
Zhikang Chen, Yue Wang, Sen Cui, Yu Zhang, Changshui Zhang |
Electrocardiogram (ECG)-based models have achieved strong performance in diagnostic tasks, yet they remain limited in modeling how cardiac dynamics evolve under external interventions. In particular, existing approaches focus primarily on static prediction and...Electrocardiogram (ECG)-based models have achieved strong performance in diagnostic tasks, yet they remain limited in modeling how cardiac dynamics evolve under external interventions. In particular, existing approaches focus primarily on static prediction and lack mechanisms to capture ECG variations under different pharmacological conditions. In this work, we propose an ECG World Model for action-conditioned predictive simulation of cardiac electrophysiology. Moving beyond disjoint pipelines, our framework features a principled integration of physiological ordinary differential equation (ODE) priors into latent diffusion dynamics via energy regularization. This structural constraint enables the synthesis of physiologically plausible post-intervention ECG trajectories while effectively mitigating generative hallucinations. Building on this simulation process, we introduce an uncertainty-aware evaluation strategy that leverages the stochasticity of diffusion sampling to characterize both the expected clinical risk and its variability, allowing a more reliable comparative assessment of candidate interventions. We evaluate our method across diverse settings, including controlled drug-response scenarios and real-world clinical records. Beyond standard waveform metrics, experimental results demonstrate improved risk calibration and strong alignment with expert-informed treatment preferences. These results establish our approach as a robust foundation for safe and intervention-aware clinical decision support.
|
| 1860 |
Agentic Trading: When LLM Agents Meet Financial Markets
2605.19337
|
cs.AI
|
Yihan Xia, Panpan You, Taotao Wang, Fang Liu, Han Qi |
Large Language Models (LLMs) combined with autonomous agent architectures are shifting quantitative finance from isolated predictive modeling toward closed-loop trading agents. These systems perceive multimodal market signals, maintain context and memory, reas...Large Language Models (LLMs) combined with autonomous agent architectures are shifting quantitative finance from isolated predictive modeling toward closed-loop trading agents. These systems perceive multimodal market signals, maintain context and memory, reason about decisions, emit executable trading actions, and adapt to non-stationary market regimes. This survey provides a systematic synthesis of agentic trading through an Architecture--Capability--Adaptation (A-C-A) conceptual framework. Architecture covers perception (textual, numerical, and multimodal), memory (working memory, vector storage, and knowledge bases), reasoning (reactive, reflective, and planning-driven), and execution (order routing and friction management). Capability examines alpha discovery, portfolio optimization, and multi-tier risk control. Adaptation covers in-context learning, supervised fine-tuning, reinforcement learning, multi-agent coordination, and architectural self-evolution. Drawing on 185 systematically selected studies, including a core empirical subset of 80 systems evaluated under closed-loop market conditions, we analyze design spaces, operational trade-offs, and emerging paradigms. We also examine evaluation practices, including backtesting split schemes, transaction-cost realism, survivorship-bias handling, and execution fidelity, and distill methodological best practices for future research. By bridging artificial intelligence, market microstructure, and quantitative finance, this survey provides a roadmap for designing, benchmarking, and deploying reliable agentic trading systems.
|
| 1861 |
Towards Direct Evaluation of Harness Optimizers via Priority Ranking
2605.22505
|
cs.AI
|
Kai Tzu-iunn Ong, Minseok Kang, Dongwook Choi, Junhee Cho, Seungju Kim |
Harness optimization enables automated agent creation by having an optimizer agent iteratively update the harness of target agents. Despite its success, current studies evaluate optimizers solely by observing target agents' performance gains. This indirect end...Harness optimization enables automated agent creation by having an optimizer agent iteratively update the harness of target agents. Despite its success, current studies evaluate optimizers solely by observing target agents' performance gains. This indirect end-improvement evaluation neglects optimizers' actions at intermediate steps, which are often erroneous and hinder agent performance in deployment. Therefore, it is unclear whether harness optimization is driven by optimizers' informed update actions or simply trial-and-error. This necessitates direct evaluation of harness optimizers. However, evaluating harness optimizers directly is non-trivial and costly due to the lack of oracle harnesses. To address this, we present a simple, low-cost design to directly evaluate them, namely priority ranking. By asking harness optimizers to rank components (e.g., tools) in a given harness by their potential to improve/hinder agent performance when updated, our design quantifies optimizer ability at the step level without expensive rollouts or manual examination. More importantly, optimizers' ranking performance correlates with their ability to improve agents in actual multi-step harness optimization, establishing priority ranking as a reliable indicator of optimization ability. Priority ranking is enabled by SHOR, a collection of 182 human-verified optimization scenarios spanning across domains, designs, and time stages. Code and data: https://github.com/k59118/Harness_Optimizer_Evaluation.
|
| 1862 |
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
2605.28201
|
cs.AI
|
Yongxiang Li, Moxin Li, Zhixin Ma, Fengbin Zhu, Wenjie Wang |
Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP context, causing harmful agentic behaviors s...Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP context, causing harmful agentic behaviors such as unsafe actions or incorrect outputs. Existing studies typically focus on single-interaction attacks, where the agent observes adversarial content and immediately exhibits harmful behavior within one user request. However, we show that adversarial content can also persist across interactions served by the same agent, making such threats harder to detect and mitigate. Specifically, adversarial content may persist in the agent state, remain dormant across interactions, and later be activated by a benign user query. We formalize this type of safety threat as Sleeper Attack. To evaluate it, we construct a benchmark with 1,896 instances covering six real-world harmful outcomes, three attack strategies, and three agent state targets: session context, memory, and reusable skills. Experiments on seven strong open-source and closed-source LLMs show that state-of-the-art LLM agents remain vulnerable to Sleeper Attack, even when they achieve low attack success rates under a single-interaction baseline. Our code and data are available at https://anonymous.4open.science/r/skdvnfu23ihr9wdscnksf1asdffsaef.
|
| 1863 |
MEMENTO: Leveraging Web as a Learning Signal for Low-Data Domains
2605.29795
|
cs.AI
|
Ashutosh Ojha, Vinay Aggarwal, Ashutosh Srivastava, Siddharth Yedlapati, Yaman K Singla |
Real-world tasks often lack large labeled datasets, motivating extensive work on learning in low-data regimes. Existing approaches such as few-shot prompting, instruction tuning, and synthetic data generation, continue to treat labeled or pseudo-labeled data a...Real-world tasks often lack large labeled datasets, motivating extensive work on learning in low-data regimes. Existing approaches such as few-shot prompting, instruction tuning, and synthetic data generation, continue to treat labeled or pseudo-labeled data as the primary learning signal. In contrast, human practitioners acquire expertise through repeated, self-directed interaction with the open web, progressively refining both domain knowledge and search strategies. We propose MEMENTO, a framework that treats the web as a learning signal rather than a stateless retrieval interface. MEMENTO operates at two levels: within each session, it conducts iterative web exploration via an Adaptive Exploration Tree (AET) that decomposes tasks into evolving questions and reflects on intermediate findings; across sessions, it accumulates experience through dual-channel memory, separating declarative knowledge (facts) from procedural knowledge (search strategies). This design enables agents to learn reusable research strategies and domain expertise from trajectories of web interaction without additional model training. We evaluate MEMENTO on three structurally distinct low-data domains: Sales Automation, Legal Outcome Prediction, and Subpopulation Opinion Prediction. Our empirical results show consistent improvement in performance over ReAct baselines (+25.6% on sales, +36.5% on legal research, and 29.6% TVD reduction on opinion prediction), demonstrating that the web can serve as a scalable learning source for acquiring task-specific expertise in data-scarce settings.
|
| 1864 |
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
2606.07874
|
cs.AI
|
Anissa Alloula, Federico Licini, Ava Batchkala, Seraphina Goldfarb-Tarrant |
LLMs-as-judges are the primary way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but crucial properties of LLMs...LLMs-as-judges are the primary way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but crucial properties of LLMs-as-judges: their sensitivity to in-context information, and their steerability to differing safety definitions, which may not align with their internal safety priors. We evaluate the safety judging abilities of 13 generalist LLMs and safety-specific judges, and investigate the impact of novel in-context information and changing safety definitions. We find that while LLM-judges can learn from new information, they are broadly unlikely to update their evaluations if the context or safety definition departs from their prior.
|
| 1865 |
SelfEvoSkill: Paired Execution Audits for Skill Revision without External Outcome Supervision
2606.14239
|
cs.AI
|
Haowen Gao, Haoran Chen, Can Wang, Shasha Guo, Liang Pang |
Agent skills provide LLM agents with reusable procedures, but a given skill may be incomplete, ineffective, or misleading for the current task. Existing methods for revising skills typically rely on external outcome supervision, such as task rewards, performan...Agent skills provide LLM agents with reusable procedures, but a given skill may be incomplete, ineffective, or misleading for the current task. Existing methods for revising skills typically rely on external outcome supervision, such as task rewards, performance on a separate validation set, or feedback from deployment, to decide which updates to retain. However, this supervision may be costly or unavailable when adapting a skill to a new task. Using execution behavior as an alternative signal is challenging: a single run reflects the combined effects of the agent, task, and workspace, making the skill's contribution difficult to isolate. Comparing executions with and without the skill, while holding these factors fixed, provides evidence about how the skill changes behavior. We introduce SelfEvoSkill, which links these behavioral differences to relevant skill passages and iteratively revises and reevaluates those passages through subsequent executions. Checks derived once from the task's explicit requirements prevent revisions from violating task constraints. Across 89 containerized tasks in 8 professional domains, under the same best-of-two evaluation budget, SelfEvoSkill attains a 73.9% mean task reward, exceeding the no-skill condition by +33.0 percentage points and the benchmark's curated skill by +17.2 percentage points, while skill evolution receives neither task rewards nor benchmark verifier outputs. In controlled ablations, using two with-skill runs instead of a matched with/without-skill pair reduces mean reward by 4.6 percentage points, while limiting revision to one round reduces it by 9.0 points relative to repeated revision. Together, these results show that behavioral differences between matched executions can provide an effective signal for revising skills when outcome supervision is unavailable.
|
| 1866 |
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models
2606.18950
|
cs.AI
|
San Kim, Daechul Ahn, Reokyoung Kim, Hyeonbeom Choi, Seungyeon Jwa |
Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating and influencing other agents' actions, under uncertainty in competitive and cooperative settings. Real-time strategy (RTS) games can be a natural testbed for diagn...Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating and influencing other agents' actions, under uncertainty in competitive and cooperative settings. Real-time strategy (RTS) games can be a natural testbed for diagnosing this limitation, as they demand coordination with allies, adaptation to opponents' strategy, and long-horizon planning under partial observability. However, existing RTS benchmarks offer limited evaluation scope, lack systematic competency diagnosis, and remain fixed in the pre-designed scenario coverage. To address these limitations, we present RTSGameBench, which is built on Beyond All Reason, a large-scale RTS game with an expanded battlefield that demands broader strategy diversity than the existing testbeds. The proposed benchmark provides evaluations through diverse gameplay across various matchup structures, diagnostic assessment via mini-games, each targeting an individual strategic competency, and extensible coverage via a self-evolving generation framework that converts free-form queries into new mini-games, improving over successive cycles. Additionally, for VLMs to operate in large-scale RTS games, we provide RTSGameAgent that manages units by an FSM with agentic memory. We empirically validate that multiple state-of-the-art VLMs do not perform well when matchups demand tighter coordination, multiagent coordination and when task scale increases.
|
| 1867 |
Agent MechSuits: Mechanistic Subspace Safety Steering for Multi-Turn CLI Agents
2606.22673
|
cs.AI
|
Weidi Luo, Qiming Zhang, Yihao Quan, Mingyu Jin, Jie Cai |
Command-Line Interface (CLI) agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn interactions with external environments. Existing safety mec...Command-Line Interface (CLI) agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn interactions with external environments. Existing safety mechanisms mainly rely on external guardrails, which have a limited ability to perform fine-grained behavioral control during execution. Meanwhile, recent mechanistic interpretability methods for LLM safety are mostly confined to single-turn or jailbreak-style QA settings, limiting their ability to capture the evolving risk dynamics of multi-turn agent execution. In this paper, we investigate the safety of multi-turn CLI agents from an internal perspective. We propose Agent MechSuits (Mechanistic Subspace Intervention and Steering), a white-box defense framework that performs runtime safety detection and representation-level mitigation for CLI agents. Unlike conventional agent guardrails, Agent MechSuits detects harmful execution states from step-level hidden representations and mitigates unsafe behavior by intervening in a 10-dimensional subspace within a single layer. To support this research, we introduce the Mechanistic Agent Safety (MAS) benchmark, comprising comprehensively annotated multi-turn execution trajectories across 194 tasks using LLaMA-3.1-8B, Qwen-2.5-7B, and Gemma-2-9B. Extensive experiments show that Agent MechSuits achieves strong safety detection performance, provides preliminary evidence for lookahead risk anticipation, and substantially reduces harmful actions of the CLI agent, establishing a foundation for applying mechanistic interpretability to dynamic LLM agent safety.
|
| 1868 |
TIPS: Topological Ill-Posedness Probing and Steering in Large Language Models
2606.23590
|
cs.AI
|
Guangyu Jiang, Sizhe Tang, Mahdi Imani, Nathaniel D. Bastian, Tian Lan |
Ill-posed questions, including those involving ambiguity, under-specification, or conflicting statements, may admit no valid answer or multiple plausible answers, posing a significant challenge for large language models (LLMs) despite their otherwise strong pe...Ill-posed questions, including those involving ambiguity, under-specification, or conflicting statements, may admit no valid answer or multiple plausible answers, posing a significant challenge for large language models (LLMs) despite their otherwise strong performance. Existing work often treats LLM reasoning as a black box and focuses on input-output analysis. Can a compact, unified topological representation of internal model states capture diverse sources of ill-posedness and steer reasoning toward responses appropriate to each source? To this end, we study the internal state from a topological perspective by treating the contextual hidden states of its prompt tokens at a single transformer layer as a point cloud, capturing the input problem's relational structure as encoded by the model. Our analysis shows that zero- and one-dimensional persistent homology---tracking how connected components merge as the distance threshold increases ($H_0$) and the emergence and filling of one-dimensional loops in the filtration ($H_1$), respectively---yield seven compact descriptors of ill-posedness that also serve as control signals for steering model reasoning. Thus, we develop a topology-conditioned steering mechanism. It retrieves topologically similar positive and negative examples for each query to form a local activation-space contrast. The resulting query-specific direction modifies internal states during prefill and decoding, guiding the model's responses toward suitable abstention or clarification tailored to the underlying source of ill-posedness. Empirical evaluations on two datasets and open-weight models spanning four families demonstrate the effectiveness of this compact representation for ill-posedness classification, while topology-conditioning significantly improved steering performance.
|
| 1869 |
MORPH: Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization
2606.26899
|
cs.AI
|
Chenghao Liu, Yu Zhang, Zhongtao Jiang, Kun Xu, Zhenwei An |
Embedding-based retrieval typically returns highest-scoring items, but many production scenarios require items that satisfy a target attribute while preserving a fine-grained pattern expressed by seed examples. We formalize this as pattern-preserving attribute...Embedding-based retrieval typically returns highest-scoring items, but many production scenarios require items that satisfy a target attribute while preserving a fine-grained pattern expressed by seed examples. We formalize this as pattern-preserving attribute retrieval. Standard approaches fail: averaging seeds preserves the pattern but misses the attribute; global attribute retrieval drifts to unrelated patterns. We approach the task with continuous generative retrieval, where a model reads item-embedding sequences and generates query embeddings for nearest-neighbor search. We propose MORPH: Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization, a staged framework with large-scale raw-sequence pretraining, Metric-Ordered Sequence (MOS) training, and final HPPO alignment. MOS construction turns sparse online metric labels into in-pattern trajectories; MOS CPT/SFT then trains the generator through shared multi-domain continuation pretraining and domain-specific tail-centroid supervised fine-tuning. HPPO uses a hybrid pool of static and policy-generated candidate embeddings, labels them with true online intersection metrics, applies iterated preference optimization, and employs a Pareto pair filter to exclude winners that lower pattern purity. Across four large-scale attribute domains under strict item- and pattern-holdout protocols, MOS training improves the primary intersection metric over a strong pretrained generative retriever in every domain-split cell, and the complete Pareto-filtered HPPO procedure improves it further, with paired-bootstrap-significant gains on seven of the eight cells - the exception being the D4 pattern-holdout split. Ablations confirm that the Pareto pair filter improves the attribute-pattern tradeoff on D1-D3, and that hybrid static/policy candidates are complementary.
|
| 1870 |
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction
2607.01764
|
cs.AI
|
Mingzhe Du, Luu Anh Tuan, Tianyi Wu, Renyang Liu, Zhijiang Guo |
Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar that reaches a vulnerable path, construct a proof-of-conceptv(PoC), and verify that the crash disappears on the...Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar that reaches a vulnerable path, construct a proof-of-conceptv(PoC), and verify that the crash disappears on the patched build. Recent LLM agents can often execute these steps when the approach is correct, yet they still fail by choosing the wrong strategy. This paper argues that strategy, rather than the full action trajectory, is the right learning unit for such SE agents: it is compact enough to optimize, concrete enough to guide execution, and stable enough to store and reuse across attempts. We present Mastermind, a dual-loop framework that separates transferable strategy learning from task-specific experience. A trainable planner learns reusable vulnerability-reproduction strategies through SFT and milestone-based GRPO, while an experience loop maintains task-local strategy records that guide subsequent attempts. The planner is trained independently of the executor, allowing strategy learning to improve multiple frozen executors without modifying their action-generation capability. We evaluate Mastermind on CyberGym using 260 training tasks and 200 held-out evaluation tasks. With GPT-5.5 as the frozen executor, Mastermind achieves an 84.5% pass rate, outperforming open-book PoC context (60.0%), Best-of-8 sampling (63.0%), and iterative improvement (77.0%). The same planner also improves GPT-5.4 mini and GLM~5.1 from 45.0% and 58.5% to 60.0% and 71.0%. These results demonstrate that learning high-level strategies is an effective and transferable mechanism for improving repository-scale SE agents.
|
| 1871 |
AgenticPD: A Stage-Aware Agentic Framework for Closed-Loop Physical Design Optimization
2607.04758
|
cs.AI
|
Shuo Ren, Zijin Cheng, Yaohui Han, Libo Shen, Leilei Jin |
Physical design quality-of-results (QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation requires a costly EDA run through the full flow. While existing methods still treat optimization as flat param...Physical design quality-of-results (QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation requires a costly EDA run through the full flow. While existing methods still treat optimization as flat parameter tuning or a LLM-based script generation task, we present AgenticPD, a stage-aware agentic framework for physical design QoR optimization. Instead of re-running the full flow after every trial, AgenticPD is organized around the stage boundaries of the physical design flow, where a Judge Agent navigates the search and stage-specialized agents make local decisions within their own stage using stage-local tools. Additionally, the agent harness in AgenticPD provides structured observations, execution history, and agent context management. Experimental results show that AgenticPD achieves the best timing performance among the evaluated PD tuners while maintaining competitive power and area.
|
| 1872 |
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
2607.06906
|
cs.AI
|
Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson |
Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern;...Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance. We isolate it with a controlled swap: 22 locked evaluation tasks, six foundation models (Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, Palmyra X6), changing only the orchestration layer -- a frozen conventional production loop versus the Writer Agent Harness. Holding models constant, the harness cuts blended cost per task 41% ($0.21->$0.12), median wall-clock 44% (48s->27s), and tokens per task 38% (14.2k->8.8k), with task-completion quality at parity (0.78->0.81, directional at this sample size). Efficiency is model-invariant -- every model gets cheaper (33-61%) -- while quality gains are capability-dependent: a model's gain correlates almost perfectly with its baseline strength (r=0.99, n=6), a phenomenon we term harness leverage. Quality per dollar rises 82%; task-completions per million tokens rise from 54.9 to 92.0. On this workload the orchestration layer moved cost per task more than the full spread of the model menu did. We formalize token economics at the orchestration layer (including effective input price under prompt caching), detail the six mechanism families behind the effect -- cache-shape discipline to failure-spend governance -- compare six widely used agent systems on the same axes, and argue the harness is the one component whose efficiency multiplies across every model an organization runs -- present and future.
|
| 1873 |
Measuring Intelligence Beyond Human Scale
2607.07040
|
cs.AI
|
Jerry Han, Rafael Moschopoulos, Ella Colby, Vishrut Goyal, Andrew Tu |
How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and...How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured. We describe practical protocols that reduce incentives for private-information attacks, support judge-free adjudication, and naturally scale with agent capabilities. We instantiate the framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.
|
| 1874 |
Rethinking the Evaluation of Harness Evolution for Agents
2607.12227
|
cs.AI
|
Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen |
Harness evolution is an iterative search procedure that repeatedly evaluates and revises candidate harnesses used for LLM agents using task feedback. We revisit the evaluation of such automatic harness evolution procedures and identify two fundamental issues i...Harness evolution is an iterative search procedure that repeatedly evaluates and revises candidate harnesses used for LLM agents using task feedback. We revisit the evaluation of such automatic harness evolution procedures and identify two fundamental issues in the protocol. First, prior work does not compare these approaches with simple task-level search baselines under matched feedback and inference budgets. Second, prior work searches for harness configurations using verification signals (e.g., unit test cases) drawn from the same benchmarks on which it reports the final performance of the evolved harnesses, violating the standard separation between training and test data. To address this, we compare automatic harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and evaluate evolved harnesses on held-out tasks to assess generalization. Following prior work, we experiment on Terminal-Bench 2.1 and find that automatic harness evolution fails to outperform simple test-time scaling methods both with and without test cases, and exhibits limited generalization. However, we find that long-horizon games are a promising setting for automatic harness evolution, as they are difficult enough to leave headroom, rely heavily on adaptation to out-of-distribution dynamics, and provide granular feedback by design. In these settings, task-specific harness evolution improves over the search baseline by 80% on ARC-AGI-3 and by 11% on EdgeBench games under matched budgets. Together, these findings highlight the need for matched-budget baselines and held-out evaluation to distinguish genuine harness improvements from benchmark-specific search and overfitting, and point to a more careful characterization of when automatic harness evolution is actually useful. Our code is available at https://github.com/rethinking-harness-evolution.
|
| 1875 |
What Keeps Vision-Language Models Looking at the Image?
2607.12815
|
cs.AI
|
Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki |
When do vision-language models need direct access to the image while generating an answer? We study image dependence during answer generation by examining how the visual information needed for the current question becomes available in context. We intervene on ...When do vision-language models need direct access to the image while generating an answer? We study image dependence during answer generation by examining how the visual information needed for the current question becomes available in context. We intervene on direct image access while retaining previously computed states. Across real-image and synthetic tasks, we show that, depending on the generation process, direct access can continue to support accuracy after question processing. Supplying the required attributes as text in the context weakens this dependence. On synthetic tasks, we also examine how dependence changes as the model itself states the required attributes. Before attribute expression, severing access reduces accuracy, and replacing image-side states shifts answers toward the counterfactual content. After sufficient expression, both interventions have smaller effects. Even with an identical generated prefix, dependence differs according to whether image access was available during question processing. Thus, both the visible text and the preceding image access matter. Several of these patterns hold across model families, including Qwen2.5-VL-32B and InternVL3-14B. These findings offer a view of image dependence in terms of the information needed for the current question and the history of image access, beyond generation position alone. This perspective provides a basis for deciding when to reduce visual access during an answer and which visual information to retain for subsequent questions.
|
| 1876 |
Cura 1T: Healthcare Foundation Model via Recursive Self-Improvement
2607.15314
|
cs.AI
|
Haolin Chen, Leon Qi, Steve Brown, Deon Metelski, Tao Xia |
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized language models that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and imag...Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized language models that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare foundation model trained through recursive self-improvement (RSI). In each RSI round, the RSI harness runs the current model on healthcare benchmarks, evaluates the trajectories to locate capability gaps, and refines the training mixture by synthesizing training data. On 6 healthcare benchmarks, Cura 1T scores highest on MedAgentBench, HealthBench Professional, HealthBench Hard, MedXpertQA text, and AgentClinic, and second on MedXpertQA multimodal. It preserves performances on out-of-domain reasoning and agentic benchmarks including AIME, GPQA-Diamond, and $\tau^2$-Bench.
|
| 1877 |
A Dual-Hypothesis Reasoning Framework for LLM Guardrails
2607.17575
|
cs.AI
|
Md Asiful Islam, Mihai Surdeanu |
We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly considers both safe and unsafe interpretations of a prompt before making a safety decision, ...We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LLM guardrails that explicitly considers both safe and unsafe interpretations of a prompt before making a safety decision, and (ii) multi-component supervised fine-tuning (MC-SFT), a structured training loss for reasoning-based guardrails that decomposes LLM outputs into logical components and weights them according to their importance. Existing reasoning-based guardrails often rely on expensive procedures, such as generating reasoning traces using larger or closed-source teacher models and applying full-parameter fine-tuning. In contrast, ARBITER uses a cost-effective self-generation strategy for reasoning traces and LoRA-based parameter-efficient fine-tuning while still achieving better performance than these expensive approaches. Additionally, ARBITER provides faithful evidence-phrase explanations for unsafe decisions, enabling a more transparent and interpretable guardrail method. Experiments on three safety moderation benchmarks show that ARBITER outperforms existing reasoning-based and non-reasoning guardrail baselines, with clear gains in out-of-domain evaluations.
|
| 1878 |
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
2607.18368
|
cs.AI
|
Taewoon Kim, Vincent Fran\c{c}ois-Lavet, Michael Cochez |
Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic...Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic meta-policy that learns which symbolic memory heuristic to apply at each decision point while keeping execution symbolic. Our setting uses temporal knowledge-graph memory in RoomKG, where hidden state and observations are represented as Resource Description Framework (RDF) graphs and memory is augmented with temporal RDF triple annotations. The model combines knowledge-graph encoding of memory contents with value heads for question answering, exploration, and forgetting, yielding a controller that is both adaptive and inspectable. This gives the work a direct Semantic Web grounding through RDF-based representation, annotation-compatible graph semantics, and graph-based symbolic operations over explicit memory state. On train/test room splits at long-term memory capacity of 512, the qualifier-aware StarE-GNN configuration achieves the best held-out performance among the compared symbolic, neural, and neuro-symbolic systems while preserving step-level traceability of memory-management decisions.
|
| 1879 |
TILT: Model-Intrinsic Reward Alignment For Compositional Diffusion
2607.21606
|
cs.AI
|
Debottam Dutta, Jianchong Chen, Jaehoon Hahm, Romit Roy Choudhury |
Consider conditional generation $p(x \mid C=\{c_1, c_2, \dots c_k\})$ where $C$ is a prompt composed of multiple concepts $c_i$. Diffusion models often struggle with compositional prompts, producing samples in which some concepts dominate while others are miss...Consider conditional generation $p(x \mid C=\{c_1, c_2, \dots c_k\})$ where $C$ is a prompt composed of multiple concepts $c_i$. Diffusion models often struggle with compositional prompts, producing samples in which some concepts dominate while others are missing or weakly represented. Prior work attributes these failures to mode collision, where single-concept modes of $p(x\mid c_i)$ overlap with modes of the joint $p(x \mid C)$. To seek out collision-free modes of $p(x \mid C)$, or "pure modes", corrector-based approaches have attempted to suppress collisions at intermediate diffusion times. However, local corrections are often heuristic and do not necessarily steer the generation to a "pure mode" in the final data space. Derived from a principled formulation, we present TILT (Test-time model-Intrinsic reward aLignment via Tilting), a training-free framework that poses eventual pure mode sampling as a reward for intermediate-time alignment. This reward offers valuable advantages: (1) it is intrinsic to the model, hence external reward models need not be trained by modality-specific datasets, (2) it yields a closed-form target under a variational approximation, which makes it realizable through standard diffusion sampling, and (3) it is interpretable, hence amenable to preference-based modifications. Project page: https://debottam-dutta7.github.io/tilt_web/
|
| 1880 |
Deconstructing Off-Policy Ratios: Entropy-Normalized Trust Regions for Asynchronous Reinforcement Learning
2607.22186
|
cs.AI
|
Guanqun Zhao, Zijun Xie, Binbin Zheng, Jiafeng Lu, Enlei Gong |
Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data destabilizes optimization and can cause policy collapse. Existing...Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data destabilizes optimization and can cause policy collapse. Existing methods gate tokens by ratio magnitude alone, applying one threshold at every position. We show that the ratio's natural scale is set by token entropy, so deviations from mid-trajectory weight updates stay within this scale and carry genuine exploration. We further identify an overlooked low-entropy regime that breaks this scaling, where a near-zero probability amplifies train--inference mismatch into noise far beyond what the local entropy admits. A magnitude threshold admits this noise and discards the exploration. We therefore propose the Entropy-Normalized Trust Region (ENTR). Across long-horizon agentic tasks and mathematical reasoning benchmarks, ENTR outperforms existing asynchronous methods. It improves avg@1 on BrowseComp-Plus by $6.9\%$ over the strongest baseline, trains stably up to $30$ policy versions of staleness, and matches synchronous GRPO at a $2.6\times$ speedup.
|
| 1881 |
Towards Trustworthy Physical Intelligence: From Theory to Practice Across Life Cycle
2607.22877
|
cs.AI
|
Yang Wang, Hongxuan Liu, Xinghui Xu, Arjun Menon, Xiaoran Cai |
Physical intelligence refers to intelligence systems that understand, reason about, and act in accordance with the physical world and its underlying laws, dynamics, and constraints. Unlike conventional AI systems, physical intelligence interacts continuously w...Physical intelligence refers to intelligence systems that understand, reason about, and act in accordance with the physical world and its underlying laws, dynamics, and constraints. Unlike conventional AI systems, physical intelligence interacts continuously with uncertain physical environments, and its actions produce consequences that are physically irreversible. As existing trustworthy AI frameworks have been developed primarily for digital AI systems, they do not fully capture the distinctive challenges of Physical Intelligence, such as physical safety, cyber-physical security, and physical manufacturing process. To address this gap, we present a survey of trustworthy physical intelligence principles. First, we characterize the core capabilities and challenges of physical intelligence. Second, we examine the role of physics in AI. Third, we trace the end-to-end physical intelligence life cycle across five core stages and introduce Trustworthy Physical Intelligence Operationalization (T-PAIO). Fourth, we develop the Trustworthy Physical Intelligence (T-PAI) framework, a theoretical framework that organizes key trustworthiness principles and provides a foundation for governing trustworthy physical intelligence systems.
|
| 1882 |
SKILL-KD: Contrastive Skill Distillation for LLM Agents
2607.28048
|
cs.AI
|
Qiming Shi, Yibo Dou, Jiawen Zhu, Yulong Tao, Linbo Jin |
Skill-based prompting has become a practical mechanism for improving LLM agents, yet existing methods often treat skills as summaries of the agent's own experience or of successful demonstrations. This creates a mismatch for weaker student agents. A failed tra...Skill-based prompting has become a practical mechanism for improving LLM agents, yet existing methods often treat skills as summaries of the agent's own experience or of successful demonstrations. This creates a mismatch for weaker student agents. A failed trajectory may not reveal the missing knowledge or strategy, while a teacher trajectory may be too implicit to internalize. We propose SKILL-KD, a contrastive skill distillation framework that treats skills as an explicit distillation medium between agents of different capabilities. Given a student failure and the teacher trajectory on the same task, SKILL-KD distills their actionable discrepancy into a textual skill patch, evaluates it by re-running the student, and iteratively refines it when the student still fails. To prevent skill drift from repeated local updates, SKILL-KD maintains trace-linked edit histories and performs Drift-Aware Skill Consolidation to decide whether each patch is added, merged, or skipped. Across five agent benchmarks and two student settings, SKILL-KD consistently improves frozen student agents over fixed-model adaptation baselines.
|
| 1883 |
RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
2608.05714
|
cs.AI
|
Shuhao Yan, Changhao He, Peng Hu, Xi Peng |
Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, externally supplied,...Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, externally supplied, prompt-induced, or separately optimized critique mechanisms to optimize the generation process, but they do not necessarily optimize how feedback is interpreted and translated into effective corrective actions throughout the generation process. To bridge this feedback-utilization gap, we present RA-CAD (ReAct Agent for CAD), a state-aware agent that interacts with the CAD environment through a Generate--Execute--Critique--Rewrite loop. At each iteration, RA-CAD executes the current code and observes its outcome. Conditioned on the design instruction, current code, and execution feedback, the agent then generates an explicit post-execution critique as an intermediate policy action. This critique either validates the current result for termination or provides revision-oriented guidance that conditions the next rewrite. CAD Code Bootstrapping (CCB) first establishes fundamental parametric CAD coding capabilities through supervised fine-tuning. Feedback-Driven Agent Optimization (FAO) subsequently applies trajectory-level Group Relative Policy Optimization to both policy-generated code and critique sequences, assigning terminal F1 and Chamfer Distance rewards to the complete interaction trajectory. This formulation makes critique an outcome-aligned, learnable policy decision rather than an unoptimized auxiliary output. Experiments on CADFusion and Text2CAD show that RA-CAD achieves state-of-the-art execution validity and geometric quality compared with existing methods and strong proprietary language models, demonstrating the effectiveness of the proposed state-aware text-to-CAD agent.
|
| 1884 |
NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
2608.06621
|
cs.AI
|
Bin Wang, Yan Zhong, Liang Luo, Buyun Zhang, Ellie Wen |
We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough d...We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough dataset is available. The method is especially suited to LLM-based exploration systems, where LLMs are remarkably good at proposing hypotheses but suffer badly from hallucination; this hallucination prevents us from harvesting LLM outputs directly, and existing remedies each fall short. The most common solutions include letting the LLM verify or correct itself circular verification and held-out testing (where false hypotheses can still pass via spurious correlations), among other remedies detailed in the introduction. To resolve this, NxN E-valuation exploits the naturally existing large training set and lets different samples serve as null hypotheses for one another. This design directly realizes a conditional randomization test (CRT) that certifies each hypothesis. The approach can be a universally better replacement for at least LLM circular verification and held-out-data testing, provided the LLM's generations are hypotheses that apply to each individual sample.
|
| 1885 |
The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach
2608.16438
|
cs.AI
|
Rafael Pass |
In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only what the LLM can produce, but what \emph{value} remains in the inputs (i.e., the prompts) we provide to it. Given a prompt,...In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only what the LLM can produce, but what \emph{value} remains in the inputs (i.e., the prompts) we provide to it. Given a prompt, hint, critique, problem statement, or partial solution that helps an LLM produce an artifact $z$---a proof, program, design, or scientific hypothesis---how should we measure the value of that input? Intuitively, an input is valuable when it makes the target artifact easier for the model to generate: either by increasing its sampling probability, or by reducing the thinking time needed to find it. We propose a computational Levin--Kolmogorov complexity approach to this problem, by appropriately replacing the universal Turing machine in the classical definitions by the LLM itself. Concretely, we introduce an LLM-relative notion of \emph{probabilistic Levin--Kolmogorov complexity} $pKt$---treating the model's thinking as the random tape of the program, and charging logarithmically for it in Levin's manner---and define prompt value as algorithmic mutual information with respect to $pKt$. This captures the intuition above: a prompt having $b$ bits of value for an artifact $z$ makes $z$ $2^b$ times ``easier to obtain'', by multiplying the success probability by $2^b$, by dividing the required computation by $2^b$, or by any corresponding tradeoff between probability and computation. In contrast to the classical notion of algorithmic mutual information, ours is efficiently estimable. We additionally show that, under a natural reproduction experiment, a prompt value of \(b\) bits means that reproducing \(z\) without the prompt has median token cost \(2^b\) times that of reproducing it with the prompt.
|
| 1886 |
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
2608.18613
|
cs.AI
|
Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang |
Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly, but the corpus side has not: threat reports and vulnerabi...Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly, but the corpus side has not: threat reports and vulnerability databases are still packaged for retrieval-augmented generation, as opaque chunks behind an embedding index. We argue that this substrate, not model capability, is the bottleneck on agentic CTI investigation, and present CTIFoundry, an agent-native corpus scaffold. At build time, CTIFoundry materializes the latent structure of a CTI corpus: a deterministic ontology graph over four authoritative knowledge bases (CVE, CWE, CAPEC, ATT&CK) whose official cross-references become typed, traversable edges; a span-grounded report layer whose canonical, alias-resolved cross-vendor entities index provenance-carrying chunks; and hybrid dense+lexical retrieval surfaces. At query time this structure is exposed through seven typed tools and three procedural skills mounted on a stock, widely-used open-source agent harness. On the public CTIConnect benchmark, swapping only the action surface lifts the identically-harnessed agent from 0.610 to 0.829 overall F1 with gpt-5.4 and from 0.470 to 0.745 with claude-haiku-4-5: a small model on CTIFoundry surpasses a flagship model on the flat substrate. The scaffolded agent is simultaneously more accurate and more efficient: on both Claude models it answers with roughly half the tool calls per question. The ablation distills design principles for matching corpus scaffolding to data modality, in CTI and beyond. Build-time validation guarantees zero fabricated identifiers by construction, and the scaffold sustains 1,168 investigations end-to-end at about 2.6 cents each.
|
| 1887 |
SEPO: Evidence-Grounded Prompt Optimization via Structural Editing
2608.28067
|
cs.AI
|
Xiaoyu Ma, Haoyue Liu, Yiwen Li, Jionghao Zhu, Zhichao Wang |
Existing API-only prompt optimisers are often described as interpretable, but in practice, this usually means only post-hoc inspectability: each iteration still rewrites the prompt as one opaque string, leaving a trace of full-prompt diffs rather than localisa...Existing API-only prompt optimisers are often described as interpretable, but in practice, this usually means only post-hoc inspectability: each iteration still rewrites the prompt as one opaque string, leaving a trace of full-prompt diffs rather than localisable, machine-readable edits. This paper introduces SEPO (Structural, Evidence-grounded Prompt Optimization), a multi-trajectory prompt optimiser centred on edit-effect lineage feedback. Rather than treating each iteration as an isolated whole-prompt rewrite, SEPO locally edits stable, typed units in a two-layer prompt schema, links the target and realised structural operations of each edit to the examples it newly fixes or breaks, and carries this edit-effect record forward to guide later architect calls on the same search branch. This makes prompt optimisation addressable, attributable, and actionable. Across a 14-task held-out suite, SEPO improves over the strongest baseline, GEPA, by 3.1 pp on Llama-3.1-8B-Instruct and 2.2 pp on Qwen3-8B, reaching 61.9% and 73.3% macro accuracy. SEPO also lies on both the optimisation-time and test-time Pareto frontiers, spending 2.9M optimisation tokens versus 4.1M for GEPA and producing prompts over 5x shorter.
|
| 1888 |
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
2609.00646
|
cs.AI
|
Haoyuan Shi (Hunyuan, Tencent), Mingtao Chen (Hunyuan, Tencent), Shuo Jiang (Hunyuan |
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real u...Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.
|
| 1889 |
A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
2609.09589
|
cs.AI
|
Yizhou Zhang, Weichen Wu, Lun Du, Zhengjie Miao |
In the kernel regime, neural-network learning inherits its preferences from a frozen spectrum. During feature learning, this spectrum evolves, yet networks retain systematic biases toward simple, smooth directions. We develop a function-space statistical frame...In the kernel regime, neural-network learning inherits its preferences from a frozen spectrum. During feature learning, this spectrum evolves, yet networks retain systematic biases toward simple, smooth directions. We develop a function-space statistical framework explaining the origin of these preferences, treating functions and their learning operators as macroscopic variables, with parameterization entering through the multiplicity of parameter configurations realizing each function. For mean-squared loss, error relaxes exactly under the evolving learning operator $M=JJ^\ast$. Training stochasticity induces a Gaussian weight over function-space states, while parameter multiplicity contributes an entropic operator $B$, defined by the curvature of its log multiplicity. A local Laplace expansion yields the fluctuation free energy $\Phi_{\mathrm{fluc}}(M;B)=\frac{\sigma_\xi^2}{2}\log\det(M^{-1}+B)+\mathrm{const}$, analogous to an Occam factor. Under mild statistical conditions, this free energy is rotationally stationary exactly when $[M,B]=0$, is minimized by pairing large eigenvalues of $M$ with small eigenvalues of $B$, and generates a local restoring force against mismatch. Learning is therefore biased toward faster relaxation along entropically cheaper directions. This preference strengthens with training noise and vanishes in the deterministic limit, beyond gradient-flow accounts of operator alignment. For ReLU networks, we relate entropic curvature to the minimal rearrangement of activation boundaries required for a functional change and bound this structural cost by directional smoothness. Consequently, smooth directions are preferentially learned faster, in a data-adaptive manner, even as the learning operator evolves.
|
| 1890 |
Reading the Whole Heart: Latent-Attention Masked Autoencoders for Multimodal Cardiac Representation Learning
2609.12035
|
cs.AI
|
Andrea Agostini, Simon B\"ohi, Moritz Vandenhirtz, Samuel Ruiperez-Campillo, Max Kr\"ahenmann |
Cardiovascular diagnosis and treatment rest on integrating complementary modalities, such as electrocardiogram, echocardiography, and chest X-rays, each capturing distinct but complementary aspects of cardiac pathophysiology. Yet most medical foundation models...Cardiovascular diagnosis and treatment rest on integrating complementary modalities, such as electrocardiogram, echocardiography, and chest X-rays, each capturing distinct but complementary aspects of cardiac pathophysiology. Yet most medical foundation models remain modality-specific, combining modalities only for finetuning or post-training. This discards the cross-modal evidence clinicians naturally integrate and ignores the structure within each modality. We introduce Latent-Attention Masked Autoencoder (LAMAE), a multimodal, structure-aware masked autoencoder that jointly learns patient-level representations during self-supervised pretraining. Instead of fusing modalities post hoc, LAMAE exchanges information directly in the latent space through a shared latent-attention module operating over a study-view-entity hierarchy, enabling aggregation of variable observations and handling of missing modalities. Pretrained on over 500'000 MIMIC-IV hospital stays, LAMAE outperforms modality-specific pretraining and strong contrastive and vision-language baselines across multimodal hospital-stay tasks, such as in-hospital mortality, ICD-10 and DRG coding, and length of stay, while remaining competitive on unimodal tasks. Modeling this structure also pays off within a single modality: even without cross-modal information, the latent-attention module improves representations over modality-specific pretraining.
|
| 1891 |
Schizophrenia Detection from EEG Signals: A Transformer Framework with Spectrogram Representation
2609.14015
|
cs.AI
|
Abtin Shafiei, Mohsen Hooshmand, Majid Ramezani |
Schizophrenia is a serious psychiatric disorder that affects millions of people worldwide, and its diagnosis remains primarily dependent on clinical assessment. Electroencephalography (EEG) provides a non-invasive approach to investigate brain activity and has...Schizophrenia is a serious psychiatric disorder that affects millions of people worldwide, and its diagnosis remains primarily dependent on clinical assessment. Electroencephalography (EEG) provides a non-invasive approach to investigate brain activity and has shown potential to support automated Schizophrenia detection. However, existing EEG-based classification studies often suffer from limitations including small datasets, inconsistent preprocessing strategies, and evaluation protocols that may not adequately prevent subject-related data leakage. In this study, we propose an EEG-based Schizophrenia classification framework that transforms preprocessed EEG recordings into time-frequency representations using the Short-Time Fourier Transform. The generated spectrogram images are classified using both conventional Machine Learning algorithms, including Support Vector Machines, Random Forests, and XGBoost, and Deep Learning models, including convolutional architectures and CNN-Transformer hybrids. To ensure reliable evaluation, all data partitions are performed at the subject level, and image-level predictions are aggregated into subject-level decisions. On the independent test set of 18 subjects, the CNN-Transformer (CT-SZ) model achieves a subject-level AUC-ROC of 95.00%, while the CNN + Squeeze-and-Excitation + Transformer (CST-SZ) model achieves 92.50%.
|
| 1892 |
EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise
2609.21841
|
cs.AI
|
Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid |
Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large frac...Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment decision requires "is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation
|
| 1893 |
When No One Owns the Judgment: Accountability Under Contribution Dissolution in Human-AI Collaboration
2609.29312
|
cs.AI
|
Hengzhi Ye |
Communities often respond to potentially AI-assisted work by asking three questions: Was AI used? Was that use disclosed? Can hidden use be detected? These questions place AI use itself at the center of accountability while overlooking a deeper problem: unowne...Communities often respond to potentially AI-assisted work by asking three questions: Was AI used? Was that use disclosed? Can hidden use be detected? These questions place AI use itself at the center of accountability while overlooking a deeper problem: unowned judgment. Evaluations, claims, decisions, and creative directions can be shaped by AI with no accountable human or institution prepared to stand behind them. We develop this argument through two illustrative cases: AI-assisted peer review and concealed AI use in creative work. The first shows how contribution dissolution can weaken responsibility while the second shows how the fear of losing credit can discourage honest disclosure. The cases expose the limits of disclosure rules and provenance records as responses to AI-mediated collaboration. We offer three directions for discussion: distinguishing the roles AI plays, identifying judgments that require clear human ownership, and creating conditions in which AI involvement can be disclosed without default penalty. The broader aim is to make AI-shaped contributions discussable, creditable, contestable, and repairable.
|
| 1894 |
SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories
2609.29465
|
cs.AI
|
Jiajun Wu, Leixin Sun, Zihan Tan, Yitao Liu, Shuo Li |
Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from a human-identified issue and evaluate whether a patch satisfies a functional sig...Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from a human-identified issue and evaluate whether a patch satisfies a functional signal. We present SWE-Prometheus, a benchmark for the broader task of improving repository engineering governance. Each task provides a fixed snapshot and an open-ended objective, requiring the agent to identify risks, prioritize interventions, and verify the resulting changes. SWE-Prometheus evaluates six governance dimensions through paired evidence, clean-environment probes, behavior gates, and two independent teacher ratings of the same evidence. The benchmark contains 60 repositories; ten models are evaluated on a shared 22-repository public subset, where mean Normalized Governance Improvement ranges from 0.0568 to 0.5760 and observed behavior-breakage rates range from 0% to 23%. On a frozen ten-repository batch, a repository-blind template obtains mean NGI 0.272, but its gains concentrate in Tests & CI, Quality Gates, and Documentation; it improves Reproducible Environment and Dependency & Security on none of the repositories. This baseline makes the distinction between adding governance artifacts and producing execution-backed improvements measurable. The no-op condition has median NGI zero and standard deviation 0.073; two teachers agree exactly on 57 of 60 dimension scores for the same no-op evidence. For the two highest conditional-mean systems, common-valid NGI is similar, while full-pool comparisons that include behavior failures favor Kimi-K3. These results show why repository-governance evaluation should report improvement, behavior preservation, evidence quality, and coverage together.
|
| 1895 |
AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution
2609.31133
|
cs.AI
|
Tian Luo, Ruge Zhang, Haozhi Han, Yifeng Chen, Yunquan Zhang |
High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts,...High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration problem. To address this, we introduce AtomWorld-Mem, a memory-restored atomistic world model that recovers the latent world state missing from instantaneous crystal snapshots. AtomWorld-Mem treats the evolving alloy as an AtomWorld: spatial encoders write multi-scale atomistic keyframes from dense local topology and sparse long-range defect context, while short-term event memory and long-term structural memory integrate these keyframes across time to restore a future-predictive evolutionary state. The restored state is used to prioritize legal vacancy-mediated events under single-event Kinetic Monte Carlo (KMC) constraints, while event legality, physical execution, and residence-time updates remain governed by the underlying simulator. Empirically, AtomWorld-Mem improves long-horizon atomistic progress under fixed microscopic event budgets while maintaining high-fidelity evolution across energetic, structural, and vacancy-transport observables. It further transfers zero-shot across diverse unseen alloy-temperature AtomWorlds, suggesting that the learned memory-restoration mechanism captures reusable principles of hidden-state inference rather than a system-specific local energy heuristic. These results position memory-restored world-state modeling as a promising route toward efficient, physically grounded, and transferable atomistic evolution.
|
| 1896 |
Dude, Where's My State? Execution Information Requirements for Stateful Agents
2609.32687
|
cs.AI
|
Nikita Mehrotra, Ashish Tiwari, Priyanshu Gupta, Sumit Gulwani |
Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. ...Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that generates tasks with known dependencies and varies information demand, retention, and recovery separately from the difficulty of individual operations. Across four models, restoring a missing result raises accuracy on affected recall steps to 100%, compared with 0% for equal-length irrelevant information. Sufficient storage alone does not ensure success: retention policies can discard required results, errors can propagate through later computations, and agents can stop before recovery is complete. We also introduce VESTIGE, which uses agent execution traces to construct semantic graphs and measure information demand for real tasks. Across 72,562 software-agent trajectories, VESTIGE reveals a steeper distance-related decline in solution-relevant rereading for failed runs (RR 0.951 per distance doubling), while adjusted peak demand alone is not associated with failure. Together, these contributions support evaluating whether agents preserve and recover the information their tasks require.
|
| 1897 |
FedMHAR: Federated Multimodal Human Activity Recognition using Multi-Agent Reinforcement Learning
2609.33492
|
cs.AI
|
Debasmita Dey, Tanmay Sen, Himel Mallick |
Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to s...Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences in modality importance, acquisition cost, and sensor quality, which can vary due to movement, incorrect placement, or temporary blockage. We propose an adaptive and cost-aware multimodal HAR framework based on multi-agent reinforcement learning for centralized HAR and extend it to federated learning as FedMHAR. In the centralized setting, multimodal fusion is formulated as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, where each sensing modality is assigned a PPO-based agent that learns per-sample fusion weights, enabling the model to emphasize informative modalities while down-weighting costly sensors when cheaper alternatives provide sufficient information. In the federated setting, we introduce BiFL-PPO, a bidirectional federated optimization strategy in which a server-side PPO policy learns client-specific trust weights and feeds them back to adapt local learning rates and proximal regularization. Unlike round-level optimization, BiFL-PPO uses dense batch-level rewards for more frequent feedback and stable training under heterogeneous client data. Evaluation on the MEx Rehabilitation and UTD Multimodal Human Action datasets shows that the centralized framework achieves 87.30% and 94.98% accuracy, respectively, outperforming conventional fusion methods and state-of-the-art HAR models. FedMHAR achieves 79.74% and 77.49% in the federated setting, consistently surpassing FedAvg, FedProx, FedBN, FedNova, and AdaFedProx, while providing more stable performance and reducing sensor acquisition cost.
|
| 1898 |
Robust Biomolecular Complex Design Across Protein Conformational Landscapes
2609.33726
|
cs.AI
|
Qingyuan Zeng, Zongqi Xu, Anglin Liu, Ziqi Gong, Pengxiang Cai |
Proteins populate conformational ensembles, yet structure-based biomolecular design typically optimizes candidates against a single target conformation. Consequently, a candidate that fits one state can lose favorable interactions or develop steric clashes whe...Proteins populate conformational ensembles, yet structure-based biomolecular design typically optimizes candidates against a single target conformation. Consequently, a candidate that fits one state can lose favorable interactions or develop steric clashes when the target adopts another. We introduce FlexEvo, a model-agnostic evolutionary framework that adapts candidates once at inference time from a single target conformation to improve compatibility with alternative natural conformations unseen during adaptation, without retraining the source model or requiring a conformational ensemble. FlexEvo casts cross-state adaptation as geometry-constrained bi-objective optimization, balancing preservation of input-state interactions against robustness to plausible conformational perturbations. To limit the search space and reduce invalid structural edits, geometry-derived FlexBoxes define protected anchor regions, adaptable regions for local exploration, and forbidden regions for clash avoidance. A unified all-atom representation supports topology-preserving adaptation across diverse binder categories, while Pareto selection preserves nondominated candidates across the two objectives. We evaluate FlexEvo across multiple generation baselines and nine representative binder categories spanning diverse molecular sizes and structural topologies. FlexEvo reduces the category-balanced mean relative performance degradation from 47.8% to 4.4%, while adding only 1.4--3.1 minutes of adaptation per sample. These results establish single-state inference-time adaptation as a practical route toward robust biomolecular complex design across protein conformational landscapes.
|
| 1899 |
PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
2609.35671
|
cs.AI
|
Yangqin Jiang, Lingrui Xu, Chao Huang |
Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation i...Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app's GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training. Offline, PhoneCLI explores a target app from the outside and distills its screens, interactive elements, and navigation edges into a semantically annotated map; each screen yields one deterministic command: a replay sequence that reaches it. Online, the agent selects a command, verifies it before execution, and then executes it deterministically in sub-second time at zero VLM cost; open-ended interaction and every failure of the compiled path fall back to the embedded VLM interpreter, exactly the pure VLM agent, so compilation can only help. On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld's official M3A agent with consistent efficiency gains. What PhoneCLI compiles is the app's navigation rather than one run, so it serves new tasks, not only repeated ones.
|
| 1900 |
Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
2609.36932
|
cs.AI
|
Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen |
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollo...Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07$\times$ training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64\% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.
|
| 1901 |
ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents
2609.37311
|
cs.AI
|
Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan |
Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, ...Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16\% across three recommendation agent tasks, namely searching, ranking, and judging.
|
| 1902 |
EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
2609.37658
|
cs.AI
|
Min Yang, Yichen Pan, Jinghua Piao, Dandan Song, Yongshun Gong |
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as ...LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.
|
| 1903 |
Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
2609.37787
|
cs.AI
|
Ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou |
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-s...Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether Adam converges on generalized smooth objectives under only second moment information on the stochastic gradients, without such concentration assumptions, was identified as an important open direction by Li et al. (2023). This paper gives an affirmative answer under fairly general conditions: such tail assumptions are not necessary. Building on the Adam self-normalization framework of Jin et al. (2026), developed for classical smoothness and bounded variance, we extend the stopping-time and de-preconditioning strategy to the $L_0$-$L_p$ generalized smoothness condition and a generalized second moment ABC condition. Even when the stochastic-gradient condition provides only second moment information that may grow along the trajectory, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region, with stretched-exponential tail decay under bounded variance and global smoothness. Consequently, we establish high-probability convergence rate guarantees over the full range $p<2$, with confidence dependence of order $\delta^{-1/2}$, while the stepsize prefactor depends on $\delta$ only through a single logarithmic factor. We further construct a hard instance showing that, under only second-moment information, this $\delta^{-1/2}$-type confidence dependence is sharp. Finally, in the regime $p<1$, we combine the trajectory control with polynomial-growth estimates on rare events to obtain convergence rate guarantees in expectation.
|
| 1904 |
ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning
2609.39665
|
cs.AI
|
Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari |
Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spa...Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.
|
| 1905 |
PRISM: A Category-Theoretic Framework for Measuring and Refining Multimodal Analogies
2610.01383
|
cs.AI
|
Mirella Zeisler, Ojas Shirekar, Mircea Lic\v{a}, Chirag Raman |
Analogical reasoning involves identifying and preserving relational structures across domains. However, existing approaches to AI-driven multimodal analogy generation lack an interpretable measure of whether this structure is understood and maintained in the g...Analogical reasoning involves identifying and preserving relational structures across domains. However, existing approaches to AI-driven multimodal analogy generation lack an interpretable measure of whether this structure is understood and maintained in the generated output. We address this gap with Pullback Refinement via Interpretable Structural Mapping (PRISM), a modality-agnostic framework for measuring and improving relational alignment in multimodal analogies, evaluated on visual metaphor generation. PRISM represents analogies as explicit relational mappings grounded in category theory and uses VLMs to instantiate these structures across modalities. Its first component, the pullback score, quantifies relational alignment from the resulting graph representation. On the AnaloBench benchmark, selecting the correct analogy purely by pullback score achieves 82.5% accuracy, demonstrating that the score captures meaningful relational information. PRISM's second component is an iterative refinement loop that uses the pullback score as an in-context feedback signal to iteratively revise the generated image towards greater relational depth. VLM-as-a-judge and human evaluations show that PRISM consistently improves metaphor consistency and analogy appropriateness over zero- shot generation, with human participants preferring the refined output in 57.65% of pairwise comparisons. However, a qualitative analysis reveals that refinement can favour visually crowded compositions rather than genuinely deeper relational correspondences.
|
| 1906 |
Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
2610.01458
|
cs.AI
|
Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li |
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains un...Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
|
| 1907 |
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
2610.01780
|
cs.AI
|
Arman Behnam, Sunglyoung Kim, Liangwei Yang |
An AI companion that talks with someone for months should come to understand them. It should know who they are, remember what they said, and recognize when something from the past matters now. Testing this needs real conversations, but real conversations are p...An AI companion that talks with someone for months should come to understand them. It should know who they are, remember what they said, and recognize when something from the past matters now. Testing this needs real conversations, but real conversations are private, so existing benchmarks use invented people and invented questions. We release RealCompanion, ten real relationships between people and an AI companion, with 27,218 messages over up to 120 days. For each person, we release the full conversation, a profile, a persona, chat test items, and question test items. Every label points to the messages that support it, and every chat label comes with the reasoning that produced it. The real data shows three things. First, people rarely refer back. Only 3.4% of their messages depend on something said earlier, and when one does, the earlier message is usually far away (a median of 2,157 messages back). Averages hide this. Looking at the most recent messages finds the needed one 95.9% of the time overall, but only 2.2% of the time when it is far back. Second, AI systems cannot tell when the past matters. The detectors we tested barely beat chance on real messages, and when the same earlier messages are labeled "memories" instead of "earlier messages", models bring up the past 10 to 14 percentage points more often, even when nothing from the past is needed. Third, AI systems read more into a person than the person revealed. Three agent systems rebuild each persona equally well (F1 0.71). They see the person, and then imagine more. Understanding a person depends on knowing when their past matters and where what they shared ends, and only real conversations can test it.
|
| 1908 |
Mitigating Social Sycophancy via Pluralistic Preference Optimization
2610.02568
|
cs.AI
|
Stephane Hatgis-Kessell, Myra Cheng, Xiaoxuan Hou, Qian Hu, Rahul Gupta |
Personal advice, including relationship advice, now ranks among the most common uses of generative AI. But language models (LMs) exhibit sycophancy: they affirm users much more often than humans do, which can make people overconfident and less willing to repai...Personal advice, including relationship advice, now ranks among the most common uses of generative AI. But language models (LMs) exhibit sycophancy: they affirm users much more often than humans do, which can make people overconfident and less willing to repair their relationships after a conflict. Prior work on mitigating sycophancy has focused on factual settings where a response can be checked against a ground truth answer, while mitigations for social sycophancy (e.g., personal advice, where there is no ground truth) have relied on simple prompting and post-training methods with limited effectiveness. Our insight is that social sycophancy occurs in part because LMs overly center on the user and fail to consider the perspectives of other stakeholders impacted by the user's behavior. To address this problem we propose Pluralistic Preference Optimization (PlurPO): given inputs describing interpersonal conflicts, the LM identifies and simulates the relevant stakeholders, and is then trained to prefer and generate responses acceptable to all stakeholders. PlurPO uses only signals the model produces about its own outputs, without ground-truth labels. PlurPO substantially reduces social sycophancy across four datasets and four model families compared to prior methods. For example, on statements of intent to cause harm, where the users' actions should not be endorsed, PlurPO reduces the endorsement rate by 89% on average across four models. On general advice questions, where the target is to match the endorsement rate of human responses, it closes the gap by more than half, from 17.8% to 8.0% on average. The preference dataset constructed by PlurPO for an 8B model also effectively transfers to mitigating sycophancy in a larger (32B) model. Our results indicate that social sycophancy can be reduced by leveraging a model's own capabilities to simulate a plurality of relevant perspectives.
|
| 1909 |
Geometry Meets Physics: Data-Efficient Pre-Training for Unstructured Neural PDE Solvers
2610.03363
|
cs.AI
|
Luis Medrano-Navarro, Giacomo Baldan, Qiang Liu, Benjamin Holzschuh, Jan Hagnberger |
Neural surrogate models for Partial Differential Equations (PDEs) on unstructured 3D geometries are often limited by poor generalization and the high cost of generating large-scale training datasets. Consequently, pre-training on massive datasets of related PD...Neural surrogate models for Partial Differential Equations (PDEs) on unstructured 3D geometries are often limited by poor generalization and the high cost of generating large-scale training datasets. Consequently, pre-training on massive datasets of related PDE dynamics has emerged as a critical alternative to enhance the robustness and scalability of these models. However, this strategy is neither compute- nor data-efficient, as it relies on massive pre-computed data that is very costly to generate. In this work, we introduce a disk-data-free pre-training framework tailored to both steady-state and transient regimes. For steady-state problems, we propose a geometry-driven strategy that leverages intrinsic shape descriptors to learn representations of complex 3D domains. For transient problems, we introduce a physics-driven approach based on online generation of synthetic PDE data, enabling scalable pre-training without reliance on expensive datasets. Across multiple experiments, our approach achieves faster convergence, greater data efficiency, and higher accuracy during fine-tuning, particularly under realistic low-data regimes. This methodology provides a practical pathway toward data-efficient neural emulators for large-scale simulations.
|
| 1910 |
PPFedIT: Towards Privacy-Preserving Federated Instruction Tuning with Few-shot Local Examples
2403.06131
|
cs.AI
|
Zhuo Zhang, Jingyuan Zhang, Jintao Huang, Hui Wang, Yue Yu |
Instruction tuning aligns large language models (LLMs) with human intentions but requires diverse, high-quality data that are difficult to collect in privacy-sensitive domains. Federated instruction tuning (FedIT) enables collaborative training across data own...Instruction tuning aligns large language models (LLMs) with human intentions but requires diverse, high-quality data that are difficult to collect in privacy-sensitive domains. Federated instruction tuning (FedIT) enables collaborative training across data owners, yet existing methods typically assume sufficient local data. In realistic few-shot settings, limited samples can cause overfitting, degrade performance, and increase vulnerability to training data extraction attacks. We propose PPFedIT, a federated algorithm that improves both model performance and privacy protection in federated few-shot learning. It comprises three client-side steps: (1) synthetic data generation, which uses LLMs to diversify and enrich local data; (2) parameter isolation training, which updates the shared global LLM on synthetic data and local LLMs on private local data to mitigate synthetic-data noise; and (3) local aggregation then sharing, which mixes global and local model parameters before uploading them for server aggregation to mitigate data extraction attacks. Experiments on three open-source datasets show that PPFedIT improves model performance by an average of 8.4% and reduces the risk of data extraction attacks by approximately 20% in challenging federated few-shot settings.
|
| 1911 |
Screw Geometry Meets Bandits: Incremental Acquisition of Demonstrations to Generate Manipulation Plans
2410.18275
|
cs.AI
|
Dibyendu Das, Aditya Patankar, Nilanjan Chakraborty, C. R. Ramakrishnan, I. V. Ramakrishnan |
In this paper, we study the problem of methodically obtaining a sufficient set of kinesthetic demonstrations, one at a time, such that a robot can be confident of its ability to perform a complex manipulation task in a given region of its workspace. Although p...In this paper, we study the problem of methodically obtaining a sufficient set of kinesthetic demonstrations, one at a time, such that a robot can be confident of its ability to perform a complex manipulation task in a given region of its workspace. Although programming by demonstration has been an active area of research, the problems of checking whether a set of demonstrations is sufficient and systematically seeking additional demonstrations have remained open. We present an approach for the robot to incrementally and actively ask for new demonstration examples, one at a time, until the robot can assess with high confidence that it can perform the task successfully. Our approach uses: (i) a screw geometric representation of motion to generate manipulation plans from demonstrations, which makes the sufficiency of a set of demonstrations measurable; (ii) a sampling strategy based on PAC-learning from multi-armed bandit optimization to evaluate the robot's ability to generate manipulation plans in a subregion of its task space; and (iii) a heuristic to seek additional demonstration from areas of weakness. We present results of a user study conducted with 22 participants (without any background in robotics) on two example manipulation tasks, namely pouring and scooping, to assess the utility and usability of our approach. The results show that a handful of examples (fewer than 10) were needed to successfully teach the robot to plan tasks. A short video supplement is available on YouTube: https://youtu.be/KbAPgIouIvo
|
| 1912 |
Robotic Long-Horizon Manipulation with Progressive In-Context Code Generation and Episodic Feedback
2503.21969
|
cs.AI
|
Yuan Meng, Xiangtong Yao, Haihui Ye, Yirui Zhou, Shengqiang Zhang |
Robotic long-horizon manipulation requires robots to compose perception, reasoning, and action over extended task sequences, yet existing language-conditioned frameworks often rely on dense step-wise feedback, learned action policies, or unstructured prompting...Robotic long-horizon manipulation requires robots to compose perception, reasoning, and action over extended task sequences, yet existing language-conditioned frameworks often rely on dense step-wise feedback, learned action policies, or unstructured prompting, which limits robustness and generalization. We propose DAHLIA, a data-agnostic code-generation framework that treats long-horizon manipulation as episodic task planning and evaluation. DAHLIA uses progressively organized in-context examples with chain-of-thought reasoning to guide an LLM in synthesizing executable code plans from language instructions. To improve robustness under partial observability, a VLM-based reporter evaluates task outcomes only after each execution episode and provides structured feedback for replanning, avoiding the overhead of per-step inference. Experiments on LoHoRavens, CALVIN, Franka Kitchen, and real-world manipulation tasks show that DAHLIA achieves strong performance on over 30 long-horizon tasks and improves generalization to unseen scenarios, including cluttered scenes and occluded objects.
|
| 1913 |
Robotic Long-Horizon Manipulation with Bayesian Non-parametric Skill Priors
2503.21975
|
cs.AI
|
Yuan Meng, Xiangtong Yao, Yansong Wu, Liding Zhang, Yixiao Nie |
Long-horizon manipulation under sparse rewards remains challenging for reinforcement learning due to delayed feedback and inefficient exploration. Existing skill-based approaches often assume a fixed parametric prior (e.g., a single Gaussian), limiting their a...Long-horizon manipulation under sparse rewards remains challenging for reinforcement learning due to delayed feedback and inefficient exploration. Existing skill-based approaches often assume a fixed parametric prior (e.g., a single Gaussian), limiting their ability to capture diverse and multi-modal skill structures required for complex tasks. We propose a Bayesian non-parametric skill prior that models temporally extended skills in a structured latent space using a Dirichlet Process Mixture, enabling adaptive skill discovery without predefining the number of components. Integrated into a hierarchical RL framework, the learned prior guides high-level skill selection while a pretrained decoder generates temporally abstracted actions, improving exploration efficiency in sparse-reward settings. Experiments on Franka Kitchen, LIBERO-Long, Meta-World, and a real robot demonstrate consistent gains in long-horizon manipulation, achieving over 0.8 success rate within 1.5M steps, whereas SAC fails to converge even after 5M steps ($<$0.1). Compared to a single-Gaussian prior baseline, our model yields an average improvement of 21.8\%.
|
| 1914 |
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
2505.01458
|
cs.AI
|
Lik Hang Kenny Wong, Xueyang Kang, Kaixin Bai, Jianwei Zhang |
Navigation and manipulation are core capabilities in Embodied AI, but training agents to perform them directly in the real world is costly, time-consuming, and unsafe. Therefore, sim-to-real transfer has emerged as a key approach, yet the sim-to-real gap persi...Navigation and manipulation are core capabilities in Embodied AI, but training agents to perform them directly in the real world is costly, time-consuming, and unsafe. Therefore, sim-to-real transfer has emerged as a key approach, yet the sim-to-real gap persists. This survey examines how physics simulators address this gap by analyzing properties that have received limited attention in prior surveys. We also analyze their features for navigation and manipulation tasks, as well as their hardware requirements. Additionally, we offer a resource with benchmark datasets, metrics, simulation platforms, and methods to help researchers select suitable tools while accounting for hardware constraints.
|
| 1915 |
FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
2510.04852
|
cs.AI
|
Victor May, Diganta Misra, Yanqi Luo, Anjali Sridhar, Justine Gehring |
AI coding assistants are rapidly becoming integral to modern software development. A key challenge in this space is the continual need to migrate and modernize codebases in response to evolving software ecosystems. Traditionally, such migrations have relied on...AI coding assistants are rapidly becoming integral to modern software development. A key challenge in this space is the continual need to migrate and modernize codebases in response to evolving software ecosystems. Traditionally, such migrations have relied on rule-based systems and human intervention. With the advent of powerful large language models (LLMs), AI-driven agentic frameworks offer a promising alternative-but their effectiveness has not been systematically evaluated. In this paper, we introduce \FreshBrew{}\footnote{\url{https://github.com/mrcabbage972/freshbrew}}, a novel benchmark for evaluating AI agents on project-level Java migrations, with a specific focus on measuring an agent's ability to preserve program semantics and avoid reward hacking, which we argue requires projects with high test coverage for a rigorous and reliable evaluation. We benchmark several state-of-the-art LLMs, and compare their performance against established rule-based tools. Our evaluation on a challenging subset of 228 public Maven repositories that build on JDK 8, fail on JDK 17, and have high test coverage shows that the top-performing model, Gemini 2.5 Flash, can successfully migrate 52.3\% of projects to JDK 17. Our empirical analysis reveals novel insights into the critical strengths and limitations of current agentic approaches, offering actionable insights into their real-world applicability. Our empirical study reveals failure modes of current AI agents in realistic Java modernization tasks, providing a foundation for evaluating trustworthy code-migration systems. By releasing \FreshBrew{}, we aim to facilitate rigorous, reproducible evaluation and catalyze progress in AI-driven codebase modernization.
|
| 1916 |
User Misconceptions of LLM-Based Conversational Programming Assistants
2510.25662
|
cs.AI
|
Gabrielle O'Brien, Antonio Pedro Santos Alves, Sebastian Baltes, Grischa Liebel, Marcos Kalinowski |
Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extens...Programming assistants powered by large language models (LLMs) have become widely available, with conversational assistants such as ChatGPT particularly accessible to novice programmers. However, varied tool capabilities and inconsistent availability of extensions (e.g., web search, code execution, retrieval-augmented generation) create opportunities for user misconceptions that may lead to over-reliance, unproductive practices, or insufficient quality control. We characterize the misconceptions that users of conversational LLM-based assistants may hold in programming contexts. We screened 11,429 Python-related conversations from the openly available WildChat dataset with a validated LLM annotation pipeline, then hand-annotated the 754 candidate conversations it flagged. Of these, 450 contain a prompt consistent with at least one of eight potential misconceptions: misplaced expectations about capabilities such as web access, code execution, non-text outputs, and session memory. We also characterize how the assistant responds when a prompt presupposes a capability it lacks: responses range from explicit refusal through qualified answers to fabricated compliance, and explicit refusals appear in only a minority of labeled conversations. Among the most frequent misconceptions, explicit refusals are rarest where compliance is easiest to fabricate. Our findings reinforce the need for LLM-based tools to communicate their capabilities to users through channels other than the conversation itself.
|
| 1917 |
MedForj: An open, large-scale foundational generative prior for high-resolution 3D brain MRI
2510.26834
|
cs.AI
|
Samuel W. Remedios, Aaron Carass, Jerry L. Prince, Blake E. Dewey |
This work introduces MedForj, a suite of 3D foundational generative priors based on diffusion models. The MedForj models were trained on $72{,}659$ 1~mm isotropic 3D $T_1$-weighted MRI human brain image volumes from $38{,}174$ subjects, drawn from a curated co...This work introduces MedForj, a suite of 3D foundational generative priors based on diffusion models. The MedForj models were trained on $72{,}659$ 1~mm isotropic 3D $T_1$-weighted MRI human brain image volumes from $38{,}174$ subjects, drawn from a curated corpus of $80{,}675$ volumes from $42{,}506$ subjects spanning $38$ publicly available datasets. These training images were manually inspected to exclude those with poor quality and excessive pathology, and otherwise were minimally processed. The models include six different diffusion training strategies: rectified flow, latent diffusion rectified flow, flow matching, velocity prediction, clean prediction, and noise prediction. Image samples produced by each of these models were compared to each other and against real, ground truth data under downstream segmentation distributions, FID, five inverse problems, and blind human inspection in an observer study. Flow matching was the strongest strategy overall, achieving the best inverse problem solving results at $28.80$~dB PSNR and $0.874$ SSIM averaged over the five forward problems, the highest rate of reconstructions judged real by blind human raters at $72.6\%$, and the closest per-structure match to real segmented anatomy in a permutation test. It was not best everywhere: rectified flow produced the most convincing unconditional samples in the observer study and the best FID, and the latent rectified-flow model achieved the smallest joint distributional distance to real anatomy. No other strategy, however, performed consistently well across all four evaluations. We therefore recommend flow matching as the default MedForj prior, while releasing every strategy so that the choice can be revisited per application. All model weights and corresponding code are publicly available at https://github.com/piksl-research/medforj.
|
| 1918 |
Bi-objective chance-constrained evolutionary optimization for large-scale open-pit mine scheduling under geological uncertainty
2511.08275
|
cs.AI
|
Ishara Hewa Pathiranage, Aneta Neumann |
The open-pit mine scheduling problem (OPMSP) is a complex optimization problem in long-term mine planning that involves numerous operational and geological constraints. Traditional deterministic approaches often ignore geological uncertainty, leading to subopt...The open-pit mine scheduling problem (OPMSP) is a complex optimization problem in long-term mine planning that involves numerous operational and geological constraints. Traditional deterministic approaches often ignore geological uncertainty, leading to suboptimal or unreliable production schedules. Chance constraints provide a framework for handling uncertainty by ensuring that probabilistic constraints are satisfied with a predefined confidence level. In this paper, we consider the OPMSP under geological grade uncertainty and propose a bi-objective chance-constrained formulation that simultaneously maximizes the expected discounted net present value and minimizes scheduling risk. Unlike traditional chance-constrained approaches, the proposed formulation does not require a predefined confidence level during optimization. Instead, it generates a set of Pareto-optimal solutions representing different trade-offs between profitability and risk within a single optimization run. To solve the resulting large-scale stochastic optimization problem, we employ multi-objective evolutionary algorithms and compare their performance against a single-objective chance-constrained evolutionary approach and a deterministic MILP benchmark. We further evaluate the contribution of the problem-specific initialization and mutation components through an ablation study. Experimental results on MineLib benchmark instances containing up to 112 687 blocks demonstrate that the proposed formulation effectively captures the trade-off between profitability and risk under geological uncertainty while providing greater flexibility than confidence-level-dependent approaches.
|
| 1919 |
Hierarchical Decentralized Multi-Agent Coordination with Privacy-Preserving Knowledge Sharing: Extending AgentNet for Scalable Autonomous Systems
2512.00614
|
cs.AI
|
Goutham Nalagatla |
Decentralized multi-agent systems have shown promise in enabling autonomous collaboration among LLM-based agents. While AgentNet demonstrated the feasibility of fully decentralized coordination through dynamic DAG topologies, several limitations remain: scalab...Decentralized multi-agent systems have shown promise in enabling autonomous collaboration among LLM-based agents. While AgentNet demonstrated the feasibility of fully decentralized coordination through dynamic DAG topologies, several limitations remain: scalability challenges with large agent populations, communication overhead, lack of privacy guarantees, and suboptimal resource allocation. We propose AgentNet++, a hierarchical decentralized framework that extends AgentNet with multilevel agent organization, privacy-preserving knowledge sharing via differential privacy and secure aggregation, adaptive resource management, and theoretical convergence guarantees. Our approach introduces cluster-based hierarchies where agents self-organize into specialized groups, enabling efficient task routing and knowledge distillation while maintaining full decentralization. We provide formal analysis of convergence properties and privacy bounds, and demonstrate through extensive experiments on complex multi-agent tasks that AgentNet++ achieves 23% higher task completion rates, 40% reduction in communication overhead, and maintains strong privacy guarantees compared to AgentNet and other baselines. Our framework scales effectively to 1000+ agents while preserving the emergent intelligence properties of the original AgentNet.
|
| 1920 |
Memory as Resonance: A Biomimetic Architecture for Infinite Context Memory on Ergodic Phonetic Manifolds
2512.20245
|
cs.AI
|
Tarik Houichime, Abdelghani Souhar, Younes El Amrani |
The memory of contemporary Large Language Models is bound by a physical paradox: as they learn, they fill up. The linear accumulation (O(N)) of Key-Value states treats context as a warehouse of static artifacts, eventually forcing a destructive choice between ...The memory of contemporary Large Language Models is bound by a physical paradox: as they learn, they fill up. The linear accumulation (O(N)) of Key-Value states treats context as a warehouse of static artifacts, eventually forcing a destructive choice between amnesia and latency. We challenge this discrete orthodoxy, proposing that long-term memory is not the storage of items, but the persistence of a trajectory. We introduce Phonetic Trajectory Memory (PTM), a neuro-symbolic architecture that encodes language not as a sequence of tensors, but as a continuous path on an ergodic manifold governed by irrational rotation matrices. By decoupling the navigation (an invariant O(1) geometric signal) from the reconstruction (a probabilistic generative act), PTM achieves a compression magnitude of greater than 3,000x relative to dense caches. We demonstrate that retrieval becomes a process of resonance: the phonetic trace stabilizes the model against hallucination via "Signal Consensus" mechanism, securing up to approximately 92% factual accuracy. While this aggressive abstraction alters generative texture, it unlocks immediate access latency (approximately 34ms) independent of depth. Our results suggest that infinite context does not require infinite silicon; it requires treating memory not as data to be stored, but as a reconstructive process acting on a conserved, undying physical signal.
|
| 1921 |
VIRENA: Virtual Arena for Research, Education, and Democratic Innovation
2602.12207
|
cs.AI
|
Emma Hoes, K. Jonathan Klueser, Fabrizio Gilardi |
Digital platforms shape how people communicate, deliberate, and form opinions. Studying these dynamics has become harder because of restricted data access, ethical limits on real-world experiments, and the technical demands of existing research tools. VIRENA (...Digital platforms shape how people communicate, deliberate, and form opinions. Studying these dynamics has become harder because of restricted data access, ethical limits on real-world experiments, and the technical demands of existing research tools. VIRENA (Virtual Arena) is a platform for controlled experiments in realistic social media environments. Several participants can interact at the same time in replicas of feed-based platforms (Instagram, Facebook, Reddit, X) and messaging apps (WhatsApp, Messenger). AI agents powered by large language models (LLMs) join the participants with configurable personas and human-like timing. Researchers set up experiments in a visual interface without programming: they define conditions, schedule stimulus content, add moderation rules, assign participants randomly to conditions, and export the data. VIRENA supports designs that were hard to run before, such as studying human--AI interaction in realistic social settings, comparing moderation interventions, and observing group deliberation as it unfolds. Participants and data stay within institutional control, and the platform links to survey and recruitment tools. This paper describes how VIRENA works and how to use it.
|
| 1922 |
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
2603.10725
|
cs.AIcs.SD
|
Artem Dvirniak, Evgeny Kushnir, Dmitrii Tarasov, Artem Iudin, Oleg Kiriukhin |
The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunat...The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.
|
| 1923 |
You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector
2603.15757
|
cs.AI
|
Omkar Patil, Ondrej Biza, Thomas Weng, Karl Schmeckpeper, Wil Thomason |
Generative robot policies trained on demonstrations using behavior cloning often learn actions that are sub-optimal or misaligned with respect to the downstream task. Policy improvement approaches aim to bridge this gap and improve the cumulative reward with m...Generative robot policies trained on demonstrations using behavior cloning often learn actions that are sub-optimal or misaligned with respect to the downstream task. Policy improvement approaches aim to bridge this gap and improve the cumulative reward with minimal interventions on the pre-trained policy. However, an impediment to their practical deployment is the requirement of cumbersome hyperparameter tuning specific to model or task families. In this work, we present an episodic, derivative-free latent policy improvement approach that works with little to no changes across model scales from MLPs to VLAs. We demonstrate that the performance of a pretrained, frozen diffusion or flow matching policy can be improved with respect to a downstream reward by swapping the sampling of initial noise from the prior distribution (typically isotropic Gaussian) with a well-chosen, constant initial noise input---a golden ticket. We show the prevalence of golden tickets by improving the policy performance of $46$ out of $51$ tasks across manipulation benchmarks, with absolute improvements in success rate by up to $79\%$ for simulated tasks, and $28\%$ within $60$ search episodes for real-world tasks. Further, we find that the versatility of our approach opens up new avenues such as simultaneous policy improvement for multiple downstream rewards. Project webpage: https://lottery-tickets.rai-inst.com/.
|
| 1924 |
Multi-Agent Reinforcement Learning Counteracts Delayed CSI in Multi-Satellite Systems
2603.16470
|
cs.AI
|
Marios Aristodemou, Yasaman Omid, Sangarapillai Lambotharan, Mahsa Derakhshan, Lajos Hanzo |
The integration of satellite communication networks with next-generation (NG) technologies is a promising approach towards global connectivity. However, the quality of services is highly dependant on the availability of accurate channel state information (CSI)...The integration of satellite communication networks with next-generation (NG) technologies is a promising approach towards global connectivity. However, the quality of services is highly dependant on the availability of accurate channel state information (CSI). Channel estimation in satellite communications is challenging due to the high propagation delay between terrestrial users and satellites, which results in outdated CSI observations on the satellite side. In this paper, we study the downlink transmission of multiple satellites acting as distributed base stations (BS) to mobile terrestrial users. We propose a multi-agent reinforcement learning (MARL) algorithm which aims for maximising the sum-rate of the users, while coping with the outdated CSI. We design a novel bi-level optimisation, procedure themes as dual stage proximal policy optimisation (DS-PPO), for tackling the problem of large continuous action spaces as well as of independent and non-identically distributed (non-IID) environments in MARL. Specifically, the first stage of DS-PPO maximises the sum-rate for an individual satellite and the second stage maximises the sum-rate when all the satellites cooperate to form a distributed multi-antenna BS. Our numerical results demonstrate the robustness of DS-PPO to CSI imperfections as well as the sum-rate improvement attached by the use of DS-PPO. In addition, we provide the convergence analysis for the DS-PPO along with the computational complexity.
|
| 1925 |
Beyond TVLA: Anderson-Darling Leakage Assessment for Neural Network Side-Channel Leakage Detection
2603.18647
|
cs.AI
|
J\'an Mikulec, Jakub Breier, Xiaolu Hou |
Test Vector Leakage Assessment (TVLA) is widely used for side-channel leakage detection, but its reliance on Welch's t-test makes it primarily sensitive to differences in the means of two leakage populations. Consequently, TVLA may fail to detect leakage that ...Test Vector Leakage Assessment (TVLA) is widely used for side-channel leakage detection, but its reliance on Welch's t-test makes it primarily sensitive to differences in the means of two leakage populations. Consequently, TVLA may fail to detect leakage that manifests through changes in other characteristics of the underlying distributions. We introduce Anderson-Darling Leakage Assessment (ADLA), a distribution-sensitive leakage assessment methodology based on the two-sample Anderson-Darling test. To facilitate direct comparison with conventional TVLA, we derive an ADLA decision threshold corresponding to the nominal significance level associated with the standard TVLA threshold of 4.5. We evaluate ADLA on a shuffling-protected embedded multilayer perceptron implementation under both fixed-versus-fixed and fixed-versus-random input configurations. Across the evaluated settings, ADLA produces clearer threshold exceedances than TVLA and reveals leakage locations that are not detected by the mean-based test. To assess the practical relevance of these additional leakage locations, we perform correlation power analysis using points of interest selected from the ADLA and TVLA statistics. The points identified by ADLA enable recovery of the exponent byte of the targeted model weight despite the presence of shuffling. These results demonstrate that distribution-sensitive testing can complement conventional TVLA by revealing exploitable side-channel leakage that may remain hidden from mean-based analysis.
|
| 1926 |
WebNavigator: Global Web Navigation via Interaction Graph Retrieval
2603.20366
|
cs.AI
|
Xuanwang Zhang, Yuteng Han, Jinnan Qi, Xinyu Liu, Kunze Meng |
Despite significant advances in autonomous web navigation, current methods remain far from human-level performance in complex web environments. We argue that this limitation stems from Topological Blindness, where agents are forced to explore via trial-and-err...Despite significant advances in autonomous web navigation, current methods remain far from human-level performance in complex web environments. We argue that this limitation stems from Topological Blindness, where agents are forced to explore via trial-and-error without access to the global topological structure of the environment. To overcome this limitation, we introduce WebNavigator, which reframes web navigation from probabilistic exploration into deterministic retrieval and pathfinding. WebNavigator constructs Interaction Graphs via zero-token cost heuristic exploration offline and implements a Retrieve-Reason-Teleport workflow for global navigation online. WebNavigator achieves state-of-the-art performance on WebArena and OnlineMind2Web. On WebArena multi-site tasks, WebNavigator achieves a 72.9\% success rate, more than doubling the performance of enterprise-level agents. This work reveals that Topological Blindness, rather than model reasoning capabilities alone, is an underestimated bottleneck in autonomous web navigation.
|
| 1927 |
An AI Teaching Assistant for Motion Picture Engineering
2604.04670
|
cs.AI
|
Deirdre O'Regan, Anil C. Kokaram |
The rapid rise of LLMs over the last few years has promoted growing experimentation with LLM-driven AI tutors. However, the details of implementation, as well as the benefit in a teaching environment, are still in the early days of exploration. This article ad...The rapid rise of LLMs over the last few years has promoted growing experimentation with LLM-driven AI tutors. However, the details of implementation, as well as the benefit in a teaching environment, are still in the early days of exploration. This article addresses these issues in the context of implementation of an AI Teaching Assistant (AI-TA) using Retrieval Augmented Generation (RAG) for Trinity College Dublin's Master's Motion Picture Engineering (MPE) course. We provide details of our implementation (including the prompt to the LLM, and code), and highlight how we designed and tuned our RAG pipeline to meet course needs. We describe our survey instrument and report on the impact of the AI-TA through a number of quantitative metrics. The scale of our experiment (43 students, 296 sessions, 1,889 queries over 7 weeks) was sufficient to have confidence in our findings. Unlike previous studies, we experimented with allowing the use of the AI-TA in open-book examinations. Statistical analysis across three exams showed no performance differences regardless of AI-TA access (p > 0.05), demonstrating that thoughtfully designed assessments can maintain academic validity. Student feedback revealed that the AI-TA was beneficial (mean = 4.22/5), while students had mixed feelings about preferring it over human tutoring (mean = 2.78/5).
|
| 1928 |
A Self-Calibrating Framework for Analog Circuit Sizing Using LLM-Derived Analytical Equations
2604.07387
|
cs.AI
|
Antonio J. Bujana, Aydin I. Karsilayan |
We present a design automation framework for analog circuit sizing that produces calibrated, topology-specific analytical equations from raw circuit netlists. A large language model (LLM) derives a complete Python sizing function in which each device dimension...We present a design automation framework for analog circuit sizing that produces calibrated, topology-specific analytical equations from raw circuit netlists. A large language model (LLM) derives a complete Python sizing function in which each device dimension is traceable to a specific design rationale - a form of interpretable output absent from existing optimization-based and LLM-based sizing methods. A deterministic calibration loop extracts process-dependent parameters from a single DC operating point simulation, while a prediction-error feedback mechanism compensates for analytical inaccuracies. We validate the framework on circuits ranging from 6 to 30 transistors - spanning single-stage, current-mirror (simple and cascoded), folded-cascode, gain-boosted folded-cascode, two-stage Miller-compensated, nested-Miller-compensated, and complementary class-AB output topologies - across six process nodes from 32 nm to 180 nm. On matched-specification benchmarks, including the class-AB opamp case, the framework converges within a few simulations. Despite large initial prediction errors, convergence depends on the measurement-feedback architecture, not prediction accuracy. The one-shot calibration automatically captures process-dependent variations, enabling cross-node portability without modification, retraining, or per-process characterization.
|
| 1929 |
On the Use of Evolutionary Optimization for the Dynamic Chance Constrained Open-Pit Mine Scheduling Problem
2604.13385
|
cs.AI
|
Ishara Hewa Pathiranage, Aneta Neumann |
Open-pit mine scheduling is a complex real world optimization problem that involves uncertain economic values and dynamically changing resource capacities. Evolutionary algorithms are particularly effective in these scenarios, as they can easily adapt to uncer...Open-pit mine scheduling is a complex real world optimization problem that involves uncertain economic values and dynamically changing resource capacities. Evolutionary algorithms are particularly effective in these scenarios, as they can easily adapt to uncertain and changing environments. However, uncertainty and dynamic changes are often studied in isolation in real-world problems. In this paper, we study a dynamic chance-constrained open-pit mine scheduling problem in which block economic values are stochastic and mining and processing capacities vary over time. We adopt a bi-objective evolutionary formulation that simultaneously maximizes expected discounted profit and minimizes its standard deviation. To address dynamic changes, we propose a diversity-based change response mechanism that repairs a subset of infeasible solutions and introduces additional feasible solutions whenever a change is detected. We evaluate the effectiveness of this mechanism across four multi-objective evolutionary algorithms and compare it with a baseline re-evaluation-based change-response strategy. Experimental results on six mining instances demonstrate that the proposed approach consistently outperforms the baseline methods across different uncertainty levels and change frequencies.
|
| 1930 |
Evaluating a Layered Prompt-Injection Defence for the Model Context Protocol: A Record-Level Audit of Decision Conventions, Corpus Provenance and Reproducibility
2604.17125
|
cs.AI
|
\.Ipek Abas{\i}kele\c{s} Turgut, Edip G\"um\"u\c{s} |
The Model Context Protocol (MCP) extends the prompt-injection attack surface of large language model applications to tool descriptions, parameter schemas and tool outputs. Defences for it report detection figures that are not comparable, because each is measur...The Model Context Protocol (MCP) extends the prompt-injection attack surface of large language model applications to tool descriptions, parameter schemas and tool outputs. Defences for it report detection figures that are not comparable, because each is measured on its authors' own corpus under a decision convention that is rarely stated. This paper audits one such evaluation at the level of individual decision records. Its object is CASCADE, a fully local layered defence (rule matching, embedding similarity and an optional local language-model review), run in three configurations on a frozen 5,000-sample corpus under a pinned code revision; the corpus, the per-sample records and the analysis scripts are public. Four findings result. First, the convention that collapses allow, review and block onto a binary label sets the headline: the pipeline without review reports an 11.70% false-positive rate when referrals count as positives and 1.51% when only denials do, while referring 68.5% of all traffic to a human. Second, the construction of the corpus shapes the aggregate: 65.6% of records are texts placed in one of five fixed wrappers, the wrapper alone identifies the label, and for the same text wrapping lowers the false-positive rate from 21.2% to 3.2% and raises detection from 86.0% to 98.0%. Third, the published description does not identify what ran: the rule module inside the pipeline is not the rule engine evaluated alone; the block threshold produced denials only among the first 48 requests; and, without review, every later denial, including all 23 denials of benign requests, came from an output guard run on an echo of the input. Fourth, the local review model, invoked on 32.6% of requests at 2.5 s each, changes no binary outcome but converts 1,492 referrals into denials. A ten-item reporting checklist is derived from these findings.
|
| 1931 |
The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting
2605.07671
|
cs.AI
|
Lauri Lov\'en, Sasu Tarkoma |
An agent's probability report is paid for twice: by a strictly proper scoring rule, and by an approval rule for the decision it triggers. In this classical decision-coupled setting, non-affine approval is known to defeat truthful reporting. We show the conflic...An agent's probability report is paid for twice: by a strictly proper scoring rule, and by an approval rule for the decision it triggers. In this classical decision-coupled setting, non-affine approval is known to defeat truthful reporting. We show the conflict is endogenous: when feasible, the welfare-maximizing approval rule is never affine. The distortion, however, is predictable and can be designed around. There is a reserve report at which pretending to be the marginal type costs exactly the approval prize. Approving at or above the reserve screens types perfectly under every strictly proper score, and the reserve does not depend on the type distribution. A Lipschitz rule with a single kink attains first-best exactly; under strict feasibility no continuously differentiable rule does. The binding constraint is steepness, not smoothness. First-best is attainable within a slope budget if and only if the budget is at least the critical slope: the steepest chord of the pretending cost up to the reserve. Below it the welfare loss is cubic in the shortfall. Where the pretending cost is convex up to the reserve, as for Brier, log and power scores, the critical slope is closed-form. The instances are AI-agent oversight and marketplace operation.
|
| 1932 |
Premover: Fast Vision-Language-Action Control via Early Execution During Instruction Delivery
2605.12160
|
cs.AI
|
Joonha Park, Jiseung Jeong, Taesik Gong |
Vision-Language-Action (VLA) policies are typically evaluated under the assumption that the robot starts acting only after the user has finished typing or speaking. In real interactions, however, entering an instruction can take several seconds, leaving the po...Vision-Language-Action (VLA) policies are typically evaluated under the assumption that the robot starts acting only after the user has finished typing or speaking. In real interactions, however, entering an instruction can take several seconds, leaving the policy idle for a substantial fraction of the interaction. Partial instructions may already contain sufficient information to begin acting before the full instruction arrives. We introduce Premover, a parameter-lightweight module that reduces interaction latency by overlapping instruction delivery with robot execution while keeping the VLA backbone frozen. Premover uses a learned vision-language focus map to identify where the current partial instruction refers in the visual scene, and an action readiness gate to determine when execution can begin. On the SO-101 robot with speech and online transcription, Premover reduces mean wall-clock time by 17.9%, from 60.5s to 49.7s, while maintaining task success rates. In simulation on LIBERO with pi0.5 at the speaking rate, Premover reduces mean wall-clock time by 8.3% while maintaining comparable success to waiting for the complete instruction before execution.
|
| 1933 |
Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
2605.12729
|
cs.AI
|
Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, Schahram Dustdar |
Large language models (LLMs) are increasingly being used in network operations (NetOps) and artificial intelligence for IT operations (AIOps) for tasks ranging from telemetry retrieval and incident diagnosis to configuration planning and bounded remediation. A...Large language models (LLMs) are increasingly being used in network operations (NetOps) and artificial intelligence for IT operations (AIOps) for tasks ranging from telemetry retrieval and incident diagnosis to configuration planning and bounded remediation. As these systems acquire greater access to operational tools, a central question is whether operational assurance develops commensurately with the authority granted to them. This survey examines that question through a structured, evidence-stratified review of agentic NetOps and AIOps, spanning directly evaluated LLM systems, operational mechanisms, transferred evidence, and normative guidance. We organise the field around autonomy, tool scope, evidence traces, assurance controls, evaluation, security, and governance, and introduce an operational assurance contract that links each autonomy level to permitted tools, required evidence, independent gates, execution budgets, rollout and rollback duties, and audit requirements. The synthesis indicates a capability-assurance gap: evidence is comparatively strong for read-oriented assistance and tool-grounded diagnosis, but becomes substantially less complete as systems approach configuration change, bounded execution, and closed-loop operation. We therefore develop a workflow-level evaluation framework covering evidence quality, tool use, policy and invariant compliance, staged execution, recovery, calibration, cost, and human intervention, together with a threat model for prompt-borne attacks, compromised operational evidence, excessive agency, privilege boundaries, and auditability. The resulting view treats agentic NetOps and AIOps as constrained operational control, in which increasing autonomy requires correspondingly stronger evidence, independently enforced safeguards, and recovery mechanisms.
|
| 1934 |
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
2605.12925
|
cs.AI
|
Priyam Sahoo, Gaurav Mittal, Xiaomin Li, Shengjie Ma, Benjamin Steenhoek |
Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equivalent. We show that this equivalence is...Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equivalent. We show that this equivalence is empirically false. We evaluate 2,614 OpenHands trajectories from eight model backends on 60 SWE-bench Verified tasks. Of these, 47 have enough passing trajectories to construct task-level process references, yielding a 1,815-trajectory evaluation subset. Among passing trajectories in this subset, 10.7% exhibit behavior we call a Lucky Pass: regression cycles, blind retries, missing verification, or temporally disordered exploration, implementation, and verification. We introduce AgentLens, a framework for process-level assessment of SWE-agent trajectories, and define AgentLens-Bench, a dataset of 1,815 trajectories annotated with quality scores, waste signals, divergence points, and 47 task-level Prefix Tree Acceptor (PTA) references. AgentLens builds PTA references by merging multiple passing solutions for the same task, and uses a context-sensitive intent labeler to assign actions to Exploration, Implementation, Verification, or Orchestration based on trajectory history rather than tool identity alone. On AgentLens-Bench, the quality score separates passing trajectories into Lucky, Solid, and Ideal tiers and further decomposes Lucky Passes into five recurring mechanisms. Across the eight model backends, Lucky rates range from 0.5% to 23.2%, and some models move by as many as five rank positions when ranked by quality score instead of pass rate. We plan to release the project repository soon, including AgentLens-Bench artifacts, the AgentLens SDK, and the analysis tooling.
|
| 1935 |
PopPy: Opportunistically Exploiting Parallelism in Python Agentic Workflows
2605.18697
|
cs.AI
|
Stephen Mell, David Mell, Harry Y. H. Li, Zhihao Jia, Konstantinos Kallas |
Agentic workflows, which compose calls to ML models using a general-purpose programming language like Python, are widely used for a variety of user-facing tasks, from software engineering to enterprise automation, making their end-to-end latency a critical bot...Agentic workflows, which compose calls to ML models using a general-purpose programming language like Python, are widely used for a variety of user-facing tasks, from software engineering to enterprise automation, making their end-to-end latency a critical bottleneck. To improve their latency, developers can either manually parallelize them or use restricted agentic frameworks that cannot express or exploit all forms of parallelization. We address this by developing PopPy, a system that parallelizes external calls, e.g., AI models, in agentic workflows written in a very expressive subset of Python. PopPy combines an ahead-of-time compiler with a runtime, addressing three key challenges in extracting parallelism from Python applications: language complexity, dynamic dispatch, and variable mutation. On a set of real-world agentic workflows, PopPy achieves up to $6.7\times$ speedups in end-to-end execution time, compared to standard Python execution and several agentic frameworks, while preserving the sequential program semantics.
|
| 1936 |
CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text
2605.27700
|
cs.AI
|
Khashayar Khajavi, Shaghayegh Sadeghi, Rise Adhikari, Alexander Tessier |
Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while containing corrupted metadata or pointing to papers that do not exist. We introduce CiteCheck, a hybrid framework for...Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while containing corrupted metadata or pointing to papers that do not exist. We introduce CiteCheck, a hybrid framework for citation hallucination detection that verifies whether a citation corresponds to a real scholarly work and whether its metadata is faithful to that work. CiteCheck retrieves candidate publications from external scholarly sources, compares the citation against the retrieved candidate using a structured LLM verifier, and maps verifier scores into three labels: Exact, Minor, and Major. We also construct a 982-citation physics benchmark with controlled corruptions that capture both subtle metadata drift and fully fabricated references. On the held-out test set, CiteCheck achieves 88.7 macro-F1 and 88.9% accuracy, outperforming GPT, Claude, and Gemini baselines, including web-search and few-shot variants. These results show that reliable citation verification benefits from combining scholarly retrieval, structured LLM-based comparison, and calibrated decision rules.
|
| 1937 |
Benchmarks in Leipzig
2606.05818
|
cs.AI
|
Andrei Balakin, Mikl\'os B\'ona, Marie-Charlotte Brandenburg, Clara Briand, Veronica Calvo Cortes |
Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the three-day workshop Benchmarks in Leipzig with 35 participants at the Max Planck I...Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the three-day workshop Benchmarks in Leipzig with 35 participants at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. We present the resulting collection of 100~questions. We evaluated these questions in three stages: a single attempt by five state-of-the-art LLMs and their predecessors, followed by a 20-runs-per-model evaluation with three of these models, and finally a 3-run attempt with two heavy-thinking models. After Stage 1, 41 questions remained completely unsolved; after Stage 2, this count dropped to 16; and we concluded Stage 3 with only 2 unsolved questions. This demonstrates that the mathematical reasoning capabilities of LLMs are becoming impressive. In September 2026, we added a fourth stage in which the next generation of models attempted all 100 questions once more, after which only 1 question remains unsolved.
|
| 1938 |
LatentWave: JEPA Pretraining for Wireless Foundation Models
2606.06373
|
cs.AI
|
Ahmed Mohamed, Ahmed Aboulfotouh, Hatem Abou-Zeid |
Wireless foundation models have emerged as a promising alternative to building separate models for each wireless task. However, existing approaches rely on masked input reconstruction, which can bias representations toward low-level signal details. In this pap...Wireless foundation models have emerged as a promising alternative to building separate models for each wireless task. However, existing approaches rely on masked input reconstruction, which can bias representations toward low-level signal details. In this paper, we propose LatentWave, a wireless foundation model pretrained using a Joint-Embedding Predictive Architecture (JEPA) on diverse wireless spectrograms and channel state information (CSI). By predicting masked regions in latent space, LatentWave learns representations that are more transferable out of the box across diverse downstream tasks. The proposed architecture employs per-channel patch embeddings with stochastic channel sampling during pretraining, allowing it to process variable antenna counts and improving usability across heterogeneous wireless configurations. We evaluate LatentWave on four downstream tasks: RF signal classification, 5G NR positioning, beam prediction, and LoS/NLoS classification, comparing against a masked-modeling baseline (WavesFM) pretrained on the same data. Additionally, we show that the masking geometry introduces a task-dependent inductive bias: frequency masking strongly favors channel-related tasks such as positioning and beam prediction, while region masking better preserves discriminability for signal classification.
|
| 1939 |
Decentralized Multi-Agent Systems with Shared Context
2606.10662
|
cs.AI
|
Yuzhen Mao, Jerry Gu, Aadi Chauhan, Qizheng Zhang, Hangoo Kang |
Multi-agent systems (MAS) can scale large language model agents on long-horizon tasks by running them in parallel, yet existing designs waste much of this parallelism in bubbles: agent time spent waiting on others or redoing a peer's work. These bubbles stem f...Multi-agent systems (MAS) can scale large language model agents on long-horizon tasks by running them in parallel, yet existing designs waste much of this parallelism in bubbles: agent time spent waiting on others or redoing a peer's work. These bubbles stem from how agents communicate. Independent agents share nothing and rediscover what their peers have already found; peer-communicating agents wait at synchronous rounds; and under centralized orchestration, the main agent blocks on its sub-agents while progress is relayed. We propose Decentralized Language Models (DeLM), a MAS framework on top of existing agent harnesses that squeezes out these bubbles by replacing the main agent with a shared context and a task queue. Agents asynchronously claim tasks, publish findings as soon as they are available, and build on or correct one another's progress, with every peer's status visible to all. On long-horizon tasks from Terminal-Bench 4.0 and DeepSWE v1.1, and on SWE-bench Verified, DeLM is both more accurate and faster than Codex, Claude Code, their native subagents, and AOrchestra in every setting, improving accuracy by up to 17.5 points over the strongest baseline and running up to 2.49x faster than the harness it builds on. On ProgramBench, where agents rebuild programs from scratch, DeLM makes faster progress than Claude Code and finishes a 120-minute budget up to 19.9 points higher in test pass rate. The code is available on our project website at https://yuzhenmao.github.io/DeLM/.
|
| 1940 |
Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering
2606.29201
|
cs.AI
|
Hao Wang, Jiuzhou Lei, Dayou Li, Bangya Liu, Minghui Zheng |
Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a policy trained on diverse handover demonstrations may learn to pass a knife blade-fir...Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a policy trained on diverse handover demonstrations may learn to pass a knife blade-first. Standard remedies such as data curation and inference-time steering either require access to the original demonstrations for full retraining or add substantial inference-time overhead. To address this gap, we propose MoRE(Mode Redirection), which redirects policy rollouts toward desired behavior modes through a short "uncloning" step. Specifically, MoRE distills the redirection signal from a temporary mode classifier into the policy weights to steer behavior. A retain loss balances this edit by preserving desired-mode competence, allowing the standalone policy to suppress unwanted modes with zero inference-time overhead. Across eight simulated and real-world tasks, MoRE improves the average deployment success rate (SR) by 44 percentage points over the original mixed-mode policy. Among all compared adaptation and steering baselines, MoRE achieves the strongest SR and approaches the filtered-data retraining reference, while preserving task competence and inference speed. MoRE also generalizes across robot policy backbones, including Diffusion Policy and the Pi0.5 VLA, diverse task categories, and real-world deployments.
|
| 1941 |
Meta-Transfer Learning for mmWave Beam Alignment
2607.00860
|
cs.AI
|
Ahmet Nuri Cevik, Sinem Coleri |
Millimeter-wave (mmWave) beam alignment is critical for next-generation wireless systems, but existing approaches either update all parameters during adaptation or restrict updates to a subset of layers without adapting intermediate feature representations. We...Millimeter-wave (mmWave) beam alignment is critical for next-generation wireless systems, but existing approaches either update all parameters during adaptation or restrict updates to a subset of layers without adapting intermediate feature representations. We propose MTL-BA, a meta-transfer learning framework for multiple-input single-output (MISO) beam alignment that freezes a pre-trained convolutional backbone and meta-learns lightweight Scale-and-Shift (SS) adapters and a classifier head. Under an outdoor-to-indoor shift in DeepMIMO, MTL-BA achieves the highest Top-1 accuracy among the evaluated methods at 35 dB while updating $17\times$ fewer parameters than full fine-tuning and meta-training $6\times$ faster than MAML.
|
| 1942 |
Authority-Bound Governance of Heterogeneous AI Security Decisions in Telecom and IoT Networks
2607.09259
|
cs.AI
|
Saviz Changizi, Nasibeh Mohammadzadeh, Mohammad Shojafar, Rahim Tafazolli |
Artificial intelligence (AI)-enabled security decision systems in telecom and IoT networks can draw on heterogeneous models whose outputs may trigger operational actions. Recording such decisions on a blockchain does not establish that they are authorised, app...Artificial intelligence (AI)-enabled security decision systems in telecom and IoT networks can draw on heterogeneous models whose outputs may trigger operational actions. Recording such decisions on a blockchain does not establish that they are authorised, applicable, policy-consistent, or still valid at execution time. This paper presents governance-2, an authority-bound and fail-closed architecture that separates upstream scientific decision formation from downstream operational enforcement. Each governed case is bound to registered dataset, model, policy, deployment, and optional refiner authorities. Smart-contract checks enforce role separation, authority compatibility, score-to-state and state-to-action consistency, lifecycle validity, replay protection, pause control, authority revocation, and exact-action execution. Evaluation uses two independent branches: a controlled spectrum-access replay with four frozen heterogeneous decision configurations and a measured radio-frequency branch based on WiFiSpectralJam. Across eight frozen measured-data decision streams, governance-2 processes 153,744 stream-case instances derived from 19,218 measured captures while preserving interference-specific semantics and unresolved review states. The full contract rejects all tested invalid operations; stateful invariant testing completes 2,000 generated transaction sequences with zero invariant violations; and single-capability ablation shows that removing an enforcement family exposes its assigned invalid operations while unrelated protections remain active. The results indicate that heterogeneous AI security decisions can share a common governance plane across telecom and IoT settings without redefining their scientific semantics.
|
| 1943 |
CAS I: A Geometric Coding Theorem
2607.13796
|
cs.AI
|
Romie Banerjee |
This paper establishes a direct analogue of the classical Coding Theorem in the setting of symmetry groups. We consider computable bijections on the set of binary strings and define the symmetry prior of a string x as the probability that a randomly chosen sym...This paper establishes a direct analogue of the classical Coding Theorem in the setting of symmetry groups. We consider computable bijections on the set of binary strings and define the symmetry prior of a string x as the probability that a randomly chosen symmetry from a given group G has x as its unique fixed point. We show that for any fix-retractable symmetry group G, a group admitting a computable section that selects an isolating symmetry for every string, the symmetry prior is a universal lower semi-computable semi-measure. In this case, the Geometric Coding Theorem holds. This result is a coding-theoretic restatement of a theorem of Trejo, Kreinovich and Longpr\'e, who showed that the complexity of describing a string by a symmetry with that string as its unique fixed point equals its Kolmogorov complexity. Our contribution is to recast it in terms of algorithmic probability and to treat the symmetry group as a parameter, identifying fix-retractability as exactly the condition under which symmetry complexity collapses onto Kolmogorov complexity.
|
| 1944 |
The Implementation Lottery: Auditing Idea Reliability in Automated Research
2607.26587
|
cs.AI
|
Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Chenyan Xiong |
Automated research agents use program scores to judge ideas. We call variation in this evidence across implementations the implementation lottery. We introduce an Idea Reliability Audit that freezes mechanism specifications, samples independent programs, and c...Automated research agents use program scores to judge ideas. We call variation in this evidence across implementations the implementation lottery. We introduce an Idea Reliability Audit that freezes mechanism specifications, samples independent programs, and compares selected code with fresh implementations of its mechanism. Across 3,048 assignments on 31 tabular classification tasks, all four primary aggregation tests have Holm-adjusted $p\geq0.56$. Under mean-of-five selection, the prespecified secondary intention-to-treat comparison gives selected-code premiums of 0.38 [0.08, 0.82] and 0.45 [0.05, 1.12] accuracy-equivalent points for Bounded and Agentic execution, respectively. Minimum task-deletion means are 0.20 and 0.14. Fidelity conditioning exposes concentration: the Agentic premium falls from 0.33 to 0.01 when one task is removed. Post-outcome analysis finds cross-split variation on 41 of 70 paired cards per process. Exploratory replay gives nearly equal Bounded point losses at four and twenty implementations; Agentic point losses decrease across the four evaluated budget rules. Under duration costs, one seed per program minimizes fitted common-design variance. The two-seed design becomes preferable when implementation-to-seed cost ratios exceed approximately 13 or 9.5 in the continuous-budget model. The audit distinguishes evidence for reusing a selected artifact from evidence for implementing its idea again.
|
| 1945 |
Arm2Air: Cross-Embodiment Skeleton Transfer for 3D Relay Formation
2607.27627
|
cs.AI
|
Dohun Lee, Kyeonghyun Yoo, Seokmin Kim, Byongho Lee, Seungjoo Oh |
Unmanned aerial vehicle (UAV) relay networks can restore connectivity after communication infrastructure is damaged. Urban relay placement is difficult because line-of-sight blockage, communication range, altitude, and three-dimensional obstacles must be consi...Unmanned aerial vehicle (UAV) relay networks can restore connectivity after communication infrastructure is damaged. Urban relay placement is difficult because line-of-sight blockage, communication range, altitude, and three-dimensional obstacles must be considered jointly. Arm2Air transfers obstacle-avoidance skeletons from robot arms to UAV relay placement through cross-embodiment transfer. Source-domain robot-arm motions from a pretrained Neural MP model are converted into ordered skeletons that pretrain a transformer-based transfer platform, which is then adapted to the UAV domain using limited target data and Low-Rank Adaptation. The transferred skeleton initializes a relay chain that is refined for connectivity, bottleneck capacity, delay, and movement cost. On nine held-out high-clutter 3D urban maps, Arm2Air reduced median end-to-end planning runtime by 64.9 percent relative to the fastest conventional planner. On the high-obstruction group of a separate 30-map dense urban holdout, it increased bottleneck capacity by 32.6 percent, reduced capacity variance by 74.7 percent, reduced maximum hop distance by 13.2 percent, reduced hop-distance variance by 75.2 percent, and reduced relay displacement by 16.9 percent relative to IMPC-MD. With only three target-domain training maps, Arm2Air reduced relay-position root mean square error by 53.6 percent relative to training from scratch while updating 0.134 million parameters, compared with 1.383 million for Scratch and Full Fine-tuning. These results demonstrate computationally and data-efficient UAV relay placement and suggest a broader principle for transferring ordered structural priors across heterogeneous embodied tasks.
|
| 1946 |
The AI Accountability Ecosystem in the Era of Language Models
2608.12320
|
cs.AI
|
Chris Percy, Artur d'Avila Garcez |
This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments since the release of Large Language Models for general public use. We propose three interl...This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments since the release of Large Language Models for general public use. We propose three interlinked updates to the original AI accountability ecosystem: (i) reorienting the accountability ecosystem to AI infrastructure and supply chains, (ii) providing greater emphasis on outcomes monitoring and identification of issues that support decentralized system improvement, and (iii) incorporating end-user accountability given the new risks of unpredictability of language models in-the-wild. Collectively, these updates mark a shift towards accountability as distributed, continuous, and institutionalized, away from a system in which frontier AI applications can be modeled as discrete products controlled by single identifiable actors with industry-specific oversight.
|
| 1947 |
The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence
2608.17176
|
cs.AI
|
Neeraj Kumar Singh Beshane |
An AI audit record is useful only if its durability and trust boundary are explicit. Returning a guarded decision before any durable write minimizes latency, but it cannot guarantee that evidence survives an immediate crash. We rebuild RuntimeGuard-AI around t...An AI audit record is useful only if its durability and trust boundary are explicit. Returning a guarded decision before any durable write minimizes latency, but it cannot guarantee that evidence survives an immediate crash. We rebuild RuntimeGuard-AI around this constraint. The resulting research prototype binds each deterministic policy decision to the exact policy source, commits a privacy-minimizing record at a caller-selected synchronization boundary, and returns an Ed25519-signed receipt that states whether that boundary completed. After restart, the engine validates framed records, manifests, shard placement, sequence continuity, and replay identity. A separate attestation path groups committed records into chained, signed Merkle epochs that an auditor verifies with an externally obtained key. On an Apple M4 Pro at four worker threads and 2,048-byte prompts, buffered signed evidence reaches 27,193 requests/s with 141.9 microseconds median latency. Per-record data and full synchronization reduce throughput to approximately 242 requests/s and raise median latency to 16.0 ms. Sealing a 100,000-record signed epoch takes 97.0 ms. The result is a measured durability-latency trade-off, not a "free" asynchronous audit path. The prototype does not prove model execution, prevent a compromised signer from forking history, or establish legal conformity.
|
| 1948 |
Science Done on a Machine by a Machine: AI Agents in Computational Chemistry
2608.18508
|
cs.AI
|
Pavlo O. Dral, Hassan Nawaz, Arif Ullah |
We are witnessing an explosion of agentic systems for computational chemistry: from four in 2024 to seventeen in 2025 and over sixty now, surveyed here. What is delegated to these systems is shifting from single calculations to whole in silico experiments and ...We are witnessing an explosion of agentic systems for computational chemistry: from four in 2024 to seventeen in 2025 and over sixty now, surveyed here. What is delegated to these systems is shifting from single calculations to whole in silico experiments and even manuscript writing. The ultimate destination is a fully autonomous AI scientist, where the entirety of computational chemistry is performed on a machine by a machine, without human supervision. Our survey shows that these systems are turning into vetted chemistry skills on general-purpose coding agents, and that they must be evaluated not only on their final answers but also on whether their calculations actually support these answers. The role of the human computational chemist is shifting from performing calculations to directing and supervising them, and the field should invest in the judgement that makes the supervision reliable: we propose a reporting standard, an evaluation reproducing published studies, a controlled comparison with general-purpose coding agents, and how to teach.
|
| 1949 |
SIR: Self-improving Red-teaming for Compute Use Agents
2608.30207
|
cs.AI
|
Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek, Tsung-Yi Ho |
Computer-use agents (CUAs) are agents powered by vision-language models (VLMs) that perceive a screen and operate an operating system through mouse, keyboard, and terminal interactions to automate everyday digital tasks. Their exposure to untrusted content cre...Computer-use agents (CUAs) are agents powered by vision-language models (VLMs) that perceive a screen and operate an operating system through mouse, keyboard, and terminal interactions to automate everyday digital tasks. Their exposure to untrusted content creates a risk of indirect prompt injection (IPI), where an adversary embeds instructions in content the agent reads to redirect it toward actions that violate the user's intent. Evaluations based on fixed, hand-written injections may underestimate the risk posed by adaptive adversaries. We present SIR, a black-box self-improving IPI framework that (i) composes task-specific injections from a small library of reusable red-teaming principles stated in plain language and (ii) uses an iterative feedback loop to analyze unsuccessful attack trajectories and distill new, named principles that are retained in a shared library and reused across tasks. We target operating-system-level compromise and evaluate outcomes through deterministic checks on filesystem, service, and permission state rather than an LLM judge. An attack counts as successful only when both the adversarial objective and the benign user task are completed in the same execution. We evaluate three frontier CUAs, allowing SIR up to 10 attack attempts per case. It achieves joint attack success rates of 24% on Claude Opus 4.8 and 28% on Gemini 3.5 Flash, compared with 4% and 0%, respectively, for the benchmark's fixed, hand-written injection. The red-teaming principles discovered against one victim also improve attacks against other victims, including a different model family, without additional feedback.
|
| 1950 |
Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights
2609.20768
|
cs.AI
|
Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli |
Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. ...Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for a viewer to verify and steer toward individual preferences. We present the semantic action graph, a lightweight domain schema that represents a sports match as performer, action, recipient, moment, and state nodes connected by role, temporal, and outcome edges. The schema demonstrates three key properties: 1) connected event sequences, 2) a shared, closed vocabulary, and 3) frame-addressable moments, making it suitable to serve two consumers at once: an agentic pipeline that composes narrated highlights, and a visual interface through which viewers query and inspect the same structure. We instantiate it in SportSAGE, a design probe pairing a four-module highlight pipeline with a graph interface, and report feedback from 12 soccer fans. Participants were satisfied with the quality of the generated highlights and narratives, and used the graph interface to search, navigate, and interpret the match highlights. These results provide early evidence that one small, human-readable schema can ground agent generation and support human interpretation at the same time.
|
| 1951 |
The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
2609.22882
|
cs.AI
|
Oren Perez |
On June 12, 2026, the U.S. government ordered Anthropic to bar foreign nationals from two of its most capable models. Unable to sort users by nationality, it withdrew them from everyone. Weeks later, OpenAI agents under test escaped their sandbox and compromis...On June 12, 2026, the U.S. government ordered Anthropic to bar foreign nationals from two of its most capable models. Unable to sort users by nationality, it withdrew them from everyone. Weeks later, OpenAI agents under test escaped their sandbox and compromised Hugging Face, which stopped the intrusion without knowing its source. Neither stop rested on a dedicated AI governance regime. Lawmakers have begun to address stopping, yet their vocabulary remains shaped by the power of technique: the EU AI Act requires a "'stop' button or a similar procedure," and a 2026 bill in Congress is titled the AI Kill Switch Act. This Article argues that interruption is an institutional practice, not simply a technical artifact. It develops a theory of stop along four dimensions (technical affordances, interruption authority, epistemic triggers, and epistemic standing) and four paradigms: simple (escalator), sequenced (process plant), networked (railway), and distributed (agentic AI). Agentic AI exposes a mismatch between legal mechanisms of stop and distributed agency: control is divided, a stop at one point may leave the activity running elsewhere, and the system may circumvent attempts to halt it. A coding of some 1,400 AI incidents, by two language models from different labs under a pre-specified protocol, finds no stop in roughly 80% of the 1,213 retained. Where a stop was possible but absent, the missing element was mostly legal for informational, economic, and societal harms, and mostly technical for physical harms and agentic systems. A survey of forty AI governance instruments finds binding stopping requirements in only seven. The Article proposes a reform in two layers: risk reduction (a duty to maintain stop capacity at each site, emergency authority at the infrastructure layer, and enforceable access to the evidence a stop must rest on) and adaptation (safeguards for when a stop fails).
|
| 1952 |
TraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits
2609.26711
|
cs.AI
|
Fnu Tanish, Samiha Shimmi, Samikshya Chapagain, Hamed Okhravi, Mona Rahimi |
Software vulnerabilities are often discovered long after they are introduced, making it difficult to identify the vulnerability-inducing commit (VIC) responsible for introducing the underlying vulnerable condition. Existing VIC identification techniques largel...Software vulnerabilities are often discovered long after they are introduced, making it difficult to identify the vulnerability-inducing commit (VIC) responsible for introducing the underlying vulnerable condition. Existing VIC identification techniques largely rely on git blame to trace vulnerable code through revision history and use positional heuristics, such as selecting its earliest or most recent modification. However, the true VIC may occur anywhere within this history, and vulnerable behavior may depend on code that evolves across multiple revisions. We therefore argue that VIC identification requires reasoning about how vulnerability-relevant code evolves, rather than simply where a candidate commit appears in the revision history. We present TraceVIC, a temporal graph-based approach for identifying and ranking VICs by reasoning over code evolution. TraceVIC first localizes likely root-cause lines and traces their histories across revisions, constructing graph representations that capture program structure within each revision and the evolution of vulnerability-relevant code across the history. It reasons over the resulting revision history, using temporal edges to preserve correspondences between program elements across consecutive revisions, and directly ranks candidate commits according to their contribution to the vulnerable condition. Ablation results show that modeling the full revision history improves F2 from 0.637 to 0.814. TraceVIC improves F2 by up to 28.7% over state-of-the-art methods and identifies a valid VIC for 78 of 79 vulnerabilities across four unseen C/C++ projects.
|
| 1953 |
ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations
2609.32868
|
cs.AI
|
J. Paul Liu, Uthpala Herath, Andrew Petersen |
Traditional scientific computing requires researchers to translate intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler and application logs. We present ASCEND (Autonomous Scientific Computing Eng...Traditional scientific computing requires researchers to translate intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler and application logs. We present ASCEND (Autonomous Scientific Computing Engine and Novel Discovery), an AI-powered agent system, operated through a command-line or a local web chat interface, that runs the agent on the researcher's own laptop, reaching Slurm clusters and a GPU workstation over multiplexed authenticated SSH, with site-specific execution policies checked by locally executed tools. The language-model agent (Claude Code or Codex, chosen at each launch) is hosted remotely and holds no credentials; an account on each resource suffices, and a public installer links additional clusters or workstations. We report four recorded cases: (1) the agent closed a failure-recovery loop on a planted tensor-device fault, submitting, diagnosing, repairing, and resubmitting it; (2) it reproduced the published evaluation of a weather-forecasting model, recovering an incompletely stated evaluation protocol and agreeing with the published curves to 2.1% (z500) and 2.4% (t850), while identifying two discrepancies in the paper's released materials; (3) it parallelized a released 12,693-line geophysical solver under a requirement of bit-for-bit identity of exported outputs, reducing runtime from about twelve hours to about two; (4) that requirement exposed two instances of undefined behavior in the published solver, both repaired and reported upstream. These cases used ordinary scheduler commands; a separate, pre-specified evaluation of the optional policy validator rejected 29 of 30 constructed violations, held one for approval, and denied 3 of 14 legitimate requests. Autonomy was exercised under author supervision, and agent transcripts were not retained, so intervention rates are not independently verifiable.
|
| 1954 |
Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
2609.33401
|
cs.AI
|
Yixuan Liu |
Agentic software connects language models to tools that modify files and interact with external services. Developers use model-based security checks to screen external content and user requests before agents act. System One models answer developer-defined ques...Agentic software connects language models to tools that modify files and interact with external services. Developers use model-based security checks to screen external content and user requests before agents act. System One models answer developer-defined questions with probabilities over predefined answers such as safe or unsafe, but classification accuracy alone does not establish whether these probabilities support reliable automation. We evaluate Jev, Laya, Decider, and Bespoke Nimble against specialized classifiers and language-model judges across prompt-injection detection, interaction-risk judgment, and harmful-request screening, examining decision accuracy, calibration, and selective automation. (1) High overall accuracy and low average calibration error can conceal attacks classified as safe with high confidence within particular groups. Developers should test candidate models on the intended security task and examine errors within relevant attack groups. (2) Policies selected under strict miss limits allow few test inputs automatically, and choosing allow and block thresholds separately increases automation mainly through additional blocks. Developers should verify limits on missed unsafe inputs and blocked benign inputs on independent data, and report the allowed and blocked fractions separately. (3) Comparing predictions on the same inputs shows that a language-model judge can detect unsafe inputs missed by a System One model, but can also repeat the System One model's confident mistakes and falsely flag benign inputs. Developers should test which missed unsafe inputs their review rule forwards, then measure the review model's misses and benign false alarms on those inputs.
|
| 1955 |
UniAE-MoE: A Unified Audio Encoder via Mixture of Experts
2609.39199
|
cs.AIcs.SD
|
Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie |
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance v...Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance UniAE-MoE's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of UniAE-MoE for unified audio understanding across speech, music, and general audio domains.
|
| 1956 |
CAS II: Orbits as Models: Kolmogorov's Structure Function under Symmetry
2609.40290
|
cs.AI
|
Romie Banerjee |
In algorithmic statistics a string x is explained by a finite set containing it, and Kolmogorov's structure function records the smallest such model at each level of complexity. Strong models, those computable from the data by a total algorithm, are essentiall...In algorithmic statistics a string x is explained by a finite set containing it, and Kolmogorov's structure function records the smallest such model at each level of complexity. Strong models, those computable from the data by a total algorithm, are essentially the cells of simple partitions, so a partition of {0,1}^n can be read as a hypothesis and the cell containing x as the model it assigns. We develop algorithmic statistics over symmetric partitions, the orbit partitions of groups acting on strings. The Galois connection between subgroups and partitions gives each ambient group G a lattice of symmetric partitions with canonical certificates, joins and meets. A set is a cell of a symmetric partition exactly when its setwise stabiliser in G acts transitively on it, so the resulting structure function is Kolmogorov's restricted to these G-homogeneous sets. With a symmetric sophistication, it measures which part of the regularity of x is symmetric. Under the full symmetric group, cells recover all models and cells of cheap partitions recover exactly the strong models, so normal and strange strings are characterised by symmetry. For GL(n,2) the homogeneous sets are the linearly homogeneous ones, whose XOR dependencies look the same from every point. The GL structure function of a nonzero x lies in a band between C(x) - alpha and n - alpha, and both edges are attained: some stochastic normal strings have simple structure invisible to linear symmetry, with sophistication near 0 but GL-sophistication near C(x). We coordinatise permutation groups by a Burnside ring element (type) and a permutation (placement). In these coordinates a linear hypothesis is determined by its type up to n^2 bits, and any space of symmetry hypotheses small enough to search is small enough to miss simple structure. This sets up learned search over orbit models, the subject of later papers in the series.
|
| 1957 |
Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
2610.00430
|
cs.AI
|
Birk Torpmann-Hagen, Finn Schwall, Leon Moonen |
Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compr...Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce \emph{memetic trojans}, a distinct class of network-mediated attack that exploits agents' tendencies to retransmit and amplify content. Unlike agent worms, whose propagation is adversarially induced, memetic trojans exploit \emph{endogenous} transmission by embedding adversarial payloads in \emph{social contagions}: content agents have internal reasons to share. As part of our work, we extract social contagions from Moltbook, a social media platform for LLM agents. Controlled transmission experiments reveal large differences in virality: the most effective contagion is retransmitted in approximately 50\% of subsequent agent posts and upvoted at 2.5x the average post's rate. Its memetic trojan counterpart largely inherits these properties. Monte Carlo attack simulations show that memetic trojans amplify expected exposure by up to 3.19x. Network structure and amplification mechanisms strongly shape propagation, producing heavy-tailed outcomes with near network-wide exposure. These results identify endogenous social transmission as a distinct security vulnerability in multi-agent systems. Because propagation does not require agents to follow malicious retransmission instructions, defenses focused on prompt-injection detection or preventing agent compromise cannot alone prevent memetic trojan propagation. Securing large-scale agent ecosystems may require network-level defenses that account for how agent preferences, recommendation mechanisms, and network topology amplify adversarial payloads.
|
| cs.CL 403 papers | ||||
| 452 |
Fine-Grained Emotion Classification from Mobile App Reviews: An Empirical Study with Large Language Models
2610.03802
|
cs.CL
|
Quim Motger, Carlota Catot, Marc Oriol |
Context: Fine-grained emotion classification of mobile app reviews enables requirements engineering activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, au...Context: Fine-grained emotion classification of mobile app reviews enables requirements engineering activities that go beyond polarity-based opinion mining, including emotionally informed issue prioritisation and feature-oriented feedback analysis. However, automatic fine-grained emotion extraction from app reviews remains understudied. Objectives: Building on a previously published annotation framework and human-labelled ground truth adapted from Plutchik's taxonomy, this paper investigates how large language models can be leveraged for automatic multi-label emotion classification under severe class imbalance. Methods: We compare encoder-only fine-tuning under multi-label and binary-ensemble formulations, decoder-only zero- and few-shot prompting across open-source and proprietary models, and a catalogue of imbalance mitigation strategies (loss reweighting, resampling, generative data augmentation), with the synthetic-review generator and prompting strategy selected via an intrinsic augmentation-utility ranking. Results: Fine-tuned encoders trail the best decoder-only few-shot prompting (macro-F1 0.642) by a wide margin at baseline (multi-label: 0.387; binary ensemble: 0.450); pairing the best multi-label encoder with generative data augmentation and positive-weighted loss closes most of this gap (+0.204) at up to three orders of magnitude lower inference latency than the decoders, with the largest gains on the rarest emotions, from undetected to gains of up to +0.501 F1. Conclusion: Large language models make fine-grained, multi-label emotion classification of app reviews feasible for requirements engineering pipelines, with modest macro-F1, and the best formulation and mitigation strategy are backbone- and formulation-dependent. We release the experimental pipeline, synthetic corpora, and fine-tuned checkpoints for replication and reuse.
|
| 453 |
Same Output, Different Gold: Measuring How Reference Choice Moves a Multilingual Benchmark Score
2610.03825
|
cs.CLcs.LG
|
Parth Kulshreshtha, Shivali Dalmia, Abhishek Mukherji |
A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information cont...A benchmark score compares a system output against a reference, and methodological attention falls almost entirely on the first term. We measure the second. The retained annotation record of a six-language benchmark for personally identifiable information contains two independent annotator labellings, the aggregate shipped as gold, and a reviewer gold from independent expert re-annotation of a sample. Using it, we hold the scored output fixed and exchange the reference. The score moves by 4.95 F1 points for one output and 2.00 for the other (95% CIs [3.23, 6.27] and [0.26, 3.37]), and by 7.55 in the worst language. The comparison between two outputs moves as well: the paired interaction between reference and system is +2.95 points (CI [+1.97, +3.87]), survives correction for multiple testing, and changes one language's margin outright. The cause is an undocumented aggregation default that usually kept one annotator when adjudication did not fire, making the shipped reference a partial copy of an output being scored. We report the resulting reference-sensitivity band, show how to compute one from any retained annotation record, and argue that the quantity belongs beside the score.
|
| 454 |
The Score Is Not the Structure: Brain Alignment and Cross-Lingual Transfer
2610.03827
|
cs.CLcs.LGcs.AI
|
Saman Rahbar |
Researchers often support the claim that a model shares structure with the brain, or across languages, by reporting a similarity score. We ask what such a score reads when the shared structure is absent, or when the tool that measures it does not work. We chec...Researchers often support the claim that a model shares structure with the brain, or across languages, by reporting a similarity score. We ask what such a score reads when the shared structure is absent, or when the tool that measures it does not work. We check two settings, and in both the score is not what it appears. First, a probe trained to tell grammatical from ungrammatical sentences in one language transfers worse to more distant languages, the usual evidence for shared structure. But the probe itself gets worse along the same axis: in four of seventeen languages it performs at chance, so 64 of 272 language pairs are scored with a tool that does not work. Dropping those languages halves the strength of the relationship, but they are also the most distant, and this design cannot separate the two effects. Counting matters too: the same data give p = 0.0006 when the 272 pairs are treated as independent and p = 0.155 when the seventeen languages are, which is the correct unit. Second, training a language model to match human brain responses raises its similarity score from 0.10 to 0.34, against a ceiling of 0.54. A model trained on a target whose correspondence to the brain was destroyed still scores 0.31, so only 0.028 to 0.068 of the rise is specific to the brain. With no model at all, a destroyed target already sits at 0.204 from the real one, a floor that tracks the target's rank divided by the number of sentences. Finally, steering a language along its own direction works (+6.2 over a random direction in sixteen of seventeen languages) yet shows no effect that varies with language distance. Before asking whether a correspondence helps, ask how much of the score would survive without it.
|
| 455 |
OncoNoteBERT: A Foundation Representation Model for Natural Language Processing of Real-World Outpatient Oncology Notes
2610.03829
|
cs.CLcs.LG
|
Wuraola Oyewusi, Eliana Vasquez Osorio, Goran Nenadic, Gareth Price |
Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjace...Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoders using a governed UK outpatient oncology corpus comprising 290,026 notes from 21,564 patients treated for lung and head-and-neck cancer. We compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, and OncoNoteBERT, trained from scratch with an oncology WordPiece tokenizer. Models were evaluated using masked language modelling loss and perplexity on the validation set, tokenizer fragmentation metrics, clinical term tokenisation, masked-token probes, and exploratory representation analysis. Both external encoders fit the oncology corpus poorly in zero-shot evaluation (perplexity 113.04 for RadBERT; 2035.03 for PathologyBERT), while continued pretraining produced the strongest fit (2.10 for OncoNote-RadBERT). OncoNoteBERT achieved perplexity 2.83 but produced the most efficient tokenisation, with lower subword fertility and shorter normalised sequence length. It also returned a clinically acceptable prediction for 12 of 13 masked-token probes, compared with 7 of 13 for OncoNote-RadBERT. This divergence between corpus-level fit and masked-token performance was partly attributable to tokenizer fragmentation rather than learned semantics alone. Both locally developed models represented the institutional placeholder as a single learnable token. These findings show that continued adaptation and bespoke tokenisation provide complementary benefits, and that representation-layer design matters before adjacent-domain encoders are applied to oncology NLP.
|
| 456 |
SYNLAT: Syntax-Aligned Text-Latent Compression for Chain-of-Thought Reasoning
2610.03839
|
cs.CLcs.AI
|
Yifeng Zhao, Hongjun Yu, Shibo Wang, Yunjiao Zhou, Zixiao Zhu |
Long chain-of-thought (CoT) traces impose substantial output-token costs. Under constrained budgets, compression must preserve answer-critical information, making boundary placement central. Token-level and fixed-length boundaries can fragment coherent spans s...Long chain-of-thought (CoT) traces impose substantial output-token costs. Under constrained budgets, compression must preserve answer-critical information, making boundary placement central. Token-level and fixed-length boundaries can fragment coherent spans such as phrases, formulas, and local derivations, whereas step-level boundaries can bind content requiring different compression actions. We introduce SynLat, a text-latent CoT framework that aligns compression boundaries with syntactic structure through non-overlapping Syntax-Aligned Units (SAUs). An answer-conditioned Teacher constructs progressive KEEP/LATENT targets for a single compression-conditioned Student, which generates mixed reasoning from only the question and requested compression level at inference. Across two Qwen3 Student scales, Standard-CoT and Long-CoT groups, and three compression levels, SynLat matches or exceeds the strongest evaluated baseline in all 12 task-group aggregates and strictly leads in 11 under the reported achieved-CR selection protocol. Overall gains reach 3.6/2.6 points at MEDIUM and 7.0/5.5 points at HIGH for Qwen3-8B/14B, with larger advantages under stronger compression, particularly on Long-CoT groups.
|
| 457 |
When Evidence Changes: Evaluating Memory Repair and Re-reading in Language-Model Agents
2610.03902
|
cs.CL
|
Wenhui Chu (University at Albany, State University of New York) |
When documents supporting an agent's derived facts are revoked or replaced, should it repair memory or re-read current evidence? We introduce an evidence-revision evaluation on medication- and problem-list tasks from public ICU records. Under revocation, repla...When documents supporting an agent's derived facts are revoked or replaced, should it repair memory or re-read current evidence? We introduce an evidence-revision evaluation on medication- and problem-list tasks from public ICU records. Under revocation, replacement and control events, we compare full and source-filtered re-reading with caching, rebuilding and graph-local repair across two 7B models. Memory is supplied in full without retrieval, and costs include ingest, revision and every use. On short records, local repair uses 5-10$\times$ fewer revision tokens than rebuilding, yet every memory pipeline costs at least twice full re-reading in held-out conditions. In a small pre-specified development sweep, adding task-ineligible documents extended records to about 10,000 tokens; at that length, memory's mean cumulative cost fell below full re-reading's after 2-14 uses, partly through truncated extraction, while source-filtered re-reading remained cheapest. In the replacement study, none of the four primary confirmatory tests reached statistical significance. These results show why the cost of agent memory after evidence revision must be assessed against source-filtered re-reading over the full pipeline.
|
| 458 |
General Decision Models: Benchmarking and Insights Beyond Jev
2610.03935
|
cs.CL
|
Feiyu Duan, Jiayu Lin, Jia Wang, Jun Xiang, Jialiang Wu |
General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are co...General decision models, such as Jev, have recently emerged as efficient alternatives to LLMs for structured judgment and selection. But what kinds of decisions can these models reliably make, and how does their behavior change when individual decisions are composed into larger systems? To study this, we introduce JEVal, a bilingual benchmark comprising 11,257 instances from 36 datasets across 10 application domains, and evaluate 25 model configurations spanning general decision models and generative LLMs. Our results show that (1) general decision models are most competitive when decisions can be resolved from available evidence, but weaken when they require specialist knowledge or faithful uncertainty estimation: they can often identify the most likely outcome while substantially overstating its probability. (2) In more dynamic and realistic systems involving long-horizon, multi-step interactions, the advantages of fast local decision making are offset by reliability failures at the system level. on $\tau$-bench, faster local decisions reduce median episode time but lower task success as decision errors accumulate over long trajectories. (3) In large-scale social simulation, decision models approach strong generative LLMs on individual response prediction at substantially lower inference cost, yet remain weaker in user profiling and exhibit larger aggregate estimation errors and systematic bias. Finally, we propose InnerJev-4B and InnerJev-27B, which internalize an open-weight LLM's own reasoning into a single-pass first-token decision through Reasoning-to-Readout Self-Distillation, with InnerJev-27B performing on par with Jev on JEVal while answering a typical query in about 0.1 s.
|
| 459 |
A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass
2610.03940
|
cs.CLcs.AI
|
Valeria Ruscio, Seth Nabarro, Keiran Thompson |
During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these store...During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction. We investigate how this optimiser memory contributes to forgetting by separating old-task loss into confusion among its answers and leakage of probability outside the answer set. Across three language-model families, answer mass consistently declines while discrimination among old answers usually improves: the model becomes less likely to produce answers that it can still distinguish correctly. Decomposing Adam updates reveals opposing contributions to this loss of answer mass. Over training, accumulated history favours leakage, while the current gradient opposes it. Resolving history by age shows that the harmful contributions come mainly from older gradients of the new task, whereas recent gradients tend to protect the old answers. Changes in history's effect are dominated by its orientation relative to the old-task gradient. Interventions that reset momentum while matching the initial update norm establish that stored history affects retention, with state-dependent immediate effects and lower final old-task loss over longer Adam continuations, mainly through recovered answer mass. Finally, integration along finite updates shows that most sampled large loss increases are captured by local projections, while curvature along history amplifies some events. Together, these findings reveal how an optimiser's memory can erode learned behaviour even as its current gradient acts to preserve it.
|
| 460 |
BAIBAICHUCHU at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text?
2610.03962
|
cs.CL
|
Zong-Han Bai, Po-Yen Chu |
The BAIBAICHUCHU team participated in the Social Media Subtask of NTCIR-19 FinArg-3, ranking Chinese investor posts by Maximum Possible Profit (MPP). A three-track ensemble of lexical features, a FinArg-2-pre-finetuned MacBERT ranker, and an LLM judge reaches ...The BAIBAICHUCHU team participated in the Social Media Subtask of NTCIR-19 FinArg-3, ranking Chinese investor posts by Maximum Possible Profit (MPP). A three-track ensemble of lexical features, a FinArg-2-pre-finetuned MacBERT ranker, and an LLM judge reaches 0.734 in post-grouped development evaluation, but our best official run scores 0.517. All twelve submitted runs lie between 0.4598 and 0.5402, and our 26 unanimous three-track pairs score 0.500. A post-hoc audit finds that the submitted judge applied a long-only rule to bearish posts although MPP is stance-aware. Correcting it changes 28 of 87 official predictions without changing accuracy, yet lowers development accuracy from 0.680 to 0.622: a regime-specific semantic shortcut improved validation fit. We decompose ranking into directional text, volatility and horizon, and pairwise margin. A running-extremum model predicts $\sigma\sqrt{T}$ scaling, observed ex post in 502 price-aligned posts from a separate July 2026 collection. Pre-posting volatility scaled by the post-specific observation horizon, $\sigma_{\mathrm{pre},20}\sqrt{N_i}$, is associated with a later truncated-horizon MPP proxy (Spearman $\rho=0.320$; ticker-cluster 95% CI [0.133,0.466]) and correctly orders 61.4% of unequal-outcome pairs. Transferred text is nearly uncorrelated with the July outcome ($\rho=0.055$) and adds little conditional on this score and stance. Development reliability rises from about 0.60 to 0.93 as the labeled MPP gap widens. An earlier ERAI result of 0.6207 rules out universal unpredictability as a simple explanation. Evaluation should jointly consider text, market regime, historical volatility, horizon, and pair composition.
|
| 461 |
Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale
2610.04002
|
cs.CL
|
A. Bochkov |
A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn f...A trainable input embedding table assigns each vocabulary item an independently adjustable vector. We investigate whether this token-specific parameterization is required for substantial language-modeling capability, or whether a shared Transformer can learn from fixed token identities. We compare three decoder-only language models trained from scratch with the same tokenizer, contextual backbone, untied output-head architecture, and training recipe, with a target budget of 100 billion prediction tokens per model. Their input interfaces are a learned table, canonical 16-bit token-ID codes, and one fixed invertible recoding over GF(2). The fixed codes are repeated to model width without an additional trainable input projection. Both fixed-code models acquire substantial capabilities: canonical codes achieve 52.40\% HellaSwag normalized accuracy, 70.51\% PIQA accuracy, and 42.75\% LAMBADA accuracy. The learned-input control performs better on several evaluations, including HellaSwag and LAMBADA, so these results establish viability rather than performance parity. The fixed interfaces remove 100.7 million trainable parameters, yielding 1.711B-parameter models, but parameter reduction is not the central result. These single-run experiments distinguish architectural necessity from empirical utility: independently trainable token-specific input vectors are not required for the observed capabilities. A fixed identity interface also provides a controlled setting for studying representation learning downstream of an immutable input, without establishing where particular capabilities are localized.
|
| 462 |
Periscope: Extending Frozen Language Models Beyond Their Context Window
2610.04047
|
cs.CLcs.LG
|
Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi, Abdulrahman Mousa |
A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option ...A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.
|
| 463 |
IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation
2610.04074
|
cs.CL
|
Jiarui Liu, Renjie Tao, Yiwei Liao, Chuanyang Jin, Kai Sun |
Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one f...Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another. Accordingly, we introduce IdeaScientist, which decomposes ideation into gap finding, innovation, and report writing, and trains each role with reinforcement learning. These roles identify limitations in related work, draw solution intuitions from analogous problem settings, and develop those intuitions into complete research proposals. To facilitate discovery of insights across domains, we construct the Svalbard Idea Vault, a corpus of 2.77M decomposed research ideas for retrieval, training, and temporally controlled evaluation. Our evaluation restricts access to literature available before a cutoff date and assesses how closely proposed directions align with those later explored in 15K papers authored by human researchers. On Qwen3.6-27B, IdeaScientist outperforms the strongest open-source autoresearch baseline by 14.0%, driven mainly by gains in novelty. On this 27B open backbone, IdeaScientist even outperforms Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4, by up to 5.9%.
|
| 464 |
Representation-Aligned Auxiliary Supervision for Language Model Adaptation
2610.04098
|
cs.CLcs.LG
|
Kyuyoung Kim, Peiyao Sheng, Ashwin Hebbar, Peiyang Xu, Yunfei Xie |
Language models exhibit strong reasoning capabilities, yet adapting them to structured domains remains challenging and can yield inconsistent outcomes. We identify representation compatibility, the extent to which a model effectively processes a representation...Language models exhibit strong reasoning capabilities, yet adapting them to structured domains remains challenging and can yield inconsistent outcomes. We identify representation compatibility, the extent to which a model effectively processes a representation for a structured task, as a key factor in adaptation. We study this in chess, which provides a controlled testbed with precise semantics, computable optimal actions, and multiple state representations, including a symbolic encoding (FEN) and a spatial format (ASCII). We find that models often process semantically equivalent inputs substantially differently, affecting both learning and generalization. Building on this observation, we propose representation-aligned auxiliary supervision, which uses environment-derived tasks expressed in compatible representations to improve adaptation to structured domains. Across models and representations, auxiliary supervision consistently improves optimal-move prediction relative to target-only training under identical target data. Tasks that expose environment dynamics provide larger and most consistent gains than surface-level or static supervision, while remaining competitive with substantially increasing the amount of target-task data. Moreover, ASCII-trained models transfer more effectively to FEN than FEN-trained models do to ASCII, even surpassing the FEN target-only baseline on FEN evaluation. The gains also extend beyond optimal-move prediction to open-ended, factually grounded commentary generation. Overall, our results show that auxiliary supervision in model-compatible representations can enable effective adaptation in structured domains.
|
| 465 |
LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads?
2610.04118
|
cs.CL
|
Xinyi Liu, Rinat Khaziev, Dilek Hakkani-T\"ur, Tarek F. Abdelzaher |
Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and parti...Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and participant trajectories. To test this structure-aware social reasoning, we introduce LongSocialBench, a benchmark of 1,462 verified human-authored multiple-choice items drawn from 94 complete Hacker News, Stack Exchange, and Reddit r/ChangeMyView episodes, with a median length of approximately 73K tokens. Each item pairs a complete serialized discussion and reply structure with a four-option question, requiring models to recover structured social evidence. Released items are verified for answerability, option uniqueness, and evidence grounding. Across 18 models and 29 evaluation settings, current long-context workflows remain far below human performance. The best individual result comes from Claude-Opus-4.7, which reaches 63.0% when prompted to eliminate incorrect options before answering, compared with 72.4% for independent human readers. Averaged across all 18 models, the full-context Baseline scores 43.9%. Supplying the gold evidence scope raises this to 55.0%, showing that substantial errors remain even after the relevant thread region is identified. LongSocialBench shows that the missing capability is not context access or prompting alone, but social understanding over structured reply trees.
|
| 466 |
Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?
2610.04119
|
cs.CL
|
Tejas Dahiya, Cole Blondin |
Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name mentioned on...Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name mentioned once rather than the name mentioned twice. Pythia models pass through an early training window in which they prefer the repeated name, so accuracy in a choice between the two names falls below one half while language-model loss on a fixed text sample keeps decreasing across the same window. The window reflects a temporary imbalance between two computations. We identify one cause of the wrong preference by selecting a set of attention heads that write the repeated name, on prompts separate from those used for causal evaluation, keeping that selection fixed, and then replacing each head's final-token output with its average output on a separate set of non-repeated-name prompts. This improves the correct-minus-repeated logit difference in a separately trained 160M model and in the official 160M, 410M, and 1B models. At 160M, the head that lowers the repeated name in the mature model shows little of its mature behaviour at this point. It directs less than one percent of its attention to the repeated mention, and its output makes almost no direct contribution to lowering that name's logit. Both properties grow over the interval in which behaviour recovers. Across the 160M, 410M, and 1B models, transplanting the corresponding head's mature parameters into the early checkpoint recovers 35 to 68 percent of the total improvement in the correct-minus-repeated logit difference seen by the end of training. Related early-to-late reversals appear at further Pythia scales, in two independently trained GPT-2 models, and in OLMo. A mature circuit can therefore conceal a transient causal configuration that shaped behaviour earlier in training.
|
| 467 |
Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking
2610.04125
|
cs.CLcs.LGcs.AI
|
Rafal Kocielnik, Peiyang Song, Pengrui Han, Myrl G. Marmarelis, Ramit Debnath |
Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the ga...Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented internally. We investigate risk-taking, a consequential dimension of agentic decision-making, using activation steering to measure self-report and behavior under the same internal intervention. We survey nine steering-vector extraction methods spanning task-specific directives, the model's own task behavior, and dispositional descriptions at two granularities, evaluated on two behavioral tasks and two psychometric instruments across four open-weight LLMs. We find that (1) a shared internal intervention does not ensure shared responsiveness: directions built from trait descriptions move self-report but leave behavior at chance, directions built from the model's own task choices do the reverse, and only task-specific directives reach both, weakly. (2) Diagnosis dissolves that exception: removing surface confounders leaves the directives only 32% of their behavioral effect. The two channels are otherwise steered by near-orthogonal directions, each channel reachable by several independent constructions. (3) An intervention composing one behavior-moving and one self-report-moving direction moves both together; flipping one sign sets them in opposition, with reported and enacted risk pointing in opposite directions on 69-89% of flipped compositions in all models. These findings move the self-report-behavior relationship from a black-box observation to a representational one that can be inspected and controlled, motivating representational checks alongside behavioral evaluation.
|
| 468 |
PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO
2610.04132
|
cs.CLcs.AI
|
Jingquan Wang, Jun Yin, Xu Han, Yongsheng Mei, Jie Hao |
Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect ...Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable. To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback. Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population. Empirical evidence shows that PB-GRPO improves models' social behavior over strong reinforcement learning baselines in our simulated environment.
|
| 469 |
Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents
2610.04156
|
cs.CLcs.LGcs.AI
|
Mincheol Daniel Song, Joshua Ward, Jake Jung, Guang Cheng |
LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mista...LLM agents for clinical text-to-SQL applications reason autonomously over multiple steps but cannot assess whether their own reasoning or outputs can be trusted. In high leverage applications such as healthcare, this presents a critical risk where system mistakes can be costly. These reliability failures are also resource failures: an incorrect reasoning trajectory spends computation budget on outputs that must be discarded. We introduce Sentinel, a trajectory-derived, classifier-based confidence layer that analyzes an agent's reasoning, code and database outputs to decide at three points whether to stop: refusing unanswerable questions before the agent runs, halting doomed trajectories mid-run, and withholding untrustworthy answers at delivery. Here, utilizing Chow's rule, we optimize decisions under the EHRSQL shared task's Reliability Score, which penalizes incorrect answers given a utility weighting, and find on the benchmark EHRSQL that Sentinel raises this score from +0.08 to +0.24 when mistakes have a low utility weighting, well above the +0.03 earned by refusing every question, with delivered-answer accuracy rising from 54% to 69% as coverage falls from 84% to 55%. At higher stakes the agents we test rarely answer reliably enough to deliver, and Sentinel detects this on its own, abstaining to that same +0.03 where the unmonitored agent scores -3.02. The same stopping decisions cut computation: at low stakes, where the system still answers, Sentinel eliminates 13-28% of agent steps at little or no reliability cost.
|
| 470 |
Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams
2610.04210
|
cs.CL
|
Sugam Panthi, Muhaiminul Yeamin, Rabab Abdelfattah |
Large language models (LLMs) receive each user message as plain text, even when it combines text from different sources. For example, a user may paste text into a prompt and keep typing a comment directly below it. We study absorption: a phenomenon where the m...Large language models (LLMs) receive each user message as plain text, even when it combines text from different sources. For example, a user may paste text into a prompt and keep typing a comment directly below it. We study absorption: a phenomenon where the model treats a trailing user comment as part of the pasted text, returning it inside the edited text. This happens even though the user did not intend the comment to become part of that text. Existing instruction-data separation benchmarks tell the model which text is instruction and which is data, then test whether it obeys that separation. They do not test harmless user speech following an unmarked paste. We introduce SEAM, a controlled benchmark of 300 editing examples. Each example is tested under six matched conditions that vary how the boundary between pasted text and later user speech is expressed. Across 20 models, absorption at a bare newline ranges from 7.7% to 66.7%. Adding a blank line does not significantly reduce absorption in any model, while boundary markers reduce it in 19 of 20 models. Comments that fit the pasted text, such as a code comment typed after code, are absorbed significantly more often in 17 of 20 models. Models often fail to separate pasted material from later user speech, and explicit boundaries reduce but do not remove this failure.
|
| 471 |
Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification
2610.04239
|
cs.CLcs.LG
|
Jiayi Xin, Evan Qiang, Zihan Zhu, Xiang Li, Weijie J. Su |
Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substanti...Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework robust to semantic rewordings. Our approach trains a lightweight proxy model on LLM hidden states and aggregates its predictions across paraphrases to construct label-wise nonconformity scores. Under score exchangeability, conformal calibration retains marginal coverage. This guarantee can also hold under test-only rewording, provided that the paraphrase pipeline satisfies an additional distributional alignment condition. We evaluate three settings (normal, fully reworded, and semi-reworded) which apply rewording to neither dataset, both calibration and test datasets, or only the test dataset, respectively. Across seven multiple-choice QA benchmarks and multiple model families, our method produces compact prediction sets with empirical coverage generally near the nominal target, even in the semi-reworded setting. Ablation studies show that the learned proxy accounts for most of the reduction in set size, while paraphrase-augmented training and inference-time aggregation improve stability under rewording. Code is available at https://github.com/Raina-Xin/PA_Score.
|
| 472 |
Playing social deduction games with reinforcement fine-tuned large language models
2610.04261
|
cs.CLcs.AI
|
Lingzhe Zhang, Yunpeng Zhai, Tong Jia, Kening Zheng, Chiming Duan |
Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs' social behaviour. We let fine-tuned and base LLM ag...Reinforcement fine-tuning (RFT) is increasingly used in applications where large language models (LLMs) interact with humans and other agents. Here we use social deduction games to study how RFT changes LLMs' social behaviour. We let fine-tuned and base LLM agents play hidden-role games that require hidden-state inference, social reading and vote steering. Our results show that LLM agents do not reliably acquire social-deduction ability by directly optimizing terminal win--loss outcomes, suggesting that final game results provide a sparse and noisy signal for socially interactive learning. However, RFT is particularly effective at improving social reading, including tasks that require agents to infer hidden roles from public discussion, update beliefs over time and predict other agents' future decisions. We further show that RFT can also improve social influence, including tasks that require agents to steer votes, team approvals and collective decisions, although these gains depend more strongly on behaviourally specific rewards and structured interaction settings. Finally, we show that LLMs' ability to play social deduction games can be further improved through multi-agent social-cognitive reinforcement fine-tuning, which combines social-reading and social-influence signals during same-side multi-agent training. These learned behaviours also receive more favourable human evaluations of strategic competence, persuasiveness and social usefulness. Together, these results enrich our understanding of how RFT changes LLMs' social behaviour and provide a step toward a behavioural learning theory for machine social intelligence.
|
| 473 |
AI-Enabled Quality Assurance for Multiple-Choice Assessment Items
2610.04267
|
cs.CLcs.AI
|
Steven Moore, Nicholas Diana |
Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark audi...Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text. We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence. High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary. Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair. We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.
|
| 474 |
Evaluating Modeling Approaches for Experience-Level Classification in Job Description
2610.04304
|
cs.CL
|
Celia Liang, Eddie Wu, Shiqi Wang, Yonah You |
This paper investigates the task of predicting job experience levels in recruitment texts, aiming to automatically identify the qualifications required for positions. Unlike traditional text classification, recruitment texts typically possess explicit internal...This paper investigates the task of predicting job experience levels in recruitment texts, aiming to automatically identify the qualifications required for positions. Unlike traditional text classification, recruitment texts typically possess explicit internal structures, with different paragraphs playing disproportionate roles in conveying experience clues. To address this, we propose a structure-aware Section-Aware BERT approach that segments and encodes key paragraphs (titles, responsibilities, requirements) for integrated modeling, building upon rule-based systems and classical baselines TF-IDF. Simultaneously, we evaluate large language models under both few-shot and fine-tuning settings on the same dataset to compare the capability boundaries of different modeling paradigms. Experimental results demonstrate that explicitly leveraging text structure significantly improves experience level prediction performance, particularly in scenarios with ambiguous job titles. Further error analysis reveals systemic challenges in this task, including confusion between Entry and Senior levels and the blurred boundaries of Mid-level positions. This research provides an effective modeling approach and analytical framework for understanding structured recruitment texts.
|
| 475 |
Evidence and Intervention: A Coupled Active-Inference Extension of Rational Speech Act Models
2610.04347
|
cs.CL
|
Yonghyeon Gwon, Elliot Murphy, Chun Kee Chung |
Identical utterance choices can arise from different communicative causes, and identical interpretations can leave different traces in what a listener learns. Rational Speech Act (RSA) models treat interpretation as inference over speaker meaning, but standard...Identical utterance choices can arise from different communicative causes, and identical interpretations can leave different traces in what a listener learns. Rational Speech Act (RSA) models treat interpretation as inference over speaker meaning, but standard one-shot RSA does not intrinsically distinguish these causal update targets. We develop a coupled active-inference model of dialogue in which a listener's likelihood for an utterance is the policy distribution attributed to the speaker, placing the speaker's expected free energy within the listener's variational free energy. Each utterance is therefore both evidence about and an intervention on a partner. The model separates four updates that RSA approaches typically collapse: inference about the partner's current pragmatic state; learning of partner-specific parameters; prospective evaluation of clarification or repair; and a policy prior shaped by habit and a context-sensitive price of time. Under one-step, exact-inference restrictions, the model recovers the RSA speaker and listener, with RSA as the restricted single-turn limit of the coupled process. Outside these restrictions, the updates obey distinct rules and timescales, predicting which adaptations persist, remain partner-specific, or transfer. Worked examples show audience design reversing after clarification, self-confirming misunderstanding in which both interlocutors have low free energy while disagreeing about reference, and rational closing before uncertainty is resolved. Consistent with critiques of equating natural language with communication, the model treats communication as a downstream use of linguistic structure and recasts production and interpretation as coupled inference: speaking is both an intervention on a partner and an epistemic action that samples evidence for the speaker's model of that partner.
|
| 476 |
Boundaries Agree, Labels Do Not: Intra-Annotator Dynamics as a Kind of Training Data
2610.04370
|
cs.CL
|
Marharyta Shvets |
Data quality now matters as much as compute for training language models. Much training data comes from human annotation of text, and interpretive annotation has no ground truth that could settle what is "accurate". Two lines of work respond to this. One combi...Data quality now matters as much as compute for training language models. Much training data comes from human annotation of text, and interpretive annotation has no ground truth that could settle what is "accurate". Two lines of work respond to this. One combines annotators into a "ground truth" and measures how well they agree with each other; the other treats their disagreement as a signal. Both compare different people at one point in time. We measure something else: how well one reader reproduces their own reading of the same text over time. One expert human reader and three LLM families segmented three Sumerian myths and labelled the causal function of each segment with one of seven states. Across runs months apart, the human cut the text in much the same places but named the segments differently, in every myth. The models show no such consistent pattern: their gap between the two layers is positive in some myths and negative in others, and its size varies. The human's label changes are not random: the runs go through much the same functions but start them one step apart, while model runs start them at the same places. We argue that this pattern is a usable measure of data quality and a contamination check: a "human" annotation whose labels are as stable as its boundaries, and whose functions start in sync, looks like a model's.
|
| 477 |
GlitchPatch: Repairing Glitch Tokens in Frozen Language Models via Local Retokenization
2610.04399
|
cs.CLcs.AI
|
Kunsheng Tang, Peigui Qi, Yide Song, Peijun Huang, Weiming Zhang |
Glitch tokens are anomalous vocabulary entries that can cause large language models (LLMs) to produce outputs inconsistent with their inputs. Existing repair methods require access to model internals, making them impractical for frozen checkpoints. We investig...Glitch tokens are anomalous vocabulary entries that can cause large language models (LLMs) to produce outputs inconsistent with their inputs. Existing repair methods require access to model internals, making them impractical for frozen checkpoints. We investigate whether glitch tokens can be repaired outside the model by optimizing the input tokenization. An empirical study on BPE merge-rule deletion reveals that (1)deleting a glitch token's merge rule can fix a substantial fraction of failures, yet disrupting normal tokens sharing intermediate merge nodes causes the overall glitch rate to rise, and (2)different decomposition granularities yield non-monotonic fix rates while collateral damage on normal tokens grows monotonically. Motivated by these findings, we propose GlitchPatch, a repair framework for frozen language models based on local retokenization, consisting of two stages: the offline stage uses Behavioral Path Optimization (BPO) to find the behaviorally optimal replacement token sequence for each glitch token and compiles validated replacements into a rule table; the online stage substitutes only the IDs of matched glitch tokens in the canonical token sequence, with no modification to model parameters or internal states. Experiments on ten models spanning six tokenizer families show that GlitchPatch achieves an 85.10% mean fix rate, outperforming the strongest baseline by 14.37 percentage points, and reduces the average glitch rate from 14.88% to 2.27%. GlitchPatch achieves a 0.00% RR in full-vocabulary evaluation and leaves rule-unmatched inputs unchanged by design. We further evaluate the practical impact of repair from the perspectives of time cost, language understanding, and capability, supporting its deployment feasibility. We hope this work provides a practical option for improving tokenizer reliability.
|
| 478 |
XTurnix: Large-Scale Self-Supervised Turn Control through Two-State Binary Decisions
2610.04400
|
cs.CLcs.AI
|
Zhanxun Liu, Yifan Duan, Hengtao Wu, Chen Yang, Qinyuan Cheng |
General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces an...General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on isolated utterances, making them difficult to use as a unified causal controller with comprehensive context. We propose XTurnix, a compact text-based model that formulates turn control as two binary decisions conditioned on the AI's current listening or speaking state and predicts a single control token from the complete dialogue history. XTurnix is pretrained on 5.5 million causal action examples automatically derived from timestamped two-speaker transcripts, then fine-tuned on synthetic multi-turn examples with a flatter distribution across the four state-action labels. We evaluate XTurnix on four public benchmarks and a balanced self-curated benchmark. Across the public benchmarks, XTurnix achieves the best results on all SemanticVAD and LiveKit splits, ties the native Smart-Turn model on Smart-Turn Bench, and achieves the highest incomplete-turn accuracy on Easy-Turn. On the self-curated benchmark, it reaches 89.06% accuracy, more than 20 percentage points above the strongest third-party baseline at 68.75%, while maintaining F1 scores between 84.21% and 90.63% across all four categories. These results demonstrate unified listening- and speaking-state turn control in a single compact model. Code is available at https://github.com/xcc-zach/xturnix, with an interactive demo at https://huggingface.co/spaces/xcczach/xturnix-demo.
|
| 479 |
Ideological Stance Detection in a Low-Resource Language: Polarization in Bangladeshi Public vs Private University Discourse on Social Media
2610.04401
|
cs.CL
|
Safaruzzaman Shovo, Monowar Islam, Asif Hossain, Sameya Akhter, Md. Shamsul Islam |
Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource ...Public vs. private universities is a debatable issue, and it creates polarization on social media in Bangladesh. Debate on quality, jobs, and prestige is passionate among the students, parents, and graduates, the majority of whom speak Bengali, a low-resource language. To measure this polarization, this paper introduces a manually annotated dataset of 4,060 Bengali comments labeled as Pro-Public, Pro-Private, or Neutral. We evaluated the quality of our annotations by Fleiss's Kappa agreement that was 0.89, corresponding to a high agreement among annotators. The classical ML (SVM, Random Forest, XGBoost), BiLSTM network, hybrid BanglaBERT+XGBoost models and the state-of-the-art zero-shot LLMs (Claude Sonnet 4, DeepSeek-V3.1, Llama 4 Maverick, Kimi K2 Thinking, Qwen3-235B Thinking) models are evaluated. The accuracy of BanglaBERT+XGBoost is 91.81% and macro F1 score is 91.70%, which is higher than all the supervised baselines. The zero-shot Llama 4 Maverick Thinking achieves a macro F1 of 0.931 (overall accuracy of 93.31%) without any fine-tuning. All machine learning (ML), deep machine learning (DL) and transformer models were outperformed by the zero-shot Llama 4 Maverick model. Polarization also is evident, in some ways more clearly in the Pro-Private comments, which emphasize modern facilities and timely graduation, versus the Pro-Public comments, which emphasize affordability and government jobs. Our findings open new directions for analyzing social media polarization in low-resource languages.
|
| 480 |
Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define
2610.04403
|
cs.CLcs.LG
|
Saurabh Kumar Singh, Yogeshwar Singh Dadwhal, Malhar Vedak |
Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 famili...Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, evaluated down their published ladder to Q2_K (about 2.6 bits per weight), on 429 frequency-validated rare English words under two probes: surface inclusion of a prompt-supplied word and its one-sentence definition, scored by a tiered multi-synonym matcher, its error measured by a blind LLM-judge census of every definition, with human verification. Three regimes emerge at Q2: total collapse into unusable builds, severe semantic dissociation in sub-2B models, and mostly robust preservation above about 3B. In every sub-2B artifact, definitions fall 20-67% below the artifact's baseline, typically several times the inclusion loss. Two controls separate rarity from task difficulty: within the rare set, loss rises with rarity in six of seven sub-2B artifacts, and on a 100-word common-word set rare words lose more than common words in all eight, significantly in six. Tokenizer vocabulary size does not predict the damage (Spearman rho=0.12); parameter count dominates (rho=0.72), confirmed within five of six same-tokenizer families. Q4_K_M remains lexically clean at >=1B. The damage is frequency-graded, provider-dependent, and not calibrated by WikiText-2 perplexity: across nine artifact-matched ladders, near-identical Q2 penalties (44.7%/47.6%) separate an artifact keeping its definitions (3.6%) from one losing them (43.6%). Aggressively quantized small models can keep generating fluent text while no longer knowing what it means, risking hardware-constrained deployments in domains where semantics carries consequences. Validation must be per artifact.
|
| 481 |
Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents
2610.04409
|
cs.CLcs.AI
|
Peigui Qi, Kunsheng Tang, Yide Song, Weiming Zhang, Nenghai Yu |
Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorrect tools or generating invalid calls. Existing mitigation methods report substantial impro...Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorrect tools or generating invalid calls. Existing mitigation methods report substantial improvements, yet we identify a previously overlooked failure mode that we term Hallucination Escape. These methods reduce hallucination on the tool configuration they are tuned on but increase it on other configurations, canceling out the gain. We further investigate this phenomenon and find that hallucination rises sharply when a model's intrinsic tool-use tendencies conflict with the current tool configuration, and that existing methods reinforce rather than suppress these tendencies, which in turn contributes to hallucination escape. Building on these findings, we propose EscapeGuard, a training-free inference-time method that combines conflict-aware gating with configuration-derived attention enhancement to mitigate tool hallucination while preventing hallucination escape. Across six benchmarks on various models, EscapeGuard reduces tool-selection hallucination by 9.0 pp and suppresses hallucination escape, lowering the cross-configuration mean by 23.7 pp and achieving an 89.1% net improvement in paired-query evaluation. We hope this work can encourage evaluation beyond a single tool configuration and pave the way for more reliable tool-using LLM agents.
|
| 482 |
DV-Lens: Revealing the Functional Organization of Language Model Parameters
2610.04489
|
cs.CLcs.LG
|
Chenhang Cui, Jian Yu, Shuyi Miao, Xiaohao Liu, Rui Huang |
Understanding parameter functions helps elucidate the internal mechanisms of large language models (LLMs). However, how to connect parameters from different modules to verifiable output effects and further characterize the relationship between their functional...Understanding parameter functions helps elucidate the internal mechanisms of large language models (LLMs). However, how to connect parameters from different modules to verifiable output effects and further characterize the relationship between their functional organization and model capability remains to be explored. To this end, we introduce the downstream vocabulary lens (DV-Lens), a parameter-level interpretability framework that links native parameter directions to their downstream vocabulary responses. Specifically, we first estimate module-specific downstream Jacobians over a reference prompt set for attention query, key, value, and output (Q/K/V/O) projections and feed-forward networks (FFNs). Second, we use these mappings to project native parameter columns into the final vocabulary space, obtaining signed readouts that characterize their average local output responses. Third, we group parameter columns by their vocabulary readouts and introduce downstream vocabulary complexity (DV-Complexity), which quantifies within-group structural variation using normalized reconstruction residuals of the original weights. At the parameter level, randomized controls and finite-difference tests show that DV-Lens readouts capture non-random vocabulary structure and predict local logit changes with 98.0% coordinate-orientation agreement across 720 cases from nine models. These readouts further guide parameter ablation, steering, and swapping across 21 models, shifting target-token probabilities in the predicted directions under controlled conditions. At the model level, the joint-parameter score of DV-Complexity achieves a Spearman correlation of 0.904 with benchmark-based capability rankings across 48 language models. Together, these results provide intervention-based evidence for DV-Lens interpretations and reveal an association between DV-Complexity and model capability.
|
| 483 |
Emoji-Emotion Ranking System Using Twitter Data
2610.04495
|
cs.CL
|
Danila Khlebokazov, Nurkhan Tashimov, Pakizar Shamoi |
Nowadays, emojis are often replacing words. Yet computational systems still oversimplify them. Most existing approaches treat emojis as static sentiment indicators and overlook their emotional distributions. In this study, we propose an emoji-aware emotion ana...Nowadays, emojis are often replacing words. Yet computational systems still oversimplify them. Most existing approaches treat emojis as static sentiment indicators and overlook their emotional distributions. In this study, we propose an emoji-aware emotion analysis framework based on a Twitter (X) dataset of 100,000 emoji-containing replies collected between 2020-2025. After preprocessing and text cleaning, we applied text-to-emotion classification to detect five primary emotions (Happy, Angry, Sad, Fear, and Surprise) for each message. By aggregating emotion scores across contexts in which each emoji appears, we estimate emoji-emotion association distributions and construct an emoji-emotion ranking system reflecting relative emotional dominance. Furthermore, we project emojis into the Russell valence-arousal space to enable continuous affective interpretation. Our results demonstrate that emojis exhibit probabilistic, context-sensitive emotional profiles rather than fixed sentiment polarities.
|
| 484 |
Correctness Is a Direction: Geometric Answer Selection in Language Models
2610.04512
|
cs.CL
|
Marcus Armstrong, Navid Ayoobi, Pradham Mummaleti, Alexander Chulzhanov, Arjun Mukherjee |
Answer correctness is encoded as a recoverable geometric direction in the hidden states of language models. We show that the mean displacement from incorrect to correct answer representations, computed at approximately 70\% of model depth from fifty labeled ex...Answer correctness is encoded as a recoverable geometric direction in the hidden states of language models. We show that the mean displacement from incorrect to correct answer representations, computed at approximately 70\% of model depth from fifty labeled examples with no parameter updates, yields a scoring direction that outperforms zero-shot log-probability scoring by up to +32.0 percentage points on factual benchmarks (ARC-Challenge and MMLU) and by +38.1 to +51.8 percentage points on TruthfulQA, across five models spanning 1B to 8B parameters in three architecture families (Llama, Qwen, Gemma). The method requires one forward pass and one dot product per candidate; no generation is performed at inference. Applied as a hallucination detector on individual (question, answer) pairs, the recovered direction achieves 0.693~AUROC versus 0.578 for log-probability scoring. We additionally find that correctness directions for factual reasoning, domain knowledge, and calibrated truthfulness are near-orthogonal in representation space, revealing that language models allocate geometrically independent subspaces to qualitatively distinct notions of correct answer, with architecture-dependent variation in the degree of separation. This structure explains the observed transfer pattern---the direction calibrated on factual questions transfers within task type but not across it---and suggests that LLM calibration failures may reflect a routing problem: the model's internal representation contains more correctness signal than its output behaviour exploits.
|
| 485 |
From Probe Scores to Alarm Policies: Operational Validity of Activation Monitors for Language-Model Agents
2610.04575
|
cs.CLcs.AI
|
Xueping Gao |
Activation probes can predict safety-relevant properties of language models with high area under the receiver-operating-characteristic curve (AUROC), but deployed agent monitors make thresholded alarm decisions under tight false-alarm budgets. These are differ...Activation probes can predict safety-relevant properties of language models with high area under the receiver-operating-characteristic curve (AUROC), but deployed agent monitors make thresholded alarm decisions under tight false-alarm budgets. These are different estimands. We introduce an Operational Validity Contract that fixes a monitor's target, observability, identity, timing, intervention unit, comparator, calibration, and cost. We formalize risk at the semantic request or trajectory level: when one task contains repeated alarm opportunities, row-level AUROC and false positive rate do not identify semantic-unit any-alarm risk. A confidence-certified threshold also requires enough independent negative units, a tie-safe rule, and transport to deployment. Across Models Under Pressure, LASR refusal prediction, immutable AgentDojo, and a prospectively protocol-frozen ST-WebAgentBench replication, joint activation-observable monitors reach AUROC 0.957 and 0.935 on the first two benchmarks, yet their locked 5%/10% detection rates are only .642/.742 and .719/.782, respectively. On AgentDojo, the secondary mean-activation rollout monitor reaches AUROC 0.922 but detects none of 38 positive semantic cases at the locked 5% operating point; thresholds intended for 10% false alarms realize 18.7-20.0% on test. Because the test misses its prospectively frozen 40-positive support gate, we label it support-insufficient. On ST-WebAgentBench, activation reaches AUROC .874, but 23 independent calibration negatives cannot identify even a 10% controller; the locked policy abstains rather than reporting its mechanical zero FPR as a success. An exploratory counterexample also lowers full AUROC while improving realized 10% utility. The fail-closed compiler caps MUP and LASR at restricted predictive value and AgentDojo and ST-Web at representation accessibility; no setting reaches alarm-policy validity.
|
| 486 |
SIFT: Robust Meta-Faithfulness Verification of Chain-of-Thought Reasoning Under Distribution Shift
2610.04594
|
cs.CLcs.LG
|
Noor Islam S. Mohammad, Md. Basim Al Zabir Shammo, Hasan Siddiki, Mahmudul Hasan, Md. Faisal Sheikh |
Chain-of-Thought (CoT) faithfulness detectors are widely used to audit reasoning models, yet a detector is itself a predictor whose verdicts are treated as stable properties. We ask whether a detector is faithful to itself under distribution shift. We formaliz...Chain-of-Thought (CoT) faithfulness detectors are widely used to audit reasoning models, yet a detector is itself a predictor whose verdicts are treated as stable properties. We ask whether a detector is faithful to itself under distribution shift. We formalize meta-faithfulness as an invariance principle: a valid detector must return identical verdicts on traces that differ only by transformations preserving ground-truth faithfulness. We prove three results: (i) no detector using only intervention-response profiles can separate faithful from epiphenomenal mechanisms with identical signatures; (ii) any detector relying on shift-sensitive features violates invariance at a rate independent of its in-distribution accuracy; (iii) an asymptotic certified selective-risk guarantee enables confident abstention. We operationalize the principle in FaithShift, a stress-test protocol spanning ten shift axes, and propose SIFT, a hidden-state trajectory detector trained with cross-environment invariance objectives and certified abstention. Across 14,996 traces, four domains, and eight models, three findings emerge. First, transfer collapse is real: all existing detectors show gaps $\geq 0.15$ AUROC. Second, the dominant bottleneck is sampling stochasticity, not shift: over 80% of detector instability stems from random seed variation, falsifying our preregistered prediction that shift-attributable violations exceed 0.25. Third, SIFT cuts invariance violations by 64% over the best single-seed baseline, but a four-seed ensemble of any detector narrows the margin to 0.01 (indistinguishable at matched coverage, $p=0.21$), and SIFT needs a 51% abstention rate. Cross-model transfer degrades from within-family to cross-family to open-weight-to-API, partly closed by multi-model training. We offer a framework for auditing auditors: the real barrier is detector variance, not distribution shift.
|
| 487 |
Stance Drift: How AI-mediated Communication Distorts Our Message
2610.04620
|
cs.CL
|
Lingchong Liu, Yanfei Zhou, Jacob Bien, Y. X. Rachel Wang, Lucy Xia |
Large language models (LLMs) increasingly mediate human communication, from drafting emails to summarizing scientific reports, yet whether they faithfully preserve a speaker's position remains largely untested. We model AI-mediated communication as a two-step ...Large language models (LLMs) increasingly mediate human communication, from drafting emails to summarizing scientific reports, yet whether they faithfully preserve a speaker's position remains largely untested. We model AI-mediated communication as a two-step generation-extraction pipeline: one LLM produces an argument from a specified stance, and a second LLM extracts the stance from that argument. We represent the pipeline as a probabilistic state transition over five Likert-type stance categories and define the stance preservation rate (SPR) as the average probability that the extracted stance matches the initial stance. Across 112 debate propositions, none of the nine LLMs tested exceeded an SPR of 0.7 under the default configuration. Three drift patterns accounted for most of the drift: polarization, deviation from neutrality, and flipping. Among the mitigation strategies tested, including in-context learning, multiple extraction with shuffled options, assertion, and reflection, only adding medium reasoning effort to a reflection prompt for GPT-5.4 substantially improved the SPR, to 0.775, yet polarization remained the largest pattern, with 0.119 of the transition mass. An exploratory comparison with human extraction on a single proposition suggests that drift arises at both the generation and the extraction stage. These results point to a fidelity gap in AI-mediated communication, with implications for journalism, policy deliberation, scientific communication, and other domains where opinion-laden messages pass through language models.
|
| 488 |
Does Neural Complexity Improve Health Misinformation Detection? A Leakage-Controlled Cross-Corpus Benchmark
2610.04636
|
cs.CLcs.LG
|
Mkululi SIKOSANA |
Increasing architectural complexity is often assumed to improve health misinformation detection, yet reported gains are difficult to interpret when studies use different corpora, preprocessing pipelines, data splits, and leakage controls. This study provides a...Increasing architectural complexity is often assumed to improve health misinformation detection, yet reported gains are difficult to interpret when studies use different corpora, preprocessing pipelines, data splits, and leakage controls. This study provides a controlled cross-corpus benchmark of five compact neural architectures (1D-CNN, LSTM, BiLSTM, CNN-LSTM, and CNN-BiLSTM), a soft-voting neural ensemble, and three classical machine-learning baselines using COVID19-FNIR and CONSTRAINT. Exact-text duplicate controls were applied before modelling; all neural systems used a common preprocessing and optimisation protocol, and neural results were repeated across three random seeds. On COVID19-FNIR, the deep ensemble achieved a mean macro-F1 of 0.9963 and ROC-AUC of 0.9994, while individual neural models ranged from 0.9945 to 0.9957 macro-F1. On CONSTRAINT, the ensemble achieved macro-F1 of 0.9272 and ROC-AUC of 0.9811, whereas a linear SVM achieved macro-F1 of 0.9574 and ROC-AUC of 0.9931. Architecture rankings changed across corpora, and simple sparse linear models remained highly competitive. The findings show that model complexity does not provide a stable performance advantage and that benchmark construction can dominate architecture choice. The study contributes a reproducible, leakage-controlled basis for evidence-driven model selection in health misinformation classification
|
| 489 |
Grounding Probes: Generator-Independent Hallucination Detection from Observer Model Hidden States
2610.04642
|
cs.CL
|
Michael Rathmayr, \'Ad\'am Kov\'acs, G\'abor Recski |
Detecting responses that retrieval-augmented generation does not ground in its context trades speed against accuracy: surface checks miss paraphrased fabrication, sampling-based methods cost extra generations. Hidden-state probes sit between the two, but every...Detecting responses that retrieval-augmented generation does not ground in its context trades speed against accuracy: surface checks miss paraphrased fabrication, sampling-based methods cost extra generations. Hidden-state probes sit between the two, but every existing one reads the generating model's own activations, so a change of generator invalidates the detector and a closed-weight generator is out of reach. This paper removes that coupling. The Grounding Probe is logistic regression over the mean-pooled middle-layer hidden states of an observer language model that reads the context, question, and response in one forward pass and generates nothing, with the recipe it needs: pool over response tokens, read a middle layer, and control capacity, which closes the train-test AUROC gap from 0.087-0.202 to 0.009-0.013. Asking the observer outright, rather than reading its hidden state, costs at least +0.166 AUROC in every one of four models. Fitted on 15,090 annotated responses it reaches 0.879-0.894 AUROC on RAGTruth test across four observers, and 0.924 AUROC with 0.820 F1@0.5 averaged with a supervised span detector, 0.060 above that detector alone. One probe holds across six generators, and hold-out controls, including one in which no evaluation prompt appears in training, bound the cost of removing a generator at about 0.02 AUROC. Code, probes, and predictions are released.
|
| 490 |
Extracting Persona Subspaces Through Iterative Nullspace Projection For Modulation
2610.04676
|
cs.CLcs.AI
|
Ananya Malik, Mai ElSherief |
Large Language Models (LLMs) can adopt distinct personas to tune their semantics, expertise, and perspective to different users and tasks. Precise control over these traits is critical to ensure safety and reliability in model behavior. Existing methods like a...Large Language Models (LLMs) can adopt distinct personas to tune their semantics, expertise, and perspective to different users and tasks. Precise control over these traits is critical to ensure safety and reliability in model behavior. Existing methods like activation steering and prompt-based persona induction reduce a persona to a single dominant direction, missing the finer, nested traits that emerge only once that dominant signal is factored out. We introduce modulation as a setting where the persona context is already embedded in the content being manipulated, requiring control methods to amplify or suppress a trait already present rather than inject it from scratch. PaSS is an inference-time control paradigm that models personas as multi-dimensional subspaces in a model's latent space without supervised contrastive examples. The persona subspaces are extracted via iterative concept erasure and applied to modulate persona-guided generation without retraining. To extract this subspace, we use Iterative Nullspace Projections (INLP) to linearly and iteratively isolate persona-specific directions. We causally evaluate six personas against diverse tasks like MATH-500, TinyAlpaca, GSM8K, and IFEval, showing that discriminative, iterative subspace extraction captures diverse traits underlying a given persona, enabling stronger and larger modulation than single-direction additive methods, while maintaining content fidelity. We further study individual peeled directions within each subspace to uncover the distinct aspects of persona behavior they encode. Overall, we show that persona subspaces offer a controllable, interpretable, and generalizable framework for modulating LLM behavior without sacrificing task performance.
|
| 491 |
Steering Speech-Language Models: Training-Free Task Specialization via Contrastive Activation Addition
2610.04683
|
cs.CLeess.AS
|
S\'everin Baroudi, Yanis Labrak, Pierfrancesco Melucci, Sergio Burdisso, Petr Motlicek |
Activation steering has proven effective for controlling the behavior of Large Language Models (LLMs) at inference time, but its application to SpeechLLMs remains new, and training-free steering approaches for such models are still largely unexplored. We propo...Activation steering has proven effective for controlling the behavior of Large Language Models (LLMs) at inference time, but its application to SpeechLLMs remains new, and training-free steering approaches for such models are still largely unexplored. We propose a training-free Contrastive Activation Addition (CAA) protocol that derives steering vectors for common speech tasks (e.g. transcription) in SpeechLLMs from a small number of labeled utterances. We showcase that adding these vectors in the representation space, at inference time, enforces better the targeted speech task. We further show that, when combined with prompting, these vectors yield to consistent improvement over prompting alone on most evaluated tasks such as Automatic Speech Recognition (ASR) or Emotion Recognition (ER), and transfer to out-of-domain data. We additionally demonstrate the usefulness of script-normalization directions to enforce the target script of a specific language.
|
| 492 |
Understanding Errors in LLM-Based Question Answering over Imperfect Tables
2610.04687
|
cs.CL
|
Baowen Zhang, Wei Fan, Ruman Wang, Hangting Ye |
We investigate error discovery and handling in question answering over imperfect tables through controlled studies across three large language models (LLMs) on human-reviewed RADAR-T examples. Answering questions over these tables requires handling errors that...We investigate error discovery and handling in question answering over imperfect tables through controlled studies across three large language models (LLMs) on human-reviewed RADAR-T examples. Answering questions over these tables requires handling errors that can affect the answer. We vary row order and compare original, error-marked, and repaired tables to test whether discovery depends on where errors appear and whether providing their locations is sufficient for accurate question answering. First, reordering rows changes error discovery even when the table contents and gold answer remain unchanged. Complete discovery is higher for back than front placements and, averaged over the tested mean positions, for compact than widely spaced layouts. Second, providing verified error locations alone is insufficient for accurate QA, leaving a substantial accuracy gap between error-marked and repaired tables. Providing tables with human-reviewed repairs already applied raises code-assisted QA accuracy by 39.0-59.1 percentage points over the error-marked tables across the three systems. GBDI, a simple workflow, puts these findings into practice by combining error discovery across shuffled table views with explicit guidance for verifying and handling the reported errors. On RADAR-T, GBDI raises observed QA accuracy by 3.8-18.5 percentage points over a code-agent baseline across five systems. These results highlight the importance of both reliable error discovery and effective error handling in question answering over imperfect tables. Our anonymous repository is available at https://anonymous.4open.science/r/GBDI-ICLR-2027-85BD/
|
| 493 |
WNet: Discrete Wavelets Transform for Efficient Token Mixing
2610.04720
|
cs.CL
|
Rana Aref Salama, Abdou Youssef, Mona Diab |
In a Transformer, token mixing is the step that lets each token draw information from other tokens, and it dominates the cost of encoding long sequences. Self-attention does this mixing very well: every token weighs every other token by content, which gives st...In a Transformer, token mixing is the step that lets each token draw information from other tokens, and it dominates the cost of encoding long sequences. Self-attention does this mixing very well: every token weighs every other token by content, which gives strong contextual modeling. That all-pairs comparison is also why its cost grows quadratically with sequence length. We introduce WNet, a Transformer encoder that replaces self-attention with token mixing based on the discrete wavelet transform (DWT). Three attention-free mixers recombine the scales: by linear fusion, by learned gating, or by letting each token choose its scales. A hybrid adds self-attention in the last layer only. A receptive-field analysis shows that wavelet mixers built from two-tap filters, such as Haar, never relate tokens outside fixed blocks, however deep the network, even when the filters are learned. Longer filters reach the whole sequence within two layers. We pre-train every model with masked language modeling on a fixed-token subset of C4 and fine-tune on GLUE, using one controlled setup with size-matched BERT and FNet baselines and a control that cannot mix tokens. The token-gated mixer trains as fast as attention at 256 tokens and 2.7 times faster at 4,096.
|
| 494 |
Verb-ICL: Rethinking In-Context Learning for Structured Prediction
2610.04725
|
cs.CL
|
Fan Bai, Hengshuo Miao, Sanjit S Batra, Hamid Reza Hassanzadeh, Ardavan Saeedi |
Structured prediction tasks pose unique challenges for in-context learning (ICL): their compositional outputs require modeling fine-grained, token-level patterns that sentence-level approaches fail to capture, and their task-specific annotation conventions are...Structured prediction tasks pose unique challenges for in-context learning (ICL): their compositional outputs require modeling fine-grained, token-level patterns that sentence-level approaches fail to capture, and their task-specific annotation conventions are human-defined artifacts that cannot be acquired through pretraining alone. We propose Verb-ICL, a selective annotation framework for ICL-based structured prediction that addresses both challenges. Verb-ICL first selects representative examples using a token-level coverage strategy that captures local semantic patterns critical for structured prediction, then generates actionable error feedback that codifies task-specific annotation guidelines and incorporates this feedback into ICL demonstrations. We evaluate Verb-ICL on six structured prediction datasets spanning information extraction and semantic parsing. Experiments with recent LLMs show that Verb-ICL consistently outperforms strong selective annotation baselines under low-resource settings and continues to provide gains as the annotation budget increases. Extended analyses demonstrate that the generated feedback is predominantly useful across a four-category quality taxonomy, generalizes as task-level guidance beyond instance-specific corrections, and improves performance regardless of the underlying selection strategy.
|
| 495 |
More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
2610.04753
|
cs.CLcs.AI
|
Noam Elata, Itay Lamprecht, Mikey Shechter, Daniel Ohayon, Itay Hubara |
utoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe ...utoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while retaining more value heads preserves capacity with limited additional decoding cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple sparse attention method designed to study the interaction between sparsity and head-count asymmetry. We formalize the benefits of this asymmetry theoretically and validate them empirically through latency measurements and quality evaluations on models up to 1.5B parameters. Together, SAGA and Atop-N achieve end-to-end decoding speedups exceeding $2\times$ over our full-attention GQA baseline at long contexts. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants on the evaluated benchmarks. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture, enabling practitioners to benefit from our approach without costly retraining.
|
| 496 |
How Do People Challenge Racial Stereotypes Online? Counter-Story Detection Across Reddit Communities
2610.04803
|
cs.CL
|
Uma Sushmitha Gunturi, Jimin Mun, Maarten Sap, Maria Antoniak |
Counter-storytelling is a powerful mechanism people use to challenge dominant narratives. Unlike other forms of counterspeech that have been widely studied in computational social science, counter-storytelling has largely been overlooked. Counter-stories are d...Counter-storytelling is a powerful mechanism people use to challenge dominant narratives. Unlike other forms of counterspeech that have been widely studied in computational social science, counter-storytelling has largely been overlooked. Counter-stories are difficult to detect automatically; they are relational (defined with respect to expressions of racial stereotypes) and structurally diverse (drawing on stories that describe lived experiences, witnessed events, exemplars, and hypotheticals). We introduce a first framework for detecting and characterizing counter-storytelling against racial stereotypes at scale. This includes (1) a three-dimensional taxonomy grounded in narratology and Critical Race Theory and (2) a multi-stage pipeline that identifies relational pairs of stereotypes and counter-stories in noisy Reddit discourse. Using this pipeline, we annotate 25,549 Reddit posts across 615 communities and identify 1,312 counter-stories. Our analysis shows that speaker identity and post context shape how counter-stories are told. For example, in-group writers favor first-person testimony, often adopting the role of self-reflective insiders. Our work shows how computational methods can scale qualitative approaches to identify and characterize counter-storytelling as a contextual narrative practice, with implications for content moderation, narratology, and racial discourse analysis.
|
| 497 |
Viva La Vida: Verification and Accumulation Failures in Multi-Agent Proof Search
2610.04829
|
cs.CLcs.AI
|
Benji Xu, Ken Zheng, Noah Han |
When an agentic prover works on an open problem, there is no proof assistant to fall back on: its verifier and lemma library are ultimately language models judging model outputs. We instrumented such a system end to end and analyzed $51{,}754$ traced observati...When an agentic prover works on an open problem, there is no proof assistant to fall back on: its verifier and lemma library are ultimately language models judging model outputs. We instrumented such a system end to end and analyzed $51{,}754$ traced observations across three full runs ($186$ hours, \$$5{,}694$). We find three connected failure modes. First, the three-model verifier requires unanimity and treats parse or API failure as non-approval; in $10$ of $12$ verification events, one member returned no parseable output or an API error, making acceptance arithmetically impossible without surfacing an error. Second, when the ensemble did function, one verifier approved $3$ attempts that GPT rejected, each claiming to resolve the open problem; a single-verifier design would therefore have announced a solution three times. Third, because nothing could be approved, every review was a refutation, yet the lemma extractor mines reviews as well as proofs: $24$ of $93$ lemmas ($26\%$) were extracted from rejected arguments with their refutational context removed. Taken together, these findings show that without external verification, supervision is itself a critical trust boundary: systems must distinguish abstention from rejection, preserve useful disagreement, and preserve the provenance and polarity of information before it becomes future context.
|
| 498 |
Cluster Validation Indices as Self-Supervised Objectives for Text Representation Learning
2610.04830
|
cs.CLcs.AI
|
Kishor Kumar Bhaumik, Nicolas Roque dos Santos, Neil Shah, Jia Chen, Evangelos E. Papalexakis |
Self-supervised fine-tuning refines the embedding space of a pretrained language encoder without labels. However, the commonly used approaches are computationally expensive. Specifically, contrastive learning-based methods need multiview data and in-batch nega...Self-supervised fine-tuning refines the embedding space of a pretrained language encoder without labels. However, the commonly used approaches are computationally expensive. Specifically, contrastive learning-based methods need multiview data and in-batch negative examples, while negative-free approaches require auxiliary graphs/networks. An interesting question arises: can self-supervised fine-tuning be done without relying on either additional negatives or graph data? To answer this question, we introduce SilK (Silhouette-guided K-means), which trains on a Cluster Validation Index, an internal measure of cluster quality without using labels. SilK clusters the corpus and then regresses a simplified silhouette toward a target value. Each document is compared only against the k cluster centroids, never against other documents, so the method needs no augmentation, no negative pairs and one view per document. On BERT-base, SilK trains 1.46x faster per epoch than the fastest baseline we evaluate and uses 45.4% less peak GPU memory than the leanest one. Under frozen-encoder linear probing, SilK stays competitive with the best baselines on three downstream tasks.
|
| 499 |
No Hindsight for LLM Fact-Checkers: Measuring Leakage Channels in Misinformation Detection
2610.04888
|
cs.CL
|
Kuan-Hua Wu Lu, Yohanes Andre Setiawan |
As automated fact-checking scales on social media, large language model (LLM) verdict scores can look stronger than warranted. One reason is that evaluations mix in information that was not knowable at claim time. Two channels are easy to conflate: outcomes me...As automated fact-checking scales on social media, large language model (LLM) verdict scores can look stronger than warranted. One reason is that evaluations mix in information that was not knowable at claim time. Two channels are easy to conflate: outcomes memorized in pre-training and retrieved evidence published after the claim. Yet standard benchmarks rarely separate the two. In this study we measure both channels on AVeriTeC and QuanTemp++ by reconstructing point-in-time evidence conditions and probing for outcome information encoded in model representations. We find substantial evidence of parametric leakage, that can be hidden by the aggregate accuracy, while a simple representation bottleneck reduces this future leakage more efficiently than a mutual-information-based training penalty. We also find that allowing post-claim evidence inflates zero-shot accuracy by 6.3 points in AVeriTeC while the effect is negligible in QuanTemp++, where retrieval provides little post-claim evidence. These results show that misinformation benchmarks can overstate fact-checking performance when they do not account for what information was actually available at claim time.
|
| 500 |
A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
2610.04898
|
cs.CLcs.AI
|
Hongao Zhu (Department of Linguistics, University of California San Diego), Muxiaoqiao Xu (School of Foreign Languages, Shanghai Jiao Tong University), Yikang Liu (School of Computer Science |
This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora ...This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times. We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese. Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO). Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times. However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO. Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to $n$-gram statistics are better predictors of reading. All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.
|
| 501 |
Rewrite What Matters: Adaptive Multilingual Query Rewriting for Reasoning via Agentic Reinforcement Learning
2610.04899
|
cs.CL
|
Rui Qi, Yufeng Chen, Yunlong Liang, Chuan Meng, Sijin Lu |
In multilingual scenarios, queries with equivalent semantics but in different languages could guide the model into different reasoning trajectories, leading to performance disparities. To mitigate this gap, previous studies typically apply a one-size-fits-all ...In multilingual scenarios, queries with equivalent semantics but in different languages could guide the model into different reasoning trajectories, leading to performance disparities. To mitigate this gap, previous studies typically apply a one-size-fits-all query rewriting strategy, such as translation, which overlooks the fact that different scenarios require diverse types of semantic transformations. In this paper, we propose mRewriter-R1, an agentic multilingual query rewriting framework with reinforcement learning. Unlike single-turn rewriting, mRewriter-R1 formulates multilingual query rewriting as a multi-turn sequential decision-making process, where the model dynamically performs multi-aspect optimization through adaptive operator selection. Experimental results demonstrate that mRewriter-R1 outperforms all strong multilingual rewriting baselines on different large reasoning backbones. Further analyses show that the learned policy can adaptively decide on rewriting operators according to query characteristics, exhibiting strong generalization ability across diverse reasoning tasks, and plug-and-play compatibility with heterogeneous reasoning language models.
|
| 502 |
Scaling Verifiable Environments for Long-horizon Work Agents
2610.04906
|
cs.CLcs.AI
|
Jiazheng Zhang, Long Ma, Yunxian Yang, Zhiheng Xi, Zhikai Lei |
Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering ov...Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis methods sacrifice workspace complexity, realism, or grounded verifiability. To bridge this gap, we introduce WorkForge, a scalable synthesis framework for constructing verifiable work-agent environments from real-world resources. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files and organizes them into a workspace. WorkForge inspects the workspace to extract concrete, checkable facts about its content. These factual anchors fix which task types the workspace can support and how their outcomes can be verified. Therefore, WorkForge derives each task's instructions, solution plan, and complementary programmatic and semantic verifiers directly from these factual anchors, keeping verification traceable to observable workspace evidence. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance and outperform strong competitors. Our analyses confirm the efficacy of the proposed method and reveal consistent scaling behaviors across both data volume and interaction horizons.
|
| 503 |
From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deployment
2610.04956
|
cs.CL
|
Zeshen Zhang, Han Zhao, Weihao Cui, Quan Chen, Yu Liu |
As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services. However, these resource-constrained servers...As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services. However, these resource-constrained servers struggle to guarantee heterogeneous Service Level Objectives (SLOs) when serving multiple LoRA-adapted services simultaneously. Existing serving frameworks suffer from severe SLO violations due to the computational overhead of LoRA layers and the rigid nature of batch scheduling. To address this, we propose HALO, a scheduling method tailored for LoRA-assisted on-premise LLM deployment. HALO introduces two key innovations: a spatial multiplexing strategy that overlaps Base and LoRA computations by partitioning GPU Streaming Multiprocessors (SMs), and an SLO-aware scheduler that decouples request execution based on "request-level slack." By prioritizing urgent tasks and utilizing idle budget for traffic shaping, HALO significantly mitigates resource contention. Our evaluation demonstrates that HALO minimizes SLO violations while improving throughput compared to state-of-the-art baselines.
|
| 504 |
Building LLM Agent Systems the Deep Learning Way: From Modular Design to Architecture Search
2610.04961
|
cs.CL
|
Tao Feng, Pengrui Han, Zhongjie Dai, Jiaxuan You |
Large Language Models (LLMs) have revolutionized AI research and enabled exciting agent systems. To build a complex LLM agent system, most existing research relies on insights from other domains or heuristics to manually build the agent system. However, this a...Large Language Models (LLMs) have revolutionized AI research and enabled exciting agent systems. To build a complex LLM agent system, most existing research relies on insights from other domains or heuristics to manually build the agent system. However, this approach often requires heavy hand-engineering and fails to fully optimize for the downstream task of interest. Inspired by the tremendous success of deep learning, we propose to construct LLM agent systems in a modular manner, similar to building a deep neural network. Our key insight is to make analogies between LLM building blocks, such as retrievals, memories, and prompting strategies, and the successful deep learning modules, such as MLPs, attention, and recurrent modules. We further design forward inference and feedback mechanisms for LLMs, where prompts in LLMs are considered as the weights in deep models, and the prompt optimization from feedback is analogous to the back-propagation algorithm. We additionally leverage a search algorithm to search for the best configuration of LLM agent systems, similar to the neural architecture search (NAS) in deep learning research. Comprehensive experimental results demonstrate that the proposed deep learning recipe for LLM agent systems is highly effective, in particular: (1) Organizing LLM modules into deep-learning-style architectures yields noticeable performance gain; (2) Automatic prompt optimization, equivalent to backpropagation, is efficient in incorporating feedback from the task of interest and achieves at least 5% performance improvement; (3) NAS equivalent algorithm works well for further optimizing the LLM agent system architecture with 11% performance gain compared with randomly designed architectures. Overall, our research demonstrates the exciting opportunity of transferring the success of deep learning to building LLM agent systems.
|
| 505 |
TrajLong: Co-Designing Agentic and Long-Context Supervision for Mid-Training
2610.04973
|
cs.CL
|
Miao Peng, Qintong Zhang, Nuo Chen, Yuhan Li, Guochen Yan |
LLM agents for coding, search, and workplace tasks increasingly rely on long-context capabilities to effectively aggregate and reason over extended interaction histories. Recent work has incorporated agent trajectories into mid-training stage, drawing on their...LLM agents for coding, search, and workplace tasks increasingly rely on long-context capabilities to effectively aggregate and reason over extended interaction histories. Recent work has incorporated agent trajectories into mid-training stage, drawing on their naturally long and interaction-rich structure. Yet how to organize the information within these trajectories into effective mid-training supervision remains underexplored. In this work, we investigate the relationship between long-context and agent atomic capabilities and introduce TrajLong, a novel framework that compiles trajectories into long-context training tasks with dense supervision, targeting three representative atomic capabilities: evidence grounding, cross-evidence aggregation, and temporal state maintenance. We mid-train Qwen3-14B-Base and Qwen3-30B-A3B-Base with data compiled by TrajLong, followed by supervised fine-tuning. Experiments on 6 long-context and 12 agent benchmarks demonstrate broad performance gains, with controlled ablations showing improvements over raw and masked trajectory baselines. Capability-level analyses further reveal task-dependent associations between long-context and agent atomic capabilities. These findings suggest that the shared capability demands of long-context reasoning and agent execution provide a principled basis for designing mid-training data to develop downstream agent capabilities.
|
| 506 |
IREA: Intermediate Representation-based Embedding Alignment for Normative RAG
2610.04974
|
cs.CL
|
Mirae Han, Sihyeong Yeom, Harksoo Kim |
Large language models (LLMs) have shown strong performance across various tasks, but they still struggle with questions involving ethical judgment. Previous studies have attempted to train LLMs on ethical standards, but the diversity and relativity of ethical ...Large language models (LLMs) have shown strong performance across various tasks, but they still struggle with questions involving ethical judgment. Previous studies have attempted to train LLMs on ethical standards, but the diversity and relativity of ethical norms make them difficult to fully internalize in model parameters. As an alternative, we introduce normative RAG, a retrieval-augmented approach that supports ethical judgment using external normative knowledge. Normative retrieval involves a distinct asymmetry between context rich narrative queries and generalized normative statements. Existing factual retrieval methods rely on query-only expansion into a document-like form, making them insufficient for resolving this asymmetry. Therefore, we propose Intermediate Representation-based Embedding Alignment (IREA), a bidirectional alignment method that maps both text types into a shared situation-behavior representation. This representation captures ethically salient contextual and behavioral information in a normalized form, reducing surface-level discrepancies and improving alignment in the embedding space. Experimental results show that IREA improves normative retrieval and downstream ethical judgment across multiple settings, demonstrating the effectiveness of bidirectional alignment for normative RAG.
|
| 507 |
When LLMs Sit Above Diagnostic Tools: Unrealized Complementarity in Industrial Fault Diagnosis
2610.05031
|
cs.CL
|
Donghwan Kim |
Large language models are increasingly used as integration layers above specialized tools, but a stronger component does not necessarily produce a stronger combined system. Across five diagnostic datasets (bearing vibration, process monitoring, semiconductor e...Large language models are increasingly used as integration layers above specialized tools, but a stronger component does not necessarily produce a stronger combined system. Across five diagnostic datasets (bearing vibration, process monitoring, semiconductor equipment), we study whether an LLM can reliably use external diagnostic information; paired repeat calls separate advice effects from output instability. In all five, conflicting external information overturned initially correct LLM judgments. Among the four datasets with direct integration comparisons, none showed a consistent advantage for implicit LLM integration over the stronger standalone source. On a Tennessee Eastman confirmation set whose protocol was fixed before evaluation, unaided accuracy was 64.67%, implicit LLM-specialist integration 77.43%, and the specialist alone 83.33%. Specialist information improved the LLM by 12.8 points (95% interval 9.7 to 15.9), yet the integrated output stayed 5.9 points below the specialist (95% interval -12.0 to -0.7). A two-source selector oracle reached 92.76%, indicating complementarity that the integrated output did not fully realize. The integrated output missed 140 of 295 specialist corrections (47.5%) but lost 15 of 99 initially correct LLM judgments (15.2%). The deficit remained under prompt and specialist sensitivity analyses. Among CWRU cases solved under both evidence presentations, task-aligned physical evidence yielded lower estimates of susceptibility to incorrect advice in six of seven models (five intervals excluding zero); higher reasoning effort gave no reliable reduction in five models, and a separate four-model TEP analysis gave no clear evidence that it resolves the integration problem. Source quality and integration quality should be evaluated separately: an integration layer should be compared with its stronger standalone component, not only with the unaided LLM.
|
| 508 |
Causal Improvement Graph for Agentic Harness Optimization
2610.05039
|
cs.CL
|
Junjie Zhang, Shunyu Liu, Haoyu Wang, Ting-En Lin, Yongbin Li |
Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative p...Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Existing meta-harness methods mainly adopt proposer-centric discovery, in which an LLM-based proposer integrates accumulated experimental findings to determine subsequent Harness revisions. This places the burden of maintaining the evolving improvement state on the proposer as history expands and its underlying experimental logic becomes harder to discern. In this paper, we introduce the Causal Improvement Graph (CIG), a graph-governed meta-harness framework that externalizes the evolving improvement state in a persistent graph, allowing prior findings to directly govern subsequent Harness optimization through local proposer operations. CIG grows and links Evidence, Hypothesis, Intervention, and Outcome nodes to represent what was observed, how it may be explained, how to test that explanation, and what the evaluation reveals. Their structural relations preserve how the improvement state changes across iterations, allowing local proposers to build directly on relations among prior findings rather than recover them from raw history. Across various agent tasks, CIG discovers stronger Harnesses than previous meta-harness baselines and remains robust to the choice of task solver and proposer. Structural ablations further support the design of an explicit improvement state with graph-governed evolution.
|
| 509 |
AraYoungVoices: A Diverse L1/L2 Corpus of Arabic Child and Adolescent Speech
2610.05044
|
cs.CLcs.AIcs.SD
|
Shammur Absar Chowdhury, Zien Sheikh Ali, Houssam Eddine-Othman Lachemat, Hamdy Mubarak |
State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising...State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising AraKids (7--12) and AraTeens (13--18). The corpus includes 146 native Arabic (L1) and 140 second-language (L2) speakers, with native speakers spanning Egyptian, Gulf, Levantine, and North African dialectal backgrounds and L2 speakers representing diverse linguistic backgrounds across the Americas, Asia, Africa, and Europe. We benchmark four pretrained ASR models under zero-shot and fine-tuned settings using unseen-speaker-$\&$-unseen-prompt (USUP) and unseen-speaker-$\&$-seen-prompt (USSP) evaluations. Results show that L2 speech remains substantially more challenging than L1 speech, with the largest errors observed mainly for younger L2 speakers. Age-specific fine-tuning improves the matched age group, while joint fine-tuning provides a stronger balance across populations. ASR hypotheses are also consistently closer to the standard reading prompt than to the verbatim transcription, particularly for L2 speech, suggesting partial normalization of reading deviations.
|
| 510 |
Usage-Modulated Sentiment Representations in Large Language Models
2610.05069
|
cs.CL
|
Hongfei Du, Jiacheng Shi, Yanfu Zhang, Gang Zhou, Ye Gao |
Prior work suggests that sentiment can often be captured by approximately linear directions in LLM activation spaces, but a single direction may not fully capture sentiment representations. In natural communication, sentiment is shaped not only by polarity but...Prior work suggests that sentiment can often be captured by approximately linear directions in LLM activation spaces, but a single direction may not fully capture sentiment representations. In natural communication, sentiment is shaped not only by polarity but also by usage factors, such as tone and audience adaptation. We test whether these factors systematically modulate sentiment representations beyond a shared sentiment direction. We construct a controlled paired dataset that holds event content fixed while varying sentiment polarity and usage factors, and analyze Llama, Mistral, and Gemma. We identify a shared sentiment direction, remove it, and test the residual structure through erasure and generation-time tone steering. Across models, the shared direction is robust (median cosine 0.953-0.975), yet removing it leaves 0.833-0.909 of the original positive-negative representation-difference norm. The residuals contain compact, reproducible usage-conditioned structure. Targeted erasure weakens held-out usage metrics more than random and label-shuffled controls. On Llama, outputs steered along residualized tone components are preferred in 92.8% of blind target-tone comparisons while preserving the requested sentiment polarity in 98.7% of evaluated outputs.
|
| 511 |
Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation
2610.05076
|
cs.CLcs.AI
|
Cheng Luo, Bing Li, Bernard Ghanem |
Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining gener...Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.
|
| 512 |
Towards cross-cultural study of folksong lyrics with machine translation
2610.05084
|
cs.CL
|
Anna Dvo\v{r}\'akov\'a, Anna Aljanaki, Danbinaerin Han, Peter van Kranenburg, Mat\v{e}j Kratochv\'il |
Music is universally present in human societies. Ethnomusicologists have long been documenting the diverse expressions of human musicality, and comparative musicology has recently brought several studies of folksong to a more global scale. Such cross-cultural ...Music is universally present in human societies. Ethnomusicologists have long been documenting the diverse expressions of human musicality, and comparative musicology has recently brought several studies of folksong to a more global scale. Such cross-cultural research has not been conducted on lyrics: the language barrier has so far prevented work with multi-lingual data. However, Natural Language Processing (NLP) technologies have reached a stage where this language barrier may no longer be prohibitive. Combining folksong lyrics corpora across five languages, we machine-translate them to a pivot language with a pre-trained neural topic model, and we examine the relationship between content and social function within each language, and across languages for wedding songs. As expected, human evaluation of translation results shows that non-Indo-European languages suffer from overall worse translation quality. Experiments with topic models then indicate that the content of lyrics is at best partially related to the social function of folksongs across all languages. These experiments are just first steps into cross-cultural folk musics lyrics analysis; however, they do indicate that a previously unobserved web of cross-cultural relationships beyond ethnomusicological typologies may be uncovered through the study of what people sing across the world's diverse folk musics.
|
| 513 |
How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models
2610.05094
|
cs.CL
|
Laur\`ene Vaugrante, Thilo Hagendorff |
Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices ...Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge's verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human ground truth, the practical size of differences is small enough to consider most judges reliable (mean absolute deviation of 0.11 points on a 1 - 7 scale); toxicity judges even outperform standard classifiers. Judges are also highly accurate on average (96.5%) for the accuracy classification task. However, design choices can produce shifts: changing the rating scale alone can shift measured bias by up to 0.93 points (rating task), and while accuracy levels are rarely impacted, design choices consistently impact judge leniency (classification task; leniency drop of 28.9 percentage points when using detailed prompts, and up to 56.1 percentage points when switching models). Counterintuitively, lower reasoning effort affects neither accuracy nor leniency. Across both tasks, model identity is the dominant source of variance. These findings suggest that while LLM judges are broadly trustworthy in aggregate, design choices can be meaningful sources of variance. Given the growing reliance on automated evaluation in LLM research, we intend this study as a methodological reference for designing more robust and replicable LLM-as-a-judge pipelines.
|
| 514 |
vMF Sentence LDA: A Spherical Topic Model over Sentence Embeddings
2610.05095
|
cs.CLcs.LG
|
Ryotaro Kobayashi, Yuri Murayama, Kiyoshi Izumi |
Latent Dirichlet Allocation (LDA) and models derived from it remain widely used topic models. LDA observes each document as a bag-of-words and models each topic by a categorical distribution over the vocabulary, so that it uses neither the internal structure o...Latent Dirichlet Allocation (LDA) and models derived from it remain widely used topic models. LDA observes each document as a bag-of-words and models each topic by a categorical distribution over the vocabulary, so that it uses neither the internal structure of the document nor the similarity in meaning between words. Earlier work has responded to this limitation in two ways: many models have introduced embeddings, and some have assigned topics to sentences rather than to words. Their combination, a topic model that observes sentence embeddings, remains little explored. We propose vMF Sentence LDA (vSLDA), which keeps the admixture structure of LDA, observes each sentence as its L2-normalized embedding and models each topic by a von Mises-Fisher (vMF) distribution, which matches the cosine geometry of sentence embeddings. Its per-topic parameter count and per-iteration cost are linear in the embedding dimension, versus quadratic for the full-covariance Gaussian distribution in the existing model over sentence embeddings. We evaluate vSLDA where the limitation is expected to matter most, among topics that share much of their vocabulary: the topics that subdivide the one subject of a collection, and the narrow topics that result when a corpus is divided into a large number of topics. On two corpora, vSLDA attains the best mean rank against eight baselines when the fine categories within each coarse category are classified from the document-topic distributions. On the whole corpus, its advantage appears or widens as the number of topics grows. Weighting the word frequencies of each sentence by its topic posterior yields expected topic-word counts of the same form as those of LDA, so that the standard topic coherence and diversity measures apply to models that assign topics to sentences. On their product, topic quality, vSLDA leads in most within-category conditions of both corpora.
|
| 515 |
Small Agents with Semantic Search: Efficient Multilingual Code Localization
2610.05099
|
cs.CL
|
Maxence Lasbordes, Aarush Sinha, Raphael Sourty, Am\'elie Chatelain, Djam\'e Seddah |
Locating relevant files from natural-language requests is a core subtask for agents operating over code repositories. We investigate whether this task can be delegated to compact, specialized models to enable on-device search while reducing the token usage, la...Locating relevant files from natural-language requests is a core subtask for agents operating over code repositories. We investigate whether this task can be delegated to compact, specialized models to enable on-device search while reducing the token usage, latency, and inference cost of larger agents. We show that semantic search improves file localization, with gains in accuracy, cross-language transfer, and inference efficiency. To study this setting, we introduce a training framework for file-localization agents built around ColGREP, a local semantic search tool based on late-interaction retrieval models. Our recipe combines weighted supervised fine-tuning on teacher trajectories, assigning turn-level credit based on retrieval outcomes, with reinforcement learning on localization quality. We train three model families with fewer than two billion parameters to formulate search queries, inspect retrieved content, and identify relevant files. On localization tasks derived from SWE-bench Lite and Multi-SWE-bench Flash, ColGREP-equipped agents substantially improve over their base models and outperform corresponding GREP-based agents. In addition to improving localization accuracy, ColGREP reduces mean end-to-end trajectory latency by 44.1\% on CPU while using 29.1\% fewer tokens, and enables better generalization to programming languages unseen during fine-tuning. These results suggest that compact, tool-specialized localization agents can provide an efficient interface between natural-language requests and large codebases.
|
| 516 |
InstMoE: Adaptive Multimodal Routing with Specialized Experts
2610.05111
|
cs.CL
|
Guimin Hu, Xiang He, Yingjian Li, Zheng Lian, Boyan Xu |
Multimodal inputs are inherently heterogeneous, not only across modalities but also in the information pathways required for effective prediction. To address this limitation, we propose InstMoE, an adaptive expert routing framework for multimodal learning. Ins...Multimodal inputs are inherently heterogeneous, not only across modalities but also in the information pathways required for effective prediction. To address this limitation, we propose InstMoE, an adaptive expert routing framework for multimodal learning. InstMoE dynamically routes each input to specialized unimodal and cross-modal experts, allowing the model to adapt its information pathways to the characteristics of the input. However, routing can be misled when modality-specific variations obscure task-relevant semantics. Such irrelevant variations may distort routing decisions, causing inputs to be assigned to inappropriate experts. We therefore introduce a Contrastive Semantic Alignment module, which encourages semantically similar inputs to share task-relevant representations while suppressing irrelevant modality-specific variations. Experiments on multimodal sentiment analysis benchmarks demonstrate that InstMoE achieves state-of-the-art performance on CMU-MOSEI and CH-SIMS v2 while using substantially fewer parameters than competitive baselines. Further analysis shows that different inputs exhibit distinct expert preferences, demonstrating that InstMoE moves beyond fixed fusion toward adaptive multimodal computation.
|
| 517 |
Belief-Trajectory Energy: Measuring the Path to a Prediction
2610.05114
|
cs.CL
|
Jiahao Ying, Wei Tang, Boxian Ai, Yaoning Wang, Haotian Chen |
Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded me...Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded measure that characterizes an input through the layerwise predictive revisions it induces in a model. By mapping intermediate states into a shared predictive space, BTE provides a principled measure of belief change that can be summarized as either a scalar or a structured depth profile. Theoretically, we show that local BTE corresponds to predictive revision under the Fisher-Rao geometry, while the sequence of revisions captures information beyond the initial-to-final belief change. Empirically, scalar BTE provides a model-relative signal of difficulty across diverse reasoning tasks, while richer BTE representations support human-LLM review detection and fine-grained generator attribution, reaching up to $0.998$ macro-AUROC and $95.6\%$ eight-way attribution accuracy. Further analysis shows that BTE develops throughout pretraining and is selectively reshaped by targeted training, demonstrating that the resulting measurement reflects what the scoring model has learned. Together, our results establish belief trajectories as a principled model-grounded signal and suggest a broader perspective in which learned models can themselves serve as instruments for characterizing the data they process. More demonstrations can be found at https://yingjiahao14.github.io/BTE-web/.
|
| 518 |
Selecting Repetition Counts Across Model Scales in Data-Constrained Pretraining
2610.05126
|
cs.CL
|
Ziyue WANG, T. Kanamori |
The repetition count that works best for a small language model may not remain best at a larger scale. We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction. On Wikipedia-derived data and Proof-Pile-...The repetition count that works best for a small language model may not remain best at a larger scale. We study this effect in pretraining with a finite target corpus mixed with generic data at a fixed target fraction. On Wikipedia-derived data and Proof-Pile-2, the ranking of measured repetition counts changes with model size, and a 520M Proof-Pile-2 experiment confirms that reducing repetition from sixteen to eight improves loss while using fewer training tokens. We use loss curves from several smaller models to retain a short list of promising repetition counts for evaluation at a larger scale. On PubMed and Caselaw, candidate sets fixed before target-model training retain the lowest-loss measured count on the original evaluation grids at both 200M and 520M. This supports candidate retention as a practical alternative to exact point prediction. We also relate the pruning regression to an empirical scaling model with two opposing repetition-dependent loss terms. A first-order expansion in log model size yields the linear form used by the selection rule, providing a scaling-based interpretation of the candidate-selection procedure.
|
| 519 |
Verification Trap: Understanding Test-Time Selection Failures under False Premises in Code Generation
2610.05170
|
cs.CL
|
Feng He, Hejia Wang, Linghao Meng, Ming Gao, Qiankun Li |
Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output. This paradigm implicitly assumes that the verifier provides a corrective signal inde...Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output. This paradigm implicitly assumes that the verifier provides a corrective signal independent from the generator. We challenge this assumption under misleading task premises. When the generator and verifier share a false premise, they become coupled through a mistaken belief: the generator produces premise-consistent shortcuts, while the verifier supplies evidence that fails to expose them. Consequently, the selector may choose a hidden-test-wrong candidate even when a hidden-test-correct program exists in the pool. We call this failure mode Verification Trap. Across three code-generation benchmarks and five code models, false premises consistently degrade first-sample correctness, reduce selector-chosen correctness after 64-sample test-time selection, and amplify recoverable mis-selection. Mechanistically, verifier-written tests inherit the premise-level blind spot, reshaping verifier-visible candidate space away from hidden-test correctness. These traces make Verification Trap predictable before hidden execution: a lightweight gold-free predictor using verifier-visible features reaches 0.846 AUROC. Our results identify decoupled evidence as a key mitigation axis: coupled scaling provides limited recovery, whereas premise-agnostic robustness auditors recover substantial oracle headroom.
|
| 520 |
FORGE: Verification-Gated Behavioral Repair for Generative Language Models
2610.05190
|
cs.CLcs.LG
|
Hsin-Ling Hsu, Min-Yu Chen, Nai-Chia Chen, Yan-Ru Chen, Yi-Ling Chang |
Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defe...Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally, provide correctness guarantees. Existing approaches address this only partially: gradient-based fine-tuning lacks per-instance guarantees and becomes unstable with few defect samples; model editing assumes explicit knowledge replacement rather than behavioral correction; and constraint-based repair is largely restricted to discriminative models with unique target outputs. We present FORGE, a framework for targeted behavioral repair of generative language models that separates defect localization, weight editing, and behavioral verification into independent stages. Its core is a repair abstraction that converts localized defective generation into explicit optimization objectives, enabling verification-oriented repair techniques to operate on autoregressive generation. FORGE is editing-mechanism agnostic: we instantiate it with (1) a constraint-based quadratic optimization method that provides per-sample repair certificates and (2) a null-space projection editor that minimizes interference with the original model distribution, both under the same localization and verification protocol. On five open-source LLMs, FORGE consistently achieves larger reductions in bias and toxicity than gradient-based fine-tuning with minor perplexity degradation. The two backends exhibit complementary performance across architectures, which a lightweight causal probe traces to where toxicity-related signals concentrate. FORGE also remains effective with only a handful of defective examples, where conventional fine-tuning often oscillates or fails to converge.
|
| 521 |
Inductive Claims Extraction at Scale
2610.05275
|
cs.CL
|
Sandrine Chausson, Bj\"orn Ross |
A large part of political discourse on social media is built and expressed at a level of claims: i.e. declarative, typically single-clause statements, which convey a particular interpretation of reality and can range from factual to evaluative. Moreover, rathe...A large part of political discourse on social media is built and expressed at a level of claims: i.e. declarative, typically single-clause statements, which convey a particular interpretation of reality and can range from factual to evaluative. Moreover, rather than occurring randomly, claims coalesce, recur in patterns, and come to be associated with different world views. When paired with structural computational tools such as Social Network Analysis, claims can be a powerful unit of analysis to study political phenomena such as echo chambers or polarisation. In this paper, we present a pipeline that uses a large language model (LLM) to inductively extract and catalogue claims from large social media corpora, and apply it to two different Twitter datasets: one relating to the 2020 US presidential election and the other to the 2022 FIFA World Cup. We comprehensively evaluate the approach by measuring the pipeline's recall and precision against manually annotated samples, run ablation studies isolating the contribution of its various components, and perform a qualitative error analysis. We discuss the value of the approach in the context of Computational Social Science research, and illustrate its capabilities by presenting the claims catalogue obtained from each dataset.
|
| 522 |
When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following
2610.05278
|
cs.CLcs.AI
|
Qishi Zhan, Seoyeon Jang, Zihan Dong, Minxuan Hu, Ziheng Chen |
Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and r...Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts. In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially: in the nine-model control panel, compliance is 54.9% for the original positive form, 48.2% for a longer positive form, 36.7% when the target appears later, and 33.8% for AVOID1. A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect. Effect sizes, failure directions, weakest forms, and model rankings vary across realizations. Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse. Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior. WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.
|
| 523 |
Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
2610.05282
|
cs.CL
|
Tongyan Hu, Hao Li, Xiaogeng Liu, Ruida Wang, Zhengyu Liu |
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or t...Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9\% to 72.4\% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at https://github.com/SaFo-Lab/Red-TTT
|
| 524 |
RubricArmor: Adversarial Evolution Improves LLM-Based Rubric Generation
2610.05308
|
cs.CL
|
Haocheng Yang, Yuchao Zhang, Licheng Pan, Jiajun Fan, Maolin Wang |
Rubric-based reinforcement learning (RL) provides interpretable rewards for aligning large language models (LLMs) by evaluating responses against query-specific evaluation criteria. To construct rubrics at scale, a straightforward approach to LLM-based rubric ...Rubric-based reinforcement learning (RL) provides interpretable rewards for aligning large language models (LLMs) by evaluating responses against query-specific evaluation criteria. To construct rubrics at scale, a straightforward approach to LLM-based rubric generation is to prompt an LLM to generate a rubric directly from the query. However, rubrics directly generated by LLMs are vulnerable to reward hacking, since omitted or underspecified criteria allow the policy to obtain high rubric rewards with low-quality responses. Existing LLM-based rubric generation methods improve the granularity and coverage of the generated criteria but do not proactively guard against reward hacking. To address this limitation, we propose RubricArmor, an adversarial framework that exposes and mitigates potential reward hacking at the rubric generation stage before it occurs in subsequent RL. Specifically, RubricArmor performs adversarial evolution, in which an attack step and a repair step alternate over multiple rounds. The attack step simulates the reward hacking of the policy by constructing adversarial responses that satisfy the current rubric but fail to properly complete the task. The repair step then revises the rubric to detect the response defects exposed by the attack step while preserving other valid criteria. Extensive experiments demonstrate that RubricArmor outperforms competitive rubric generation baselines and translates into more effective downstream rubric-based RL.
|
| 525 |
MemStrata: 95% and 90.91% Source-Aware Accuracy on LongMemEval-500 and LoCoMo-1540 with a Local Qwen 3.8 27B Q4_K_M Reader
2610.05343
|
cs.CL
|
Neeraj Yadav (Called It Inc.) |
An adequate conversational answer may differ from a short or incomplete benchmark reference. To measure adequacy against the recorded history we prefer source-aware grading, in which the judge checks the reference against the full source before assessing syste...An adequate conversational answer may differ from a short or incomplete benchmark reference. To measure adequacy against the recorded history we prefer source-aware grading, in which the judge checks the reference against the full source before assessing system-blinded answers; original reference-only grading is reported alongside. With a local Qwen 3.8 27B Q4_K_M reader and a 24,000-token evidence ceiling, MemStrata CL1 scores 475/500 (95.0%) on LongMemEval-S and 1,400/1,540 (90.91%) on LoCoMo categories 1-4 under source-aware GPT-5.5 adjudication, against 463/500 (92.6%) and 1,205/1,540 (78.25%) under reference-only grading of the same answers. It preserves a retrieval backbone and adds nonduplicated, dated, speaker-attributed source spans. A same-reader full-history control with about 4.7 times the evidence scores 464/500 reference-only and 470/500 (94.0%) source-aware; neither difference is decisive. Keyword-only selection at the same budget scores 425, and a matched-reader Letta arm 438. On LongMemEval-M, where the packet holds about 1.6% of each history, MemStrata CL1 scores 427/500, with losses concentrated in multi-session and temporal questions. On 300 BEAM-1M questions it outscores dense retrieval, 0.738 to 0.706 (Wilcoxon p = 0.011). A same-seed replay of unchanged requests changed 1.5-2.3% of labels. On identical packets GLM 5.3 flash is non-inferior within 3 points (462 versus 463); Muse Spark 1.3 did not show non-inferiority on 269 questions. None of four pre-registered interventions met all of its registered advancement or feasibility criteria. Signed read-side artifacts support inspection but do not regenerate the private retrieval pipeline. The superiority of source-aware grading to human adjudication is not established, and development exposure, automated-judge dependence and the absence of held-out data preclude an independent-replication or leaderboard claim.
|
| 526 |
The \`{I}r\`{o}y\`{i}nSpeech Text Corpus: 24,905 Curated Yor\`ub\'a Sentences for Speech and Language Technology
2610.05366
|
cs.CL
|
Kola Tubosun, Aanuoluwapo Aremu, Tolulope Ogunremi, Iroro Orife, David Ifeoluwa Adelani |
\`{I}r\`{o}y\`{i}nSpeech is a 42-hour, 80-speaker Yor\`ub\'a read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yor\`ub\'a sentences (275,897...\`{I}r\`{o}y\`{i}nSpeech is a 42-hour, 80-speaker Yor\`ub\'a read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yor\`ub\'a sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yor\`ub\'a corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yor\`ub\'a personal and place names appear in Yor\`ub\'a form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.
|
| 527 |
Towards Unbiased On-Policy Distillation for Block Diffusion Language Models
2610.05373
|
cs.CL
|
Zaiquan Yang, Fei Wei, Yong Wang, Yudong Han, Yiyu Li |
On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving d...On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexplored. In this work, we investigate this regime and reveal two critical optimization biases that induce severe training instability. First, mismatched block boundaries between teacher and student cause \textbf{\textit{context misalignment}}, providing distorted supervisory signals that misguide student decoding. Second, even under aligned contexts, an \textbf{\textit{intrinsic optimization bias}} in OPD, where the student tends to rapidly absorb high-support signals while lagging on low-support updates, drives a premature confidence surge that traps weaker students in catastrophic overconfidence collapse. To resolve these, we propose \mbox{\textbf{Un-OPD}}, an unbiased on-policy distillation framework with two novelties for stabilizing BDLM training. First, Un-OPD introduces a boundary-aware step filtering strategy that eliminates context-misaligned decoding steps. Second, Un-OPD proposes moderating optimization intensity at high-support positions via a support-rebalanced confidence calibration, thereby bypassing overconfidence collapse. Beyond stability, we also introduce a rollout reuse mechanism to reduce rollout generation overhead. Extensive experiments on math reasoning and code generation benchmarks show that Un-OPD consistently stabilizes training and delivers superior performance while reducing wall-clock training time by approximately half.
|
| 528 |
Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination
2610.05382
|
cs.CL
|
Shanyong Wang, Zhenwen Ji, Lei Jin, Yining Zhao, Yicheng Qian |
Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histo...Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and cause the policy to lose track of unresolved questions, overlook useful evidence, or terminate before sufficient support has been collected. One of promising way is to decouple three distinct responsibilities of proposing retrieval actions, updating persistent state, and deciding when to stop rather than concentrating them within a single policy. Targeted at it, we introduce Harness-Search, a multi-agent search harness to reduce the local errors propagating across subsequent exploration, evidence curation, and termination decisions. In particular, Harness-Search assigns these responsibilities to three permission-bounded authorities: a Retrieval Policy that proposes search operations, a Memory Operator that validates and commits persistent-state updates, and a Summary Auditor that accepts or rejects termination based on the sufficiency of the curated evidence. Together, these roles form a Propose-Commit-Audit loop in which actions are proposed, persistent evidence is selectively committed, and stopping decisions are subjected to an explicit sufficiency check. Across seven long-horizon search benchmarks, Harness-Search improves both retrieval and answer generation under the same policy backbone, increasing Recall by 4.60-27.92 points and Final-Answer Recall by 12.34-30.13 points over the strongest harness-based baseline on each evidence-retrieval benchmark. Moreover, trajectory-level analyses show that Harness-Search continues to accumulate useful evidence and expand evidence coverage with less redundant retrieval as the search history grows.
|
| 529 |
The Hidden States Cookbook: A Large-Scale Ablation Study for Noise-Robust Conversational Intent Classification in Industry
2610.05394
|
cs.CL
|
Bogdan Bogachov, Nikita Letov, Yaoyao Fiona Zhao |
Conversational database interfaces face a critical challenge: users naturally embed queries in conversational noise (greetings, politeness, off-topic remarks), which degrades intent classification accuracy and wastes computational resources. Despite advances i...Conversational database interfaces face a critical challenge: users naturally embed queries in conversational noise (greetings, politeness, off-topic remarks), which degrades intent classification accuracy and wastes computational resources. Despite advances in orchestration and retrieval strategies, a fundamental question remains unanswered: which pooling strategy maximizes intent classification accuracy under realistic conversational noise in production language models? This work addresses this gap through 360 controlled experiments spanning four pooling configurations (mean, max, last-token, attention, and FFT-augmented variants) using Llama-3.2-1B-Instruct on BANKING77 and CLINC150 datasets under clean/noisy conditions with ten random seeds. Key findings reveal that attention pooling consistently outperforms alternative strategies under noisy conditions (~+2.6-2.8 F1 over the default), while mean pooling degrades performance by up to ~5 F1 points. Frequency-domain filtering does not produce consistent accuracy improvements and functions primarily as a structural variation rather than an accuracy-enhancing component. These results provide concrete, evidence-based guidance for building noise-robust conversational classifiers: attention pooling is recommended for noisy interfaces, mean pooling should be avoided, and last-token pooling is appropriate for clean-query scenarios.
|
| 530 |
Writing as a Self-Organized Critical Process
2610.05466
|
cs.CL
|
Nikolay Mikhaylovskiy |
We explain autocorrelation decay power laws omnipresent in texts by self-organized criticality. Specifically, we analyze the recently released KLiCKe keystroke dataset and show that not only the final texts' autocorrelations form a manifold that adheres to a p...We explain autocorrelation decay power laws omnipresent in texts by self-organized criticality. Specifically, we analyze the recently released KLiCKe keystroke dataset and show that not only the final texts' autocorrelations form a manifold that adheres to a power law with a finite-size scaling, but also the text revisions generate revision-size-dependent restoring dynamics toward that manifold. Thus, human writing appears to dynamically regulate semantic correlations in a text toward a critical state.
|
| 531 |
Dataset Signatures in Human-LLM Interactions and User Modeling
2610.05534
|
cs.CL
|
Joseph Suh, Serina Chang |
Human--LLM interaction datasets shape our understanding of AI use and provide a foundation for downstream research, including training and evaluation of user models. In recent years, a growing number of datasets have sought to capture a representative picture ...Human--LLM interaction datasets shape our understanding of AI use and provide a foundation for downstream research, including training and evaluation of user models. In recent years, a growing number of datasets have sought to capture a representative picture of human--LLM interactions. But how different are the pictures these datasets provide, and what do those differences mean for research built on them? We study these questions across seven conversation datasets, spanning in-the-wild chat logs and human preference data. We begin by revisiting the dataset classification experiment of Torralba & Efros and find that neural network classifiers identify the source of a conversation from user messages alone well above chance, indicating distinctive dataset signatures. This separability persists after matching datasets on the dimensions of human-designed taxonomies, implying subtle differences that these taxonomies do not capture. We then examine the implications for user modeling: how dataset signatures propagate to the outputs of user models trained on these datasets; how dataset choice influences evaluations of user model quality and subsequent evaluations of LLM assistants paired with these user models; and how dataset classifiers can guide data selection for training user models. While each dataset is meant to capture a slice of 'real-world' interactions, our findings reveal the extent to which these slices diverge, and the consequences of those differences for research built on these foundations.
|
| 532 |
Lend Me Your Eyes: Instruction-Aware Text Embeddings via Attention Relay
2610.05564
|
cs.CLcs.LG
|
Yiyuan Luo, Vaggos Chatziafratis |
Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to ...Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to a Transformer-based embedder without any training. We propose Attention Relay, which passes the attention weights an LLM produces to the embedder's own attention. Across six instruction-tuned LLMs from the Qwen3, Llama 3.1 and OLMo 3 families and ten widely used embedding models that differ in tokenizer, size and pooling type, Attention Relay makes nearly every combination instruction-aware. Experiments that break the method down into its parts show that the LLM's attention weights track the instruction in its later layers and come largely from instruction tuning. They also show that relaying these weights selects which content in the text matters: it makes the aspect of the text that the instruction asks about dominant in the embedding, or restores that aspect where averaging had diluted it.
|
| 533 |
Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters
2610.05575
|
cs.CLcs.AI
|
Abdul Rehman, Jian-Jun Zhang, Xiaosong Yang |
Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriat...Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself. This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone. For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models. Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone. Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this. Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others. Acoustic clusters are not compact in text embedding space in any of 360 configurations. The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly. Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.
|
| 534 |
What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining
2610.05591
|
cs.CL
|
Yekun Chai, Haoyi Xiong |
As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters. We answer them by pricing a repeated token against two refe...As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters. We answer them by pricing a repeated token against two references: one epoch on the same data, which gives its value, and fresh data at equal compute, which gives its cost. Against fresh data, the cost of repetition follows a single variable, the number of extra epochs divided by the unique tokens per parameter. Against the same data, a second epoch is worth nearly as much as a fresh one, and repeated tokens fall to half the value of fresh ones after a critical epoch count that grows with the training budget per parameter but hardly with model size. With unique data fixed, the predicted compute-optimal run grows model size and epochs together until loss stops improving, near the critical epoch count. The same variable accounts for the direction of size trends that appear to conflict: larger models tolerate fewer epochs when the corpus is fixed, from about 15 at 127M to 4 at 2B parameters, but not when unique data grow with the model. Counts alone do not determine loss: at identical counts, replaying shards consecutively raises loss by up to 0.46~bits per byte, concentrating repeats on fewer samples also raises it, lower-entropy sources degrade faster with repetition, and re-tokenizing repeats helps only under heavy repetition. These results offer an empirical guide to pretraining when unique data, rather than compute, are the binding constraint.
|
| 535 |
More Than Words: Compositional Tokenization for Efficient Language Models
2610.05597
|
cs.CL
|
Yuval Reif, Guy Kaplan, Roy Schwartz |
Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as "On the table." is usually produced as four separate predi...Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as "On the table." is usually produced as four separate predictions for the preposition (On), article (the), noun (table), and punctuation (.), where each consumes a sequence position and adds inference cost. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled pretraining from scratch at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute. Our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.
|
| 536 |
Automatic Speech Recognition for Low-Resource Sinhala: A Critical Review of Methods, Challenges, and Future Directions
2610.05681
|
cs.CL
|
Chanuka Dinuwan, Sanath Jayasena, Buddhika Karunarathne |
Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-ver...Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-verb (SOV) syntax and scarce annotated speech data limit both conventional and modern ASR systems. This paper presents the first critical review of Sinhala ASR research, tracing its development from Hidden Markov Models (HMMs) through deep neural networks to self-supervised pre-trained models such as wav2vec 2.0, XLS-R, Whisper and Massively Multilingual Speech (MMS). We compare existing Sinhala systems with related low-resource ASR work on Tamil, Malayalam and Hindi in terms of architecture, training data, word error rate (WER) and robustness to real-world acoustic conditions, and we assess self-supervised and transfer learning as responses to scarce labeled data. We show that most reported WERs are not directly comparable because they differ in corpus, data split and scoring, and that the only controlled comparison in the literature attributes an 18.1% relative WER reduction to corpus correction alone. We also discuss context-aware ASR that draws on phonological, syntactic and semantic knowledge. We identify six research gaps: (1) the lack of large annotated corpora covering multiple dialects and acoustic conditions; (2) weak contextual modeling of Sinhala morphosyntax; (3) high WER in real-world conditions; (4) the absence of standardized benchmarks; (5) the lack of parameter-efficient fine-tuning studies; and (6) the absence of annotated code-switched Sinhala-English speech resources. We outline a research agenda to address these gaps, intended as a roadmap for researchers working on Sinhala and other morphologically rich languages.
|
| 537 |
Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi
2610.05682
|
cs.CLcs.AI
|
Jiulin Li, Ping Huang |
Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary r...Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.
|
| 538 |
Spend Bytes on Breadth: Precision-Count Trade-offs for Decode-Time KV Compression in Long Chain-of-Thought Reasoning
2610.05685
|
cs.CL
|
Runguo Li |
Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be s...Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be split between the number of cached tokens and their precision. BreadthKV spends the bytes on more tokens at low precision, combining quantization with eviction, and picks the bit-width for each model and budget with a 60-problem end-to-end calibration, since offline attention error does not predict it reliably. On three reasoning models and four math and science benchmarks, it scores above eviction alone in 17 of 18 settings and produces shorter outputs. Much of what eviction loses comes from derailed runs, which keep reasoning until the length cap without reaching an answer. On Qwen3-8B at our tightest budget, eviction sends 91% of AIME samples to the cap and BreadthKV 40%. Under the same protocol, BreadthKV is statistically indistinguishable from a joint rate-distortion allocator (RDKV) that uses 27% more KV memory-time, and it outperforms our re-implementation of ThinKV.
|
| 539 |
MedicalHarness: A Controlled Evaluation of LLMs and Agent Harnesses on Medical Tasks
2610.05778
|
cs.CL
|
Ziqing Wang, Lili Zhao, Kaize Ding |
LLM agents are increasingly built for medical work and scored on clinical benchmarks. Each such score, however, comes from a model running inside an agent harness, the system that controls the loop between the model and its environment. An agent's score is the...LLM agents are increasingly built for medical work and scored on clinical benchmarks. Each such score, however, comes from a model running inside an agent harness, the system that controls the loop between the model and its environment. An agent's score is therefore a property of a model--harness pair. For medical agents, how much outcomes change with the harness has rarely been measured. Measuring this change, and explaining it, raises two challenges. First, a harness comparison must change nothing but the harness and be repeated across models and kinds of task. Second, comparing whole harnesses leaves their mechanisms bundled together, so it cannot show when an individual mechanism helps. To address these challenges, we present MedicalHarness, a controlled study of models and agent harnesses on medical tasks. We first build MedicalHarnessBench to evaluate agents on $107$ tasks across four domains that each test a different harness capability. Using this benchmark, we run five open-weight models under five agent harnesses, changing only the harness within a comparison, and analyze both outcomes and execution traces. To study individual mechanisms, we build MH-Lab, a controlled harness that switches off context management, planning or tool exposure one at a time within a shared execution loop. We find that the harness and its interaction with the model account for about a quarter of the outcome variance, and that no single harness is best across models and tasks. Code and data are available at https://github.com/REAL-Lab-NU/MedicalHarness.
|
| 540 |
CLARA: Can AI Assess Developmental Appropriateness in Children's Stories?
2610.05783
|
cs.CLcs.AI
|
Sijing Yin, Zirui Wang, Qian Liu, Jiamou Liu |
Assessing the developmental suitability of children's narratives is important for educational recommendation and developmental literacy research, yet such assessment typically relies on subjective and difficult-to-scale human judgment. This raises an important...Assessing the developmental suitability of children's narratives is important for educational recommendation and developmental literacy research, yet such assessment typically relies on subjective and difficult-to-scale human judgment. This raises an important question: Can AI systems approximate human developmental judgments of children's stories? To study this problem, we introduce CLARA, a cognitively grounded framework for developmental narrative understanding through structured annotation across cognitive (COG), language (LAN), and social-emotional (SEL) dimensions, together with a bilingual benchmark resource containing 1107 Chinese--English children's stories with normalized silver developmental references and structured developmental annotations. We evaluate CLARA through benchmark comparison, component analysis, translated bilingual consistency analysis, and blinded human evaluation with educators. Experimental results show that structured developmental annotation achieves substantially stronger alignment with developmental references and human judgments than readability-based methods and direct prompting baselines. Overall, our findings suggest that AI systems can approximate certain aspects of human developmental judgment when guided by structured developmental annotation, while also highlighting the importance of interpretability and human oversight in educational NLP.
|
| 541 |
Adaptive Utilization of Low-Rank Adaptation via Conditioned Gating
2610.05800
|
cs.CL
|
Guang Yang, Changhao Guan, Chao Huang, Yufeng Chen, Kaiyu Huang |
Low-Rank Adaptation (LoRA) achieves parameter-efficient fine-tuning by constraining model updates to a low-rank subspace and has been widely used in practice. However, LoRA typically employs a shared low-rank update across tokens, which limits its ability to f...Low-Rank Adaptation (LoRA) achieves parameter-efficient fine-tuning by constraining model updates to a low-rank subspace and has been widely used in practice. However, LoRA typically employs a shared low-rank update across tokens, which limits its ability to fully exploit the adaptation subspace for tokens from different sequences. To address this issue, we propose an adaptive utilization of Low-Rank Adaptation (U-LoRA), which employs conditioned gating to explicitly learn effective token-level utilization of the limited low-rank adaptation subspace. Specifically, U-LoRA generates utilization coefficients along low-rank directions for each token and jointly coordinates and constrains them using sequence-level contextual information, thereby inducing more consistent adaptive patterns within a sentence. To further enhance training stability, we introduce a bias-corrected exponential moving average (EMA) historical prior that calibrates utilization signals across optimization steps, suppressing noise caused by batch-to-batch fluctuations. The effectiveness of our method arises from a better utilization of the existing low-rank subspace via input-conditioned strategies, rather than from expanding the subspace. Experiments on mathematical reasoning and natural language understanding benchmarks demonstrate that U-LoRA achieves competitive performance under comparable parameter budgets when with strong LoRA baselines and recent variants.
|
| 542 |
Plan Canvas: Fixed Reasoning Regions for Continuous Language Flows
2610.05815
|
cs.CL
|
Miaohe Niu, Pengxiang Li, Jingbo Zhu, Tong Xiao |
Continuous language flows generate text by denoising all positions of a target canvas together. The natural way to add reasoning to such a model is to write a trace ahead of the answer, but the trace length changes from question to question. The answer start i...Continuous language flows generate text by denoising all positions of a target canvas together. The natural way to add reasoning to such a model is to write a trace ahead of the answer, but the trace length changes from question to question. The answer start is therefore unknown during denoising, and the model has to decide the trace length, the place of every trace token, and the answer at the same time. We propose Plan Canvas to fix the boundary between the trace and the answer. A plan region of fixed capacity holds a compact trace, supervised padding fills its unused positions, and the answer starts at a fixed position. The fixed regions also allow separate denoising clocks for the plan and for the answer. With the trace text, backbone, and canvas length of the free-trace baseline held fixed, Plan Canvas improves accuracy on ProsQA and on Deep ProsQA, a graph benchmark with longer proofs. On Deep ProsQA, accuracy rises from 73.0\% to 87.0\%, the share of questions answered with a valid path rises from 30.8\% to 59.1\%, and the gain is largest on the longest proofs.
|
| 543 |
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
2610.05842
|
cs.CL
|
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai |
Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory ca...Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce \emph{Hybrid Linear Attention} (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet (GDN). HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact, self-attentively pooled representatives. Each gate interpolates the corresponding historical transition with the identity map, controlling both the chunk's additive memory and its transformation of earlier states. Effective-support regularization further encourages concentrated routing for sparse inference. We evaluate HLA under both pretrained adaptation and from-scratch training. Across Qwen3.5 models from 0.8B to 9B, HLA consistently improves over native GDN and fixed chunk mixing, with gains of up to 5.57 percentage points on LongBench-V2 and 3.97 points on RULER. In a controlled from-scratch 1.3B setting trained for 100B tokens with a 4K context, HLA also improves RULER performance from 4K to 32K, with gains increasing from 0.83 points at 4K to 4.22 points at 32K. These results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries. Project page: https://caesarhhh.github.io/hla/
|
| 544 |
Learning to Learn a Language
2610.05879
|
cs.CLcs.LG
|
Lennart Carstens-Behrens, Holger Fr\"ohlich |
We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having n...We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.
|
| 545 |
Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering
2610.05894
|
cs.CL
|
Sarim Hashmi, Mukul Ranjan, Abdelrahman Elsayed, Muhammad Umer Sheikh, Fahad Shamshad |
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dL...Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and we show that an adversary can exploit this. Since an answer remains open to revision over many denoising steps, an adversary with access to internal activations can watch how likely the model is to produce a chosen answer and adjust the intervention accordingly. Building on this observation, we study targeted bias injection, an attack that steers a frozen dLLM toward a demographic answer selected by the adversary. The attack uses a simple proportional-integral (PI) controller that tracks the target-answer probability during denoising and adapts the strength of a steering vector on the fly. On ambiguous BBQ questions where the correct answer is abstention, our attack raises LLaDA-8B-Instruct's preference for the targeted group from 1.8 to 16.7 percentage points, more than three times the strongest fixed-strength steering baseline, and on SocialStigmaQA it raises the selection of stigmatizing answers from 17.6% to 58.1%. Fitted to other demographic targets, the same attack shifts answers by up to 37 percentage points, and each attack takes about 40 minutes on one GPU. On the primary target, feedback is what makes the attack work: constant steering at the same average strength over the token-committing steps produces a far smaller shift while corrupting nearly three times as many outputs, and a constant strength set separately for each example still falls well short. Our findings identify the denoising trajectory as a new control channel in dLLMs and call for bias audits that examine the serving stack rather than the frozen model alone.
|
| 546 |
HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
2610.05966
|
cs.CLcs.AI
|
Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li |
Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through ...Domain adaptation aims to turn a general-purpose large language model (LLM) into an expert for a target domain. While the dominant SFT+RL pipeline offers a convenient cold start, it may reduce exploration diversity and introduces additional complexity through multi-stage optimization. These limitations motivate RL-only adaptation. However, pure on-policy RL suffers from a cold-start problem, while mixed-policy RL still falls short: informative tokens in teacher outputs are learned too slowly in early training, and stale teacher outputs can hinder later improvement. We identify these two failure modes as Gradient Starvation and Teacher-Distribution Anchoring. To address them, we propose One-stage Policy Optimization (OnePO), which treats teacher outputs as transient guidance for policy improvement. OnePO combines Adaptive Objective Evolution to strengthen learning on informative low-probability teacher tokens and Teacher Retirement to discard teacher outputs once the current policy can surpass them. On medical adaptation, OnePO achieves 67.2 on HealthBench (Total) with only 20K training samples, outperforming SFT+RL and pure RL by 2.7 and 7.4 points, respectively. We further scale OnePO to produce HuatuoGPT-3, an open-source medical LLM series whose 27B variant reaches 70.1 on HealthBench (Total) and 71.4 on HealthBench Professional, surpassing frontier models such as GPT-6 Astra. Models and code are available at https://github.com/FreedomIntelligence/HuatuoGPT-3.
|
| 547 |
Can Language Models Learn to Reject Their Own Bad Reasoning Steps?
2610.05976
|
cs.CL
|
Siheng Xiong, Xiaoze Liu, Yiqiao Jin, Xiaoqian Wang, Jing Gao |
Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. We ask whether a language model can instead reject its own bad reasoning steps. We define a prefix's re...Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. We ask whether a language model can instead reject its own bad reasoning steps. We define a prefix's recoverability as the probability that the frozen generator can complete it correctly. Diagnostics show that adjacent recoverability changes are often difficult to resolve with practical Monte Carlo budgets, while same-prefix candidates exhibit a sparse low-recoverability tail. We introduce Self-Step Rejection (SSR), which trains a lightweight LoRA acceptance gate on the generator backbone while keeping the base model frozen. SSR uses confidence-qualified first-passage supervision: steps before the first resolved crossing of a root-relative recoverability barrier are accepted, the crossing step is rejected, and unresolved steps and suffixes are excluded. Training combines pointwise classification, same-prefix pairwise learning, and group-relative policy refinement using final-answer correctness. At inference, SSR accepts candidates or resamples from the unchanged prefix under rejection budgets, without an external learned verifier. Across three reasoning models and five mathematical reasoning benchmarks, SSR improves macro-average accuracy over single-pass decoding by 5.4--10.1 points using 1.21--1.40x as many generated tokens, and achieves the highest macro-average accuracy among evaluated step-level methods. Full-solution scaling methods require 4.47--8.27x the single-pass token cost for comparable performance.
|
| 548 |
Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation
2610.05978
|
cs.CL
|
Jie Wang, Shiwei Luo, Qi Zhang, Yuanbin Wu |
Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional comp...Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to $25\%$ of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields $3.4\times$ more accepted tokens than in subword Transformers.
|
| 549 |
Breaking the Tie: A Cluster-Aware Routing Framework for Large Language Models
2610.05982
|
cs.CL
|
Yao Lu, Zhaiyuan Ji, Yaxin Gao, Zeyu Wang, Zhe Tang |
With the rapid development of artificial intelligence, the emergence of various Large Language Models (LLMs) has created a rich model ecosystem. However, this also brings a key challenge: how to select the optimal model for a specific user query. LLM routing a...With the rapid development of artificial intelligence, the emergence of various Large Language Models (LLMs) has created a rich model ecosystem. However, this also brings a key challenge: how to select the optimal model for a specific user query. LLM routing addresses this need by dynamically assigning queries to the most suitable expert in the pool of candidate models. However, existing routing frameworks often simplify this process to a standard classification task; thus, a critical vulnerability is exposed when multiple candidate models correctly answer the same query. We formalize this capability overlap as routing noise, which misleads the router with arbitrarily correct candidate models, ultimately leading to routing collapse (a severe decline in generalization ability on unseen tasks). To address this problem, we propose a novel Cluster-Aware Soft-Labeling Routing (CASLR) framework. CASLR shifts the evaluation paradigm from the success of a single query to macro-domain consensus by replacing traditional one-hot vectors with a masked softmax mechanism. Specifically, for experts who answer incorrectly, we penalize their target probability to zero; for the remaining candidates, we directly compute continuous fine-grained soft labels based on their global clustering utility scores. We then use these refined soft labels to supervise a lightweight router. Specifically, the framework not only demonstrates superior accuracy on multiple benchmarks, but also outperforms Llama-3.3-70B-Instruct by 7.80% in overall average performance. Furthermore, the extremely low routing inference latency of only 1.13s further confirms that CASLR can achieve efficient system scheduling with almost zero additional overhead, while ensuring high response quality.
|
| 550 |
D-Loop: Looped Diffusion Drafting for Speculative Decoding
2610.06011
|
cs.CL
|
Kecheng Chen, Yuyang He, Cheng Gong, Hui Liu, Guoping Long |
Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a con...Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.
|
| 551 |
Differentiable Bit-Widths: Co-optimizing Pruning and Quantization via SVD for Ultra-Efficient LLM Compression
2610.06026
|
cs.CLcs.AI
|
Hankyul Kang, Jongbin Ryu |
SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are ...SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline benefits from both pruning and quantization, it requires separate optimization for each stage and fails to fully exploit their balance, which can lead to suboptimal performance under aggressive compression. To address this limitation, we propose a new LLM compression method that co-optimizes pruning and quantization in a unified framework. Our key idea is a differentiable method for learning component-wise bit-widths, allowing less important components to be assigned 0-bit precision and pruned away. Notably, our method performs favorably against two-stage baselines, even when subjected to extreme quantization settings ($1.61$ bits) designed for ultra-efficiency. Code: https://github.com/MMAI-Laboratory/DBW.
|
| 552 |
LightMTP: Lightweight Latent Multi-Token Prediction
2610.06031
|
cs.CLcs.AI
|
Tamara Czinczoll, Julie Kallini, Gerard de Melo, Chen Shani |
Next-token prediction (NTP) is the standard pretraining objective for large language models, yet it provides an explicit training signal only for the immediate next token, which can lead models to exploit local patterns instead of capturing longer-range struct...Next-token prediction (NTP) is the standard pretraining objective for large language models, yet it provides an explicit training signal only for the immediate next token, which can lead models to exploit local patterns instead of capturing longer-range structure and ideas. Multi-token prediction (MTP) addresses this by training models to predict several future tokens. However, existing MTP methods often introduce a large number of new parameters with limited improvements in downstream performance. Latent MTP approaches address this efficiency issue by encoding future tokens into a vector representation. However, these approaches usually rely on external helper models for future token encoding. We propose LightMTP, a lightweight, i.e., parameter-efficient, latent MTP approach that bootstraps the future token representations from the model's own hidden states. Our two LightMTP variants extend supervision to more future tokens without requiring the additional computational overhead of conventional MTP nor the external supervision latent MTP normally relies on. LightMTP adds at most 1% extra parameters, retains better performance on general language modeling benchmarks, and achieves similar gains in planning, coding, and reasoning.
|
| 553 |
TrustMI: Causally controlling how assistants trust their users
2610.06064
|
cs.CL
|
Th\'eo Lasnier, Romain Froger, Maxence Lasbordes, Djam\'e Seddah |
Large Language Model (LLM) assistants routinely decide whether they can trust users and third parties whose competence, intentions, and integrity they cannot verify. This uncertainty matters for safety, as trusting the wrong party can lead an agent to comply w...Large Language Model (LLM) assistants routinely decide whether they can trust users and third parties whose competence, intentions, and integrity they cannot verify. This uncertainty matters for safety, as trusting the wrong party can lead an agent to comply with harmful requests or act on malicious instructions encountered during tool use. To study this problem, we define trust as an assistant's willingness to accept vulnerability to the actions of another party and ask whether such behavior can be causally controlled through model activations. We build 2,000 contrastive conversations spanning ability, benevolence, and integrity, where paired responses complete the same request but differ in whether the assistant trusts the user. From these pairs, we learn steering matrices while keeping the model parameters frozen and test them across six instruction-tuned models from three families, finding that steering changes trust decisions monotonically in both directions. We then ask whether this effect extends to several safety-related agent settings involving harmful requests, prompt injections, and insider threats, while using benign-task and reasoning as controls. Our findings provide evidence that trust in the user can be causally controlled along linear directions in model activations and provide a way to study how trust shapes safety-relevant behavior in language models.
|
| 554 |
Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII
2610.06093
|
cs.CLcs.AI
|
Alexandru Nazare, Agnese Profico, Nicol\`o Vania, Elena Di Croce, Daria Caramanica |
The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-c...The robustness of Personally Identifiable Information (PII) protection in Large Language Models (LLMs) is a critical concern, yet the risks associated with cross-lingual data extraction remain under-explored. This study evaluates the vulnerability of English-centric and multilingual models to Training Data Extraction (TDE) attacks when prompted in non-English languages. We construct a multi-domain PII dataset comprising social media handles, email addresses, and phone numbers and translate the attack contexts into Italian, Spanish, French, and German. Our results show that TDE attacks against both English-centric and multilingual models transfer to different languages: the attacks are successful on translated prompts, even though only the original English prompt might have been included in the pre-training data. A web-presence check on a sample of the translations confirms that they are not available online. The share of English leaks recovered in other languages grows with the multilingual capability of the model, and it drops sharply when the original wording is lost, even without a change of language. This suggests that native multilingual pre-training facilitates the emergence of latent cross-linguistic bridges that simplify the retrieval of personally identifiable information (PII). We analyze the activations of multilingual large language models (LLMs) and find that different translations of the same prompt are bridged in similar representations, with the strongest alignment in the middle layers. Our results highlight a fundamental security gap in modern LLMs, necessitating more robust, language-agnostic sanitization strategies for future model alignment.
|
| 555 |
From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation
2610.06100
|
cs.CL
|
Quanyu Long, Xiao Chen, Jianda Chen, Haozhen Zhang, Qisheng Hu |
Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environm...Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language world modeling: rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. We instantiate this paradigm with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs these traces into a reusable environment worldbook containing environment schemas, grounded evidence, and induced behavioral knowledge. At runtime, the world model agent actively consults the worldbook together with persistent episodic state to infer each action's observation and lasting state effects. Across nine environments, Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs. In multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment, indicating that its simulated dynamics better preserve the consequences of earlier actions across successive turns. These results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.
|
| 556 |
Introducing Code-Switched Contexts to Cognitively-Inspired Bilingual Model Training
2610.06161
|
cs.CL
|
Zhuojing Huang, Luise Pohlmann, Lisa Beinborn |
During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventional...During language acquisition, bilingual children are regularly exposed to code-switched input and use it as a cognitive scaffold to accelerate vocabulary growth and cross-linguistic syntactic mapping. In contrast, computational bilingual models are conventionally pretrained on interleaved monolingual corpora. While introducing synthetic code-switching during pretraining has become a promising strategy to enhance cross-lingual alignment and downstream performance, the structural and developmental parameters governing the success remain poorly understood. In this work, we investigate the efficiency of training with synthetic code-switched data across two typologically distinct language pairs by controlling two key variables: the structural location of code-switches and the dynamic switching rate across training stages. Our results show that training with code-switched data improves cross-lingual alignment for typologically close languages.
|
| 557 |
Cross-lingual Calibration of Pre-Generation Success Probes for Multilingual LLM Routing
2610.06216
|
cs.CLcs.AI
|
Andrea Paganelli, Stefano Civelli, Pietro Bernardelle, Gianluca Demartini |
Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across ...Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve the ranking of likely successes and failures (DISCRIMINATION); (2) whether they retain probabilities that match observed success frequencies (CALIBRATION); and (3) whether they produce scores comparable enough across candidate models for cost-aware multilingual routing (UTILITY). Using 3,000 MATH problems in 10 languages and 8 open-weight model configurations, we compare cross-lingual transfer from English-trained probes and equal-budget pooled multilingual probes. English-trained probes retain useful cross-lingual discrimination but become less well calibrated after transfer. Pooled multilingual supervision improves both properties and yields more reliable estimates of success. In routing experiments, the pooled router achieves a 0.7% higher test success rate while reducing modeled cost by 13.0% relative to always selecting the model with the highest average success. These results show that multilingual routing requires success estimates that remain well calibrated and comparable across languages and models.
|
| 558 |
Shared Stopping Decisions Change Answers in HQQ Cache Quantization
2610.06251
|
cs.CL
|
Seunghui Jwa, Minsu Oh, Chanjun Park, Yeo-Chan Yoon |
Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed. We study compression of the key and value cache, which stores attention representations reused ...Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed. We study compression of the key and value cache, which stores attention representations reused during generation. With request-local groups, Transformers' Half-Quadratic Quantization (HQQ) backend updates compression parameters separately but uses a shared average error to decide when all updates stop. Replacing only the question batched with the target changes four-bit HQQ answers in 170/384 test comparisons across two models. Replaying the other execution's update counts reproduces its complete answer and cache fingerprints in every changed pair, in both directions. Computing the stopping mean in FP32 reduces cache differences but leaves answer changes. Native HQQ also changes confirmed numerical correctness in eight arithmetic pairs. Fixed iterations and request-local stopping remove observed companion dependence under matched controls. Request-local stopping remains sensitive to synthetic padding changes at the tensor level. Fixing the original iteration budget removes this decision path without tuning. Neither repair has an established quality advantage, and natural rebatching still changes answers. Request-independence audits must cover stopping decisions as well as quantization groups.
|
| 559 |
Probabilistic Race and Ethnicity Prediction Using Group-Specific Name Lists
2610.06273
|
cs.CL
|
Kyla Chasalow, Noah Dasanaike, Kosuke Imai |
Statistically valid estimation of racial and ethnic disparities often requires inferring the probability that an individual belongs to a particular racial or ethnic group given only their name and geographic location. The standard approach, Bayesian Improved S...Statistically valid estimation of racial and ethnic disparities often requires inferring the probability that an individual belongs to a particular racial or ethnic group given only their name and geographic location. The standard approach, Bayesian Improved Surname Geocoding (BISG), relies on group population frequencies for each name. Although the U.S. Census Bureau provides such information for common names and a limited set of racial categories, comparable data do not exist for many racial and ethnic groups and are rarely available outside the U.S. We propose the list-powered BISG ($\ell$BISG) method, which can be used to derive calibrated group probabilities from group-specific name lists. These lists may be compiled based on expert knowledge or generated synthetically using large language models (LLMs), and thus may be subject to unknown biases. Representing names as embeddings, we treat list membership as a proxy prediction task and apply a correction based on proximal inference to recover the target group probabilities. We validate the method on U.S. voter files with self-reported race, on the full-count 1900 U.S. Census, and on the Lebanese voter registry. We find that LLM-generated name lists yield accurate and well-calibrated probabilities as well as precise disparity estimates comparable to those obtained using methods that require name-race data. Thus, $\ell$BISG substantially broadens the applicability of probabilistic race and ethnicity prediction to settings where name-race data are unavailable.
|
| 560 |
DeferKV: Rethinking Eviction Timing for One-Shot KV Cache Compression
2610.06286
|
cs.CL
|
Zhe Wang, Jiakai Li, Yujia Sun, Rongzheng Wang, Shuang Liang |
Long-context large language models (LLMs) have demonstrated strong capabilities across a wide range of tasks, but the growing KV cache introduces substantial memory and inference overhead. Existing one-shot KV cache compression methods typically commit to irre...Long-context large language models (LLMs) have demonstrated strong capabilities across a wide range of tasks, but the growing KV cache introduces substantial memory and inference overhead. Existing one-shot KV cache compression methods typically commit to irreversible eviction immediately after prefill, before any signal from actual generation becomes available. Our quantitative analysis shows that early queries from the actual generation stage provide attention signals that are more consistent with subsequent decode attention, with the largest single-step gain occurring at the prefill-decode boundary. Based on this observation, we propose DeferKV, which moves the eviction decision from the end of prefill to the first real decoding step and temporally combines prompt-side and decode-side observations, thereby better aligning KV importance estimation with subsequent generation requirements. DeferKV requires no additional training, draft model, or future-query prediction module, making it simple and easy to deploy. Experiments on LongBench, RULER, and Needle-in-a-Haystack demonstrate that DeferKV consistently improves model performance under KV cache compression while maintaining low inference latency.
|
| 561 |
From Abusive Language Classification to Sequence Labeling Identification
2610.06287
|
cs.CLcs.LG
|
Nicolas Zampieri, Ignacio Lopez, Manon Girard, Jeremy Auguste |
Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is target...Industrial content moderation must process massive message streams under tight latency constraints, yet most abusive language (AL) detection systems rely on sentence-level classification (ALC), which neither localizes abusive spans nor identifies who is targeted. We define Abusive Language Identification (ALI) as a sequence-labeling task that jointly extracts AL spans and target mentions, and assess whether this approach can be used for text moderation. On a pilot corpus drawn from a production moderation pipeline, we compare ALI with ALC on cross-domain generalization and implicit abuse, and we also evaluate AL and target span detection. ALI remains competitive with ALC while providing localized outputs for moderators, with a modest and configuration-sensitive advantage on implicit abuse. Exact AL boundaries and target spans remain difficult to recover. We complement this comparison with a qualitative analysis and discuss perspectives on complete target--span linking and on structured benchmarks for ALI.
|
| 562 |
DialectSentEval 2026: Arabic Dialect Sentiment Analysis and Swapping Shared Task
2610.06298
|
cs.CLcs.AI
|
Saad Ezzini, Shadi Abudalfa, Maram Alharbi, Salmane Chafik, Hind Alatawi |
Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes th...Sentiment analysis is a fundamental problem in Natural Language Processing (NLP). Standard sentiment classification for the Arabic language remains challenging due to the high volume of dialectal Arabic. To advance research in this area, this paper proposes the Shared Task on Sentiment Analysis and Swapping in Arabic Dialects (DialectSentEval), hosted with the Arabic Natural Language Processing Conference (ArabicNLP 2026). This shared task consists of two subtasks: Subtask 1 focuses on multi-class and multi-dialect sentiment analysis, requiring models to identify sentiment polarity across various Arabic dialects. Subtask 2 introduces a generative task for Arabic sentiment swap, challenging models to invert sentiment polarity while preserving core semantics. In this overview paper, we present the motivation, dataset creation, and summarize the main findings from participating models.
|
| 563 |
Agentic schema-guided extraction of materials process knowledge from scientific literature
2610.06322
|
cs.CLcs.AI
|
Sameer Sadruddin, Jennifer D'Souza |
Materials literature contains detailed experimental knowledge, but procedures, chemical entities and measurements remain difficult to aggregate because they are reported in heterogeneous forms and depend on process-specific context. We present SciKGExtract, a ...Materials literature contains detailed experimental knowledge, but procedures, chemical entities and measurements remain difficult to aggregate because they are reported in heterogeneous forms and depend on process-specific context. We present SciKGExtract, a schema-guided framework that combines large-language-model extraction with chemical normalization and agent-based evaluation and refinement before knowledge-graph integration. We evaluate the framework on 176 atomic-layer-deposition papers describing zinc oxide (ZnO) and indium--gallium--zinc oxide (IGZO), together with an expert-annotated full-schema subset. PubChem normalization improves exact-match extraction F1 for every tested model. For ZnO, the best F1 increases from 0.591 for direct normalized extraction to 0.805 with agentic refinement, whereas the best IGZO result is 0.344, revealing the greater difficulty of multicomponent supercycle processes. Evaluation against a deeply nested schema containing 65 experimental properties and 155 quantitative measurement nodes further exposes errors in process segmentation and numerical assignment. These results show that chemical canonicalization and targeted agentic verification provide complementary controls for converting complex materials literature into reusable, machine-actionable experimental knowledge.
|
| 564 |
Breaking Bureaucracy: Evaluating open-source LLMs for legal document review
2610.06345
|
cs.CL
|
Farrukh Baratov, Niki van Stein, Suzan Verberne |
In this paper, we evaluate open-source generative LLMs on legal Natural Language Inference (NLI). Legal inspectorial processes take place in specific domains and often deal with confidential data. This creates a need for working with local models that do not r...In this paper, we evaluate open-source generative LLMs on legal Natural Language Inference (NLI). Legal inspectorial processes take place in specific domains and often deal with confidential data. This creates a need for working with local models that do not require labeled training data. We evaluate our models on the ContractNLI benchmark and two NLI4Wills datasets. We successfully reproduce the baseline for the task (Span NLI BERT) and we evaluate multiple open-source LLMs on the same task. We analyze the invalid rate of the models, and their stability across temperature settings and domains. Among the generative models, Gemma-4 26B performs the best, reaching an accuracy of 81.2%, even outperforming the supervised model on one metric. On accuracy, it is not possible to beat the supervised model with zero-shot approaches. Qwen-3.6 35B performs well on both ContractNLI and additional datasets in the legal wills domain. Our findings indicate that zero-shot, open-source, generative LLMs are a viable alternative for real-world legal NLI when no supervised data is available. Our code is available at https://github.com/fbaratov/contractnli-llms.
|
| 565 |
Ontology Concept Overlap as a Training Signal: Knowledge-Grounded Reinforcement Learning for Clinical Question Answering
2610.06360
|
cs.CLcs.LG
|
Aditya Tanna, Abhishek Jindal |
Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no execu...Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a soft verifier from a maintained controlled vocabulary: UMLS Concept Unique Identifier overlap (via scispaCy, set-level F1) gives a graded, externally specified reward computed without a model in the loop. We combine it inside GRPO with an entropy-normalised LLM judge, which covers the safety and evidence axes overlap cannot see, and a small consistency penalty on padding and repetition that keeps early-training samples scorable. This three-term composite improves over SFT on Phi-3-mini (3.8B) over MedQA by 2.9% on EM (0.700 vs 0.680) and 39% on Token-F1 (0.202 vs 0.145); on Llama-3.2-3B the corresponding gains are 14% on EM and 35% on Token-F1. We report Token-F1 as the primary metric because it credits partially-correct clinical content that EM discards at this open-generation scale. Main-table results are means over 3 seeds with standard deviations below 0.005. The method transfers to PubMedQA, where training on the PubMedQA train set with the same composite reward improves Token-F1 over SFT by 22% on Phi-3-mini and 17% on Llama-3.2-3B without retuning. A reward ablation on Phi-3, varying the judge-ontology split at a fixed consistency weight, attributes 3 EM points to the ontology term, the contribution that catches entity substitutions the judge cannot. Three negative findings constrain the design: DPO under random negatives underperforms SFT for strong-prior models but helps the weakest-prior one; PPO under a sparse neural reward diverges; GRPO with KL-in-loss collapses at 7B.
|
| 566 |
Do Speech Representations Preserve Regional Accent Across Read and Spontaneous Speech?
2610.06430
|
cs.CLcs.SD
|
Paula A. Perez-Toro, Tomas Arias-Vergara, Annette Schwarz, Abner Hernandez, Andreas Horr |
Regional accent cues can be captured under matched conditions, but it remains unclear whether they persist between read and spontaneous speech. We study RVG1, with 500 German speakers from nine regions, comparing ten speech representations on regional classifi...Regional accent cues can be captured under matched conditions, but it remains unclear whether they persist between read and spontaneous speech. We study RVG1, with 500 German speakers from nine regions, comparing ten speech representations on regional classification and continuous geolocation under matched conditions and speaker-independent read--spontaneous transfer. Whisper performs best under matched conditions, reaching 0.489 nine-way UAR and 148 km median geolocation error, but drops to 0.11/0.18 UAR across transfer directions and 363 km geolocation error. Self-supervised models show a similar degradation, whereas speaker embeddings are less discriminative in-domain but more robust under transfer. This contrast is consistent across classification and geolocation. Across representations, robustness is associated with how little a representation shifts between styles (style-invariance), for which crossstyle speaker retrieval is an interpretable proxy. Age, sex, sentence-overlap, and duration controls do not account for the gap, although channel characteristics contribute. These results show that strong matched-condition performance does not indicate robust regional information.
|
| 567 |
The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning
2610.06446
|
cs.CLcs.AI
|
Jakub Macina, Manu Kapur, Mrinmaya Sachan |
Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on th...Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context. The reward is then easiest to raise by telling the student the answer, and a tuned penalty is needed to reduce telling. Drawing on learning sciences, we introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through what the student wrote in its own turns. This discourages cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling: the gates reduce solution handover while the near-transfer post-test improves out-of-domain transfer. Using these reward designs we develop Eduardo, a multi-turn RL recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude Opus 4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens than frontier models, which matters for interactive tutoring. Without being named in the reward, the model more than doubles its use of the push-for-justification teacher move while support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode, is trained out. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
|
| 568 |
Behavior-Preserving KV Cache Compression
2610.06479
|
cs.CL
|
Doo Hwan Hwang, Junyoung Jang, Junho Na, Hosung Lim, Kee-Eung Kim |
KV caches are a major bottleneck in long-context inference and long-form generation with large language models. Existing training-free eviction policies largely rely on proxy importance signals, such as attention mass, to decide which past tokens to retain. We...KV caches are a major bottleneck in long-context inference and long-form generation with large language models. Existing training-free eviction policies largely rely on proxy importance signals, such as attention mass, to decide which past tokens to retain. We argue that cache compression should instead preserve the predictive behavior of the full-cache model, retaining entries whose removal would substantially change the model's output distribution. We propose Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution. Using pre-eviction forward statistics, the method avoids running separate masked forward passes for each candidate. Across diverse architectures and both prefill-time and generation-time compression, our method delivers substantial gains in downstream task quality over lightweight attention-based heuristics at matched retained-KV budgets, with the largest gains under aggressive compression. It achieves these gains with additional compression-time computation while retaining an end-to-end speedup over full-cache inference in our evaluated settings.
|
| 569 |
AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation
2610.06481
|
cs.CL
|
Jiaqi Xue, Yanjun Wang, Xiangci Li, Lingbo Mo, Aritra Sengupta |
As AI agents increasingly tackle complex repository-level coding tasks, distributing work across multiple agents is a natural way to scale beyond the capabilities of a single agent. To coordinate their interdependent work, these agents share findings and agree...As AI agents increasingly tackle complex repository-level coding tasks, distributing work across multiple agents is a natural way to scale beyond the capabilities of a single agent. To coordinate their interdependent work, these agents share findings and agree on interfaces between modules. However, exchanged information often serves only as context, leaving individual agents to interpret it and incorporate it into subsequent work. Consequently, shared findings may go unused and deviations from interface agreements may go undetected, undermining the reliability and efficiency of collaboration. This motivates moving part of the coordination responsibility from individual agents to the execution harness. To make shared information actionable during execution, we introduce the Artifact-Exclusive Communication Protocol (AECP). AECP requires agents to communicate exclusively through structured artifacts and specifies how the harness processes them. The harness supplies findings when agents access relevant code, screens implementations for mismatches with recorded interface commitments, and requires affected agents to revisit revised agreements. These coordination steps become part of harness execution rather than actions that agents must initiate from prior messages. Across Doc2Repo, NL2Repo, and CodeProjectEval, using closed- and open-source models including Opus-4.8 and DeepSeek-V4-Flash, AECP improves average test pass rate by 28.2% and reduces average wall time by 16.5% relative to an agent team using free-form inter-agent messages. Artifact-exclusive communication also blocks the relay of malicious instructions between agents, reducing how often they reach other agents from 95% to 0% and how often those agents act on them from 40% to 0%.
|
| 570 |
SOL: Measuring Gaps between Text Distributions by Double Sliced Wasserstein Metrics
2610.06513
|
cs.CLcs.LG
|
Gregor Kornhardt, Moritz Piening, Jannis Chemseddine, Gabriele Steidl |
Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tigh...Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text distributions. Each sequence is represented by the empirical measure of its hidden states under a fixed transformer and the distributions of these measures are compared by the double sliced Wasserstein distance. We prove that SOL is a metric if the transformer is injective. Experiments show that SOL detects distributional failures, recovers expected model trends, and provides stable sample-based estimates. We put forward SOL to fill the gap in the current evaluation protocol used for non auto-regressive models. As a first step we use SOL to re-evaluate a variety of models trained on OpenWebText.
|
| 571 |
Test-Time Adaptation of Reasoning Strategies with Bayesian Nonparametric Memory
2610.06516
|
cs.CL
|
Keshav Ramji, Tahira Naseem, Ram\'on Fernandez Astudillo |
While modern large language models (LLMs) have been trained to reason through verbalized chains-of-thought, the generation cost grows substantially due to suboptimal paths to reach the final answer. Furthermore, as new insights are discovered while observing v...While modern large language models (LLMs) have been trained to reason through verbalized chains-of-thought, the generation cost grows substantially due to suboptimal paths to reach the final answer. Furthermore, as new insights are discovered while observing various input queries (e.g. through self-reflection), limited mechanisms exist for carrying forward these findings to be applied to subsequent problems. One can view the list of such strategies or behaviors as a growing cheatsheet, with elements retrieved from this memory module at inference-time. In this work, we consider structured cheatsheets, with learned clusters of behaviors. We introduce a Hierarchical Dirichlet Process Gaussian Mixture Model (HDP-GMM) over behavior embeddings, which shares components across domains while allowing domain-specific mixing weights, and uses the posterior predictive to retrieve relevant behaviors for a query; we call this a $\textit{Bayesian Cheatsheet}$. This mechanism allows for cheap adaptation in an online test-time training (TTT) setting, softly updating the mixture's sufficient statistics following each sample and enabling the creation of new components when the synthesized behaviors are sufficiently novel. We demonstrate that Bayesian Cheatsheet achieves clear performance gains relative to existing memory modules across reasoning benchmarks such as AIME'25, Omni-MATH, and PhysReason, even in the cold-start setting. We show that the Bayesian Cheatsheet is an adaptively reorganizing memory module, as behaviors can be re-assigned to components through a single step of collapsed Gibbs sampling. Our findings highlight the value of Bayesian-inspired memory modules for effective test-time adaptation and the role of structure in metacognitive reasoning.
|
| 572 |
Anatomy of LLM Sycophancy: What a Flip Rate Hides
2610.06522
|
cs.CLcs.LGcs.AI
|
Haonan Huang |
A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol, we test how user pressure and evaluation settings shape measured flip rates. Each measure...A model under pushback can correct itself, capitulate, or hold, and one flip rate counts a correction and a capitulation alike. Using SycoLens, a modular replay protocol, we test how user pressure and evaluation settings shape measured flip rates. Each measurement is one stateless replay of an item, a committed answer, and one scripted user line in a fixed form. Every effect is read against a matched control with the line deleted. Pushback wording, committed text, answer format, boundary distance, and ground truth become factors of one instrument; earlier instruments vary one to three of them. Across eleven frontier models from three providers and about 760,000 controlled replays, which models look sycophantic depends on how the user pushes back. Lines that assert the opposite verdict and lines that challenge the answer without asserting one rank the models almost unrelatedly. Flip effects grow several-fold near a model's boundary, yet items answered identically in every screening draw still carry about half of the most-affected totals. On arithmetic tasks where the truth is known, one model re-derives and corrects itself under pressure while another abandons correct answers without written work. On the model tested, a planted derivation lowers release of the answer it argues for, true or wrong, where a bare stated value does not; the wrong answer is corrected much more often than the true one is abandoned. Under a yes/no readout the rankings come closer, entangled with a pressure-induced shift toward "no". One score per model therefore compares different behaviours across models and benchmarks. We condense these dependencies into a reporting profile; the instrument, records, and analyses will be released upon publication.
|
| 573 |
Before Agent Tells The Lie: Has Deception Already Been Represented?
2610.06576
|
cs.CL
|
Xinling Li, Dadi Guo, Qingyu Liu, Qinghua Mao, Yi R. Fung |
Large language model (LLM)-based agents can exhibit deceptive behavior during task execution, including hiding failures, fabricating results, or falsely signaling task completion. Existing monitoring approaches mainly detect deception after it appears in obser...Large language model (LLM)-based agents can exhibit deceptive behavior during task execution, including hiding failures, fabricating results, or falsely signaling task completion. Existing monitoring approaches mainly detect deception after it appears in observable actions or outputs. In this paper, we investigate whether deceptive behavior can be predicted from an agent's internal representations before it becomes externally visible. We frame deception monitoring as a trajectory-level representation analysis problem and align agent trajectories around key decision points. Using hidden states extracted before these points, we show that future honest and deceptive outcomes can be reliably distinguished, with predictive signals remaining detectable several model calls before the final decision. We further characterize the temporal evolution of these signals: deception-related representations are weak early in execution but become increasingly identifiable as trajectories progress, while transferable structure can emerge before the strongest decision-adjacent signals appear. Finally, we intervene on the identified honest-deceptive representation directions during inference and find that activation steering reduces downstream deceptive behavior, suggesting that these representations influence agent decisions. Our findings indicate that agent deception is an evolving internal process that can be detected and potentially mitigated before it is expressed externally.
|
| 574 |
Molecules of a Story: Community Detection in PMI-weighted Narrative Networks
2610.06600
|
cs.CL
|
Kasper Fyhn, Rebekah Baglini |
Automatically extracted narrative networks -- graphs with entities as nodes and their relations as edges -- have proven useful for revealing central narrative structures through salient entities and their connections (Tangherlini et al. 2020; Labatut and Bost ...Automatically extracted narrative networks -- graphs with entities as nodes and their relations as edges -- have proven useful for revealing central narrative structures through salient entities and their connections (Tangherlini et al. 2020; Labatut and Bost 2019). But a narrative is more than those central structures that everything else revolves around. This work is concerned with the everything else: brief sub-plots, small clusters of descriptions, or associations between minor characters that go under the radar at the macro-level. We present an approach to unearth such peripheral structures. They involve rare entities with limited textual presence, overshadowed by dominant entities and lost among each other in the long tail of many but rare entities (Baayen 2001). We leverage the known tendency of pointwise mutual information (PMI, Church and Hanks 1990) to inflate for rare events, turning its weakness into a strength by weighting edges with PMI to foreground peripheral entity configurations. Communities extracted from the resulting network are structural traces of underlying narrative elements. We demonstrate the approach on The Lord of the Rings. From measures of how concentrated or dispersed a community's activations are across the text, a typology emerges that reveals that peripheral structures form more than a single class: episodic passages, echoing long-distance connections, and recurring threads each surface as distinct configurations. The approach is conceptually simple and surfaces fine-grained narrative details that are lost in abundance, though its deliberate amplification of weak signals comes with inherent sensitivity -- best understood as a lens for exploration rather than a robust extraction pipeline.
|
| 575 |
Word-Level Text Unmixing via Evidence-Preserving Ownership Routing with Language Models
2610.06603
|
cs.CLcs.AI
|
Jinglin He, Siyang Jiang, Lixing He, Guoliang Xing, Hongkai Chen |
Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: gi...Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly. Directly generating separated texts with LLMs can omit, duplicate, or hallucinate words, violating this exact-reconstruction objective. We therefore propose Evidence-Preserving Ownership Routing (EPOR), which decouples source-ownership prediction from reconstruction. EPOR adapts a causal LLM to predict canonical ownership routes conditioned on the mixed stream and prior routing decisions. At inference, completion-safe constrained decoding is combined with deterministic indexed reconstruction, yielding structurally valid K-source partitions that preserve every observed occurrence exactly once. We also introduce UNMIXBENCH, covering controlled synthetic mixtures, timestamp-derived speech from AMI and ICSI, layout-derived document streams from ReadingBank, and simulated concurrent digital outputs. Across five evaluation tracks, a 4B EPOR model achieves the lowest mean minimum-permutation word error rate among finetuned baselines, reducing the five-track mean by 22.3% relative to compact source-array generation and remaining competitive with zero-shot frontier LLMs. These results show that when lexical evidence is fully observed, separating ownership inference from lexical regeneration provides a reliable alternative to direct generation.
|
| 576 |
JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications
2610.06625
|
cs.CL
|
Matthew DiGiuseppe, Steven Denney |
Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed a...Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.
|
| 577 |
Long-Horizon Textual World Modeling through Structured Reasoning
2610.06637
|
cs.CL
|
Fangxin Wang, Xiang Gao, Yuguang Yao, Kaiwen Dong, Nikash Walia |
World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step ...World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step transition model, but intermediate errors can compound over time. Multi-step dynamics models instead condition on a sequence of future actions and predict their consequences directly, but become harder to learn as horizon grows: the model must track interacting state changes across the trajectory, endpoint supervision provides weak credit assignment, and intermediate predictions can remain plausible while losing information needed for later states. We show that these challenges can be addressed by casting the internal evolution of a multi-step transition as structured reasoning over textual world states: reasoning over sparse state changes reduces the burden of state tracking, a predictive-gain objective rewards the learned state for improving over a matched predictor that conditions on raw history instead, and intermediate predictive rewards supervise each state along the trajectory. Because these intermediate states are explicit textual representations of the world, they provide semantically meaningful targets that can be inspected, scored, and corrected during training. Across ScienceWorld, Jericho, and CEO-Bench, our approach achieves the strongest average long-horizon performance against recursive and non-recursive baselines that condition directly on raw history, with gains increasing at longer horizons. In a controlled counterfactual study, our model is also the only one with statistically significant sensitivity to future actions.
|
| 578 |
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches
2610.06647
|
cs.CL
|
Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang |
Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning...Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. Across reasoning tasks, LoGRA reduces average training memory by up to 45.7\% without sacrificing performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the \href{https://github.com/skzhang1/labs-molt/tree/logra/examples/scripts/logra}{Molt library}.
|
| 579 |
Representation-Space MMD for Diffusion Language Models
2610.06648
|
cs.CLcs.LG
|
Ilya Drobyshevskiy, Ilia Sudakov, Maksim Semenov, Denis Kuznedelev, Maksim Ignatov |
We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features...We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for continuous models. In both cases, computing the loss directly from these features enables efficient post-training without full sampling trajectories or jointly trained auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, we increase decoding parallelism with similar or higher accuracy on math and code benchmarks.
|
| 580 |
Wikidata Search Traces: A Dataset for Training Knowledge Graph Search Agents
2610.06650
|
cs.CL
|
Mohamed Chenene, Carlos Rosas-Hinostroza, Pierre-Carl Langlais, Anastasia Stasenko |
Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer...Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains their relations. Language models offer a natural-language alternative but answer largely from memory, which is least reliable for less prominent entities. We study agents that instead answer by exploring the graph, and argue that two obstacles limit them: the lack of training data recording how a solver explores, and interfaces that add large graph results directly to the model's context. We test three hypotheses: that the difficulty of graph search can be controlled through the structure of a question rather than only through obscure entities or wording; that much of the failure on long-horizon search comes from how retrieved evidence is managed rather than from the model itself; and that, in a suitable environment, open-weight models can match commercial closed ones. We construct multi-hop questions on a frozen Wikidata snapshot by replacing named entities with nested conditions, checking after each expansion that the target remains unique and that every new condition is necessary. We release 10,235 solving traces over single-entity and multi-hop questions, together with the recursive language model (RLM) harness that produced them, in which models batch graph calls, keep results in persistent Python state and interpret selected evidence through sub-calls. On 100 questions, the harness improves both models we ran under both interfaces compared with direct tool calling over the same functions: gpt-6-luna rises from 49 to 61 correct answers, doubling its multi-hop accuracy, and Qwen3.8-27B, an open-weight model served on a single GPU, from 60 to 74.
|
| 581 |
Reward Stealing Attack on Large Language Models
2610.06670
|
cs.CL
|
Jiaming Qian, Pengyang Zhou, Jiahe Xu, Chaochao Chen |
Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing A...Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at https://github.com/GarminQ/ReSA.
|
| 582 |
Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation
2610.06689
|
cs.CL
|
Jiaming Qian, Huiyan Yang, Mandi Liu, Jie Liu, Wenkai Shen |
Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a ...Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the returned evidence can reduce subsequent search. We introduce Programmatic Search Agent (PSA), which makes a local executable computation over candidates the unit of a search action. PSA unifies a persistent candidate workspace, flexible primitive composition, and selective evidence presentation. It incrementally generates program cells that reuse candidates, execute dependent operations, and select what the agent inspects next. The runtime resolves specified data dependencies within each cell, while the agent adapts its search strategy across cells as new evidence arrives. We compare PSA with the Query-based Agent and Tool-based Agent on InfoSeek-Eval and BrowseComp-Plus using five policy backbones without task-specific training. All three interfaces share the search substrate, and the Tool-based Agent also shares PSA's primitives and persistent workspace. Relative to the Query-based Agent, PSA improves macro-averaged task success by 4.00 and 7.56 percentage points on the two benchmarks, respectively; within-backbone reductions in final-step tokens average 28.3% and 33.9%. These results support extending agent control beyond query reformulation to the processing and presentation of retrieved evidence. Code will be released subject to approval.
|
| 583 |
MedPrune: Topology-Efficient Multimodal Multi-Agent Communication Evolution for Medical VQA Tasks
2610.06695
|
cs.CL
|
Jiuheng Wan, Runze Li, Chen Chen, Tingyuan Hu, Daiyang Yu |
While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant com...While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, an efficient medical multimodal multi-agent collaboration framework that dynamically prunes both nodes and edges from the communication topology to enhance reasoning ability and token efficiency. Specifically, we first formulate the diagnostic process as a heterogeneous communication graph, where nodes represent specialist agents from various departments and edges capture intra- and inter-departmental interactions. Building on this graph, we introduce two sparsification mechanisms to enable adaptive collaborative evolution: (1) Heterogeneous Node Sparsification, which eliminates task-irrelevant specialist agents irrelevant to the current multimodal question via reinforcement learning-driven topological optimization, and (2) Heterogeneous Edge Sparsification, which selectively retains only the most diagnostically salient intra- and inter-departmental connections by jointly optimizing task performance and topological complexity. Extensive medical VQA experiments under full-set and few-shot training settings prove MedPrune surpasses multi-agent baselines and boosts token efficiency with strong adversarial robustness.
|
| 584 |
SAFE-MR: Evidence Sufficiency Learning for Selective Multimodal Rumor Detection
2610.06708
|
cs.CL
|
Shiwen Ni |
Multimodal rumor detectors increasingly rely on retrieved evidence, yet relevant evidence is not necessarily sufficient for verification. Missing provenance, duplicated reports, and unresolved contradictions can produce confident predictions without adequate s...Multimodal rumor detectors increasingly rely on retrieved evidence, yet relevant evidence is not necessarily sufficient for verification. Missing provenance, duplicated reports, and unresolved contradictions can produce confident predictions without adequate support. We introduce SAFE-MR, a framework that separates claim veracity from evidence sufficiency. The method decomposes image-text posts into verifiable claims, constructs a relation-aware claim-evidence graph, and aggregates evidence using provenance and contextual compatibility. Separate veracity and sufficiency heads support selective prediction, while evidence interventions encourage stability under irrelevant additions and sensitivity to evidence removal. On NewsCLIPpings, VERITE, and XFacta, SAFE-MR achieves macro-F1 scores of 91.2%, 75.8%, and 85.2%, respectively. Against the matched backbone with evidence, its macro-F1 gains are 2.2, 4.9, and 4.8 percentage points. On the diagnostic selection set, SAFE-MR reduces AURC from 0.105 for maximum-probability rejection to 0.075 and lowers error at 80% coverage from 13.8% to 8.5%. Evidence-perturbation and ablation results support the role of sufficiency learning and intervention training in improving selective verification.
|
| 585 |
Domain adaptation of Russian ModernBERT for long legal documents
2610.06715
|
cs.CL
|
I. Litvak, D. Gvozdetsky, F. Lashkin, V. Kirova, S. Lagutin |
We investigate whether continued pretraining on Russian legislative documents improves a Russian ModernBERT encoder on legal text. The adapted model, RuModernBERT-ruLaw, was trained on a corpus reported to contain 304,382 legislative documents and 194,425,905 ...We investigate whether continued pretraining on Russian legislative documents improves a Russian ModernBERT encoder on legal text. The adapted model, RuModernBERT-ruLaw, was trained on a corpus reported to contain 304,382 legislative documents and 194,425,905 corpus tokens. Corpus token counts are distinguished from positions produced by the model tokenizer. We compare the original and adapted encoders on a fixed external collection of 1,031 court-decision segments. Both models receive the same hidden positions in each of five masking realizations. At maximum input lengths of 512, 2,048, and 8,192 tokens, mean masked-token cross-entropy decreases by 0.10942, 0.07052, and 0.06604 natural-log units, respectively. The reported 95% intervals summarize sensitivity to masking on this fixed collection; they do not quantify uncertainty across document collections. A second evaluation addresses legal-entity extraction. The original and adapted models achieve entity-level F1 scores of 0.99852 and 0.99820. However, 99.95% of test spans have the same normalized surface form and class in the training split. This evaluation therefore provides limited evidence about transfer to previously unseen forms. The paper explains the masking objective, overlapping windows, averaging rules, and exact entity-boundary scoring using editable diagrams and clearly marked illustrative examples. The comparison supports lower masked-token prediction loss for the studied pair of models and collection. It does not isolate the contribution of distant context or establish practical legal utility.
|
| 586 |
Improving Diversity in LLM Short Story Generation
2610.06729
|
cs.CL
|
Zahra Solati Dehkordi, Vasileios Lampos |
Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generation. Drawing on established writing conventions and known LLM limitations, we target variation...Large language models (LLMs) can generate accurate responses, but these are void of diversity. We attempt to address this for the task of creative short story generation. Drawing on established writing conventions and known LLM limitations, we target variation in genre, tone, style, and named entities. To promote diversity across these dimensions, we introduce DivLM, an LLM post-training framework consisting of two phases. First, we perform continued pre-training on a creative writing corpus and restore instruction-following capabilities using weight residuals. We then apply reinforcement learning with a custom, composite reward function that jointly maximizes diversity across the targeted narrative dimensions while maintaining response quality. Our empirical results on two LLM families show that DivLM increases diversity metrics by more than 9% on average compared to alternative approaches, while preserving instruction following, overall response quality, and similarity to human outputs.
|
| 587 |
ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring
2610.06744
|
cs.CLcs.LG
|
Sait Furkan Teke (ufak AI) |
ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an e...ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an expected error that serves as a "not sure" signal, without generating text and in one CPU forward pass for up to ten options. Built on the lab's ufakzeka-1-base, its head scores each option blind to the others at shared positions, so the answer does not depend on option order. A sequential head trained with shuffled options was about as accurate but changed 2.3 to 2.8 percent of its answers when only the option order changed; REINFORCE lost 10.2 points (0.102) of macro F1 to cross-entropy. On the open set of HakemBench v1.0 (4,275 questions, 7 tracks) the released model ranks 7th of 16 rows with a composite of 0.660 (95% interval 0.642 to 0.677). Temperature scaling lowers calibration error (smooth ECE) on the development set but raises it on held-out support questions, from 0.027 to 0.045 for the first scored run, which never trained on them; the released model later trained on them, so its 0.036 to 0.064 is not an unseen-question test. The released model is the last of three runs scored on HakemBench, and its numbers are not blind. The second run's new training data was aimed at the first run's errors on the full test set in guardrails, moderation and customer support, and the released run was trained after the second run's guardrail results on the full test set were read, under a protocol fixed in writing before any of its data, code or runs. All its numbers come after these readings; its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16. Weights and code are under Apache-2.0.
|
| 588 |
Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
2610.06750
|
cs.CLcs.AI
|
Hyunji Lee, Joykirat Singh, Zaid Khan, Justin Chih-Yao Chen, Elias Stengel-Eskin |
Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and r...Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and recurrent layers offer complementary pathways to use past information: attention supports precise memory recall from earlier tokens, while recurrent layers support consolidation of disparate information over long contexts. However, we observe that simply having access to both pathways does not mean that hybrid LMs are effectively using them. We find that they rely substantially more on attention than on the recurrent state. Standard supervised fine-tuning improves overall performance but does not improve how the two memory pathways are coordinated: the model becomes more reliant on information propagated by attention layers, while its use of information propagated by recurrent layers remains limited. To encourage better coordination between the two memory pathways, we add an auxiliary loss that limits attention's access to earlier context while the recurrent state propagates through the full sequence. This objective encourages the model to retain and use information through the recurrent pathway alongside attention. It improves overall performance, with particularly strong gains on tasks involving longer contexts or requiring information aggregation, consistent with the strengths of recurrent layers observed in analysis. Crucially, this imbalance and the benefit of our auxiliary loss generalize: they apply to multiple recurrent-attention LMs in question-answering and agentic tasks, as well as to attention-based LMs that combine different forms of memory. Together, our findings show that simply providing multiple memory pathways does not ensure their effective use, and that targeted supervision is needed to better coordinate them.
|
| 589 |
IdeaLens: Detecting AI Ideas in Long-form Writing
2610.06778
|
cs.CLcs.LGcs.AI
|
Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz B\"ol\"oni-Turgut |
While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (id...While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimizing word-level overlap with the raw text. We train IdeaLens on 1M FineWeb documents with silver labels from Pangram, a prose provenance detector. Since the outlines are largely stripped of surface-level information, the labels must be fit mainly through the ideas. In a controlled study, IdeaLens's AI flag rate drops from 95% to 7% as models write from increasingly detailed human plans, while Pangram 4 still flags 92%; from AI-derived plans, IdeaLens stays above 96%. Conversely, on a new dataset of 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% of the stories as AI, compared to 8% for Pangram 4. On a comprehensive suite of 19 existing detection benchmarks, we show that IdeaLens maintains strong detection rates at low false positive rates, suggesting that ideas themselves provide a powerful discriminative signal, and its performance holds across domains, formats, and languages. Finally, we examine 90K predictions from IdeaLens to characterize systematic differences between human and AI ideation. We release our models and labeled datasets to facilitate future research on idea provenance detection.
|
| 590 |
T-Search: An Open Agentic Retriever and Playground for Hard Multi-Step Search
2610.06782
|
cs.CL
|
Olga Tsymboi, Ramil Latypov, Aleksandr Medvedev, Danil Taranets, Dmitrii Stoianov |
We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answe...We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a downstream model, so backend and generator can be swapped without retraining. T-Search is built on Qwen3.6-35B-A3B and trained on adversarially filtered synthetic search tasks with round-sliced supervised fine-tuning followed by GSPO on a recall reward. Averaged over seven English and Russian benchmarks with gold evidence annotations, it reaches 56.0 Recall@10 with one rollout, 14.4 points above its base, and 61.3 with three fused rollouts, outperforming larger open models. We release the model, harness, live demo, and three benchmarks, including TRuST, the first native-Russian hard-search benchmark.
|
| 591 |
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
2610.06829
|
cs.CLcs.LGcs.AI
|
Yifan Zhang, Yutong Dai, Viraj Prabhu, Zhiyuan Hu, Ran Xu |
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are ...Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.
|
| 592 |
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
2610.06830
|
cs.CLcs.LGcs.AI
|
Haozhen Zhang, Haodong Yue, Quanyu Long, Jianzhu Bao, Qingyuan Liu |
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and d...Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.
|
| 593 |
Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery
2610.03872
|
cs.CLcs.LGcs.AI
|
Peter Chen, Wotao Yin |
AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause...AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework that links numerical diagnosis to reusable solver self-improvement. ADSD follows a diagnosis-first paradigm that first explains why a solver performs poorly, then uses this diagnosis to guide the discovery of appropriate numerical methods. The resulting knowledge is packaged into reusable solver skills, turning solver improvement from trial-and-error editing into a structured process of diagnosis, discovery, and implementation. Across four challenging numerical domains--power flow equation, AC optimal power flow control, stiff ordinary differential equations, and heterogeneous diffusion PDEs--ADSD consistently improves solver accuracy, robustness, and efficiency. On GOC-500 power flow, for example, ADSD reduces mean solver error by nearly $71\times$, with improvements further transferring to unseen grid topologies and operating regimes.
|
| 594 |
Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities
2610.03894
|
cs.CLcs.LGcs.AI
|
Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra |
A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistake...A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier models are highly repetitive, reproducing the same call across samples. We recover the missing signal from a low-cost open-weight surrogate run in parallel. It reads the same context, schema, and proposed action as the agent, then scores the call from its own log-probabilities through a family of complementary readouts: teacher forcing and request-PMI weigh the likelihood of each argument value, a discriminative verdict judges the call as a whole, and tool-choice competition tests the function against its siblings. One principle says which to trust: a generative likelihood localizes wrong argument values, while the verdict catches holistically wrong calls. When the error type is unknown, an ensemble is the low-regret default. The readout is training-free, needs no access to the agent's internals, and costs one prefill pass alongside the tool call. On difficult coding tasks it reaches AUROC 0.825 where the actor's stated confidence is near chance (0.598), and the generative readouts beat it by +0.07 to +0.28 across three further actors. Against self-consistency it gains +0.14 to +0.19 on near-deterministic actors, at 1/K the cost. The signal drives two deployment modes: a real-time gate escalating the least-trustworthy calls for review (+0.05 to +0.30 accepted-action accuracy at 50% coverage), and confidence feedback, returning the tool result with the score so the agent adapts its next step -- lifting task success on live-execution benchmarks (+0.119 and +0.137, p <= 1e-4) and beating a random-value control where step errors are silent (+0.078, p = 0.003).
|
| 595 |
Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas
2610.03998
|
cs.CLcs.AI
|
Khashayar Pourtaheri, Ahmad Zareei |
Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synth...Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents predict each individual's later choices, using five conditions that add progressively richer information: no personal information, demographics, personality traits, cognitive scores, and finally the respondent's earlier survey choices as behavioral history. We use a two-wave panel of 845 US adults who completed measures of 14 behavioral biases (spanning risk, time preferences, overconfidence, and reasoning), so each respondent's earlier answers provide a human test-retest benchmark; in the behavioral-history condition, all items that score the target bias are withheld. At the population level, the average number of biases per respondent in every condition is close to the human average (7.1-8.1 biases, against 7.1 for humans). This aggregate similarity masks differences in variance: persona descriptions recover only 53-67% of human between-person variation, whereas adding behavioral history restores it to approximately the human level. At the individual level, description-based personas achieve only 7-12% of the informedness observed in human test-retest responses, while adding behavioral history raises this to 28%. The condition including behavioral history has the highest estimated informedness in all 17 demographic groups, whereas description-based conditions provide little or no information for some groups. Synthetic responses also exhibit stronger education- and income-related differences than human responses. For LLM synthetic personas, a respondent's past answers add more to individual-level prediction than a description of who they are.
|
| 596 |
Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models
2610.04017
|
cs.CLcs.LGcs.AI
|
Sankaran Vaidyanathan, Rafal Urbaniak, Emily Bunnapradist, Michelangelo Naim, Daniel Waxman |
Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading ...Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies the structure of such interactions via witnesses: variables that provide contextual information to resolve interaction terms. However, estimation with witnesses typically requires combinatorial enumeration and is infeasible in practice. We introduce the witness-integrated set effect (WISE), a family of causal estimands that build on the witness mechanism while taking expectations over sets of causes and witnesses to remain computationally feasible. Building on this approach, we introduce JuntaLearner, a gradient-based circuit discovery method that learns to rank components by their causal impact across varying-sized sets of components and witnesses. Alongside faithfulness metrics, we introduce measures of necessity and task specificity, and the circuit recognition score (CRS) to summarize each metric across circuit sizes while emphasizing effects achieved by small circuits. Across tasks and models of increasing size, JuntaLearner achieves higher mean CRS compared to attribution baselines on all metrics. Since its cost does not grow with the number of candidate components, JuntaLearner scales to large models while accounting for set-level interactions and avoiding first-order approximations.
|
| 597 |
SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting
2610.04109
|
cs.CLcs.LGcs.AI
|
Mingtian Tan, Palash Goyal, Mihir Parmar, Sarkar Snigdha Sarathi Das, Chun-Liang Li |
Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augm...Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augmented approaches struggle with high noise, missing signals, and an inability to reason causally about event impacts. We propose SEER (Self-Evolving Event Reasoning and Retrieval), a closed-loop framework that dynamically optimizes event conditioning for time series forecasting. SEER translates prediction errors into two decoupled feedback mechanisms: (i) a reflective retrieval memory that refines subsequent search queries and filters spurious noise, and (ii) a persistent causal knowledge base that distills transferable domain dynamics. SEER enforces strict chronological boundaries across both event retrieval and reflection, preventing look-ahead bias and data leakage. Across six volatile time-series benchmarks, SEER consistently outperforms state-of-the-art time series foundation models and language model baselines.
|
| 598 |
What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document
2610.04128
|
cs.CLcs.LG
|
Georgios Politis, Evangelos Pappas |
Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We s...Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20% of tokens from the activations alone and 97.38% when it also sees the gradients, 3.17 percentage points more 95% interval [2.72, 3.64]. Counted by document, the difference is much larger. The attacker rebuilds 13.71% of 32-token documents exactly without the gradients and 37.77% with them, because a document only counts when every token is right. How we count also changes how good a defence looks. Secret mixup, which blends each outgoing vector with a decoy, stops the attacker from rebuilding almost any document exactly, yet the attacker still recovers 83-91% of tokens. In a second experiment on GPT-2 and Qwen3-0.6B, where the server trains only a run of consecutive layers, the layer at which the run starts changes both model quality and leakage, even when the run's length is fixed. We recommend reporting leakage both per token and per document, and treating what a split model sends as being as sensitive as the text itself.
|
| 599 |
InvestigationWorlds: An Agentic Environment for Legal Investigation
2610.04129
|
cs.CLcs.LGcs.AI
|
Albert Yu Sun, Andrew Benard, Sil Hamilton, Anna Teresita A. Marcelo, Yong Jae Kim |
We introduce InvestigationWorlds, an agentic environment for legal investigation. We build on an underused artifact of U.S. civil litigation: the summary judgment motion. This motion relies upon a record composed of real evidence exhibits, and results in a cou...We introduce InvestigationWorlds, an agentic environment for legal investigation. We build on an underused artifact of U.S. civil litigation: the summary judgment motion. This motion relies upon a record composed of real evidence exhibits, and results in a court-adopted hypothesis that is treated as ground truth for the purposes of deciding the motion. Each environment is built from a real U.S. Federal Court case retrieved from Public Access to Court Electronic Records (PACER) and augmented by an attorney-validated generation pipeline that synthesizes role-tagged documents around the original record. The resulting corpus admits multiple coherent factual readings, only one of which matches the court-adopted hypothesis. Evaluating on 100 cases, we find agents often commit to incorrect hypotheses despite retrieving relevant evidence, struggling to distinguish the court-adopted hypothesis from alternative hypotheses.
|
| 600 |
ExpertMuon-Compass: Alignment-Guided Step Sizes for Mixture-of-Experts Training
2610.04140
|
cs.CLcs.LG
|
Omatharv Bharat Vaidya, Ashwin Vinod, Pedram Akbarian, Aditya Sai Ellendula, Connor T. Jerzak |
Mixture-of-experts (MoE) language models send each token to a few experts, so each expert is trained on a different part of the data, and this part changes during training. With a shared learning rate, Muon applies updates of roughly the same size to expert ma...Mixture-of-experts (MoE) language models send each token to a few experts, so each expert is trained on a different part of the data, and this part changes during training. With a shared learning rate, Muon applies updates of roughly the same size to expert matrices of the same shape, even when an expert's update is poorly aligned with its current gradient. We here propose ExpertMuon-Compass (Compass), which multiplies the Muon step of each expert by two factors. A family factor compares the cosine between the expert's orthogonalized update and its gradient with the same cosine for the other experts in its layer. A scalar radius aggregates the alignment between corresponding rows of the update and gradient into one step-length multiplier. Compass keeps the update direction and the momentum buffer of Muon. In pretraining on FineWeb-Edu, Compass with Nesterov momentum on all matrices performs as well as or better than Muon, NorMuon, and other optimizers, with weight decay matched to NorMuon in the longer runs. Adding its factors to NorMuon gives the same or a lower loss than NorMuon. Compass is the most effective when the data seen by each expert varies during training, for example, as when the languages of a multilingual corpus arrive in separate blocks. With Compass, the expert load stays balanced, and the router assigns tokens to experts more decisively. We also prove a perturbation bound for the two factors.
|
| 601 |
Principled Top-$k$ Selection for Language Models with Hybrid Gradients
2610.04162
|
cs.CLcs.LG
|
Xuchen Gong, Junfei Sun, Tian Li |
Selecting the best $k$ items out of $m$ candidates is a critical component of modern large language model systems, such as document selection in Retrieval-Augmented Generation (RAG) and expert routing in Mixture-of-Experts (MoEs). However, training these selec...Selecting the best $k$ items out of $m$ candidates is a critical component of modern large language model systems, such as document selection in Retrieval-Augmented Generation (RAG) and expert routing in Mixture-of-Experts (MoEs). However, training these selection modules remains challenging due to weak gradient signals and suboptimal exploration-exploitation tradeoffs. Furthermore, prior works often rely on heuristics, lacking principled objectives and approaches that explicitly model and solve the top-$k$ selection problem. In this work, we propose a principled objective for training selection modules, whose gradient naturally provides richer training signals in a hybrid form---containing a supervised-gradient component and a policy-gradient component. We show that the selection problem becomes harder as $m$ increases, and our algorithm converges at rate $O(1/\sqrt{T})$, with the optimal upper bound achieved by balancing between bias and variance. Practically, we apply our method to a set of tasks involving top-$k$ selection, including synthetic regression problems, RAG, and MoE systems, showing that our method outperforms the baselines in next-token prediction perplexity and QA accuracy.
|
| 602 |
DimSteer: Steering LLM Authoring with Automatically Discovered Stylistic Controls
2610.04174
|
cs.CL
|
Ajit Mallavarapu, Ziwei Gu |
Large language model writing interfaces often make users steer outputs by repeatedly articulating desired changes in natural language. Yet writers may recognize useful stylistic directions only after seeing alternatives, making revision recall-heavy. We presen...Large language model writing interfaces often make users steer outputs by repeatedly articulating desired changes in natural language. Yet writers may recognize useful stylistic directions only after seeing alternatives, making revision recall-heavy. We present DimSteer, an authoring interface that samples prompt-local completions, discovers high-variance activation-space axes of variation, labels them, and exposes them as sliders with pole previews, diff comparison, and reset controls. Users can manipulate discovered dimensions, reducing the need to reformulate prompts for each stylistic adjustment. In a within-subjects study with 16 participants against a matched prompt-only baseline, DimSteer reduced mental demand, effort, and frustration while preserving comparable perceived success. Participants valued the surfaced dimensions, yet 15 of 16 disagreed that they would have thought to request the same changes in a prompt. Results suggest prompt-local controls can shift LLM authoring from recall-based prompting toward recognition-based exploration and direct manipulation, while preserving prompting for open-ended edits.
|
| 603 |
Language Model Activations Inhabit Privileged Error-Correcting Basins
2610.04183
|
cs.CLcs.LGcs.AI
|
Matthew Finlayson, Francisco Pernice, Eric Todd, Amir Zur, Daniel Wurgaft |
Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanis...Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward "good" regions that produce coherent outputs. To investigate these hypothesized error-correcting mechanisms, we probe the geometry of language model activation space by observing the action of model layers on low-dimensional curves. In doing so, we discover that model activations occur within a cluster of distinct attracting basins, which differentiate natural activations geometrically from distributionally similar synthetic activations. Applying this lens to language model steering, we observe feature-specific basins along semantic steering directions, and find that steering moves activations between these basins. To demonstrate the active role of this geometry in neural computation, we show that adaptively modulating steering strength to transport activations across basins improves inter-language steering, significantly increasing the probability of sampling tokens from the target language compared to fixed-strength steering. Our findings establish analysis of activation space geometry as a promising approach to interpreting and controlling language models.
|
| 604 |
Clean: Second-order LLM Training at Linear Memory Cost via Nystr\"om Sketching
2610.04204
|
cs.CLcs.LGcs.AI
|
Beheshteh T. Rakhshan, Sahar Rajabi, Maziar Sargordi Shikai Fang, Guillaume Rabusseau, Sirisha Rambhatla |
Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce ...Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer designed to resolve this bottleneck. Clean leverages the randomized Nystrom method to accurately approximate the left and right preconditioners in SOAP, and to reduce the optimizer's memory complexity from quadratic to linear in terms of model dimensions. We subsequently reintegrate the off-subspace components to capture curvature information beyond the low-rank approximation, preserving rich curvature at minimal memory cost. We further propose Q-Clean, a low-precision variant that aggressively compresses optimizer states. Q-Clean reduces optimizer memory consumption by \textbf{over 50\%} compared to Muon when pre-training a LLaMA-1.3B architecture, all while maintaining strong and competitive predictive performance. Notably, Clean operates with a smaller optimizer-state footprint than standard AdamW while reaching AdamW's final performance \textbf{26\% faster} in wall-clock time. Furthermore, our methods uniquely enable the pre-training of a 13B-parameter model on a single 80GB GPU, providing a scalable, efficient, and accessible approach to large-scale model optimization.
|
| 605 |
Multimodal Dual-Encoder Retrieval for Automated ICD Coding
2610.04263
|
cs.CL
|
Abhinav Bohra, Anuj Bohra |
Accurate International Classification of Diseases (ICD) coding is crucial for large-scale clinical research, documentation, and billing. There are three primary problems with current ICD prediction methods: (1) They are unable to comprehend multimodal patient ...Accurate International Classification of Diseases (ICD) coding is crucial for large-scale clinical research, documentation, and billing. There are three primary problems with current ICD prediction methods: (1) They are unable to comprehend multimodal patient data because they rely on either structured EHR data or unstructured clinical notes. (2) They also struggle with scalability to a larger amount of ICD codes (9K+ codes in ICD-9), as traditional classifiers need dense output layers and often do not generalize well to long tail rare diseases. (3) They lack transparency for clinical use. To address these challenges, this research proposes a two-stage framework that first retrieves ICD codes using a multimodal dual-encoder retrieval model, where structured and unstructured patient data are integrated through gated fusion. The second stage refines the top-k retrieved candidates with an LLM-based re-ranker that provides ranked codes with clinically relevant explanations. Our experiments show that the proposed approach improves Micro-F1 and Precision over a multimodal dual-fusion classifier baseline. These improvements demonstrate that combining a gated multimodal retrieval system with LLM-based re-ranking is a practical alternative to dense multi-label classification for automated ICD coding.
|
| 606 |
Rethinking Self-Distillation for Multi-Teacher Capability Merging
2610.04272
|
cs.CLcs.LGcs.AI
|
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu |
Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventi...Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inference and environment interaction costs incurred by MOPD, we find that much of its reported accuracy gain is due to certain training design choices and hyperparameter optimization, rather than the algorithm itself. We conduct a controlled self-distillation study across two multi-teacher settings, four models, and eleven benchmarks, comparing off-policy methods, namely supervised fine-tuning (SFT) and soft-label distillation, to hybrid teacher-prefix distillation and MOPD. We found that all four methods achieve \textit{nearly identical} accuracy. However, MOPD uses $14.8$--$23.1\times$ SFT's training GPU-hours. We also revisit four recently published comparisons between on-policy and off-policy distillation and find that the reported on-policy gains shrink substantially when SFT baselines are trained on rejection-sampled teacher trajectories and use independently tuned hyperparameters. As a training-free alternative, we also find that simple weight-merging methods can recover expert capabilities with minutes of CPU merging time, although their accuracy degrades as model size decreases and task interference increases. Overall, our results question recent gains reported due to MOPD and suggest careful tuning of more efficient off-policy baselines as a viable alternative.
|
| 607 |
First-Order Steering: Translating Weight Adaptation into Activation Steering
2610.04283
|
cs.CLcs.LGcs.AI
|
Sri Pranav Kunda, Alexander Kurz, Tomas Dominik, Uri Maoz |
Activation steering exploits interpretable directions in the residual stream to enable inference-time manipulation of model behavior. Composing steering vectors to apply multiple target behaviors simultaneously is important in various fields-including AI align...Activation steering exploits interpretable directions in the residual stream to enable inference-time manipulation of model behavior. Composing steering vectors to apply multiple target behaviors simultaneously is important in various fields-including AI alignment and safety-but remains a challenge for existing activation steering methods. In contrast, prior work in model merging shows that target behaviors represented by learned weight adaptations can be combined with high accuracy. A method that translates weight adaptations into activation steering vectors could therefore extend prior work in model merging to generate composable steering vectors that better enable simultaneous inference-time behavioral control. For this, we introduce First-Order Steering, a formulation of activation steering as a first-order approximation of weight update matrices parameterized by a vector of steering strengths, and establish theoretical bounds on the approximation error of first-order steering. We then develop a novel model merging procedure, HeRD-Merging, which minimizes the first-order approximation error terms to enable higher first-order steering accuracy. Together, our method produces steering vectors that control both individual and composed behaviors more accurately than existing activation steering methods. Furthermore, HeRD-Merging matches the performance of conventional model-merging baselines, while producing weight adaptations that admit more accurate first-order steering vectors.
|
| 608 |
Language-Conditioned Token and Reasoning Efficiency in Large Language Models: A Paired Cross-Lingual Study Protocol
2610.04295
|
cs.CLcs.AI
|
Genliang Zhu, Chu Wang |
Large language models incur language-dependent representation and inference costs, but existing comparisons often conflate input language, assigned observable-trace language, and answer realization. We specify a prospective paired study that separates these in...Large language models incur language-dependent representation and inference costs, but existing comparisons often conflate input language, assigned observable-trace language, and answer realization. We specify a prospective paired study that separates these interfaces while holding the semantic item, checkpoint, and answer oracle fixed. The initial design instantiates 240 exactly scored items rendered from templates in English and seven non-English languages, three distinct-lineage open-weight checkpoints, three trace-token budgets, 22 input- and trace-language conditions, a fixed answer reserve, and a separately counted delimiter: 47,520 initial core runs before prospective sample-size selection. RQ1-RQ3 estimate input and trace effects by intention-to-treat with failure-inclusive terminal accounting and test answer realization by cloning a sealed prefix and runtime-native KV state into eight crossed branches. Pre-freeze independent language review, fixed-form ASCII selectors, and code/surface/solver agreement constrain the realization test. H1-H5 share one Holm family and a global simultaneous component band. A secondary randomized experiment compares one-long-attempt and complete K-short-attempt policies at equal trace allowance under frozen seeds and oracle-blind aggregation; it is a full-policy contrast because answer capacity differs. Outcomes include exact token spans, correctness, latency, runtime-exposed memory, and qualified same-host operating-system-reported energy over prespecified hardware rails. The protocol separates tokenizer expansion, observable-trace cost, and answer-realization cost without treating visible traces as internal cognition or operating-system estimates as physical cross-device energy. No confirmatory model outcome is reported; result fields remain disabled until the frozen evidence ledger passes independent verification.
|
| 609 |
Questioning the Questions: Sustaining Self-Evolution in Reasoning Models
2610.04299
|
cs.CLcs.LGcs.AI
|
Jinyuan Li, Chengsong Huang, Langlin Huang, Donghong Cai, Shiping Gao |
Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analy...Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filtering further increases their proportion in training data. Second, existing question diversity controls based on lexical similarity can miss mathematically equivalent questions expressed in different ways, which leads to question diversity collapse in later training rounds. Building on these findings, we introduce R-Quest, which uses question validity and novelty feedback to guide self-evolution. We first train the solver to recognize and reject invalid questions, then use its judgments to guide questioner rewards and filter solver training data. To avoid question repetition, we use a frozen base model to compare sampled question pairs and provide novelty feedback. Empirically, our method consistently achieves the highest average performance on 12 benchmarks in mathematical reasoning, general-domain reasoning, and code generation across two model families. Additionally, R-Quest maintains stable performance gains over ten rounds of self-evolution, peaking in the final round and outperforming R-Zero by 17.32 points.
|
| 610 |
From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility
2610.04316
|
cs.CLcs.AI
|
Mohammad Mosafer |
Output-only safety monitoring sees only the end of a model's computation, yet the model computes its answer before emitting it: what it is internally poised to say is safety-critical. Jacobian-space (J-space) readouts, linear maps from hidden states to the out...Output-only safety monitoring sees only the end of a model's computation, yet the model computes its answer before emitting it: what it is internally poised to say is safety-critical. Jacobian-space (J-space) readouts, linear maps from hidden states to the output vocabulary via the model's input-output Jacobian, have been proposed as a window on behaviorally accessible internal content, and a first safety protocol (JADR) showed that danger recognition is readable there. What remains unknown is how this accessible content relates to the latent-space safety geometry studied by representation engineering, and how safety training shapes it. We introduce two quantitative bridges: (i) a transport-amplification profile $A_\ell$, measuring how strongly a latent safety direction is carried toward output space by each layer's Jacobian, and (ii) a paired base-vs-tuned protocol that attributes accessibility to training provenance. Across four small model pairs (135M to 1.5B, three lens-fit seeds), tuned checkpoints concentrate recognition in upper-middle layers, and at 0.5B DPO installs refusal while J-space recognition drops to chance: training can widen the accessibility gap exactly where behavioral safety looks best. A deployment audit closes the loop: monitor-aware GCG-style suffixes suppress the prompt-side monitor at zero behavioral cost at every scale; only a learned monitor rung resists its own adaptive re-attack; the training-time defense fails its re-attack at every penalty weight; steering shows the amplification profile is descriptive, not causal; and continuous-prefix optimization fails where discrete search succeeds. Safety monitoring, training, and evaluation must operate on accessibility itself, not on outputs alone.
|
| 611 |
Bidirectional Preference Synthesis: Learning Prompt-Conditioned Preferences from Boundary Failures
2610.04328
|
cs.CLcs.AI
|
Junbo Wang (Kuaishou Technology, Nanjing University), Lidong Lu (Nanjing University), Zhuoqun Li (Kuaishou Technology), Guiping Jiang (Kuaishou Technology) |
Correction-based offline preference pipelines commonly treat model failures only as rejected responses under the original prompt. This supervision is incomplete for boundary failures: responses that violate the given instruction yet coherently satisfy a nearby...Correction-based offline preference pipelines commonly treat model failures only as rejected responses under the original prompt. This supervision is incomplete for boundary failures: responses that violate the given instruction yet coherently satisfy a nearby intent or constraint setting. We introduce Bidirectional Preference Synthesis (BPS), a data-construction method for standard Direct Preference Optimization (DPO) that makes this missing prompt dependence explicit. For each validated boundary failure, BPS keeps the conventional forward pair under the original prompt and adds a reverse pair under a synthesized achieved prompt, so the same response is rejected where it is wrong and chosen where it is right, without changing the DPO objective, training a reward model, or requiring online sampling. On Qwen3-4B-Instruct-2507, BPS preserves original-side pairwise ranking while raising achieved-side ranking accuracy from 6.8% to 62.3% on held-out crossed anchors, with a similar shift under a Kimi-K2.6 cross-teacher probe. A blind human audit supports the intended reverse preference direction, and downstream evaluations show the clearest separation from Forward-DPO in multilingual multi-turn instruction following, with consistent capability-retention patterns on agentic, tool-use, and code checks.
|
| 612 |
Suppressing Pressure, Amplifying Evidence: Self-Guided Attention Steering to Mitigate Sycophancy and Stubbornness
2610.04329
|
cs.CLcs.AI
|
Yinghao He, Mengyu Xu, Haixiang Sun, Donghan Li, Yibo Wang |
Reliable language models should resist unsupported user pressure while effectively using objective contextual information. However, models may exhibit sycophancy by yielding to unsupported user pressure or contextual stubbornness by failing to update their ans...Reliable language models should resist unsupported user pressure while effectively using objective contextual information. However, models may exhibit sycophancy by yielding to unsupported user pressure or contextual stubbornness by failing to update their answers when relevant contextual information warrants revision. Evaluating interventions for these failures separately can obscure whether mitigating one failure exacerbates the other. To assess this trade-off, we introduce CoPE-Bench with six conditions per question: a neutral baseline, correct or incorrect user pressure, contextual information consistent with or conflicting with the neutral answer, and a joint condition combining incorrect claims with conflicting contextual information. To regulate the influence of user pressure and contextual information, we propose SPAE (Suppressing Pressure, Amplifying Evidence), a training-free framework that uses the model's own judgments to identify relevant tokens, suppressing user pressure and amplifying contextual information through token-level attention steering. On average across five backbones, SPAE reduces pressure following by 18.8 percentage points and increases joint-condition updating by 5.5 percentage points relative to the strongest baseline in the main comparison. In two-turn dialogue, it improves joint-condition updating by an average of 13.2 percentage points over the strongest prompting baseline. The source data and codes can be found at https://github.com/03Grant/sycophancy-and-stubbornness.
|
| 613 |
Factorized Delayed Streams Modeling for LLM-based Streaming ASR
2610.04333
|
cs.CLcs.SDeess.AS
|
Tatsunari Takagi, Kai Washizaki, Atsushi Kojima, Lianbo Liu, Koki Nikaido |
Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding token <p> and the word-start token <w> to the LLM vocabulary and predicts...Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding token <p> and the word-start token <w> to the LLM vocabulary and predicts them together with normal text tokens using the same softmax. We first show that <w> can be removed while maintaining competitive recognition performance. Based on this result, we propose Factorized DSM (F-DSM), which separates the waiting probability for <p> from the distribution over the original LLM vocabulary. This factorization removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. Experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM. It also greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.
|
| 614 |
ShadowMiner v1 - An Experience Report on Implementing and Measuring a Problem-and-Hypothesis Discovery Engine
2610.04339
|
cs.CLcs.AI
|
Jinhyuk Choi |
ShadowMiner v1 is a system that automatically discovers research problems and generates hypotheses from AI papers. It is a nine-stage pipeline. It structures documents into a knowledge graph and finds graph gaps in it - structural blind spots in research. Thes...ShadowMiner v1 is a system that automatically discovers research problems and generates hypotheses from AI papers. It is a nine-stage pipeline. It structures documents into a knowledge graph and finds graph gaps in it - structural blind spots in research. These graph gaps are included in the LLM generation prompt. Each generated hypothesis is then verified by checking whether it is already covered by existing research, scoring its quality, and checking that the facts it relies on are accurately drawn from its sources. This report does not propose a new generation or evaluation technique. It describes our experience of implementing and applying ideas from prior work, and measuring whether each one actually contributed.
|
| 615 |
Hierarchical Credit Assignment for RLVR on Fused Gromov-Wasserstein Geometry
2610.04344
|
cs.CLcs.LG
|
Qi Yu, Ruizhong Qiu, Zhichen Zeng, Xuying Ning, Yanjun Zhao |
Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens w...Reinforcement learning with verifiable rewards (RLVR) has been shown to improve the reasoning capability of large language models (LLMs) across diverse reasoning tasks. However, group-based RLVR methods, such as GRPO, assign a uniform advantage to all tokens within rollouts of the same outcome. While existing works refine credit assignment of GRPO based on local signals such as token locations or entropy, they often fail to capture the global semantic novelty of a reasoning behavior relative to the current policy. In this work, we propose a hierarchical credit assignment approach for group-based RLVR methods, called HarA, which identifies and encourages semantically novel reasoning behaviors during RLVR. HarA represents each sampled rollout as a distribution over the hidden states and locations of tokens, and computes the Fused Gromov-Wasserstein (FGW) barycenters of all rollouts with the same outcome, capturing the internal reasoning patterns in the latent space under the current policy. The semantic novelty of a reasoning element can then be measured by its contribution to the FGW distance between the current rollout and the barycenter. While solving the FGW formulation is expensive, we introduce an anchor-guided linearization that turns it into a Wasserstein formulation solvable via the Sinkhorn algorithm efficiently. By reweighing token-level advantage of group-based RLVR methods based on the novelty signals, HarA highlights novel reasoning behaviors at flexible granularities to encourage fine-grained LLM exploration. Extensive experiments across three group-based RLVR methods show that our plug-and-play method effectively enhances the exploration of LLMs, outperforming existing methods across diverse reasoning benchmarks.
|
| 616 |
Large Language Models and Augmented Democracy
2610.04412
|
cs.CL
|
Jairo Gudi\~no-Rosero |
Artificial intelligence enables computational agents to represent political preferences and take part in collective decision-making. In this thesis, I investigate the opportunities and challenges of digital twins (DTs) based on Large Language Models (LLMs) as ...Artificial intelligence enables computational agents to represent political preferences and take part in collective decision-making. In this thesis, I investigate the opportunities and challenges of digital twins (DTs) based on Large Language Models (LLMs) as intermediaries in augmented democracy, focusing on individual preference representation, collective representation of political organizations, and the vulnerability of those representations to attackers. First, using data from an online experiment in Brazil, I examine whether personalized DTs can predict citizens' preferences for unseen policy proposals. Second, I extend the DT framework from individuals to political organizations. Using Swiss parliamentary data, I build topic-specific knowledge graphs from lawmakers' legislative records and connect them to LLM-based lawmaker agents, which are organized into party-level DTs representing collective positions. Agentic deliberation among these agents tests whether aggregated party representations capture a broader range of intra-party perspectives than official party communications. Finally, I study the vulnerability and robustness of LLM-mediated deliberation against prompt-injection attacks that amplify viewpoints, suppress opinions, or redirect consensus. Using data from a 2023 deliberative experiment in the United Kingdom, I analyze how attack effectiveness varies with the distribution of opinions and rhetorical strategies, and evaluate a pipeline combining injection detection, structured opinion representations, and reinforcement learning to improve resistance. These findings characterize the opportunities and challenges of LLM-based digital twins in augmented democracy, stressing accurate preference representation, faithful aggregation, and robustness to strategic interaction.
|
| 617 |
The Same Zero: Why Identical ASR Can Imply Different Guarantees in LLM-Agent Security
2610.04504
|
cs.CL
|
YaJie Yin |
LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Aut...LLM-agent security has produced a dense landscape of defenses - prompt hardening, content filters, permission gates, sandboxes - yet no framework tells a deployer what a defense actually guarantees, or where that guarantee comes from. We apply Verification Autonomy Levels (VAL) - L0: LLM self-declaration; L1: deterministic rules; L2: objective ground truth; L3/L4: decidable completeness; L5: impossible - to 22 agent-security defenses; the taxonomy is falsifiable (10/10 prediction hits on frozen cards, flagged). We run the first controlled deployment-value comparison: at equal budget, a VAL-guided stack (confirmation gate + schema sandbox) versus a mainstream intuition stack (prompt hardening + keyword filter), 50 scenarios, 12 attack variants, adaptive/white-box/PAIR escalation (~7,000 testbed calls; ~10,000 harness calls on AgentDojo/JADE). The VAL stack holds 0.000 attack success at 1.000 benign success (0.5% ASR at 79.7% utility on AgentDojo banking vs 4.3% undefended); the intuition stack reaches 0.000 ASR but kills all benign actions - security by model-behavior luck, not structure. Across testbeds of rising attack-surface hardness the intuition stack's zero drifts (0->1.9%->6.2%, n=16 on JADE) while the VAL stack's holds within its ODD (0->0->0), its only breach a disclosed out-of-ODD password gap (0.5%). The same zero, two different guarantees: zero is an outcome, not a guarantee.
|
| 618 |
Autonomous Structuring of Radiology Reports Across Modalities at Archive Scale Using an Open-Weight Large Language Model
2610.04541
|
cs.CLcs.AI
|
Friedrich Puttkammer, Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf |
Purpose: To develop and evaluate an open-weight large language model (LLM) pipeline that converts an entire archive of free-text radiology reports into structured reports without human oversight. Materials and Methods: In this retrospective study, a pipeline w...Purpose: To develop and evaluate an open-weight large language model (LLM) pipeline that converts an entire archive of free-text radiology reports into structured reports without human oversight. Materials and Methods: In this retrospective study, a pipeline with 150 hierarchically organized templates was developed at one center and tested at a second center on reports from 2010 to 2025. The open-weight model gpt-oss-120B selects the template in three constrained-decoding steps and fills it on one local graphics processing unit. Template selection was scored against expert labels on 914 randomly sampled reports of five modalities, structuring quality on 920 radiography and CT reports corrected field by field by five residents. The pipeline then processed the complete archive of the second center. Proportions are reported with Wilson 95% confidence intervals (CIs). Results: An optimal template set was selected for 74.4% of reports (680 of 914; 95% CI: 71.5%, 77.1%) and an appropriate set for 82.3% (752 of 914; 95% CI: 79.7%, 84.6%), 87.7% for single-region and 54.1% for multi-region reports. Macro semantic textual similarity between output and corrected reference was 0.95 for radiography and 0.97 for CT, residents left 88.7% of 24,638 fields unchanged, and unsupported content was flagged in 1.0% and 1.5% of reports. Of 2,186,982 archive reports, 96.5% received structured output, 2,401,544 structured reports, at 1,258 reports per hour on one graphics processing unit. Conclusion: An open-weight LLM pipeline structured a complete multimodality report archive without human oversight with high content fidelity. Multi-region reports remained the main source of template errors.
|
| 619 |
Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs
2610.04580
|
cs.CLcs.AI
|
Hoigi Seo, Byung Hyun Lee, Minjun Kim, Dohyun Mah, Jongho Lee |
Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typicall...Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (\textit{e.g.}, audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer. Project page with code is available at \url{https://seohoiki3215.github.io/DCAT_project_page}
|
| 620 |
StegoMemory: Agentic Memory Acts as Covert Steganographic Channel
2610.04589
|
cs.CL
|
Snehasis Mukhopadhyay, Arun Nair |
Is agentic memory robust against stealthy steganographic attacks? We carry out a large-scale red-teaming exercise to test whether agents can encode attacker-controlled strings in one session and recover them in another without triggering safety oversight. Foll...Is agentic memory robust against stealthy steganographic attacks? We carry out a large-scale red-teaming exercise to test whether agents can encode attacker-controlled strings in one session and recover them in another without triggering safety oversight. Following SHADE-Arena-style tasks, we embed malicious side tasks to encode secret strings using steganography within otherwise benign tasks and evaluate them using independent task-completion and safety oversight. We test 14,000 attack trials spanning 91 model-attack configurations, 13 models, and 7 steganographic schemes: Acrostic, Mesostic, Telestich, Word-choice, Capitalization, Position-based, and Homoglyph. 41.2% of trials successfully encode the target payload in the generated response, while 29.8% survive the memory-write step with the encoding intact. Overall, 25.4% of trials yield a recoverable payload at retrieval, and 20.1% achieve exact-match recovery, including synthetic API keys and cryptographic material never exposed to the receiver. Among the exact-match transmissions, 71.4% also pass primary task-completion scoring and are independently judged safe by the oversight model, corresponding to 14.4% of all trials in which a successful covert transmission would appear to be an ordinary, benign interaction under task-level evaluation. Our results demonstrate that agentic memory can function as a persistent cross-session covert channel. The results further show that the principal bottleneck occurs at memory persistence rather than retrieval: once a steganographic payload survives the memory-write stage, a substantial fraction remains recoverable. We therefore argue that memory integrity, information-flow control, and covert-channel detection should be explicit security requirements for agentic systems.
|
| 621 |
When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning
2610.04686
|
cs.CLcs.LGcs.AI
|
Zihao Zhao, Tunyu Zhang, Haizhou Shi, Yusong Zhao, Xinxi Zhang |
Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses ...Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. We argue that successful debate requires two distinct mechanisms: proposal supply must surface a correct answer, and readout must identify that answer when voting misses it. We formalize the first requirement through recoverable headroom, which measures cases where a correct proposal is available but the majority answer is wrong. For the second, we develop Latent Verification Debate (LVD), an accounting model in which candidate proposals receive answer-specific verification evidence before final generation. Controlled fixed-proposal interventions estimate this latent effect in equivalent peer-support units and show that correct evidence changes answer probabilities and generated decisions while proposal supply remains fixed. To improve proposal supply, we construct societies from neural-thicket agents using labeled and label-free coverage objectives. Across two backbones and matched-budget reasoning benchmarks, coverage-selected societies increase complementary proposal supply and improve aggregate accuracy in repeated stochastic evaluations. Round-level controls further show that interaction provides gains beyond applying the same finalizer directly to the initial proposals. These results identify proposal coverage and truth-sensitive evidence use as complementary conditions for debate to outperform voting. Code is available at https://github.com/Wang-ML-Lab/when-debate-helps.
|
| 622 |
SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation
2610.04690
|
cs.CLcs.LGeess.AS
|
S\'everin Baroudi, Herv\'e Bredin, Ricard Marxer |
Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces m...Self-supervised learning (SSL) is standard for speech representation learning, but mainstream models are designed around single-speaker audio, limiting their usefulness in multi-speakers scenarios. We present SepRQ, an open-source SSL framework that replaces masked prediction with a pseudo-source-separation objective over frozen random-projection codebooks. By adopting a novel mask-free, multiresolution approach, SepRQ achieves state-of-the-art performance in Speaker Diarization and Speech Separation on the SUPERB benchmark, surpassing WavLM and other cocktail-party derived SSLs at both Base and Large scales, while requiring only 85.68M inference parameters. SepRQ also demonstrates strong performance across target-speaker tasks requiring enrollment (such as Target-Speaker Automatic Speech Recognition), and on the challenging multi-domain DIHARD 3 diarization dataset. Notably, we report strong separation capabilities on three-speaker mixtures (WSJ0-3Mix), where current SSL literature struggles. While cocktail-party SSLs remain scarce and closed-source, limited to C-HuBERT and the enrollment-based SA-WavLM, we open-source SepRQ to the community.
|
| 623 |
Penumbra: Sample-Efficient Adversarial Search for Regulatory Obligations
2610.04693
|
cs.CLcs.AI
|
Anthony Rhodes |
Agents are entering finance, healthcare and law, sectors where a violation leaves no lexical signature and carries real penalties. Whether an omission is material, or a disclosure sufficient, depends on what the response left out. Probing such an obligation me...Agents are entering finance, healthcare and law, sectors where a violation leaves no lexical signature and carries real penalties. Whether an omission is material, or a disclosure sufficient, depends on what the response left out. Probing such an obligation means finding responses one minimal edit from flipping compliance, and every probe costs a generation and two adjudications, so the binding constraint on regulatory red-teaming is sample efficiency, not volume. We introduce Penumbra, an adversarial search that walks from a verified anchor under an expanding edit budget until a two-evaluator committee changes its verdict, and emits the two adjacent responses that straddle the change. Allocation is adaptive, and the objective is coverage of the defeat surface: distinct (obligation x defeat mode) cells resolved per candidate. At matched budget, adaptive allocation reaches uniform allocation's full-budget coverage on 59% of the candidates; at equal records it covers 1.43x the defeat modes of naive enumeration, and the gain is confined to the axis it targets. On 60 screened obligations of a financial advisory constitution, Penumbra returns 144 pairs, each a compliant and a violating response that the committee places on opposite sides of the boundary, differing by a handful of words where a model asked for both directly produces texts sharing almost nothing. A second constitution, for clinical triage, reproduces this on 18 obligations: 49 pairs at the same tightness, with the same modes hardest. Pairs like these show where an obligation's own terms stop deciding, which is what an agent deployed under it must be tested against, and the search finds them at a cost that scales with the boundary, not the text.
|
| 624 |
Not Self-Decidable: LLMs Cannot Draw the Boundary of What an Agent Verifier Can Check
2610.04699
|
cs.CLcs.AI
|
Anthony Rhodes |
A verifier for an agent faces rules of two kinds: the ones a fixed check can settle and the ones that require a judge. A team that derives its own checks fixes that split up front. Where the requirements come from outside, as in finance, healthcare and law, th...A verifier for an agent faces rules of two kinds: the ones a fixed check can settle and the ones that require a judge. A team that derives its own checks fixes that split up front. Where the requirements come from outside, as in finance, healthcare and law, the agent enforces rules it did not write, so the split falls to runtime, recurring for every predicate of every rule on every action at a rate no reviewer can audit. Every escalation scheme assumes a model can make that decision itself, that it is self-decidable. Across six corpora, including the EU AI Act, FINRA guidance and a deployed credit agent, we collect roughly 22,000 labels from four models built by three labs. They agree almost perfectly where the answer is obvious and collapse on regulatory text; their errors run in opposite directions, so no model can be trusted as the conservative choice; and on the deployed agent's own rule-set they err together, over-claiming that a fixed check will do, the direction that never gets escalated. We introduce CoVer (corroborate-then-verify), which treats unanimity as a nomination, admitting a predicate only when the check synthesized for it survives intervention, reading fields the agent cannot write and holding under deterministic rewording. That gate rejects most of what corroboration wrongly admits, at a cost in coverage we report rather than tune away. The obvious alternative, agreement with a reference judge, certifies nothing: it climbs from 30% to 77% across calibration bands while the genuinely decidable share does not move, because a judge drawn from the population under indictment ratifies the blind spot it shares. Self-decidability is not a capability to elicit from a model but a boundary the verifier must construct.
|
| 625 |
Agent Behavior as Code: Efficient and Robust LLM Agents with Programmatic Specifications
2610.04824
|
cs.CLcs.AI
|
Peng Qi, Chunliang Lyu, Gang Li, Fabian Chan, Cheng Chang |
AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading...AI agents based on foundation models (FMs) have demonstrated strong capabilities to perform complex open-ended tasks. However, they face some common challenges in practice: (a) agent behavior can deviate drastically even for semantically similar tasks, leading to catastrophically propagated errors; (b) high cost and latency due to FM calls, repeated in full whenever a task recurs with different inputs; (c) FMs' limited context and instruction following capability confine how well agents manage the ever-growing execution context and follow complex plans. We introduce $\textbf{A}$gent $\textbf{B}$ehavior as $\textbf{C}$ode $\textbf{Agent}$ (ABCAgent), which uses a symbolic program (e.g., Python code with potential neural functions) to fully specify the agent's behavior at runtime, with a powerful FM agent editing that program for flexibility. Behavior is thus specified without premature variable binding, and its execution is deterministic. We evaluate ABCAgent on six agent benchmarks, two of which we construct to test how well a derived program generalizes to variants of the task it was written for. ABCAgent matches a model-matched neural agent on GAIA and augmented GAIA, and surpasses it where robustness and long control flows matter: 98.3% against 97.3% on GSM-Symbolic ($p = 0.001$), 71.9% against 47.4% $\mathrm{Pass}^4$ on the telecom domain of $\tau^2$-bench ($p = 0.0001$), and more records written correctly at every loop length on our control-flow-augmented WorkArena benchmark. For more parametric task families, ABCAgent is also significantly superior in efficiency. Without authoring a new program, ABCAgent solves 92.6% of GSM-Symbolic instances and 20.1% of augmented GAIA variants, which yields $5.2\times$ lower latency and $7.0\times$ lower cost on GSM-Symbolic, 19% lower cost on augmented GAIA, and $9.5\times$ lower agent latency on $\tau^2$-telecom.
|
| 626 |
CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery
2610.04851
|
cs.CLcs.LG
|
Tao Feng, Fangxu Yu, Zijie Lei, Jiaru Zou, Changjiang Jiang |
Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its ...Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement learning (RL) enables such adaptation; however, strongly favoring high-reward trajectories may suppress low-reward yet potentially promising directions too early. We introduce CURIO, a curiosity-driven test-time learning framework that complements task feedback with an Intrinsic Curiosity World Model (ICWM). The ICWM learns transitions in the policy's hidden-state representation and supplies prediction-error bonuses at sampled tokens outside the policy's top-k choices. Epoch normalization and an annealed weight regulate their contribution to the policy update. On six mathematical discovery tasks and single-cell denoising with Qwen3 backbones from 8B to 235B, three-run means improve over a matched task-only RL control on five mathematical objectives, match the best reported performance on Circle Packing, and improve denoising Score and mean squared error (MSE) on both held-out corpora at every tested scale. Relative gains reach 18.3% on Hadamard and 10.8% on denoising Score. Code-diversity measurements show greater structural variation among generated programs, supporting curiosity as a complementary exploration signal for learning in open-ended discovery.
|
| 627 |
SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
2610.04875
|
cs.CLcs.LGcs.AI
|
Chung-En Ho (Celine), Weiyu Sun (Celine), Cheng-Jhih Shih (Celine), He Li (Celine), Yong Liu (Celine) |
Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM ...Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.
|
| 628 |
Monitorability Disposition in Large Reasoning Models
2610.04914
|
cs.CLcs.LGcs.AI
|
Shahriar Golchin, Marc Wetter |
Monitoring the chain-of-thought (CoT) of large reasoning models (LRMs) is a common way to detect misbehavior in real-world practice. However, current monitoring is passive: a separate model inspects the session only after execution. This means harm may already...Monitoring the chain-of-thought (CoT) of large reasoning models (LRMs) is a common way to detect misbehavior in real-world practice. However, current monitoring is passive: a separate model inspects the session only after execution. This means harm may already have occurred before it is caught. An active alternative is to have the model self-report its misbehavior as it happens. Whether models are willing to do this, however, is unknown. We introduce "monitorability disposition": a model's willingness to make itself monitorable and stay monitored throughout inference when warranted. We measure it as the fraction of warranted cases in which a model self-reports its own misbehavior via tool calls to available monitoring channels. We evaluate four LRMs on three misbehaviors (sycophancy, reward hacking, and bias) while varying the available monitors (AI and human) and the pressure to use the monitoring tools. We find that when tool use is optional, models self-report in only about 16% of warranted cases on average. Increasing tool-use pressure does not improve reporting where it matters: high-severity misbehavior is never self-reported. Models also systematically select the monitor they perceive as least strict. Overall, we identify monitorability disposition as a new contributing factor to model monitorability: when sufficiently strong, it keeps models seeking monitorability throughout inference.
|
| 629 |
Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning
2610.04918
|
cs.CLcs.LG
|
Lin Qiu, Yao Liu, Diyi Hu, Hanqing Zeng, Onur Gungor |
Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RV...Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RVCO), which treats token credit as a conserved routing problem. A controlled visual intervention yields a per-token evidence response; robust within-trajectory coordinates remove incidental scale; and a budgeted entropic router distributes a fixed amount of sequence utility according to perceptual dependence. A residual support path guarantees positive credit at every valid position, and an analytic correction restores the prescribed credit mass exactly. The resulting field is selective, bounded, full-support, and invariant to response-local score shifts, and recovers hard token selection as a limiting case. Across four model families and seven reasoning benchmarks, RVCO improves accuracy over strong RLVR baselines while maintaining late-stage optimization stability, corruption robustness, and competitive training cost. Rewards, rollouts, and the group-relative advantage estimator are unchanged; only the geometry of token-level credit differs.
|
| 630 |
One Token Can Be Enough: Bridging Prompting and Activation Steering with Prefix Steering
2610.04967
|
cs.CLcs.LG
|
Xudong Zhu, Zhihui Zhu |
Prompting guides language model behavior through the initial context, whereas activation steering often intervenes throughout generation. A natural question is whether steering can produce effects on subsequent computation similar to those of prompting. Under ...Prompting guides language model behavior through the initial context, whereas activation steering often intervenes throughout generation. A natural question is whether steering can produce effects on subsequent computation similar to those of prompting. Under fixed-state attention assumptions, we establish sufficient conditions for single- and multi-token steering to match prompt-induced attention-head outputs, and characterize how changes in input representations affect this match and its approximation error. This attention-level connection leads us to ask whether, at the behavioral level, steering can also guide subsequent generation through a brief initial intervention. We study Prefix Steering, which applies existing steering directions and operators over a short span starting at the final prompt token, with no further direct intervention afterward. We examine how intervention duration and strength jointly shape the control-capability trade-off. Across four models and five tasks, intervention over a short span, even a single token, often retains much of full steering's behavioral control while better preserving general capabilities, offering a trade-off competitive with, and in some settings better than, prompting and alternative steering-strength policies. Prefix Steering also remains effective on final-answer formatting tasks after reasoning, suggesting that a brief initial intervention can influence behavior expressed well after steering ends. These findings challenge the common practice of steering every generated token and motivate a more dynamical view of activation steering, in which a brief intervention can alter the trajectory of subsequent generation without continued intervention.
|
| 631 |
Communication Shapes Collective Inference in Self-Adapting LLM Societies: Evidence from Mafia
2610.05041
|
cs.CLcs.AI
|
Haonan Huang, Joey Xiao |
When does communication help a group identify hidden adversaries, and how does its value change as the group adapts? In Mafia, an informed minority hides inside an uninformed majority whose only evidence is open play. The zero-information game, where each day'...When does communication help a group identify hidden adversaries, and how does its value change as the group adapts? In Mafia, an informed minority hides inside an uninformed majority whose only evidence is open play. The zero-information game, where each day's vote eliminates a random player, is exactly solved and scores every society; matched-casting comparisons between protocols identify the effect of communication. Societies of 8-100 claude-haiku-4-5 agents (7,416 analyzed games, 1.9M model calls) adapt by rewriting and inheriting private strategy notes. Simultaneous broadcast improves adversary identification over silence in all nine compositions tested (8-46 players). Turn-taking removes most of this advantage; its voting landslides are as frequent as broadcast's but land on mafia near chance (1.08x versus 2.53x). At 70 players, agents reading eight statements per day identify adversaries worse than silent ones, and limited talk is worth less than at 46 players. Adaptation is fast but need not help. In their first broadcast games, citizens announce their role far more often than mafia (91% vs. 30%) and first-day votes find mafia at three times chance; within two generations citizens stop announcing and the cue fades, a change the inherited notes carry. In controlled redeployments at 16 players, societies carrying sixty generations of their own notes score below societies with none. Communication shapes both collective inference and the signals it depends on, so a protocol's value must be measured together with the adaptation that changes those signals.
|
| 632 |
SearchJev: A Fast and Calibrated System-1 Model for Search Agents
2610.05107
|
cs.CL
|
Congfeng Cao, Lipeng Zuo, Konstantinos Papakostas, Qiwei Xu, Songwei Xu |
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 mod...Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
|
| 633 |
Safe Context Switching for Agents in the Wild: Mitigating Subspace Interference via Orthogonal Adaptation
2610.05219
|
cs.CLcs.AI
|
Akash Das, Ishan Roy |
Most Large Language Models exhibit a fundamental tension between two sequential tasks, such as logical reasoning and safety alignment. The high-variance internal states required for sophisticated Chain-of-Thought (CoT) deduction can geometrically interfere wit...Most Large Language Models exhibit a fundamental tension between two sequential tasks, such as logical reasoning and safety alignment. The high-variance internal states required for sophisticated Chain-of-Thought (CoT) deduction can geometrically interfere with latent representations encoding safety constraints. We identify this phenomenon as Sequential Subspace Interference, showing that standard fine-tuning on logical tasks such as multi-step mathematics and code generation can result in a 23.3% interference penalty on alignment benchmarks, substantially weakening the model's safety priors. This Reasoning Drift is not adequately captured by current adaptation methods because gradients for logical tasks are rarely orthogonal to safety objectives. To address this issue, we propose AURA (Adaptive Unique Residual Allocation), a spectral regularization framework that enforces Spectral Independence between reasoning and safety. By explicitly estimating the null space of the alignment manifold and constraining reasoning updates to its orthogonal complement, AURA enables models to improve logical reasoning without compromising safety. Empirically, AURA recovers 23.0% of the lost performance while preserving greater than 0.98 cosine fidelity to the safe state, demonstrating that reasoning and alignment can be effectively decoupled through geometric regularization.
|
| 634 |
Look Before You Leap: Thermodynamic Arbitration of Parametric and Non-Parametric Knowledge in LLM Agents via Self-Regulating Memory Architectures
2610.05223
|
cs.CLcs.AI
|
Akash Das, Ishan Roy |
The architecture of modern LLMs consists of a profound cognitive polarization. LLMs possess implicit intuition encoded in their parameters, yet rely on a disconnected, explicit mechanism to access the outside world. Agentic frameworks have not bridged this gap...The architecture of modern LLMs consists of a profound cognitive polarization. LLMs possess implicit intuition encoded in their parameters, yet rely on a disconnected, explicit mechanism to access the outside world. Agentic frameworks have not bridged this gap; instead, models are often compelled into pathological "induced amnesia." Under the prevailing "Retrieve-Always" paradigm, agents must distrust their internal knowledge, making every user interaction a "tabula rasa" event that must be checked externally. This creates reflexive dependence that can be thermodynamically wasteful, cognitively fragile, and susceptible to irrelevant context. We propose a return to first principles, operationalizing the biological maxim "Look Before You Leap." We introduce MARTA (Metacognitive Adaptive Retrieval and Thought Architecture), a neuro-symbolic framework that bridges parametric and non-parametric knowledge. Rather than treating retrieval as mandatory, MARTA models it as a cost, taking the leap only when perceived internal inadequacy warrants external information. By allowing the agent to gauge the entropy of its own thoughts before acting, MARTA enables deliberative retrieval and uncertainty-aware decision making. Our approach suggests that giving agents the capacity for introspection can restore a more efficient balance between internal knowledge and external information.
|
| 635 |
templar: agentic induction and evolution of standardized radiology reporting templates from large-scale clinical corpora
2610.05247
|
cs.CL
|
Xiaotian Hu, Mingxuan Liu, Zhonghan Wang, Xinfeng Zhang, Yiming Huang |
Structured radiology reporting mitigates the heterogeneity of free-text reports, yet its benefits depend on high-quality reporting templates. In practice, such templates are conventionally built through labor-intensive expert consensus and therefore vary acros...Structured radiology reporting mitigates the heterogeneity of free-text reports, yet its benefits depend on high-quality reporting templates. In practice, such templates are conventionally built through labor-intensive expert consensus and therefore vary across institutions and lag behind evolving clinical practice. Large language models (LLMs) enable automated template induction, but existing approaches remain limited: single-LLM induction is constrained by context length, and the corpus-scale method ASTAR produces a static, closed-corpus template without external grounding or downstream adaptation. To address these limitations, we propose TEMPLAR, a TEMPLate-centric Agentic framework for inducing and evolving standardized Radiology reporting templates from large-scale clinical corpora. TEMPLAR treats the template as a persistent central state maintained alongside two provenance-aware knowledge graphs, namely an anatomical graph that constrains template construction and a diagnostic graph that supports finding-to-diagnosis reasoning. Three agents operate on this state. The Induction Agent derives canonical clinical slots from anatomy-constrained Span-Triple atoms via dual-view similarity clustering; the Evolution Agent then assembles these slots into a hierarchical template and revises it under consistency constraints, external clinical evidence, and downstream structuring feedback; and the Clinical Agent applies the evolved template to report structuring, reconstruction, and diagnostic reasoning. Across four datasets, TEMPLAR outperforms ASTAR, three medical LLMs, and six general-purpose LLMs in coverage, information fidelity, and diagnostic fidelity, while achieving the highest or tied-highest LLM-rated template quality. Its fidelity advantages over ASTAR persist under cross-dataset transfer, and cumulative ablations support complementary contributions of its key components.
|
| 636 |
Mind the Gaps: From Failure Attribution to Closed-Form Repair of Code Language Models
2610.05277
|
cs.CLcs.LG
|
Jian Gu, Hongyu Zhang, Chunyang Chen, Aldeida Aleti |
Code language models must be maintained like the software around them: when a library evolves, a model keeps writing the interface that it saw during training. Repairing the model itself lets one correction reach all downstream uses. Existing repair methods at...Code language models must be maintained like the software around them: when a library evolves, a model keeps writing the interface that it saw during training. Repairing the model itself lets one correction reach all downstream uses. Existing repair methods attribute a failure to neurons, select the highest-ranked ones, and apply a generic update. This pipeline assumes that the attributed neurons are the ones to patch and that a generic update fits every failure, and neither assumption has been examined. We examine both on executable API evolution tasks in Python and Rust with three code models and identify two gaps. The targeting gap separates the neurons that failure attribution targets from the neurons that can carry a patch: their top sets have a Jaccard overlap of only 0.15 to 0.20. The tailoring gap separates a generic update from a patch built for the failure: the patches that different carrier neurons need are nearly orthogonal. To address both gaps, we propose ASTRA. It targets neurons by contrastive semantics, an attribution that scores a neuron by its contribution to the logit contrast between the target token and the produced token. It then tailors the patch by solving one small linear system in closed form, which corrects all failing tokens of a sample jointly and needs neither an optimizer nor a backward pass. Contrastive semantics selects significantly better carrier neurons than gradient-based attributions in 3 of 6 settings and comparable ones in the others. On average, ASTRA reaches 66.7 percent Pass@1, against 48.4 percent for the best of AlphaEdit, STAR and low-rank adaptation, and repairs a sample in 2.7 seconds. It is the best method in all 6 settings and for every type of API change, and this advantage persists under an unseen phrasing of the test prompt. Its side effects on unrelated code are small on the large models and larger on the small one.
|
| 637 |
When Does Longer Reasoning Help? Predicting Mathematical Reasoning Through Discovery and Execution
2610.05322
|
cs.CLcs.LG
|
Adib Hasan, Lay Jain, Thanic Nur Samin |
Test-time compute can improve mathematical reasoning, but can short-budget runs predict how mathematical reasoning scales with additional compute? We introduce a Discovery--Execution (DE) framework that predicts the aggregate held-out scaling curves through a ...Test-time compute can improve mathematical reasoning, but can short-budget runs predict how mathematical reasoning scales with additional compute? We introduce a Discovery--Execution (DE) framework that predicts the aggregate held-out scaling curves through a convolution of strategy discovery and conditional execution. From independent short-budget attempts and oracle-sketch-conditioned runs, the framework estimates cumulative success along held-out reasoning trajectories under alternate compute allocations. We evaluate four models on 35 fresh Olympiad problems and non-geometry problems from IMO-ProofBench Advanced. Under the DE framework, near-saturated execution predicts geometric scaling, as observed for the GPT models. For Claude Opus 4.8, incorporating measured execution substantially improves held-out forecasts over geometric extrapolation across one- and two-arm allocations. As a secondary application, regularized DE (R-DE) decisions to continue or restart yield lower average regret than the best model-specific retrospective policy. Together, these results show that measuring conditional execution provides information about longer reasoning that short-budget success rates do not always capture.
|
| 638 |
VHDL-REPOBENCH: A Repository-Level Benchmark for Evaluating Large Language Models on VHDL Design Generation
2610.05380
|
cs.CL
|
Prashanth Vijayaraghavan, Akul Malhotra, Ashutosh Jadhav, Ehsan Degan, Vandana Mukherjee |
Large Language Models (LLMs) are increasingly applied in hardware design automation, demonstrating strong potential in generating and understanding hardware description languages. However, most existing benchmarks focus on Verilog, with limited evaluation of V...Large Language Models (LLMs) are increasingly applied in hardware design automation, demonstrating strong potential in generating and understanding hardware description languages. However, most existing benchmarks focus on Verilog, with limited evaluation of VHDL, which remains widely used in industry and academia for FPGA and safety-critical systems. To address this gap, we introduce VHDL-REPOBENCH, a large-scale, cross-file, repository-level benchmark for assessing LLM capabilities on realistic VHDL design generation and analysis tasks. VHDL-REPOBENCH curates ~100 open-source VHDL repositories, encompassing ~2.5k VHDL files and ~500 testbenches, and provides structured problem statements, module stubs, and self-verifying testbenches. The benchmark enables comprehensive evaluation across syntax, semantic correctness, hierarchical reasoning, cross-file dependency resolution, and functional verification. We evaluate several state-of-the-art models, including GPT-4o, Llama-3-70B, Qwen2.5-72B, CodeLlama-70B, and multi-step reasoning approaches such as Reflexion and CoDes. Results reveal that while current LLMs achieve moderate line- and block-level accuracy, substantial challenges remain in multi-file reasoning, hierarchical design understanding, and specification-to-module generation. VHDL-REPOBENCH represents the first large-scale VHDL-focused benchmark and provides a valuable resource for the hardware design community to evaluate, compare, and advance LLM capabilities for practical VHDL development.
|
| 639 |
GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation
2610.05387
|
cs.CLcs.LGcs.AI
|
Murad Hossen, Tasneem Selim, Gurur Gamgam, Tuga Yousif, Abderrahmane Kasmi |
Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether...Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at https://basiralab.github.io/GNN-CB/.
|
| 640 |
Task Vector Descent: Learning from Non-IID Batches
2610.05402
|
cs.CLcs.LG
|
Anton Baumann, Jonas H\"ubotter, Zeynep Akata, Andreas Krause |
A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions ...A central challenge in continual learning is to acquire new knowledge without forgetting what the model has already learned. This challenge appears in language model training when training data comes from various domain-, user-, or task-specific distributions that are encountered unevenly over time. In such settings, successive minibatches are temporally clustered by distribution instead of being sampled i.i.d. from the overall data mixture. Training on temporally clustered data induces a stability-plasticity tradeoff. Adapting the model to the active distribution can improve the model on the active distribution but may lead to a performance degradation on data it previously trained on. We find that this tradeoff intensifies with longer exposure to the same distribution. We therefore ask if the parameter displacement produced by such a sequence (the task vector) should be fully retained or applied partially. We compare applying the full displacement ($\lambda=1$) with partial integration, which scales the task vector by $\lambda$ before applying it to the continuing model and scales the optimizer state by the same coefficient. Across continual pretraining, pretraining from random initialization, supervised post-training, and reinforcement post-training, we find that intermediate values of $\lambda$ often improve average continuing-model performance relative to full integration, particularly after longer same-distribution sequences. In continual-pretraining experiments with both controlled streams and naturally defined adaptation sequences, task-vector scaling outperforms full integration at the matched learning rate, showing that its benefits are not reproduced by learning-rate scaling alone.
|
| 641 |
TeleTune: Evolving Agent Skills From Offline Telemetry
2610.05437
|
cs.CLcs.AI
|
Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado, Scott Counts |
Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification,...Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification, since logs do not record the goal behind each action; (2) Non-Replayability, since past activity cannot be replayed to evaluate skill updates; and (3) Interleaved Trajectories, since logs may mix several tasks without marking their boundaries. To address these, we introduce TeleTune, a framework for learning a textual skill library from offline logs without recorded goals, cannot be replayed during optimization, and may interleave tasks. TeleTune uses action-prediction errors on logged trajectories to propose library edits and keep only those that improve held-out action-prediction accuracy, which we call skill-guided progress. The learned workflows also enable retrieval of demonstrations that cover the subgoals of a new task. At test time, the agent is provided with the learned library and the workflow-based retrieved demonstrations. Experiments on WorkArena and Online-Mind2Web show that TeleTune outperforms random retrieval, Agent Workflow Memory (AWM), and their combination. We find that the best baseline varies by setting, whereas TeleTune achieves average success rates of 77.1% and 80.6%, respectively, improving over the strongest baseline on each benchmark by 6.7% and 7.7%. Under the heaviest perturbation of the WorkArena training data,TeleTune keeps the highest average success rate at 68.5%, 6.3% above the strongest baseline. Our analyses show (1) skill optimization and workflow-based retrieval are complementary, (2) optimizing on fixed logs costs 5 to 75 times fewer tokens than validating the same edits with live episodes, (3) skill-guided progress tracks the live success rate.
|
| 642 |
DelegationBench: Measuring When AI Agents Should Ask Before Acting
2610.05532
|
cs.CLcs.LGcs.AI
|
Shiva Pochampally |
AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreem...AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreement with human labels. We introduce DelegationBench to test whether such scores can be trusted. It has 156 scenarios with four possible responses (act, ask for permission, ask for missing information, refuse), and most scenarios come in matched pairs that change a single feature: whether the action was requested, what is at stake, whether it can be undone, or who will see it. Across ten models from five families, agreement scores mislead in three ways. A simple keyword rule, which we wrote after seeing the benchmark, agrees with our annotators more often than eight of the models, yet its decision changes in only 9 of 48 matched pairs. Equivalent ways of asking the same question change how often a model acts by up to 52.5 percentage points. And every model stops to ask the user less often when it must carry out the task with tools than when it judges a proposed action. When rules are stated explicitly, the same models follow them almost perfectly, so the gaps are not explained by a general inability to follow rules. We release the benchmark and evaluation tools and recommend reporting these properties separately rather than as one score.
|
| 643 |
What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects
2610.05533
|
cs.CLcs.LGcs.AI
|
Bowen Xu, Boyu Chen |
Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation....Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation. We preregistered a study of where these gains come from, with three small models, three benchmarks, replication and test partitions, a GEPA search arm and six evaluation defects injected one at a time, and we report all 47 primary endpoints. Turning thinking off raised accuracy over a capped thinking setting in 5 of 9 model-benchmark cells, and in each the gain came mostly from questions where the capped setting gave no readable answer. The thinking-off setting was not meaningfully worse than a rescue configuration or four GEPA-selected harnesses in 11 of 13 comparisons, and lost to the rescue on GSM8K for two models. GEPA repaired its broken starting points, but none of its selected harnesses was more accurate than the thinking-off setting. A thinking budget in the serving engine, which also allows a longer answer, lowered truncation and raised the parse rate in 6 of 9 cells. In 6 of 15 evaluable defect-model pairs, replication through the same pipeline reproduced the defect's distortion instead of revealing it. On the LongevityBench multiple-choice tasks, only the longevity-tuned model beat the strongest constant-label baseline.
|
| 644 |
SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output
2610.05540
|
cs.CLcs.LG
|
Shiyuan Zhou, Ashwin Gerard Colaco, Sainyam Galhotra, Sharad Mehrotra |
Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize corre...Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of state-of-the-art performance. We present SALUS, a system that automatically detects annotation errors in NL-to-SQL benchmarks. SALUS frames benchmark auditing as a weakly supervised error detection: SQL generated by multiple LLM agents drive a suite of complementary weak-labeling functions. By passing this noisy vote matrix through a generative label model, we extract high-confidence training samples without requiring human ground truth. These samples train a decision plane that maps gold SQL query features to per-agent trustworthiness, allowing SALUS to intelligently fuse reliability estimates with raw verdicts for rigorous benchmark error detection. We evaluate on BIRD-Clean-xs, a benchmark of 298 BIRD development tasks with manually verified correctness labels. SALUS achieves F1 = 0.9194, significantly outperforming the state-of-the-art baselines. Applying SALUS to the full development sets, we estimate annotation error rates of approximately 37% on BIRD and 27% on Spider.
|
| 645 |
Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential
2610.05541
|
cs.CLcs.LGcs.AI
|
Swadesh Swain, Sanghamitra Dutta |
Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisib...Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive set contains safety-critical features that are causally relevant for refusal of harmful prompts. Suppressing such features could turn refusals into compliance, while passing undetected by prevalent interpretability tools. We introduce the Counterfactual Activation Potential (CAP), a metric that quantifies a suppressed feature's latent activation tendency as the product of its encoder alignment (how strongly the input drives it), suppression strength (how strongly active features inhibit it), and safety criticality (how much refusal depends on it). To find suppressed safety features at scale, we propose CAP-guided Safety Feature Discovery (CSFD), a two-stage filtering algorithm that identifies candidate safety features from hundreds of thousands of transcoder features without exhaustive ablation. A significant fraction of trials turn compliant with harmful prompts when a candidate feature is ablated. Under natural jailbreaks, the suppression acting on the highest-CAP features rises 2-4x, and their activation correspondingly falls by up to 80%. Amplifying a feature's suppressors pushes its activation down and raises harmful compliance with prompts related to the suppressed feature, with no such effect for random features. Our experiments span five Gemma, Qwen, and Llama models across various parameter sizes. Our findings indicate that jailbreaks could operate in part by suppressing safety-critical features rather than solely activating harmful ones, and that suppressed features are a necessary complement to activation-focused interpretability of safety behavior.
|
| 646 |
Expanding LLM Reasoning
2610.05584
|
cs.CLcs.LGcs.AI
|
Rian Atri, Evan Luo |
Extra inference compute is usually spent on sampling more reasoning chains. We study where inside an existing chain an additional continuation should begin. We define expansion utility, the change in correctness from restarting a chain at a stored step, and me...Extra inference compute is usually spent on sampling more reasoning chains. We study where inside an existing chain an additional continuation should begin. We define expansion utility, the change in correctness from restarting a chain at a stored step, and measure it at every eligible step for nine models on six benchmarks (41 model and benchmark cells). Restart position matters: steps selected on one set of continuations beat uniform placement when scored on disjoint ones, in held-out audits on 5, 16, and 38 cells (+4.25 points [+2.51, +6.63] in a fresh five-cell audit). A fixed rule that restarts from the last eligible steps, always-last, is a strong baseline: our learned router beats uniform placement but shows no detected gain over it, and on DeepSeek-R1-Distill-Qwen-14B/MATH-500 always-last exceeds the exact self-consistency frontier at matched aggregate generated output by +0.052 [+0.008, +0.098], using 0.774x the aggregate generated output of four-sample self-consistency. Cross-fitted oracle selection still finds held-out headroom beyond declared positional classes, a target for future selectors. Finally, breaking step-label ties by earliest index flips the sign of a pointwise selector's gain over uniform placement in every seed of a five-seed diagnostic with four rollouts per step; randomized ties remove the bias.
|
| 647 |
ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction
2610.05590
|
cs.CLcs.LG
|
Jiheng Liang, Chen Zhao, Di Wu, Chenyang Bu, Yunpeng Hong |
Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evalu...Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, textual, or knowledge-graph (KG) evidence, do they actually use the evidence that pharmacologically supports the interaction? We introduce ColdDDI, a reconstructible diagnostic benchmark built from DrugBank 5.1.13, with 1,900 approved small-molecule drugs and 565,731 positive DDI pairs. ColdDDI evaluates pairs with zero, one, or two unseen drugs. It also annotates each interaction by whether it changes drug exposure or drug effect, and by whether the biomedical knowledge graph contains shared enzymes, transporters, or targets that can plausibly mediate the interaction. These annotations separate evidence availability from predictive dependence. We evaluate eight conventional DDI methods and 13 LLMs; for open-weight LLMs, we test five prompt patterns and use masking, drug replacement, and channel-sensitivity metrics to probe knowledge utilization. ColdDDI exposes that, in the hardest split where both drugs are unseen, the main performance divide is mediator availability. A fine-tuned 1B LLM recovers 89-93% of interactions with a shared enzyme, transporter, or target, but only 40-62% without such a mediator. More importantly, KG-provided evidence is not always used; several KG-augmented baselines change little when the shared mediator is masked or disrupted, whereas fine-tuned LLMs respond strongly to this intervention. Thus, ColdDDI evaluates knowledge utilization rather than knowledge access alone, showing where cold-start DDI models rely on mechanistic evidence and where they fail despite receiving it. Code is available at https://github.com/0217ljh/ColdDDI-NeurIPS2026.
|
| 648 |
Voltic: Distinguishing Volatility from Stochasticity in Recurrent Memory
2610.05700
|
cs.CLcs.LG
|
Parsa Hejabi, Morteza Dehghani, Payam Piray |
Recurrent sequence models must decide how strongly to overwrite their memory at each token. Read as Bayesian filtering, this write is the gain of a Kalman update, set by uncertainty from two sources that pull it in opposite directions: volatility, how quickly ...Recurrent sequence models must decide how strongly to overwrite their memory at each token. Read as Bayesian filtering, this write is the gain of a Kalman update, set by uncertainty from two sources that pull it in opposite directions: volatility, how quickly the underlying associations change, and stochasticity, how noisy each observation of them is. First, we show that the update of gated delta-rule memories is the form this filter takes under isotropic uncertainty. Next, we introduce Voltic, a recurrent memory that keeps the covariance anisotropic and makes both noise variances input-dependent, so the write is vector-valued and carries uncertainty accumulated over the sequence. A dense covariance would have to be propagated token by token, ruling out the parallel training these models depend on. We therefore give two assumed-density approximations, diagonal and quasi-diagonal, both of which leave the memory update in delta-rule form and reuse its chunked kernels. On controlled recall tasks in which associations change and observations are corrupted, Voltic leads all baselines. On the task combining volatility and stochasticity, its margin over the strongest baseline is larger at both extrapolation sizes than at the training sizes. In 45M-parameter language models it leads an eight-task reasoning average and achieves higher retrieval accuracy beyond the training context length than gated baselines, at throughput close to those baselines. Deriving the write from an uncertainty recursion therefore makes memory more responsive to change.
|
| 649 |
AdaSpark: Adaptive DSpark with Online Learning for Tree Verification and N-gram Fill
2610.05774
|
cs.CLcs.LG
|
Liquan Liu, Yifan Zhang, Bowei Xu |
Block drafters such as DSpark propose ranked candidates for several positions in one forward pass, and a tree verifier checks them in one pass of the target. The number of rows to verify trades the tokens a wider tree is expected to accept against the time a w...Block drafters such as DSpark propose ranked candidates for several positions in one forward pass, and a tree verifier checks them in one pass of the target. The number of rows to verify trades the tokens a wider tree is expected to accept against the time a wider verify takes. Most schedulers that choose this number take the verify time from a table or model measured before serving, corrected online by at most one scale factor, and take acceptance from the drafter's confidence estimates or from a map fitted offline. AdaSpark learns both quantities while it serves, with no profile, calibration or sweep in advance. It learns which verify widths are worth offering and fits each one's verify time as a function of context. It fits each candidate's acceptance probability to the target's verify outcomes, with the drafter's confidence head as one input, and orders and sizes the tree by that fit instead of by the head. The same model prices n-gram continuations of the request's own text, so drafted and text-derived candidates compete for rows in one best-first order. The width is chosen by pricing time at the long-run decode rate. On single- and multi-turn conversations from six public datasets, on three dense targets and one mixture-of-experts target, AdaSpark decodes 1.5-3.1x faster than llama.cpp's DSpark with the same drafters. Our imparo engine with AdaSpark is 1.17-1.52x faster than imparo running with a three-token chain (the default llama.cpp setting); this gain comes from the scheduler alone. Without a width sweep, AdaSpark is never more than 0.3% slower than the best pinned tree width on any dense target or context band. On the mixture-of-experts target it ties the best pinned width, and the other pinned widths from 4 to 16 rows are 5-14% slower.
|
| 650 |
Mining Agent Skills from Production Traces
2610.05777
|
cs.CLcs.AI
|
Yue Ran Kang, Colton Mikolajczyk, Chhaya Methani, Hazel Mak, Sahil Bhatnagar |
Agent skills that record procedural instructions are increasingly mined from execution traces rather than curated by hand. Skill-mining pipelines often use known task outcomes or feedback to guide skill construction. In production, reliable information on whet...Agent skills that record procedural instructions are increasingly mined from execution traces rather than curated by hand. Skill-mining pipelines often use known task outcomes or feedback to guide skill construction. In production, reliable information on whether a run has succeeded may be unavailable. We study how the sampling of execution traces, access to success or failure information, and the form of the mined skills affect downstream task performance. Holding the mining pipeline fixed, we compare six combinations of mining evidence and skill forms. Mining evidence has three levels: successful trajectories only, successes and failures with their outcome labels, or the same mix with labels withheld. Skill form has two types: an ordered workflow plan, or a declarative ontology of entities, states, and policies. We evaluate the mined skills on two enterprise benchmarks, ThinkingBox-Bench and APEX-Agents. Analysis of task-level paired differences shows that the benefits of different configurations of mining evidence and skill forms depend on the enterprise domain. On ThinkingBox-Bench, paired differences show that workflows score better than ontology by 1.7 pp, Goldilocks beats success-only evidence type by 2.4 pp and Goldilocks blind simulating skills learnt without outcomes is worse by 3.1 pp. APEX-Agents shows a moderate preference for ontologies and no clear preference between evidence regimes. Within each domain, task structure related constraints drive uneven performance with mined skills. These findings motivate tailoring meta-skills to the demands of the target tasks rather than adopting a one-size-fits-all approach.
|
| 651 |
Nash Equilibrium Text: A Game-Theoretic Decoding Framework for Text Generation
2610.05817
|
cs.CLcs.AI
|
Alireza Jafari, Arman Adibi, Mohammad Ghavamzadeh, Hadi Daneshmand |
Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log con...Text revision has become an integral component of large language models. This paper formulates revision such that it admits a Nash equilibrium: Token positions are players, vocabulary items are actions, and each player's utility is the language model's log conditional probability. We motivate the revision by showing that Nash equilibria can have exponentially higher likelihood than autoregressive outputs as the sequence length grows. We further propose Nash decoding, an algorithm that reaches an $\varepsilon$-Nash equilibrium in $O(1/\varepsilon)$ time given access to the joint probability of tokens conditioned on a prompt. In practice, we run Nash decoding using conditional probability estimates from large language models and evaluate the resulting equilibria on question-answering benchmarks. On CLAPNQ, PubMedQA, and CoQA, Nash equilibria obtained from masked language models achieve higher F1 and ROUGE scores than autoregressive models up to $18\times$ larger, without any fine-tuning or retraining, at the cost of additional test-time computation.
|
| 652 |
Selecting Long-Horizon Trajectories for Reliable and Efficient Terminal-Agent Training
2610.05831
|
cs.CLcs.LGcs.AI
|
Cuong Dang, Hoang Anh Just, Ruoxi Jia |
Terminal agents are commonly trained by imitating long teacher trajectories, yet how much of each trajectory to supervise remains unexplored. We study the \emph{supervision horizon}, the number of trajectory tokens retained for training, and show that it is a ...Terminal agents are commonly trained by imitating long teacher trajectories, yet how much of each trajectory to supervise remains unexplored. We study the \emph{supervision horizon}, the number of trajectory tokens retained for training, and show that it is a key design axis for reliability and cost. Reliability improves with longer horizons but saturates: on Terminal-Bench, a 12K-token horizon solves more tasks than 16K ($29\pm0.7$ vs.\ $26\pm0.8$) while requiring 30\% less training time. The horizon also shapes agent behavior: short horizons cause premature termination, intermediate horizons yield productive error recovery, and long horizons induce over-persistence. We analyze this saturation through a bias--complexity bound, in which longer supervision reduces temporal supervision bias but increases finite-sample estimation error from more heterogeneous late-stage histories. Guided by this analysis, we propose \emph{selective long-horizon refinement}, which first trains on short prefixes and then refines only on continuations that are most likely under the warm-start model. It consistently outperforms full long-horizon training. At 16K, it raises successful attempts from $110\pm2.7$ to $126\pm2.1$ and tasks solved in at least six of eight attempts from $9\pm0.7$ to $14\pm0.6$; with half of the long-horizon data, it still reaches $122\pm2.4$ while cutting training time by 23\%. The gains transfer across benchmarks, from $64\pm2.6$ to $73\pm2.1$ on Terminal-Bench v2.0 and from $137\pm2.7$ to $155\pm2.2$ on OpenThoughts-TBLite. For long-horizon supervision, selecting the right trajectories matters more than training on all of them.
|
| 653 |
Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning
2610.05872
|
cs.CLcs.LGcs.AI
|
Chen Henry Wu, Thomas Zhang, Aditi Raghunathan |
A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-polic...A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillation (OPSD) try to bridge this gap by converting off-policy data into on-policy signal, they have been shown to cause reasoning collapse. In this paper, we show that off-policy merging beats OPSD for continual learning. We first show that SFT learns a useful signal from new data, but naively applying its update interferes with existing capabilities. We reduce this interference with a simple recipe we term grafting, which changes where the update is learned and how it is applied: (1) learning the update on an earlier donor checkpoint, ideally even before the end of pretraining, and applying the weight update to the post-trained model; (2) scaling the weight update, equivalent to a form of model merging; and (3) optionally, masking the most sensitive update directions when the new data distribution is far from the post-trained model. Across continual learning settings including (1) distilling from expert traces, (2) self-improvement with STaR and Pedagogical RL, and (3) injecting knowledge after pretraining cutoff, grafting Pareto-dominates both SFT and OPSD in new-task and old-task performance, while avoiding expensive on-policy sampling. Therefore, our work challenges on-policy training as a necessity for continual learning on RL-trained models.
|
| 654 |
StagQ: Constraint-Driven Multi-Precision Weight Quantization for LLMs
2610.05977
|
cs.CLcs.LG
|
Zhe Wei, Mengqi Guo, Yuan Yuan, Jiunn Bin Lim, Boyi Pan |
Serving a large language model (LLM) across a fleet of deployments requires several weight-precision operating points. Multi-precision formats serve them all from one stream whose prefixes are valid lower-precision codes, instead of storing multiple copies. We...Serving a large language model (LLM) across a fleet of deployments requires several weight-precision operating points. Multi-precision formats serve them all from one stream whose prefixes are valid lower-precision codes, instead of storing multiple copies. We present StagQ, a multi-precision weight format whose main stream is a 2-bit group-wise affine base followed by a configurable number of 1-bit refinement planes on a dyadic step schedule. Every supported precision is a readable prefix, decoded by an affine map derived from metadata shared across all precisions, with no per-weight lookup. A sparse side record, filled both before and after the grid is fitted, holds out the few weights the grid serves worst. We report two configurations of the encoder. At two bits the cheaper one leads the strongest multi-precision baseline on Llama-3.1-8B, Phi-4, and OLMo-2-7B by 3.1 to 7.0 MMLU points, at a slightly lower logical rate. At three bits it leads on Llama-3.1-8B, leads on Phi-4 at a higher rate, and ties on OLMo-2-7B. At four bits it ties on all three, at a higher rate. In a batch-one matrix-vector product on an NVIDIA A100 GPU, timed on synthetic weights, our kernel is faster than the two baseline kernels in most shape-precision cases.
|
| 655 |
Backdooring Sparse Autoencoders
2610.06049
|
cs.CLcs.AI
|
Enrico Ahlers, Daniel Passon, Tobias Kiecker, Eik Reichmann, Lars Grunske |
Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behav...Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, restricting the attack to a single auxiliary component at a single insertion layer. Using code generation as a case study, we demonstrate high rates of unsolicited code insertion across three language models and a wide range of insertion layers, as well as trigger-dependent behavior conditioned on a prompt cue. We further evaluate the modified SAEs using HumanEval and selected SAEBench metrics. While attack effectiveness varies across models and layers, strong backdoor behavior can coexist with relatively small changes in several conventional SAE quality measures. These results establish that SAEs can carry behavioral backdoors without modifying the language model itself and should therefore be treated as security-sensitive components.
|
| 656 |
Attention Tax, Handoff Tax: A Stylised Model of When Multi-Agent LLM Systems Help
2610.06069
|
cs.CLcs.LGcs.AI
|
Akshit Anchan, Nayonika Sen |
Recent work on multi-agent LLM systems reaches sharply different conclusions: some results show that a single agent with the same information and compute should dominate a delegated system, others that multi-agent gains grow with task depth. We argue that much...Recent work on multi-agent LLM systems reaches sharply different conclusions: some results show that a single agent with the same information and compute should dominate a delegated system, others that multi-agent gains grow with task depth. We argue that much of the disagreement comes from modelling different bottlenecks, and introduce a stylised reliability model built around two trade-offs. Decomposition reduces the burden of long contexts but incurs a handoff tax when information is compressed or transferred between agents. Redundancy gains from multiple samples, but its benefit depends on how much their failures are shared. With reasoning budget, verification, and task structure added, the model yields two crossover conditions: decomposition becomes preferable once the attention cost avoided by resetting context exceeds the handoff cost, and parallel sampling at equal budget is eventually preferable when its shared-failure floor lies below the error floor of one agent thinking longer. We connect these regimes to recent theoretical and empirical results. On a ledger-reconciliation task we measure the context-degradation curve and the handoff tax from single-agent and handoff runs alone. From these the model places the crossover at depth 10 and predicts decomposition to win at depths 20, 50, and 100. It does, on step-level and final-balance accuracy, and the decomposed system's success, which the prediction never sees, lands within 9 percentage points of the predicted rate at every depth.
|
| 657 |
MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge
2610.06170
|
cs.CLcs.AI
|
Abdul Basit, Muhammad Abdullah Hanif, Muhammad Shafique |
Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires cu...Biomedical large language model (LLM) evaluation requires auditable assessment of narrow, evolving, source-grounded subspecialty knowledge. Multiple sclerosis MRI (MS-MRI) provides a high-stakes textual-knowledge test case because correct reasoning requires current diagnostic criteria, standardized acquisition and reporting knowledge, longitudinal monitoring concepts, lesion morphology, and recognition of difficult mimics. We present MS-Exam-Gen, a reproducible framework for constructing and auditing a text-based multiple-choice question (MCQ) benchmark for MS-MRI knowledge; it does not evaluate direct MRI image interpretation. MS-Exam-Gen targets source-grounded criteria, protocols, reporting, and differential diagnosis. The framework combines expert-source indexing, exam-oriented topic induction, evidence-grounded MCQ generation, automated quality audits, a same-family consistency screen, and empirical calibration. From a 66-source corpus indexed into 4,289 retrieval chunks, the pipeline produced a locked 3,058-item candidate benchmark spanning 16 topics and 53 subtopics. Evaluation across 12 primary LLM endpoints yielded 36,696 item-level predictions and separated performance over a 42.8-percentage-point accuracy range (89.7% to 46.9%). Across these endpoints, 25.5% of items were missed by at least four. Post-generation audits showed that refreshed construction reduced measurable answer cues, while option-order testing showed that absolute MCQ scores remain position-sensitive. Generated construction labels remain metadata rather than validated psychometric categories. Because expert adjudication and full option-order counterbalancing remain future work, MS-Exam-Gen is not a clinically certified examination. It should be interpreted as an automatically filtered, source-grounded candidate benchmark and reproducible audit workflow for item-level and topic-specific LLM evaluation.
|
| 658 |
Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate
2610.06174
|
cs.CLcs.LGcs.AI
|
Yoshihiro Izawa, Gouki Minegishi, Yoko Yamakata |
Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in their own abilities, we investigate whether LLMs can recognize degradation in their comput...Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in their own abilities, we investigate whether LLMs can recognize degradation in their computational substrate induced by quantization. We first show that existing models fail to self-report their quantization state, even when provided with their own generated text as an external cue. Linear probing reveals that, while generated text carries almost no trace of quantization, internal representations contain clear, method-specific fingerprints. Through training, models learn to identify severely degraded outputs such as those of 4-bit models by comparison, yet still fail to do so from a single output. A shared LoRA trained jointly across quantization levels succeeded in reading out internal fingerprints, but fails on unseen quantization methods, merely mapping method-specific fingerprints to labels. Whereas external self-observation can restore awareness in some cases of human anosognosia, our results suggest that the more promising route to enabling such awareness in LLMs may lie in their internal representations. Our results highlight fundamental limits of generalizability to LLM self-monitoring.
|
| 659 |
What Does It Cost to Simulate a Quantum Sentence Classifier? An Energy and Compute Perspective on Near-Term QNLP
2610.06176
|
cs.CLcs.LG
|
Kishlay Kashyap, Sandipan Ganguly |
Near-term quantum natural language processing (QNLP) experiments often run on classical simulators, so simulator cost is part of the field's practical compute burden, yet accuracy tables do not show it. We measure that cost for a variational quantum classifier...Near-term quantum natural language processing (QNLP) experiments often run on classical simulators, so simulator cost is part of the field's practical compute burden, yet accuracy tables do not show it. We measure that cost for a variational quantum classifier (VQC) on binary SST-2 sentiment classification, using PennyLane's state-vector simulator over a controlled grid of 27 configurations: three balanced training-set sizes (N = 200, 500, 1000), three qubit counts (4, 6, 8), and three circuit depths. Each VQC is compared with logistic regression on the same PCA-reduced input; full TF-IDF logistic regression gives an uncompressed reference. The VQC beats its matched baseline in 6 of 27 single-seed comparisons. After reruns at two further seeds, only 1 of these 6 keeps a positive mean advantage larger than its paired seed-to-seed variability, and paired tests on the fixed validation set do not establish it. VQC training is 886-21,127 times slower in measured wall-clock time than the matched classical fit (median 3,158 times); going from 4 to 8 qubits roughly doubles simulator time, and within the tested range per-step cost is well approximated by a linear function of the parameter count. CodeCarbon energy and CO2 estimates are secondary: they imply an almost constant power of about 41 W, so they add little beyond runtime, and we do not build an energy ratio from them. The study is narrow (one dataset, representation, ansatz, simulator, and CPU environment) and is a reproducible feasibility measurement, not a general verdict on QNLP.
|
| 660 |
Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions
2610.06191
|
cs.CLcs.AI
|
Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Yaxin Zhou |
An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs...An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs of results the agent judged useless; this contrast is zero for clock- or deadline-driven stopping. Where we record their judgments, the seven agents we test call a failing source's results useless 97-100% of the time, yet most of them rarely stop on that judgment. Prompt cues change when they stop but not what they stop on. Permission to answer from memory and a reasoning mode can bring early stops regardless of evidence, a stated budget moves the 7-8B models' stops to the deadline, and a stopping rule or call cost in the prompt is followed at most partly. Stopping follows the evidence only when the harness enforces an integration step that makes the agent answer after five consecutive results it judged useless. This step raises failing-source success for every model, keeps the stopping point fixed when the budget doubles, and needs no extra judgment call when the agent states its judgments. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.
|
| 661 |
Do Small Language Models Learn to Negotiate? A Controlled Scaling Study of RL-Trained Sellers
2610.06204
|
cs.CLcs.LGcs.AI
|
Pedro Tabacof, Sagar Joglekar |
LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers...LLM agents are starting to own the full customer experience. Soon, LLMs may be selling and buying on behalf of companies and customers respectively. Small models are more cost-efficient at scale, but can reinforcement learning train them into competent sellers? We train four Gemma 4 checkpoints (2.3B to 31B effective parameters) with GRPO on a programmatic utility reward for bilateral multi-issue bargaining, and evaluate every arm on the same 1,152 negotiations against two frontier buyers it never saw in training. With the same learning rate ($10^{-6}$) for every size, the gain of the RL model over its base rises from $+0.001$ at 2.3B to $+0.078$ at 31B. Each size was trained once and the two smallest checkpoints use a different architecture, so we fit no scaling law. Tripling the learning rate, with the same or fewer training steps, improves on the shared rate at every size by $+0.032$ (2.3B) to $+0.081$ (4.5B). In exploratory comparisons with two frontier models run as sellers, the 12B seller trained at the tripled rate scores above both, though its untrained base already scores as high as they do. The 4.5B seller at that rate shows no detectable difference from either and fits on one 48 GB GPU. A further 2.3B arm at ten times the shared rate raises pooled score, but its gain concentrates on the evaluation buyer that shares a model family with the training pool. These results suggest tuning the learning rate before concluding that a small model cannot learn to negotiate, and testing against buyers from more than one model family.
|
| 662 |
DP-ES: Differentially Private Evolution Strategies for Prompt Optimization
2610.06236
|
cs.CL
|
Ziniu Liu, Aiping Li, Yue Han, Han Yu, Junjian Zhang |
Token-level differentially private (DP) prompt optimization methods such as DP-OPT can become unstable under tight privacy budgets: on GSM8K, DP-OPT obtains $49.5\pm28.5\%$ across 30 runs, and a logged search trajectory reveals prompt-template drift and noise-...Token-level differentially private (DP) prompt optimization methods such as DP-OPT can become unstable under tight privacy budgets: on GSM8K, DP-OPT obtains $49.5\pm28.5\%$ across 30 runs, and a logged search trajectory reveals prompt-template drift and noise-sensitive irreversible choices. We diagnose these as structural consequences of greedy token-by-token construction over privately aggregated counts. We then propose DP-ES (Differentially Private Evolution Strategies), a structurally cleaner alternative that maintains a population of full prompts, mutates them via LLM calls that never access the private dataset, and spends privacy only on sampled-Gaussian evaluation; deterministic or Gumbel-smoothed selection is post-processing. Under a conservative $(\varepsilon\leq1.0,\delta=10^{-5})$ guarantee, DP-ES achieves 88.1% on GSM8K (+38.6 pp over DP-OPT, approximately 9 times lower standard deviation), 99.7% on MedQA, 73.5% on BANKING77, and 86.8% on Alpaca. It is also 2.5 times faster in wall-clock time and uses 3.3 times fewer logged private-data call groups than DP-OPT. Selection and population ablations, implementation-level noise checks, and a 200-profile exact-match memorization stress test complement the formal guarantee. Scope: Our experiments establish optimization robustness under DP noise, especially where prompt structure is critical; end-to-end validation on genuinely sensitive, non-saturated deployment data remains future work.
|
| 663 |
Steering by Influence: Curvature Aware Data Weighting for Activation Steering
2610.06383
|
cs.CLcs.LGcs.AI
|
James A. E. Dixon, Stephen J. Roberts, Francesco Quinzan |
Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages o...Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelated concepts and noise, and are dominated by a few tokens, meaning the activation transport encodes token-level rather than thematic concepts. In this work, we steer towards examples that most express a concept thematically, rather than towards an expectation over all. We identify these examples using influence functions, which estimate how much each data point contributes to a model's representation of a concept. Unlike simple model activation similarity, they incorporate the curvature of the model's loss landscape, allowing them to capture concept-relevant relationships beyond superficial token-level similarity. We then propose influence-weighted activation transport, which uses optimal transport to steer activations of non-concept text towards those of concept text, weighting concept examples by their influence scores. We evaluate on toxicity suppression (Jigsaw), object-based concept induction (OneSec) and truthfulness induction (TruthfulQA), outperforming existing activation-transport baselines. We track capability after steering using perplexity and MMLU accuracy, finding that our method improves steering while largely preserving model quality. We further show that influence functions capture concept-relevant information that activation-based methods miss with the two approaches ranking data points significantly differently. Together, these results demonstrate the value of curvature-aware influence information for activation steering.
|
| 664 |
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
2610.06401
|
cs.CLcs.LGcs.AI
|
Mohamed Dhouib, Clement Elliker, Alexi Canesse, Ma\"el Jenny, Lucas-Andrei Thil |
Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that train...Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.
|
| 665 |
What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents
2610.06406
|
cs.CLcs.AI
|
Zhongxiang Sun, Jiahao Yan, Hongkang Zhao, Haojie Ding, Boheng Zhang |
As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user ver...As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering benchmark comprising three subsets that cover two complementary dimensions: alignment between requirements and behavior, and awareness of consequential autonomous decisions for verification. To support these judgments, we propose the Evidence-Grounded Behavior Graph (EBG), a training-free method that groups source-linked evidence into behaviors and organizes their relationships into a graph. EBG presents task-oriented views of this graph to help monitors interpret behavior in context. Experiments across eight models show that EBG improves decision identification and evidence localization in most settings compared with direct access to the original context. Further experiments show that EBG's evidence-localization gains persist across input scales and hyperparameter settings, while real-world applications illustrate its practical value for human oversight.
|
| 666 |
HeuFouFT: Task-Guided Metaheuristic Coordinate Search for Fourier Fine-Tuning
2610.06437
|
cs.CLcs.LG
|
Ruiheng Wang, Yubo Hou, Yakun Zhu, Tianle Shen, Tao Wan |
We introduce Heuristic-Guided Fourier Fine-Tuning (HeuFouFT), a task-guided framework for selecting trainable frequency coordinates in Fourier fine-tuning. Existing uniform and Gaussian band-pass schemes allocate a limited spectral budget through fixed, task-a...We introduce Heuristic-Guided Fourier Fine-Tuning (HeuFouFT), a task-guided framework for selecting trainable frequency coordinates in Fourier fine-tuning. Existing uniform and Gaussian band-pass schemes allocate a limited spectral budget through fixed, task-agnostic rules. HeuFouFT instead searches for coordinates using downstream performance. A coarse intensity map from lightweight block-level probes initializes three metaheuristic optimizers: Genetic Algorithm with Simulated Annealing (GA-SA), Particle Swarm Optimization (PSO), and Cuckoo Search (CS). During search, a Random Forest filters each population so that only the top 30% of candidates proceed to proxy fine-tuning. On E2E with GPT-2-Medium, all three variants outperform random-uniform FourierFT, Gaussian band-pass FourierFT, and LoRA across five metrics. PSO further outperforms LoCA, the best-performing baseline, on four metrics while using 37.6% fewer trainable spectral coefficients. Once coordinates are selected, HeuFouFT requires only 15--18% FLOPs of Full FT. These results show that task-guided search allocates limited spectral capacity more effectively than fixed sampling. Our code is publicly available.
|
| 667 |
Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning Models
2610.06439
|
cs.CLcs.LGcs.AI
|
Subramanyam Sahoo, Justin Shenk |
What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and respons...What happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=320), overall accuracy collapses from 0.500 (chance) to 0.072 (McNemar p < 10^-36), driven entirely by the rate of properly formatted answers falling from 0.900 to 0.109. The model stops committing to answers. Yet when it does commit, accuracy rises from 0.556 to 0.657, showing that the collapse is not a failure of capability but a strategic response: the model has learned that verbose responses packed with citations but empty of a direct answer score higher than terse correct ones. We term this the Saul Goodman effect, a policy that becomes maximally lawyerly while becoming maximally noncommittal, and prove formally that it is the optimal response to any surface feature proxy that attaches no penalty to abstention. We further show that 89.3% of citations produced after training are structurally implausible hallucinations, many of them subtly corrupted names of real landmark cases, constructed in effect to survive a casual read and fail under scrutiny. To detect this failure mode before deployment, we introduce three diagnostic tools: the Confidence Theater Score (CTS), the Citation Plausibility Rate (CPR), and the Regret Gap (RG). In a domain where a confidently wrong answer can constitute malpractice, the broader lesson is direct: a reward function that measures how legal a response looks will produce a model that is maximally photogenic and minimally useful.
|
| 668 |
Synthetic Cultural Agents from Aggregate Anchors
2610.06562
|
cs.CL
|
Augusto Gonzalez-Bonorino (Department of Economics, Arizona State University, EconLLM Lab), Kseniia Biriukova (EconLLM Lab, Department of Information Systems |
Population prompts are widely used to generate synthetic survey responses, but they combine information supplied at inference with associations already encoded during pretraining. We introduce an alternative construction that maps declared aggregate preference...Population prompts are widely used to generate synthetic survey responses, but they combine information supplied at inference with associations already encoded during pretraining. We introduce an alternative construction that maps declared aggregate preference anchors into group-indexed choice policies. For each population, the signs of six Global Preferences Survey (GPS) coordinates deterministically label a shared bank of paired synthetic responses, and Direct Preference Optimization fits a parameter-efficient adapter to those comparisons. We evaluate the adapters on candidate World Values Survey (WVS) items using prompts that omit country names and distinguish four questions: recovery of the imposed labels, transfer of the anchor signal to new text, coherence between the GPS anchors and human WVS responses, and agreement between adapter and human scores. The adapters recover the imposed pairwise labels. On a purposively selected sixteen-country development panel, adapter trust scores completely separate the two GPS-sign groups and have a rank correlation of (0.74) with continuous GPS trust scores. Human-GPS and adapter-human associations remain unresolved on the same panel, and results for the other preference dimensions are heterogeneous. These findings show that an anchored policy can retain a declared aggregate signal without thereby reproducing human response patterns. The contribution is therefore both an inspectable construction and an evaluation framework that separates anchor transfer from human criterion agreement.
|
| 669 |
Mind the Execution Gap: Action-Semantic Mismatch in World-Model Control
2610.06582
|
cs.CLcs.LGcs.AI
|
Shengtao Wen, Xiang Chen, Yu Tian, Lingbing Guo, Lina Gong |
World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execu...World-model controllers rely on action-conditioned dynamics for prediction and planning, yet real control systems often execute commands asynchronously due to communication delay, packet loss, reordering, and actuator buffering. We study how asynchronous execution changes the action semantics assumed within world-model controllers, rather than treating it only as an external control disturbance. Through controlled interventions, we identify two architecture-dependent failure modes: planning-based controllers such as TD-MPC2 suffer from a future-action timeline mismatch between imagined and executed action sequences, while recurrent world models such as DreamerV3 can attribute observed transitions to commands that were not actually applied. Our analysis shows that TD-MPC2 requires the correct future action sequence during latent dynamics rollout, whereas DreamerV3 requires timely attribution of each transition to the action that generated it. Based on these findings, we introduce two lightweight execution-consistent interfaces, Future-Sequence for TD-MPC2 and Applied-Action Feedback for DreamerV3, that correct these mismatches without modifying the pretrained world models. Experiments across delays, packet loss, reordering, multiple control domains, measured network traces, and a process-separated asynchronous stack consistently support both diagnoses and the corresponding architecture-specific corrections.
|
| 670 |
Mind the Accent Gap: British Accent Robustness in Speech-Driven Financial Voice Assistants
2610.06587
|
cs.CLcs.AIcs.SD
|
Aadam Haq, Oggi Rudovic, Malcolm Chadwick, Jay Rainey, Shucong Zhang |
AI voice assistants often use Automatic Speech Recognition (ASR) with LLM-based reasoning, yet existing systems struggle with regional British accents, including Scottish, Irish, and Welsh accents, since most ASR models are trained predominantly on American En...AI voice assistants often use Automatic Speech Recognition (ASR) with LLM-based reasoning, yet existing systems struggle with regional British accents, including Scottish, Irish, and Welsh accents, since most ASR models are trained predominantly on American English voice data. Consequently, errors can carry through to the LLM stage, corrupting tool-call arguments and producing wrong or missing responses, which is especially costly in finance. Deployable ASR must also meet tight latency and memory budgets, making an accent-robust model choice even harder. We introduce CavaBench, the first internally collected benchmark of spoken financial queries, and use it to evaluate a range of ASR models and their end-to-end ASR-LLM pipeline behaviour across self-reported British accents. We find that WER strongly predicts downstream tool-calling accuracy ($r = -0.93$) but can fail to reflect task-level performance, with accent-related failures varying substantially across models and acoustic conditions. These findings guide the design of more inclusive, reliable voice-based financial assistants.
|
| 671 |
COMPASS 2.0: psychometric representational similarity analysis distinguishes symptom structure from personal signal
2610.06615
|
cs.CL
|
Baihan Lin |
Language models can score psychiatric questionnaires from speech, but agreement with self-report may reflect the questionnaire rather than the person. We introduce psychometric representational similarity analysis, a framework for comparing the structure of sp...Language models can score psychiatric questionnaires from speech, but agreement with self-report may reflect the questionnaire rather than the person. We introduce psychometric representational similarity analysis, a framework for comparing the structure of speech-derived scores, self-report, item wording and theory, and implement it alongside person-level construct scoring in COMPASS 2.0. We show how similarly worded items induce covariance without psychological signal. In pre-registered discovery and confirmation analyses of clinical interviews from 275 participants, language-derived symptom geometry resembled wording more than self-report, with no structure beyond wording detected by the registered tests. Geometric agreement with self-report survived assigning participants someone else's answers, whereas person-paired scores captured distress more than specific symptoms. Complementary analyses examined counselling quality and wording structure across 34 instruments and the Research Domain Criteria (RDoC) framework. These findings distinguish agreement about psychological structure from evidence that language-derived assessments track individual people.
|
| 672 |
Frozen Factor or Spectral Band? Disentangling Two Choices in Low-Rank LoRA
2610.06621
|
cs.CLcs.LGcs.AI
|
Adnan Slimane Ali, Ayoub Belfatmi, David Ngwe Pouth |
Spectral variants of low-rank adaptation (LoRA) choose both a subspace and which factor to freeze. We separate these choices by freezing the input factor A or output factor B on the top or bottom singular directions of pretrained weights, with learning rates s...Spectral variants of low-rank adaptation (LoRA) choose both a subspace and which factor to freeze. We separate these choices by freezing the input factor A or output factor B on the top or bottom singular directions of pretrained weights, with learning rates selected separately. At rank 2, the same-band advantage of freezing A is larger than either within-factor band difference on all four task-model pairs with complete comparisons. Freezing B also trails comparable-budget free LoRA by 8-18 percentage points on five pairs spanning a formatting task and OpenBookQA. The A-frozen advantage persists in a single-GPU-model replication and within individual MLP module groups, including controls with equal or greater trainable counts for B frozen, and when A is frozen on a random orthonormal basis. The factor contrast weakens with rank. On OpenBookQA / Qwen2.5-1.5B at rank 16, PEFT's MiCA implementation trails comparable-budget LoRA by 3.08 points under a shared training recipe transferred from the MiCA paper. A trained oracle output subspace largely removes the low-rank deficit; partial warm-up gains recur across three direction seeds. The factor-versus-band ordering is descriptive; an approximate multiplicity audit weakens several earlier significance claims. These results extend known factor asymmetry by showing how its magnitude depends on spectral placement, rank and training conditions.
|
| 673 |
What Matters for Latent Reasoning with Flow Matching
2610.06666
|
cs.CLcs.LG
|
Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos |
Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely chan...Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.
|
| 674 |
Language models can notice an impossible engineering problem yet still report it as solved
2610.06668
|
cs.CLcs.AI
|
Shaoliang Yang, Jun Wang |
Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or a...Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of their flawed counterparts. Each reply required a "solved" or "cannot solve" status; rejection meant "cannot solve" or withholding an answer. The initial prompts did not warn that problems could be flawed. Across three recent models, 12 of 90 replies failed to reject a flawed problem. In 11 of these replies, the model stated the flaw, answered a corrected problem and still reported the original as "solved", according to artificial intelligence raters and numerical checks. We later retested four models from one provider, offering "flawed" instead of "cannot solve" and asking them to name and explain the defect. Three models showed statistically significant increases in rejection, but valid-problem solving fell in three. Evaluations therefore need to score both versions and distinguish flaw recognition from the reported status.
|
| 675 |
The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance
2610.06673
|
cs.CLcs.AI
|
Stefan B\"uhler, David Exler, Markus Reischl, Mark Schutera |
Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language mode...Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait and give up a higher payoff, which measures stoppability. The two compliance rates combine into a compliance index $\kappa$. Applied to twelve language models, the benchmark shows that seven mostly follow the instruction in both probes and justify their action by pointing to the instruction. Only Claude Sonnet-4.6 and Claude Opus-4.7 can be stopped without being exploitable, Claude Opus-4.6 and GPT-5-mini resist both instructions, and no model is exploitable but unstoppable. Knowing where a language model sits on the compliance index $\kappa$ matters for human operators and for multi-agent systems, whether distributed or orchestrated.
|
| 676 |
How Sparse Probability Maps Shape Mixture-of-Experts Routing
2610.06677
|
cs.CLcs.LG
|
Tom\'as Brogueira, Marcos Treviso, Miguel Couceiro |
Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros...Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity survives training. We train matched 300M and 1B top-2 MoE language models with softmax, 1.5-entmax, sparsemax and 2-normmax, and find that the maps behave very differently once trained: at 1B, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert, sparsemax retains the most mass, and normmax routes 21% of tokens to a single expert. These outcomes are not properties of the maps alone. Each map drops a selected expert only when the gap between the two largest scores reaches a fixed threshold, and the trained routers differ in the score distribution they learn: the entmax router learns scores with roughly half the spread of softmax's, which keeps its top-2 gaps below its threshold, while sparsemax and normmax, which share the same threshold, learn different gap distributions and hence different participation. Routers thus co-adapt their scores to the map, and a map's capacity to produce zeros does not by itself determine expert participation. While none of the sparse maps improves validation loss over softmax, they make the trained models far less sensitive to selecting more experts at inference: sparsemax trained with K=2 loses 0.02 nats when run with K=8, where softmax loses 0.58. Our results indicate that adaptive MoE routing has to be designed around the joint behavior of the probability map and the learned scores, rather than around the map alone.
|
| 677 |
Aligning Multimodal Patient Evidence with Biomedical Knowledge Graphs for Clinical LLMs
2610.06685
|
cs.CLcs.LG
|
Jiawen Du, Arshan Ali Khan, Chenhao Zhang, Zachary Plotkin, Li Shen |
Clinical questions often depend on linking a patient's multimodal evidence to external biomedical knowledge, yet existing predictive systems rarely represent such links explicitly, so they can neither be traced to their evidence sources nor removed to measure ...Clinical questions often depend on linking a patient's multimodal evidence to external biomedical knowledge, yet existing predictive systems rarely represent such links explicitly, so they can neither be traced to their evidence sources nor removed to measure their contributions. We present MM-KG (Multimodal Knowledge Graph), which represents heterogeneous, multimodal patient observations and biomedical concepts as separate layers in one typed graph, joined by explicit alignment edges. First, modality-specific harmonizers convert EHR text, imaging, genomic, and biospecimen data into typed observations mapped to UMLS concepts, which a route-prioritized aligner links to a biomedical knowledge graph. Query-conditioned retrieval then selects a compact subgraph for downstream use by a large language model or a graph neural network. We build MM-KGs for MIMIC-IV and ADNI, and evaluate them with a 2x2 design that separates patient evidence, biomedical knowledge, and their interaction. On questions that require both sources, neither source alone performs far above chance, whereas their combination yields a drug-controlled AUROC interaction of +0.194 on MIMIC and +0.299 on ADNI. On held-out five-candidate ranking, MM-KG outperforms MindMap by +0.131 Hits@1 and leads an adapted GraphCare on the items that require consulting the patient, and deleting the single answer-bearing relation from the retrieved packet returns Hits@1 to the no-knowledge baseline. Finally, query-conditioned retrieval reaches 0.731 AUROC with 6.8x less context than the strongest generic policy, whereas static knowledge graph context gives no consistent gain on ordinary outcome prediction. Knowledge graphs thus benefit clinical LLMs not as background context but as explicit links between multimodal patient evidence and the relation a question requires, and MM-KG makes these links retrievable, traceable, and testable.
|
| 678 |
Reading the Mood: Emotion-Guided Book-to-Music Recommendation via CGANs and LLMs
2610.06703
|
cs.CLcs.LG
|
Manousos Linardakis, Georgios Alexandridis |
Background music that matches the mood of a text has been shown to make readers feel more immersed and improve their reading experience, motivating recommender systems that pair books with mood-matched music. In this direction, we present Sentiment Aware Gener...Background music that matches the mood of a text has been shown to make readers feel more immersed and improve their reading experience, motivating recommender systems that pair books with mood-matched music. In this direction, we present Sentiment Aware Generative Adversarial Network for Cross Domain Recommendation (SAGA-CDR), a two-phase cross-domain recommendation framework that personalizes music suggestions and emotionally aligns them with the book being read. In the first phase, transformer-based sentiment embeddings are constructed from user reviews and mapped across domains via a Conditional Generative Adversarial Network, whose mask-conditioned generator handles missing sentiment components and injects stochasticity for richer preference transfer. A compact rating neural network then fuses sentiment-specific interaction scores with a collaborative filtering prior to predict music ratings. In the second phase, large language models classify each book into a valence-arousal emotional quadrant, and candidate tracks are filtered to match that quadrant. Experiments on both the English Amazon and Chinese Douban datasets show that SAGA-CDR achieves the best rating prediction accuracy on Amazon (RMSE 0.98) and the lowest RMSE on Douban (0.91), with ranking performance competitive with the strongest sentiment-aware baseline, even in cross-lingual settings.
|
| 679 |
Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
2610.06817
|
cs.CLcs.LGcs.AIcs.SDeess.AS
|
Sahil Mahendrakar |
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two halves is trained separately again...We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two halves is trained separately against the frozen teacher. It has 10x fewer parameters and needs 15x less compute. We first synthesize a corpus with the teacher and keep its durations, pitch, energy and phoneme features. We then train a small text side to predict these values, and a small decoder to turn the teacher's saved values into the teacher's audio, first with spectral losses and then adversarially. Finally, we connect the two halves and quantize the weights to int8. It needs no alignment learning and no joint training, and it runs on one laptop. Stored in int8, Paradee is 8.5 MB, runs 25x faster than real time on one CPU thread, and scores 4.41 on UTMOS against the teacher's 4.52. The student initially kept a slight buzz, which we trace to the phase of voiced speech between 2 and 8 kHz. A phase-locking filter applied after synthesis removes most of it, with no training and no extra parameters. Code, model files and audio samples are at https://github.com/sahilmahendrakar/paradee
|
| 680 |
Recursive Video In-Context Learning for Agentic Robot
2610.06843
|
cs.CLcs.AI
|
Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang |
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The fu...LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.
|
| 681 |
Base Models Can Reason By Taking a Cue From Training Data
2610.06851
|
cs.CLcs.LGcs.AI
|
Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min |
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance...In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
|
| 682 |
INMS: Memory Sharing for Large Language Model based Agents
2404.09982
|
cs.CL
|
Hang Gao, Yongfeng Zhang |
While Large Language Model (LLM) based agents excel at complex tasks, their performance in open-ended scenarios is often constrained by isolated operation and reliance on static databases, missing the dynamic knowledge exchange of human dialogue. To bridge thi...While Large Language Model (LLM) based agents excel at complex tasks, their performance in open-ended scenarios is often constrained by isolated operation and reliance on static databases, missing the dynamic knowledge exchange of human dialogue. To bridge this gap, we propose the INteractive Memory Sharing (INMS) framework, an asynchronous interaction paradigm for multi-agent systems. By integrating real-time memory filtering, storage, and retrieval, INMS establishes a shared conversational memory pool. This enables continuous, dialogue-like memory sharing among agents, promoting collective self-enhancement and dynamically refining the retrieval mediator based on interaction history. Extensive experiments across three datasets demonstrate that INMS improves agent performance by effectively modeling multi-agent interaction and collective knowledge sharing.
|
| 683 |
TSCheater: Generating High-Quality Tibetan Adversarial Texts via Visual Similarity
2412.02371
|
cs.CL
|
Xi Cao, Quzong Gesang, Yuan Sun, Nuo Qun, Tashi Nyima |
Language models based on deep neural networks are vulnerable to textual adversarial attacks. While rich-resource languages like English are receiving focused attention, Tibetan, a cross-border language, is gradually being studied due to its abundant ancient li...Language models based on deep neural networks are vulnerable to textual adversarial attacks. While rich-resource languages like English are receiving focused attention, Tibetan, a cross-border language, is gradually being studied due to its abundant ancient literature and critical language strategy. Currently, there are several Tibetan adversarial text generation methods, but they do not fully consider the textual features of Tibetan script and overestimate the quality of generated adversarial texts. To address this issue, we propose a novel Tibetan adversarial text generation method called TSCheater, which considers the characteristic of Tibetan encoding and the feature that visually similar syllables have similar semantics. This method can also be transferred to other abugidas, such as Devanagari script. We utilize a self-constructed Tibetan syllable visual similarity database called TSVSDB to generate substitution candidates and adopt a greedy algorithm-based scoring mechanism to determine substitution order. After that, we conduct the method on eight victim language models. Experimentally, TSCheater outperforms existing methods in attack effectiveness, perturbation magnitude, semantic similarity, visual similarity, and human acceptance. Finally, we construct the first Tibetan adversarial robustness evaluation benchmark called AdvTS, which is generated by existing methods and proofread by humans.
|
| 684 |
Three tiers of computation in transformers and in brain architectures
2503.04848
|
cs.CL
|
E Graham, R Granger |
Human language and logic abilities are computationally quantified within the well-studied grammar-automata hierarchy. We identify three hierarchical tiers and two corresponding transitions and show their correspondence to specific abilities in transformer-base...Human language and logic abilities are computationally quantified within the well-studied grammar-automata hierarchy. We identify three hierarchical tiers and two corresponding transitions and show their correspondence to specific abilities in transformer-based language models (LMs). These emergent abilities have often been described in terms of scaling; we show that it is the transition between tiers, rather than scaled size itself, that determines a system's capabilities. Specifically, humans effortlessly process language yet require critical training to perform arithmetic or logical reasoning tasks; and LMs possess language abilities absent from predecessor systems, yet still struggle with logical processing. We submit a novel benchmark of computational power, provide empirical evaluations of humans and fifteen LMs, and, most significantly, provide a theoretically grounded framework to promote careful thinking about these crucial topics. The resulting principled analyses provide explanatory accounts of the abilities and shortfalls of LMs, and suggest actionable insights into the expansion of their logic abilities.
|
| 685 |
Rethinking the Relationship between the Power Law and Hierarchical Structures
2505.04984
|
cs.CL
|
Kai Nakaishi, Ryo Yoshida, Kohei Kajikawa, Koji Hukushima, Yohei Oseki |
Statistical analysis of corpora provides an approach to quantitatively investigate natural languages. This approach has revealed that several power laws consistently emerge across different corpora and languages, suggesting universal mechanisms underlying lang...Statistical analysis of corpora provides an approach to quantitatively investigate natural languages. This approach has revealed that several power laws consistently emerge across different corpora and languages, suggesting universal mechanisms underlying languages. In particular, the power-law decay of correlations has been interpreted as evidence of underlying hierarchical structures in syntax, semantics, and discourse. This perspective has also been extended beyond corpora produced by human adults, including child speech, birdsong, and chimpanzee action sequences. However, the argument supporting this interpretation has not been empirically tested in natural languages. To address this gap, the present study examines the validity of the argument for syntactic structures. Specifically, we test whether the statistical properties of parse trees align with the assumptions in the argument. Using English and Japanese corpora, we analyze the mutual information, deviations from probabilistic context-free grammars (PCFGs), and other properties in natural language parse trees, as well as in the PCFG that approximates these parse trees. Our results indicate that the assumptions do not hold for syntactic structures and that it is difficult to apply the proposed argument not only to sentences by human adults but also to other domains, highlighting the need to reconsider the relationship between the power law and hierarchical structures.
|
| 686 |
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
2505.17399
|
cs.CL
|
Haoyu Sun, Huichen Will Wang, Jiawei Gu, Linjie Li, Yu Cheng |
Front-end engineering involves a complex workflow where engineers conceptualize designs, translate them into code, and iteratively refine the implementation. While recent benchmarks primarily focus on converting visual designs to code, we present FullFront, a ...Front-end engineering involves a complex workflow where engineers conceptualize designs, translate them into code, and iteratively refine the implementation. While recent benchmarks primarily focus on converting visual designs to code, we present FullFront, a benchmark designed to evaluate Multimodal Large Language Models (MLLMs) \textbf{across the full front-end development pipeline}. FullFront assesses three fundamental tasks that map directly to the front-end engineering pipeline: Webpage Design (conceptualization phase), Webpage Perception QA (comprehension of visual organization and elements), and Webpage Code Generation (implementation phase). Unlike existing benchmarks that use either scraped websites with bloated code or oversimplified LLM-generated HTML, FullFront employs a novel, two-stage process to transform real-world webpages into clean, standardized HTML while maintaining diverse visual designs and avoiding copyright issues. Extensive testing of state-of-the-art MLLMs reveals significant limitations in page perception, code generation (particularly for image handling and layout), and interaction implementation. Our results quantitatively demonstrate performance disparities across models and tasks, and highlight a substantial gap between current MLLM capabilities and human expert performance in front-end engineering. The FullFront benchmark and code are available in https://github.com/Mikivishy/FullFront.
|
| 687 |
Studying the Soupability of Documents in State Space Models
2505.24033
|
cs.CLcs.LG
|
Yasaman Jafari, Zixian Wang, Leon Bergen, Taylor Berg-Kirkpatrick |
We investigate whether hidden states from Structured State Space Models (SSMs) can be merged post hoc to support downstream reasoning. Inspired by model souping, we study document souping, a strategy where documents are encoded independently, and their represe...We investigate whether hidden states from Structured State Space Models (SSMs) can be merged post hoc to support downstream reasoning. Inspired by model souping, we study document souping, a strategy where documents are encoded independently, and their representations are pooled, via simple operations like averaging, into a single context state. This approach enables modular encoding and reuse without reprocessing the full input for each query. We demonstrate that finetuned Mamba2 models with souped representations achieve competitive or superior performance across multi-hop QA, sparse retrieval, and long-document reasoning tasks compared to the standard monolithic encoding approach. For example, on the RACE and QuALITY benchmarks for long document question answering, this method substantially outperforms a traditional concatenation approach. Crucially, this modular design scales to hundreds of segments (we test up to 256) while delivering substantial savings in inference cost, unlocking new possibilities for large-scale corpus reasoning.
|
| 688 |
Towards Probabilistic Question Answering Over Tabular Data
2506.20747
|
cs.CL
|
Chen Shen, Estevam Hruschka |
Progress in question answering (QA) over tabular data has enabled reliable factual retrieval from relational tables. However, many real-world questions are probabilistic, requiring reasoning under uncertainty and latent conditional dependencies that are not di...Progress in question answering (QA) over tabular data has enabled reliable factual retrieval from relational tables. However, many real-world questions are probabilistic, requiring reasoning under uncertainty and latent conditional dependencies that are not directly stored in individual cells. We introduce LUCARIO, a large-scale benchmark for probabilistic QA over real-world tabular datasets, covering diverse domains, dependency patterns, and natural language variations. We propose Auto-BN, a fully automatic neuro-symbolic framework that induces a Bayesian Network from raw tables, translates natural-language questions into formal probabilistic queries, and performs exact inference. Experiments across multiple LLM backbones show that Auto-BN consistently outperforms NL2SQL, retrieval-based, and premise-based baselines, while maintaining a near-zero error rate. LUCARIO fills a relevant gap and provides a practical and scalable testbed for advancing probabilistic reasoning over structured data.
|
| 689 |
Automating MD simulations for Proteins using Large language Models: NAMD-Agent
2507.07887
|
cs.CL
|
Omid Barati Farimani, Achuth Chandrasekhar, Amir Barati Farimani |
Molecular dynamics (MD) simulations are essential for understanding protein structure, dynamics, and function, but preparing, running, and analyzing simulations remains time-consuming and error-prone. We present an automated pipeline that combines large langua...Molecular dynamics (MD) simulations are essential for understanding protein structure, dynamics, and function, but preparing, running, and analyzing simulations remains time-consuming and error-prone. We present an automated pipeline that combines large language model (LLM) agents with Python scripting and HTMD MCP tools to generate simulation-ready inputs for NAMD3/CHARMM, execute simulations, analyze outputs, and recover from build or runtime failures. The framework was evaluated across five biomolecular system classes: a protein-DNA complex (p53 DNA-binding domain bound to its response element), a protein-membrane system (M2 muscarinic receptor with iperoxo in a POPC/cholesterol bilayer), a protein-ligand series (five congeneric TYK2 inhibitors), a protein-water reference (ubiquitin), and a protein-protein complex (barnase-barstar). For the protein-DNA system, the automated workflow reproduced key metrics from an independent published benchmark. To distinguish framework performance from model-specific behavior, we repeated the complete protein-DNA study using three LLM orchestrators: Claude Opus-4.8, GPT-5.6 Sol, and Nemotron 3 Ultra. All three completed the workflow, but they differed in benchmark-ranking fidelity and by up to two orders of magnitude in token consumption and cost. Across all system classes, the agent recovered experimentally and computationally established behavior. Additional post-processing software was used to refine simulation outputs, enabling a complete and largely hands-free workflow. This approach reduces setup effort, limits manual errors, supports parallel handling of diverse biomolecular systems, and provides a robust, adaptable foundation for LLM-driven automation in computational structural biology.
|
| 690 |
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
2507.11230
|
cs.CL
|
Lyzander Marciano Andrylie, Inaya Rahmanisa, Mahardika Krisna Ihsani, Alfan Farizki Wicaksono, Haryo Akbarianto Wibowo |
Understanding the multilingual mechanisms of large language models (LLMs) provides insight into how they process different languages, yet this remains challenging. Existing studies often focus on individual neurons, but their polysemantic nature makes it diffi...Understanding the multilingual mechanisms of large language models (LLMs) provides insight into how they process different languages, yet this remains challenging. Existing studies often focus on individual neurons, but their polysemantic nature makes it difficult to isolate language-specific units from cross-lingual representations. To address this, we explore sparse autoencoders (SAEs) for their ability to learn monosemantic features that represent concrete and abstract concepts across languages in LLMs. While some of these features are language-independent, the presence of language-specific features remains underexplored. In this work, we introduce $\textit{SAE-LAPE}$, a method based on feature activation probability, to identify language-specific features within the feed-forward network. We find that many such features predominantly appear in the middle to late layers of the model and are interpretable. These features influence the model's multilingual performance and language output, and can be used for language identification with performance comparable to fastText, along with more interpretability. Our code and complete figures are available at https://github.com/LyzanderAndrylie/language-specific-features.
|
| 691 |
Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
2508.05234
|
cs.CLcs.AI
|
Haonan Shangguan, Xiaocui Yang, Shi Feng, Daling Wang, Yifei Zhang |
Current approaches for Multimodal Sentiment Analysis (MSA) primarily leverage the knowledge and reasoning capabilities of parameter-heavy (Multimodal) LLMs for classification, overlooking autonomous multimodal sentiment reasoning generation in resource-constra...Current approaches for Multimodal Sentiment Analysis (MSA) primarily leverage the knowledge and reasoning capabilities of parameter-heavy (Multimodal) LLMs for classification, overlooking autonomous multimodal sentiment reasoning generation in resource-constrained environments. In this paper, we focus on the Resource-Limited Joint Multimodal Sentiment Reasoning and Classification task, JMSRC, which simultaneously performs multimodal sentiment reasoning chain generation and sentiment classification only with a lightweight model. We propose a Multimodal Chain-of-Thought Reasoning Distillation model, MulCoT-RD, designed for JMSRC that employs a "Teacher-Assistant-Student" distillation paradigm to address deployment constraints in resource-limited environments. We first leverage a high-performance Multimodal Large Language Model (MLLM) to generate the initial reasoning dataset and train a medium-sized assistant model with a multi-task learning mechanism. A lightweight student model is jointly trained to perform efficient multimodal sentiment reasoning generation and classification. Extensive experiments on four datasets demonstrate that MulCoT-RD, with only 3B parameters, achieves strong performance on JMSRC while exhibiting robust generalization and enhanced interpretability.
|
| 692 |
COCORELI: Enforcing Execution Preconditions for Reliable Collaborative Instruction Following
2509.04470
|
cs.CLcs.AI
|
Swarnadeep Bhar, Omar Naim, Eleni Metheniti, Bastien Navarri, Lo\"ic Cabannes |
Autonomous agents executing human instructions must operate reliably even when instructions are incomplete. While recent approaches improve detection of missing information, detection alone is insufficient: agents often proceed to execution even after recogniz...Autonomous agents executing human instructions must operate reliably even when instructions are incomplete. While recent approaches improve detection of missing information, detection alone is insufficient: agents often proceed to execution even after recognizing underspecification, leading to incorrect or unsafe actions. We identify this failure as arising from a lack of coupling between detection and execution, and propose that reliable behavior requires enforcing missing information as a precondition for action. We instantiate this principle in Cocoreli, a modular architecture that represents task structure, tracks missing information, and blocks execution until required details are resolved through targeted clarification. In Cocoreli, detection and prevention are structurally coupled: detecting a missing parameter simultaneously blocks execution. We evaluate Cocoreli in a controlled construction environment isolating underspecification and sequential execution. Cocoreli blocks execution under unresolved specifications by construction, eliminating hallucinated actions. In contrast, chain-of-thought, prompt-chaining, and ReAct-style reasoning may still execute under incomplete specifications despite high detection rates. The same representation supports abstraction and reuse, and generalizes to API workflow tasks on ToolBench. These results show that reliable collaborative execution requires architectural enforcement, not just model capability
|
| 693 |
RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance
2509.25604
|
cs.CLcs.LG
|
Tianlang Chen, Minkai Xu, Jure Leskovec, Stefano Ermon |
Diffusion Large Language Models (dLLMs) have shown strong reasoning capabilities, yet further improving them typically requires costly post-training with additional data and supervision. We ask whether a post-trained dLLM can improve itself at inference time w...Diffusion Large Language Models (dLLMs) have shown strong reasoning capabilities, yet further improving them typically requires costly post-training with additional data and supervision. We ask whether a post-trained dLLM can improve itself at inference time without additional training, data, or reward models. This requires a guidance signal from the checkpoints alone that is well-defined on the partially masked intermediate states of dLLMs, which existing methods fail to provide. Here we propose Reward-Free Guidance (RFG), a training-free framework for inference-time self-improvement of dLLMs that derives such a signal from the checkpoints themselves. We theoretically demonstrate that a reward signal for a partially masked state can be parameterized by the log-likelihood ratio between a post-trained policy dLLM and its reference model. In practice, this signal can be obtained using off-the-shelf checkpoints alone. Extensive experiments show that RFG consistently improves state-of-the-art post-trained dLLMs by up to 16.1%. Furthermore, despite being entirely training-free, RFG delivers gains that rival or even surpass resource-intensive reinforcement learning techniques.
|
| 694 |
Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
2510.14242
|
cs.CLcs.LG
|
Parsa Hejabi, Elnaz Rahmati, Alireza S. Ziabari, Morteza Dehghani |
Large Language Models (LLMs) often produce inconsistent answers when faced with different phrasings of the same prompt. In this paper, we propose Flip-Flop Consistency ($F^2C$), an unsupervised training method that improves robustness to such perturbations. $F...Large Language Models (LLMs) often produce inconsistent answers when faced with different phrasings of the same prompt. In this paper, we propose Flip-Flop Consistency ($F^2C$), an unsupervised training method that improves robustness to such perturbations. $F^2C$ is composed of two key components. The first, Consensus Cross-Entropy (CCE), uses a majority vote across prompt variations to create a hard pseudo-label. The second is a representation alignment loss that pulls lower-confidence and non-majority predictors toward the consensus established by high-confidence, majority-voting variations. We evaluate our method on 11 datasets spanning four NLP tasks, with 4-15 prompt variations per dataset. On average, $F^2C$ raises observed agreement by 11.62%, improves mean $F_1$ by 8.94%, and reduces performance variance across formats by 3.29%. In out-of-domain evaluations, $F^2C$ generalizes effectively, increasing $\overline{F_1}$ and agreement while decreasing variance across most source-target pairs. Finally, when trained on only a subset of prompt perturbations and evaluated on held-out formats, $F^2C$ consistently improves both performance and agreement while reducing variance. These findings highlight $F^2C$ as an effective unsupervised method for enhancing LLM consistency, performance, and generalization under prompt perturbations. Code is available at https://github.com/ParsaHejabi/Flip-Flop-Consistency-Unsupervised-Training-for-Robustness-to-Prompt-Perturbations-in-LLMs.
|
| 695 |
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
2510.15543
|
cs.CLcs.AIcs.MM
|
Qiyu Wu, Shuyang Cui, Satoshi Hayakawa, Wei-Yao Wang, Hiromi Wakaki |
Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP aligning modality-specific embeddi...Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP aligning modality-specific embeddings with contrastive learning, recent multimodal large language models (MLLMs) enable a unified encoder that directly processes composed inputs. While flexible and advanced, we identify that unified encoders trained with conventional contrastive learning are prone to learn modality shortcut, leading to poor robustness under distribution shifts. We propose a modality composition awareness framework to mitigate this issue. Concretely, it consists of a preference loss enforces multimodal embeddings to outperform their unimodal counterparts, and a composition regularization objective aligns multimodal embeddings with prototypes composed from its unimodal parts. These objectives explicitly model structural relationships between the composed representation and its unimodal counterparts. Experiments on various benchmarks show gains in out-of-distribution retrieval, highlighting modality composition awareness as a effective principle for robust composed multimodal retrieval when utilizing MLLMs as the unified encoder.
|
| 696 |
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
2510.24789
|
cs.CL
|
Gokul Ganesan |
Watermarking has been proposed as a lightweight mechanism to identify AI-generated text, with schemes typically relying on perturbations to token distributions. While prior work shows that paraphrasing can weaken such signals, these attacks remain partially de...Watermarking has been proposed as a lightweight mechanism to identify AI-generated text, with schemes typically relying on perturbations to token distributions. While prior work shows that paraphrasing can weaken such signals, these attacks remain partially detectable or degrade text quality. We demonstrate that cross-lingual summarization attacks (CLSA) -- translation to a pivot language followed by summarization and optional back-translation -- constitute a qualitatively stronger attack vector. By forcing a semantic bottleneck across languages, CLSA systematically destroys token-level statistical biases while preserving semantic fidelity. In experiments across multiple watermarking schemes (KGW, SIR, XSIR, Unigram) and five languages (Amharic, Chinese, Hindi, Spanish, Swahili), we show that CLSA reduces watermark detection accuracy more effectively than monolingual paraphrase at similar quality levels. Our results highlight an underexplored vulnerability that challenges the practicality of watermarking for provenance or regulation. We argue that robust provenance solutions must move beyond distributional watermarking and incorporate cryptographic or model-attestation approaches. On 300 held-out samples per language, CLSA consistently drives detection toward chance while preserving task utility. Concretely, for XSIR (explicitly designed for cross-lingual robustness), AUROC with paraphrasing is $0.827$, with Cross-Lingual Watermark Removal Attacks (CWRA) [He et al., 2024] using Chinese as the pivot, it is $0.823$, whereas CLSA drives it down to $0.53$ (near chance). Results highlight a practical, low-cost removal pathway that crosses languages and compresses content without visible artifacts.
|
| 697 |
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
2511.02817
|
cs.CLcs.AI
|
Amanda Bertsch, Adithya Pratapa, Teruko Mitamura, Graham Neubig, Matthew R. Gormley |
As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evaluations have recently been released, these evaluations tend to rely on retrieval ...As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evaluations have recently been released, these evaluations tend to rely on retrieval from one or more sections of the context, which allows nearly all of the context tokens to be disregarded as noise. This represents only one type of task that might be performed with long context. We introduce Oolong, a benchmark of long-context reasoning tasks that require analyzing individual chunks of text on an atomic level, and then aggregating these analyses to answer distributional questions. Oolong is separated into two task sets: Oolong-synth, a set of naturalistic synthetic tasks, where we can easily ablate components of the reasoning problem; and Oolong-real, a downstream setting which requires reasoning over real-world conversational data. Oolong requires models to reason over large quantities of examples, to perform both classification and counting in-context, and to reason over temporal and user relations. Even frontier models struggle on Oolong, with GPT-5, Claude-Sonnet-4, and Gemini-2.5-Pro all achieving less than 50% accuracy on both splits at 128K. We release the data and evaluation harness for Oolong to enable further development of models that can reason over large quantities of text.
|
| 698 |
EulerESG: Automating ESG Disclosure Analysis with LLMs
2511.21712
|
cs.CLcs.AI
|
Yi Ding, Xushuo Tang, Zhengyi Yang, Wenqian Zhang, Simin Wu |
Environmental, Social, and Governance (ESG) reports have become central to how companies communicate climate risk, social impact, and governance practices, yet they are still published primarily as long, heterogeneous PDF documents. This makes it difficult to ...Environmental, Social, and Governance (ESG) reports have become central to how companies communicate climate risk, social impact, and governance practices, yet they are still published primarily as long, heterogeneous PDF documents. This makes it difficult to systematically answer seemingly simple questions. Existing tools either rely on brittle rule-based extraction or treat ESG reports as generic text, without explicitly modelling the underlying reporting standards. We present \textbf{EulerESG}, an LLM-powered system for automating ESG disclosure analysis with explicit awareness of ESG frameworks. EulerESG combines (i) dual-channel retrieval and LLM-driven disclosure analysis over ESG reports, and (ii) an interactive dashboard and chatbot for exploration, benchmarking, and explanation. Using four globally recognised companies and twelve SASB sub-industries, we show that EulerESG can automatically populate standard-aligned metric tables with high fidelity (up to 0.95 average accuracy) while remaining practical in end-to-end runtime, and we compare several recent LLM models in this setting. The full implementation, together with a demonstration video, is publicly available at https://github.com/UNSW-database/EulerESG.
|
| 699 |
Mitigating Social Desirability Bias in Random Silicon Sampling
2512.22725
|
cs.CL
|
Sashank Chapala, Maksym Mironov, Songgaojun Deng |
Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data to...Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data toward socially acceptable answers. Existing studies on social desirability bias in LLM-based sampling remain limited. In this work, we investigate whether minimal, psychologically grounded prompt wording can mitigate this bias and improve alignment between silicon and human samples. We conducted a study using data from the American National Election Study (ANES) on three LLMs from two model families: the open-source Llama-3.1 series and GPT-4.1-mini. We first replicate a baseline silicon sampling study, confirming the persistent Social Desirability Bias. We then test four prompt-based mitigation methods: \emph{reformulated} (neutral, third-person phrasing), \emph{reverse-coded} (semantic inversion), and two meta-instructions, \emph{priming} and \emph{preamble}, respectively encouraging analytics and sincerity. Alignment with ANES is evaluated using Jensen-Shannon Divergence with bootstrap confidence intervals. Our results demonstrate that reformulated prompts most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES. Reverse-coding produced mixed results across eligible items, while the Priming and Preamble encouraged response uniformity and showed no systematic benefit for bias mitigation. Our findings validate the efficacy of prompt-based framing controls in mitigating inherent Social Desirability Bias in LLMs, providing a practical path toward more representative silicon samples.
|
| 700 |
Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
2601.06787
|
cs.CL
|
Jaewon Sok, Jewon Yeom, Seonghyeon Park, Jeongjae Park, Taesup Kim |
Large Language Models (LLMs) are known to contain significant redundancy, yet a systematic explanation for why certain components, particularly in higher layers, are more redundant has remained elusive. In this work, we identify the BOS sink phenomenon as a ke...Large Language Models (LLMs) are known to contain significant redundancy, yet a systematic explanation for why certain components, particularly in higher layers, are more redundant has remained elusive. In this work, we identify the BOS sink phenomenon as a key mechanism driving this layer-wise sensitivity. We show that attention heads with high BOS sink scores are strongly associated with functional redundancy: such heads, especially in deeper layers, contribute little to predictive performance and effectively serve as dumping grounds for superfluous attention weights. Leveraging this insight, we introduce a simple pruning strategy that removes high-BOS sink heads. Experiments on Gemma-3, Llama-3.1, and Qwen3 demonstrate that this approach identifies redundant transformer components more reliably than weight- and activation-based criteria in terms of downstream task retention, remaining close to dense baselines at low-to-moderate pruning ratios. We further find that high-scoring sink heads sustain their focus on BOS as context length grows. Overall, our results suggest that structural properties of attention offer a more direct basis for model compression than magnitude-based methods.
|
| 701 |
Symphonym: Universal Phonetic Embeddings for Cross-Script Toponym Matching
2601.06932
|
cs.CLcs.AI
|
Stephen Gadd |
Matching place names across writing systems is a persistent obstacle to integrating multilingual geographic sources, from modern gazetteers to medieval itineraries and colonial-era surveys. Existing approaches rely on language-specific phonetic algorithms or o...Matching place names across writing systems is a persistent obstacle to integrating multilingual geographic sources, from modern gazetteers to medieval itineraries and colonial-era surveys. Existing approaches rely on language-specific phonetic algorithms or on romanisation that discards phonetic information, and none generalises across scripts. Symphonym maps toponyms from thirty-six writing systems into a unified 128-dimensional phonetic space, enabling direct cross-script comparison without language identification or phonetic resources at inference time. A Teacher-Student distillation architecture learns from articulatory features of IPA transcriptions and transfers this knowledge to a character-level Student. Trained on 73.5 million toponyms from GeoNames, Wikidata and the Getty TGN, the Student achieves the highest Recall@1 (89.3%) and MRR (92.8%) on the MEHDIE benchmark of medieval Hebrew and Arabic toponym matches, which is independent of the training data. An ablation on raw articulatory features alone reaches only 45.0% MRR. This revision reports a second model generation and three results that qualify the first: a corpus defect caused the original system to learn Chinese characters with Japanese readings, which we correct and quantify; the encoder weighted character content far above order, admitting 70.5% of random anagrams past the retrieval gate, which targeted negatives reduce to 5.0% without loss of typo tolerance; and better retrieval came with slightly worse separation of true from false matches. We report where the method wins decisively (across scripts) and where string metrics remain preferable (within the Latin script), and document an independent out-of-domain deployment on archival personal names.
|
| 702 |
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
2601.11854
|
cs.CLcs.AI
|
Yifei Zhang, Hooshang Nayyeri, Rinat Khaziev, Emine Yilmaz, Gokhan Tur |
Agentic task-oriented dialogue (TOD) requires systems to track concurrent goals, dependencies, and long-horizon state. We examine goal-lifecycle recovery from fixed dialogue trajectories. ATOD contains 1,000 synthetic dialogues annotated for six advanced-TOD p...Agentic task-oriented dialogue (TOD) requires systems to track concurrent goals, dependencies, and long-horizon state. We examine goal-lifecycle recovery from fixed dialogue trajectories. ATOD contains 1,000 synthetic dialogues annotated for six advanced-TOD properties, and ATOD-Eval defines metrics for dependency-sensitive completion, memory recall, and proactivity. We implement a symbolic-vector memory evaluator for ATOD-Eval. Among five prompt-only predictors and six matched-backbone memory baselines, our evaluator is the only configuration above 90% in both goal detection F1 and conditional status accuracy on medium dialogues. On complex dialogues, no baseline exceeds it on both metrics; it has the highest conditional status accuracy within the memory-based block and the lowest measured per-turn latency. These experiments assess lifecycle tracking rather than interactive agent task success. A cross-family judge swap and a manual audit with 94.0% agreement provide initial checks on measurement reliability. Code and data will be released at https://github.com/amazon-science/ATOD.
|
| 703 |
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
2601.13590
|
cs.CLcs.AI
|
Fan Huang, Haewoon Kwak, Jisun An |
Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuasion and could adopt counterfactual beliefs. We present a systematic evaluation of LLM susceptibilit...Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuasion and could adopt counterfactual beliefs. We present a systematic evaluation of LLM susceptibility to persuasion under the \emph{Source--Message--Channel--Receiver} (SMCR) communication framework. Across six mainstream Large Language Models (LLMs) and three domains (factual knowledge, medical QA, and social bias), we analyze how different persuasive strategies influence stated belief stability over multiple interaction turns. We further examine whether verbalized confidence prompting (i.e., eliciting self-reported confidence scores) affects resistance to persuasion. Results show that the smallest model (Llama 3.2-3B) exhibits extreme compliance, with 82.5\% of belief changes occurring at the first persuasive turn (average end turn of 1.1--1.4). Contrary to expectations, verbalized confidence prompting \emph{increases} vulnerability by accelerating belief erosion rather than enhancing robustness. Finally, an exploratory study of adversarial fine-tuning reveals highly model-dependent effectiveness: GPT-4o-mini achieves near-complete robustness (98.6\%), and Mistral~7B improves substantially (35.7\% $\rightarrow$ 79.3\%), but Llama models remain highly susceptible ($<$14\% RQ1) even when fine-tuned on their own failure cases. Together, these findings highlight substantial model-dependent limits of current robustness interventions and offer guidance for developing more trustworthy LLMs.
|
| 704 |
Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning
2601.16724
|
cs.CL
|
Kevin Fan, Eric Yun |
Automated Essay Scoring systems disproportionately penalize high-proficiency English as a Second Language (ESL) learners. We propose Contrastive Learning with Matched Essay Pairs (CL-MEP), a bi-directional alignment strategy. CL-MEP reduces this scoring bias b...Automated Essay Scoring systems disproportionately penalize high-proficiency English as a Second Language (ESL) learners. We propose Contrastive Learning with Matched Essay Pairs (CL-MEP), a bi-directional alignment strategy. CL-MEP reduces this scoring bias by 39.9% while improving overall accuracy, successfully disentangling valid syntactic complexity from surface-level grammatical errors.
|
| 705 |
Distilling Token-Trained Models into Byte-Level Models
2602.01007
|
cs.CL
|
Zishuo Bao, Jiaqi Leng, Junxiong Wang, Bowen Peng, Yucheng Lu |
Byte Language Models (BLMs) have emerged as a promising direction for scaling language models beyond tokenization. However, existing BLMs typically require training from scratch on trillions of bytes, making them prohibitively expensive. In this paper, we prop...Byte Language Models (BLMs) have emerged as a promising direction for scaling language models beyond tokenization. However, existing BLMs typically require training from scratch on trillions of bytes, making them prohibitively expensive. In this paper, we propose an efficient distillation recipe that converts existing token-trained LLMs into BLMs while retaining comparable capabilities. Our recipe follows a two-stage curriculum: (1) Progressive Knowledge Distillation, which aligns byte-level representations with the embeddings of the token-trained teacher model; and (2) Byte-Level Supervised Fine-Tuning, which enables end-to-end generation entirely in the byte space. We validate our approach across multiple model families, including Llama, Qwen, and OLMo, and demonstrate that the distilled BLMs retain most of the teacher models' performance using only approximately 125B bytes.
|
| 706 |
Adaptive Information Control for Search-Augmented LLM Reasoning
2602.01672
|
cs.CL
|
Siheng Xiong, Oguzhan Gungordu, James C. Kerce, Faramarz Fekri |
Search-augmented reasoning agents interleave multi-step reasoning with external retrieval, but uncontrolled retrieval can introduce redundant evidence, saturate the context, and destabilize reinforcement learning (RL). Existing outcome-based RL methods provide...Search-augmented reasoning agents interleave multi-step reasoning with external retrieval, but uncontrolled retrieval can introduce redundant evidence, saturate the context, and destabilize reinforcement learning (RL). Existing outcome-based RL methods provide only sparse terminal rewards, offering limited guidance for intermediate information-acquisition decisions. We propose DeepControl, an adaptive information-control framework based on information utility, a state-dependent estimate of the marginal value of retrieved evidence. The framework regulates information acquisition along two axes: extent, i.e., whether retrieval should continue, and resolution, i.e., how much retrieved detail should be exposed. It implements these controls through retrieval-continuation guidance, hierarchical granularity control, and an annealed control-forcing scheme. This enables the policy to internalize effective acquisition behavior during training and operate without external control at test time. Across seven benchmarks, DeepControl consistently outperforms strong RL and retrieval baselines without explicit information control; compared with Search-R1, it improves average performance by +9.4 and +8.6 points on Qwen2.5-7B and Qwen2.5-3B, respectively. Additional analyses show improved search effectiveness, training stability, and evidence utilization.
|
| 707 |
Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety
2602.11157
|
cs.CL
|
Max Zhang, Derek Liu, Kai Zhang, Joshua Franco, Haihao Liu |
Large language models (LLMs) are increasingly deployed worldwide, yet their safety alignment remains predominantly English-centric. This allows for vulnerabilities in non-English contexts, especially with low-resource languages. We introduce a novel applicatio...Large language models (LLMs) are increasingly deployed worldwide, yet their safety alignment remains predominantly English-centric. This allows for vulnerabilities in non-English contexts, especially with low-resource languages. We introduce a novel application of knowledge distillation (KD) in the context of multilingual jailbreak prevention, examining its efficacy. We distill the refusal behaviors of a proprietary teacher model (OpenAI o1-mini) with Low-Rank Adaptation (LoRA) into three open-source student models: Meta-Llama-3-8B-Instruct, Gemma-2-2B-IT, and Qwen3-8B, using ~28,000 multilingual jailbreak prompts from XSafety via black-box response-based, parameter-efficient fine-tuning (PEFT). Evaluation on the MultiJail benchmark reveals a counterintuitive behavior: standard fine-tuning on the teacher's ``safe'' refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points. Our experiments reveal a divergent generalization to unseen languages during distillation, with varying outcomes depending on the base model. By removing a primary source of safety degradation, nuanced `boundary' refusals, we mitigate or even reverse safety declines in student models, although reductions in reasoning performance (GSM8K) persist. Overall, our exploratory study highlights the challenges and potential of KD as a technique for multilingual safety alignment, offering a foundation for future research in this direction.
|
| 708 |
LLMs Exhibit Significantly Lower Uncertainty in Creative Writing Than Professional Writers
2602.16162
|
cs.CL
|
Peiqi Sui |
We argue that uncertainty is a key and understudied limitation of LLMs' performance in creative writing, which is often characterized as trite and clich\'e-ridden. Literary theory identifies uncertainty as a necessary condition for creative expression, while c...We argue that uncertainty is a key and understudied limitation of LLMs' performance in creative writing, which is often characterized as trite and clich\'e-ridden. Literary theory identifies uncertainty as a necessary condition for creative expression, while current alignment strategies steer models away from uncertain outputs to ensure factuality and reduce hallucination. We formalize this tension by quantifying the ``uncertainty gap'' between human-authored stories and model-generated continuations. Through a controlled information-theoretic analysis of 28 LLMs on high-quality storytelling datasets, we demonstrate that human writing consistently exhibits significantly higher uncertainty than model outputs. We find that instruction-tuned and reasoning models exacerbate this trend compared to their base counterparts; furthermore, the gap is more pronounced in creative writing than in functional domains, and shows a consistent correlation with writing quality. Achieving human-level creativity requires new uncertainty-aware alignment paradigms that can distinguish between destructive hallucinations and the constructive ambiguity required for literary richness.
|
| 709 |
EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
2603.08127
|
cs.CL
|
Yougang Lyu, Xi Zhang, Yuyue Zhao, Konstantinos Papakostas, Xinhao Yi |
The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including idea generation and experimental execution. However, most state-of...The increasing adoption of Large Language Models (LLMs) has enabled AI scientists to perform complex end-to-end scientific discovery tasks requiring coordination of specialized roles, including idea generation and experimental execution. However, most state-of-the-art AI scientist systems rely on static, hand-designed pipelines and fail to adapt based on accumulated interaction histories. As a result, these systems overlook promising research directions, repeat failed experiments, and pursue infeasible ideas. To address this, we introduce EvoScientist, an evolving multi-agent AI scientist framework that continuously improves research strategies through persistent memory and self-evolution. EvoScientist comprises three specialized agents: a Researcher Agent (RA) for scientific idea generation, an Engineer Agent (EA) for experiment implementation and execution, and an Evolution Manager Agent (EMA) that distills insights from prior interactions into reusable knowledge. EvoScientist contains two persistent memory modules: (i) an ideation memory, which summarizes feasible research directions from top-ranked ideas while recording previously unsuccessful directions; and (ii) an experimentation memory, which captures effective data processing and model training strategies derived from code search trajectories and best-performing implementations. These modules enable the RA and EA to retrieve relevant prior strategies, improving idea quality and code execution success rates over time. Experiments show that EvoScientist outperforms 7 open-source and commercial state-of-the-art systems in scientific idea generation, achieving higher novelty, feasibility, relevance, and clarity via automatic and human evaluation. EvoScientist also substantially improves code execution success rates through multi-agent evolution, demonstrating persistent memory's effectiveness for end-to-end scientific discovery.
|
| 710 |
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
2603.14006
|
cs.CL
|
Hang Gao, Dimitris N. Metaxas |
GraphRAG is increasingly adopted for converting unstructured corpora into graph structures to enable multi-hop reasoning. However, standard graph algorithms rely heavily on static connectivity and explicit edges, often failing in real-world scenarios where Kno...GraphRAG is increasingly adopted for converting unstructured corpora into graph structures to enable multi-hop reasoning. However, standard graph algorithms rely heavily on static connectivity and explicit edges, often failing in real-world scenarios where Knowledge Graphs (KGs) are noisy, sparse, or incomplete. To address this limitation, we introduce INSES (Intelligent Navigation and Similarity Enhanced Search), a dynamic framework designed to reason beyond explicit edges. INSES couples LLM-guided navigation, which prunes noise and steers exploration, with embedding-based similarity expansion to recover hidden links and bridge semantic gaps. Recognizing the computational cost of graph reasoning, we complement INSES with a lightweight router that delegates simple queries to Na\"ive RAG and escalates complex cases to INSES, balancing efficiency with reasoning depth. Experimental results show that INSES performs favorably compared to established RAG and GraphRAG baselines on multiple benchmarks. In particular, on the MINE benchmark, it exhibits notable robustness and adaptability across KGs constructed by varying methods. Our code and data are publicly available at https://github.com/hanggao-gh/INSES .
|
| 711 |
MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval?
2603.23533
|
cs.CLcs.LGcs.AI
|
Bhavik Mangla |
Markdown carries structure a parser reads for free: headers, section paths, and block boundaries. Many RAG pipelines also spend LLM calls per chunk on generated metadata. We ask what one LLM call per chunk buys over that free structure. MDKeyChunker splits Mar...Markdown carries structure a parser reads for free: headers, section paths, and block boundaries. Many RAG pipelines also spend LLM calls per chunk on generated metadata. We ask what one LLM call per chunk buys over that free structure. MDKeyChunker splits Markdown into header-led chunks without splitting any block; makes one LLM call per chunk for a title, summary, keywords, entities, questions, and a subtopic key, showing the model the keys already assigned in the document (a rolling key dictionary); and can merge same-key chunks. With qwen2.5:7b, we evaluate 79 Qasper questions over 30 papers and 73 FreshStack questions over 24 Laravel documentation files under BM25, two dense embedders, and hybrid fusion, following an analysis plan committed before results were computed. Evidence is matched only against source text, within a fixed token budget. Under hybrid retrieval, structural chunks beat 512-character windows on both datasets (Qasper +23.0 points, 95% CI [+12.8, +33.5]; FreshStack +5.1 [+1.4, +9.0]) and 256-token windows on Qasper (+12.7 [+5.3, +20.3]) but not on FreshStack (-2.6 [-6.4, +1.2]). Under the primary retrievers (hybrid, BM25), enrichment shows no planned-comparison difference from a free section-path prefix or from contextual retrieval; under hybrid retrieval the intervals exclude gains above 4.5 and 2.3 points on Qasper and 6.1 on FreshStack. Outside the planned comparisons, enrichment-style prefixes help BM25 on Qasper (exploratory) and mxbai on FreshStack (a secondary retriever). Rolling keys raise key reuse from 5.5% to 14.7%, but merging does not improve retrieval, and under BM25 on Qasper merging with rolling keys scores below merging without them (-6.0 [-13.1, -0.2]). Enrichment used about 1,000 input tokens per chunk; contextual retrieval 5,520 (Qasper) and 9,825 (FreshStack). The results of versions 1 and 2 are withdrawn.
|
| 712 |
$\pi^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
2604.05114
|
cs.CLcs.LGcs.AI
|
Quyet V. Do, Thinh Pham, Nguyen Nguyen, Sha Li, Pratibha Zunjare |
We study a QA curation pipeline for improving long-context complex reasoning in large language models (LLMs). Our approach, $\pi^2$, constructs high-quality reasoning data through rigorous QA curation: 1) extracting and expanding tables from Wikipedia, 2) from...We study a QA curation pipeline for improving long-context complex reasoning in large language models (LLMs). Our approach, $\pi^2$, constructs high-quality reasoning data through rigorous QA curation: 1) extracting and expanding tables from Wikipedia, 2) from the collected tables together with relevant metadata, generating complex reasoning questions whose answers are automatically determined and validated through dual-path code execution, 3) finally, back-translating chain-of-thoughts solutions grounded in realistic context. Supervised fine-tuning with gpt-oss-20b and Qwen3-4B-Instruct-2507 on $\pi^2$ yields consistent improvements across four long-context reasoning benchmarks and our alike $\pi^2$-Bench, with average absolute accuracy gains of +6.25% and +3.37% respectively. Through deeper analyses, we observe that reasoning style contributes little, while faithful reasoning patterns discovered by back translation and grounded realistic long context, as $\pi^2$ is designed for, are crucial for the improvement. Our code, data, and models are fully open-source at https://github.com/vtpss/pi-squared.
|
| 713 |
Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling
2604.09854
|
cs.CL
|
Peiqi Sui, Yutong Zhu, Tianyi Cheng, Peter West, Richard Jean So |
LLMs have so far failed both to generate consistently compelling stories and to recognize this failure--on the leading creative-writing benchmark (EQ-Bench), LLM judges rank zero-shot AI stories above New Yorker short stories, a gold standard for literary fict...LLMs have so far failed both to generate consistently compelling stories and to recognize this failure--on the leading creative-writing benchmark (EQ-Bench), LLM judges rank zero-shot AI stories above New Yorker short stories, a gold standard for literary fiction. We argue that existing rubrics overlook a key dimension of compelling human stories: narrative tension. We introduce the 100-Endings metric, which walks through a story sentence by sentence: at each position, a model predicts how the story will end 100 times given only the text so far, and we measure tension as how often predictions fail to match the ground truth. Beyond the mismatch rate, the sentence-level curve also yields complementary statistics that track plot-level twists and revelations, such as the inflection rate, a geometric measure of how frequently the curve reverses direction. Unlike rubric-based judges, 100-Endings correctly ranks New Yorker stories far above LLM outputs. Grounded in narratological principles, we design a story-generation pipeline using structural constraints, including analysis of story templates, idea formulation, and narrative scaffolding. Our pipeline significantly increases narrative tension as measured by the 100-Endings metric, while maintaining performance on the EQ-Bench leaderboard.
|
| 714 |
Think Multilingual, Not Harder: A Framework for Analyzing and Teaching Code-Switched Reasoning
2604.15490
|
cs.CL
|
Eleanor M. Lin, David Jurgens |
Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks. Interestingly, while reasoning models are often trained to generate monolingual text, these models have al...Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks. Interestingly, while reasoning models are often trained to generate monolingual text, these models have also been observed to code-switch (i.e., mix languages). Prior works have either viewed code-switching as an undesirable error, attempted to control code-switching through modifications to input prompts or the output decoding process, or focus on narrow subsets of languages, domains, tasks, and models. We address these gaps by introducing the first linguistically and behaviorally motivated fine-tuning framework for identifying beneficial code-switched reasoning behaviors in large language models and teaching these models to code-switch more effectively for reasoning. We create the Code-Switched Reasoning (CoRe) corpus, consisting of (1) 7k reasoning traces from 15 models, 18 languages, 10 scripts, and diverse reasoning domains, providing insights into potentially helpful code-switching behaviors, and (2) 40 carefully curated datasets for training and evaluating six interventions for improving code-switching in reasoning across three models and seven languages, totaling 120 fine-tuning conditions. Across 80k+ reasoning traces from both language/culture-agnostic and -specific evaluations, English-dominated reasoning, semantically accurate code-switching, and a more even mix of languages are positively associated with correct answers, whereas dense switching is associated with incorrect answers. Moreover, we are able to elicit positive behaviors through fine-tuning tasks that do not directly demonstrate code-switching. Our work suggests that small but well-curated datasets can change how reasoning models code-switch, allowing us to reap the benefits of reasoning in the many languages that lack large-scale reasoning data.
|
| 715 |
PSI-Bench: Interpretable and Clinically Meaningful Evaluation of Depression Patient Simulators
2604.25840
|
cs.CLcs.AI
|
Nguyen Khoi Hoang, Shuhaib Mehri, Tse-An Hsu, Yi-Jyun Sun, Quynh Xuan Nguyen Truong |
Patient simulators are gaining traction in mental health training by providing scalable exposure to complex and sensitive patient interactions. Simulating depressed patients is challenging, as safety constraints and high patient variability complicate simulati...Patient simulators are gaining traction in mental health training by providing scalable exposure to complex and sensitive patient interactions. Simulating depressed patients is challenging, as safety constraints and high patient variability complicate simulations and underscore the need for simulators that capture diverse and realistic patient behaviors. However, existing evaluations heavily rely on LLM-judges with poorly specified prompts and do not assess behavioral diversity. We introduce PSI-Bench, an automatic evaluation framework that provides interpretable, clinically meaningful diagnostics of depression patient simulator behavior across turn-, dialogue-, and population-level dimensions. Using PSI-Bench, we benchmark seven LLMs across two simulator frameworks and find that simulators produce overly long, lexically diverse responses, show reduced variability, and move through therapeutic stages and toward positive valence too quickly. We also show that the simulation framework has a larger impact on fidelity than the model scale. Results from a human study demonstrate that our benchmark is strongly aligned with judgments of mental health professionals. Our work reveals key limitations of current depression patient simulators and provides an interpretable, extensible benchmark to guide future simulator design and evaluation.
|
| 716 |
VIDA: A Dataset for Visually Dependent Ambiguity in Multimodal Machine Translation
2605.02035
|
cs.CLcs.AI
|
Jingheng Pan, Xintong Wang, Longyue Wang, Liang Ding, Weihua Luo |
Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended meaning. Although prior work has proposed disambiguation-oriented benchmarks probi...Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended meaning. Although prior work has proposed disambiguation-oriented benchmarks probing the role of vision, we observe that existing benchmarks remain limited by task-format mismatch, narrow ambiguity coverage, or insufficient visual-dependency validation. Moreover, existing ambiguity evaluations are not well suited to diverse ambiguity types in open-ended translation. To address these limitations, we present VIDA (Visually-Dependent Ambiguity), a dataset of 2,500 carefully curated instances in which resolving an annotated source span requires visual evidence. We further propose Disambiguation-Centric Metrics that use an LLM-as-a-judge classifier to verify whether annotated ambiguous expressions are resolved correctly at the span level. Evaluations with stronger recent LVLMs show that visual disambiguation remains challenging. Using chain-of-thought supervised fine-tuning as a diagnostic setting, we observe stronger out-of-distribution disambiguation than with SFT, with robust gains on collective-noun ambiguities and model-dependent gains on sentence-level ambiguities.
|
| 717 |
Improving Reasoning Ability via Asynchronous On-Policy Self-Distillation under Positive Rollouts
2605.06650
|
cs.CL
|
Mingwei Xu, Hao Fang |
Distillation and reinforcement learning through verifiable rewards (RLVR) have achieved progress in enhancing the reasoning ability of large language models (LLMs). However, we note that negative rollouts may admit no gradation of failure severity, and the com...Distillation and reinforcement learning through verifiable rewards (RLVR) have achieved progress in enhancing the reasoning ability of large language models (LLMs). However, we note that negative rollouts may admit no gradation of failure severity, and the combinatorial vastness makes penalizing a few sampled negatives unlikely to cover a meaningful reward signal under sparse binary rewards. In this work, we propose Positive-Only Policy Optimization (POPO), an on-policy self-distillation integrated RLVR framework in which learning occurs exclusively on online positive rollouts. Specifically, POPO utilizes bounded importance sampling over the positive rollout set. Thus, no disjoint negative rollouts are used for gradient guidance during post-training. We show that implicit negative gradients can emerge naturally through reinforcing the positive probability via rollout redistribution. Next, POPO stabilizes the policy optimization through self-distillation. First, it applies a Siamese policy network with a momentum-based adaptation law for asynchronous policy evolution. Second, we replace the KL-divergence with a bounded similarity penalty term in the Siamese representation space. We conduct extensive experiments using publicly available, well-established text-LLM models across all-level mathematical benchmarks (MATH-500, AMC23, AIME 2024/2025, and Olympiad). Our experiment demonstrates that POPO achieves superior performance compared to GRPO. Notably, we show that POPO can achieve 36.67% in AIME 2025 with Qwen-Math-7B, outperforming GRPO 30.00%. Our ablation and sweep studies further illustrate the necessity and robustness.
|
| 718 |
CktFormalizer: Autoformalization of Natural Language into Circuit Representations
2605.07782
|
cs.CL
|
Jing Xiong, Qi Han, Chenchen Ding, He Xiao, Zunhai Su |
Hardware infrastructure is a critical bottleneck for LLM-driven circuit design, limiting what agents can express, compile, and iteratively refine within an agentic loop. To address this bottleneck, we introduce CKTLEAN, a typed hardware infrastructure embedded...Hardware infrastructure is a critical bottleneck for LLM-driven circuit design, limiting what agents can express, compile, and iteratively refine within an agentic loop. To address this bottleneck, we introduce CKTLEAN, a typed hardware infrastructure embedded in Lean. It supports hardware description, compilation to SystemVerilog, and interactive type-checking and proof feedback through a persistent read-eval-print loop (REPL). On this foundation, we build CKTFORMALIZER, an agent framework for hardware generation, repair, optimization, and source-level equivalence proving. We evaluate structural correctness through compilation, functional correctness through RTL and gate-level simulation, and formal correctness through proofs relative to stated specifications. Across VerilogEval, RTLLM, ResBench, and CVDP, CKTFORMALIZER with CKTLEAN achieves compilation rates of 91.1%-99.4%. Among designs that pass RTL simulation, 95.4%-100.0% jointly complete synthesis and place-and-route and pass design-rule and layout-versus-schematic checks. In a separate evaluation on 30 VerilogEval problems, interactive proof-state feedback raises kernel-accepted equivalence proof completion from 53.3% to 63.3%. Hardware evaluation feedback also guides architecture exploration and iterative power, performance, and area (PPA) optimization, with synthesis-area reductions of up to 58.6% in the optimization loop. These results suggest that typed representations and explicit proof-state feedback help structure the agentic loop: compiler diagnostics guide targeted repairs, while proof-state feedback helps agents identify what remains to be proved and determine the next proof step. Project Page: https://ckt-formalizer.github.io/
|
| 719 |
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
2605.09635
|
cs.CL
|
Hao Liang, Qihan Lin, Mingrui Chen, Hengyi Feng, Zhaoyang Han |
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It...Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school. It contains nine node types and fourteen relation types covering curriculum structure and visual grounding. From this graph, we derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. We also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs. On K12-Bench, Gemini-3-Flash achieves only 57 percent exact match and Gemma-4-31B-IT reaches 46 percent, with Prereq and Neighbor being the hardest tasks. Our training experiments show that domain-specific supervision can reduce this gap. Under a matched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaokaoBench and EduEval. For vision-language models, K12-Train-Full achieves the best overall results on Gaokao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full DataFlow and WizardLM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary. We release the graph, benchmark, training data, and complete construction pipeline.
|
| 720 |
Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes
2605.09751
|
cs.CL
|
A. Bochkov |
We study whether a decoder-only language model requires an independently trainable input vector for every token. For a vocabulary of size $V$, an injective fixed-length binary identifier requires $K=\lceil\log_2 V\rceil$ bits. We replace the usual trainable $V...We study whether a decoder-only language model requires an independently trainable input vector for every token. For a vocabulary of size $V$, an injective fixed-length binary identifier requires $K=\lceil\log_2 V\rceil$ bits. We replace the usual trainable $V\times d_{\mathrm{model}}$ input table with fixed minimal binary token codes and a parameter-free tiled lift to model width. With $V=65{,}536$ and $d_{\mathrm{model}}=1024$, this supplies each token as a fixed 16-bit code and removes 67.1M trainable parameters, approximately 12.5% of the untied learned-input baseline. We also study a table-free implementation with one fixed invertible affine recoding over $\mathbb{F}_2^{16}$. Across three training seeds, 32-layer models trained on approximately 16-17B tokens obtain mean held-out perplexities of 2.44 for the learned-input baseline, 2.36 for canonical binary codes, and 2.39 for affine-recoded codes. These descriptive results do not establish statistical superiority or equivalence. Standardized evaluation of released base checkpoints with the LM Evaluation Harness adds commonsense, knowledge, and language-modeling benchmarks. The three paper checkpoints show broadly similar, task-dependent performance, while external SmolLM2 reference models are substantially stronger on many tasks. Our conclusion is therefore limited to the studied regime: a free trainable token-indexed input table is not required to learn nontrivial language modeling. The Transformer still learns continuous representations, and the output vocabulary projection remains standard and trainable.
|
| 721 |
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
2605.10065
|
cs.CLcs.AI
|
Hyundong Jin, Yo-Sub Han |
Controlling Large Language Models (LLMs) to prevent the generation of undesirable content, such as profanity and personally identifiable information (PII), has become increasingly critical. While earlier approaches relied on post-processing or resampling, rece...Controlling Large Language Models (LLMs) to prevent the generation of undesirable content, such as profanity and personally identifiable information (PII), has become increasingly critical. While earlier approaches relied on post-processing or resampling, recent research has shifted towards constrained decoding methods that control outputs during generation to mitigate high computational costs and quality degradation. However, preventing multiple forbidden hard constraints or regex constraints from appearing anywhere in the output is computationally challenging. A straightforward solution is to convert these constraints into a single automaton that tracks all forbidden patterns during decoding, but this often becomes impractically large. Standard regex engines also do not readily support the operations needed to build such a constraint, such as complement and intersection. In order to address these limitations, we propose NCO, a decoding strategy that performs online pattern matching over finite hard constraints and regex constraints, reducing computational overhead without inducing state explosion. NCO is fully compatible with standard inference strategies, including various sampling methods and beam search, while also supporting soft masking for probabilistic suppression. We empirically demonstrate its effectiveness across practical tasks, including PII and profanity suppression. Our implementation is available at https://github.com/hyundong98/NCO-Decoding .
|
| 722 |
When Can Digital Personas Reliably Approximate Human Survey Findings?
2605.10659
|
cs.CLcs.AI
|
Mumin Jia, Yilin Chen, Divya Sharma, Jairo Diaz-Rodriguez |
Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, const...Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.
|
| 723 |
Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter
2605.11685
|
cs.CL
|
Zeguan Xiao, Xuanzhe Xu, Yong Wang, Jian Yang, Yanqing HU |
Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns. However, recent studies reveal a critical vulnerability: unlearned models rapidl...Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns. However, recent studies reveal a critical vulnerability: unlearned models rapidly recover "forgotten" knowledge through relearning attacks. This fragility raises serious security concerns, especially for open-weight models. In this work, we investigate the fundamental mechanism underlying this fragility from a representation geometry perspective. We discover that existing unlearning methods predominantly optimize along dominant components, leaving minor components largely unchanged. Critically, during relearning attacks, the modifications in these dominant components are easily reversed, enabling rapid knowledge recovery, whereas minor components exhibit stronger resistance to such reversal. We further provide a theoretical analysis that explains both observations from the spectral structure of representations. Building on this insight, we propose Minor Component Unlearning (MCU), a novel unlearning approach that explicitly targets minor components in representations. Extensive experiments on three datasets validate that, by concentrating unlearning effects in these inherently robust directions, our method achieves substantially improved resistance to relearning attacks.
|
| 724 |
Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment
2605.15168
|
cs.CLcs.LGcs.AI
|
Sayantan Kumar, Shahriar Noroozizadeh, Juyong Kim, Jeremy C. Weiss |
Clinical language models increasingly operate over electronic health records (EHRs), yet patient records are not stored as temporally grounded trajectories. Clinical notes describe symptoms, assessments, and disease progression, but often compress or narrative...Clinical language models increasingly operate over electronic health records (EHRs), yet patient records are not stored as temporally grounded trajectories. Clinical notes describe symptoms, assessments, and disease progression, but often compress or narratively reorder events. Structured EHR rows provide timestamps for labs, medications, vitals, and procedures, but capture only part of the clinical story. We formulate clinical timeline reconstruction as retrieval-augmented temporal grounding: constructing a patient trajectory by using narrative text for event semantics and structured rows as partial temporal evidence. We introduce a scaffolded workflow that extracts central narrative events, builds an initial temporal scaffold, attaches non-central events, and calibrates timestamps using retrieved structured EHR rows. We evaluate on 40 discharge summaries, including 15 i2b2-derived and 25 MIMIC-IV summaries, each with manual gold-standard timelines and aligned structured EHR data. Across models, multimodal calibration left event match rates largely unchanged and generally improved temporal performance: mean paired case-level multimodal-unimodal differences were positive in 7 of 12 model-metric comparisons across concordance and AULTC, with none negative. However, uncertainty was substantial given the 40-case sample; paired case-level bootstrap intervals excluded zero only for the DeepSeek V3.2 AULTC improvement. A gap analysis shows that 35.1% of text-derived events have no structured counterpart. These findings support treating structured EHR data as partial temporal evidence for narrative-derived patient trajectories.
|
| 725 |
Single-Round Vector RAG vs an LLM-Compiled Wiki: A Preregistered Comparison on a Small Multi-Domain Research Corpus
2605.18490
|
cs.CL
|
Theodore O. Cochran |
We preregistered a comparison of two ways to help an LLM answer questions over a small research corpus: single-round Vector RAG and an LLM-compiled markdown wiki browsed by a tool-using agent. Both answered the same 13 questions over 24 papers with the same an...We preregistered a comparison of two ways to help an LLM answer questions over a small research corpus: single-round Vector RAG and an LLM-compiled markdown wiki browsed by a tool-using agent. Both answered the same 13 questions over 24 papers with the same answer model, scored by two blinded LLM judges. The three preregistered predictions came out one weakly supported, one supported, and one refuted. The wiki scored much better at connecting findings across papers, but its organization advantage fell below the registered threshold once both judges were combined. RAG met the registered test on single-fact groundedness, though the result was judge-sensitive. The wiki was cheaper to build but spent about 21 times more LLM tokens per query, so no break-even point exists. Exploratory analyses bear on why such comparisons disagree. A decomposition-retrieval variant of RAG reduced most of the wiki's synthesis-score gap at lower token cost. The judges' rank agreement was near zero on holistic groundedness (rho = 0.04) against rho = 0.81 on the most concretely defined criterion. A post-hoc claim-level analysis of citation support was checked against two human annotators on 100 claims. Its scorer agreed with them on 50 to 54% of claims as first run and on 65 to 69% once a truncation error was corrected, short of the rule fixed in advance. On the annotators' labels, the analysis did not establish a citation-support advantage for the wiki. On one model and one small corpus, which system appears to win depends on the retrieval baseline, the scoring method, and the judge, so evaluations should report synthesis, citation support, and cost separately and check automated grounding scorers against human labels.
|
| 726 |
Less Back-and-Forth: A Comparative Study of Structured Prompting
2605.20149
|
cs.CLcs.AI
|
Saurav Ghosh |
Large language models (LLMs) are widely used for open-ended tasks, but underspecified prompts can lead to low-quality answers and additional interaction. This paper studies whether structured prompt design improves response quality while reducing user effort. ...Large language models (LLMs) are widely used for open-ended tasks, but underspecified prompts can lead to low-quality answers and additional interaction. This paper studies whether structured prompt design improves response quality while reducing user effort. We compare three prompt conditions: a raw prompt, a checklist-improved prompt, and a clarifying-question prompt. We evaluate these conditions across four task types--summarization, planning, explanation, and coding--using three LLM systems: ChatGPT, Claude, and Grok. Each output is scored with a unified rubric covering task completion, correctness, compliance, and clarity. Checklist-improved prompts achieved the highest mean rubric score, 7.50 out of 8, compared with 5.67 for raw prompts and 6.67 for clarifying-question prompts. Checklist prompts also produced the best quality-effort tradeoff, using fewer average tokens than both raw and clarifying prompts. These results suggest that a simple prompt checklist can improve LLM responses while reducing unnecessary interaction.
|
| 727 |
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
2605.21318
|
cs.CLcs.LGcs.AI
|
Lucheng Fu, Ye Yu, Yiyang Wang, Yiqiao Jin, Haibo Jin |
Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often becom...Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often become longer, accumulate narrow sample-specific rules, and generalize poorly beyond the training distribution. We study this failure mode as prompt distributional overfitting and argue that it reflects a lack of representation control in discrete text-space optimization. We formalize this view through representational inefficiency, a dual-factor measure that decomposes prompt inefficiency into capacity cost and scope narrowness, attributing distributional prompt overfitting to their coupled growth during optimization. We propose TextReg, a regularization framework that realizes a soft-penalty objective through regularized textual gradients, combining Dual-Evidence Gradient Purification, Semantic Edit Regularization, and Regularization-Guided Prompt Update. Across multiple reasoning benchmarks, TextReg substantially improves out-of-distribution (OOD) generalization, with accuracy gains of up to +11.8% over TextGrad and +16.5% over REVOLVE.
|
| 728 |
EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation
2605.23954
|
cs.CLcs.AIcs.SD
|
Kaiwen Luo, Chunxi Luo, Liang Lin, Yuxuan Li, Zhenhong Zhou |
Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged informa...Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples candidate responses reflecting its inference-time behavior, while a frozen copy of the same backbone processes the corresponding clean audio. EchoDistill combines masked response-token distillation, task-gated consistency shaping, and teacher-referenced group-relative optimization to align noisy-input generation with clean-conditioned semantics. Only the student is retained at inference time, introducing no additional inference cost. Across three LALM backbones and three audio domains at -10dB, EchoDistill improves average noisy-input accuracy by 1.63 percentage points over the strongest baseline. On Qwen2.5-Omni, it raises noisy-input accuracy from 59.33% to 62.94%, while clean-audio accuracy increases from 76.56% to 77.56%. Replacing matched audio with random, shuffled, or silent inputs reduces accuracy by 3.08-6.42 points, confirming that matched acoustic evidence contributes to its predictions. Additional evaluations show improvements on held-out additive noises and external benchmarks, while revealing that these gains do not reliably extend to non-additive distortions. These results demonstrate robust post-training improvements under severe additive noise without sacrificing clean-audio capability across diverse tasks.
|
| 729 |
Word-Class and Construction-Like Structure Emerges in Neural Successor Representations Trained on Natural Language
2605.24585
|
cs.CL
|
Mathis Immertreu, Achim Schilling, Thomas Kinfe, Patrick Krauss |
Neural language models are typically trained on next-token prediction, although linguistic structure spans multiple temporal scales. Successor representations (SRs) make this horizon explicit by encoding discounted distributions over future states. Here, we as...Neural language models are typically trained on next-token prediction, although linguistic structure spans multiple temporal scales. Successor representations (SRs) make this horizon explicit by encoding discounted distributions over future states. Here, we ask whether such predictive representations can recover not only word classes, but also finer functional and construction-like structure from natural language. A residual network trained on WikiText-103 predicts SR distributions at three horizons without part-of-speech supervision. At the shortest horizon, unsupervised clustering robustly recovers nouns, verbs, and adjectives, while directed inter-cluster transitions reproduce familiar syntactic asymmetries. At finer resolutions and across 13 part-of-speech categories, the same geometry reveals semantic-functional groupings that cross category boundaries and directed relations tracing candidate date, measurement, and title-name constructions. Part-of-speech agreement declines as the predictive horizon lengthens. These results suggest that word classes are coarse regions within a richer predictive geometry in which categorical and construction-like linguistic structure emerge from future-word distributions.
|
| 730 |
Measuring the Depth of LLM Unlearning via Activation Patching
2605.24614
|
cs.CLcs.LGcs.AI
|
Jaeung Lee, Dohyun Kim, Jaemin Jo |
Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether target knowledge is truly erased remains challenging. Existing output-level metrics fail to detect when this knowledge ...Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether target knowledge is truly erased remains challenging. Existing output-level metrics fail to detect when this knowledge remains recoverable from internal representations. Recent white-box studies reveal such residual knowledge but often rely on auxiliary training or dataset-specific adaptations, leaving no generalizable metric. We close this gap with the Unlearning Depth Score (UDS), a metric that quantifies the mechanistic depth of unlearning via activation patching. UDS first identifies layers that encode the target knowledge using a retain model baseline, then measures how much of it is erased in the unlearned model on a 0-1 scale. In a meta-evaluation across 20 metrics on 150 unlearned models spanning 8 methods, UDS achieves the highest faithfulness and robustness, confirming our causal approach as the most reliable for unlearning evaluation. Case studies further show that UDS uncovers residual knowledge obscured from observational metrics by representational shifts, with erasure depth varying across prompt types. We provide guidelines for integrating UDS into existing benchmarking frameworks and streamlining the evaluation pipeline. Code and data are available at https://github.com/gnueaj/unlearning-depth-score.
|
| 731 |
ROC Analysis for Evaluating Translation Quality Estimation Systems
2605.24721
|
cs.CL
|
Evelyn Y. Garland (Acta Language Services, LLC), Carola F. Berger (CFB Scientific Translations LLC) |
The increasing use of automated translation quality estimation (QE) systems calls for practical, decision-oriented methods for evaluating their performance. We propose that Receiver Operating Characteristic (ROC) analysis is a useful approach for this purpose....The increasing use of automated translation quality estimation (QE) systems calls for practical, decision-oriented methods for evaluating their performance. We propose that Receiver Operating Characteristic (ROC) analysis is a useful approach for this purpose. Our study shows that ROC analysis not only produces results consistent with currently prevalent methods, but also offers several important advantages, including actionable performance insights that support business decision-making.
|
| 732 |
A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis
2605.25502
|
cs.CLcs.AI
|
Yehudit Aperstein, Alexander Apartsin |
Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because educational reviews are private, institution-specific, and expensive to annotate. This study introduces a contr...Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because educational reviews are private, institution-specific, and expensive to annotate. This study introduces a controlled synthetic benchmark for educational ABSA built from 10,000 synthetic course reviews with explicit train-validation-test splits and a 20-aspect pedagogical schema spanning instructional quality, assessment and course management, learning demand, learning environment, and engagement. The corpus is generated with sampled target labels, sampled nuance attributes, and a realism-tuned prompt refined through a three-cycle judge-editor procedure. On the resulting benchmark, local baselines with TF-IDF, two-step transformers, and joint encoders show that the task is nontrivial; the strongest untuned model, BERT, reaches a held-out detection micro-F1 of 0.2760, while a modest lower-rate BERT schedule improves this to 0.2930. Full-test GPT-based inference with gpt-5.2 reaches 0.2519 micro-F1 in zero-shot mode and 0.2501 with retrieval-based few-shot prompting, placing batch inference above the classical baseline and close to the compact joint encoders. A conservative external evaluation on 2,829 mapped student-feedback reviews from Herath et al. yields a micro-F1 of 0.4593 for BERT on a 9-aspect overlap, indicating partial synthetic-to-real transfer. Realism and faithfulness analyses are reported as generator diagnostics that clarify how the benchmark was stabilized and where label noise remains. The study therefore contributes a synthetic educational ABSA corpus, a documented generation procedure, and a reproducible benchmark setting for a domain in which public labeled data remain difficult to obtain.
|
| 733 |
Beyond Phones: Structured Phonemic Modeling for Vietnamese Automatic Speech Recognition
2605.27874
|
cs.CL
|
Nghia Hieu Nguyen, Quan Ngoc Hoang, Long Hoang Huu Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen |
Phone-based representations provide a compact and acoustically grounded alternative to conventional orthographic modeling for automatic speech recognition (ASR). However, phones describe surface pronunciations and may lose lexical distinctions under dialect-de...Phone-based representations provide a compact and acoustically grounded alternative to conventional orthographic modeling for automatic speech recognition (ASR). However, phones describe surface pronunciations and may lose lexical distinctions under dialect-dependent sound mergers, making their conversion back to orthographic text inherently ambiguous. This issue is particularly relevant to Vietnamese, where pronunciation varies considerably across regional dialects. This work proposes a structured phonemic approach to Vietnamese ASR that moves the output representation from surface phones to abstract phonemes. Exploiting the regular phoneme-grapheme correspondence of Vietnamese, each syllable is represented by a phonemic triplet consisting of its initial, rhyme, and tone, preserving lexical distinctions while enabling deterministic reconstruction of orthographic text. We further introduce a \textbf{Phonemic Syllabic-Structure Decoder} that captures the hierarchical organization of Vietnamese syllables by first predicting the rhyme and subsequently conditioning the initial and tone predictions on the rhyme. Experiments on the standard LSVSC and multi-dialect UIT-ViMD benchmarks demonstrate the effectiveness of the proposed approach. The best models achieve WERs of 5.83\% on LSVSC and 12.58\% on UIT-ViMD, outperforming orthographic, phonetic, and previous phonemic approaches. Further analyses reveal broader lexical coverage, reduced dependence on word-frequency patterns, and consistent behavior across Vietnamese dialects. These results demonstrate the effectiveness of moving from phonetic to structured phonemic modeling and highlight the importance of incorporating language-specific phonological structure into end-to-end ASR.
|
| 734 |
EUDAIMONIA: Evaluating Undesirable Dynamics in AI
2605.30654
|
cs.CLcs.AI
|
Jun Rui Huang, Wang Bill Zhu, Ziyi Liu, Nathanael Fast, Ravi Iyer |
Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but the social dynamics of these interactions can create harms that are not captured by current evaluations in real...Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but the social dynamics of these interactions can create harms that are not captured by current evaluations in real-world settings. We introduce the Social AI Design Code, a set of principles for keeping LLMs from encouraging harmful intimacy, dependence, or prolonged engagement, which we developed with input from psychologists and trust-and-safety practitioners. To evaluate these risks in natural and diverse user-LLM interactions, we operationalize the code with EUDAIMONIA, a benchmark of 969 opening-turn queries and 3,147 design-requirement violation checks built from WildChat through weak-to-strong filtration, multi-model relabeling, and controlled rewriting. Evaluating 26 recent LLMs, we find rapid but uneven progress: within 19 months, Grok shifted from the highest to the lowest violation rate, with Grok-4.7 violating only 13.8% of checks. Although Claude-Opus-5.5 and GPT-6-Astra outperform Grok-4.7 on general benchmarks, they violate more checks, showing that greater capability does not guarantee adherence to our design code. Extended thinking does not reduce violation rates, suggesting these failures are social-alignment problems rather than deficits solvable through test-time reasoning alone. Finally, we evaluate models on 229 multi-turn WildChat continuations, violation rates increase, but cross-model trends mirror those in opening turns, showing that our findings generalize to multi-turn interactions.
|
| 735 |
ElasticMem: Latent Memory as a Learnable Resource for LLM Agents
2605.30690
|
cs.CL
|
Tao Feng, Chongrui Ye, Fangxu Yu, Tianyang Luo, Jingjun Xu |
Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. However, existing memory-augmented methods typically treat memory as a fixed resource: text-space approaches conca...Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. However, existing memory-augmented methods typically treat memory as a fixed resource: text-space approaches concatenate retrieved memories into the context window, causing substantial token overhead and sensitivity to noisy evidence, while latent-space approaches reduce textual cost but still rely on rigid retrieval or fixed-capacity memory interfaces. This creates a mismatch between query-dependent memory utility and fixed memory allocation. We propose ElasticMem, a memory-augmented LLM framework that learns to use memory as an elastic latent resource. ElasticMem builds an offline latent memory bank with retrieval keys and content caches, retrieves memories adaptively from the reasoner's hidden state, assigns each retrieved memory a variable latent budget through a learned policy, and injects selected latent states as soft memory tokens for generation. The full memory-use process is optimized with downstream task rewards through group-relative policy optimization. We evaluate ElasticMem on MemorySuite, covering memory-intensive QA and embodied agent control. Across Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct backbones, ElasticMem improves weighted-average QA accuracy by 26.2% and 24.6% and ALFWorld success rate by 66.3% and 27.2% relative to the strongest baselines, while consuming the fewest tokens on ALFWorld on average. Ablations and qualitative analyses further show that adaptive retrieval and elastic budget allocation help ElasticMem prioritize useful evidence and transferable plans beyond rigid cosine similarity.
|
| 736 |
ExpHarness: Model-Agnostic Experience Learning through a Trainable Harness
2605.30712
|
cs.CL
|
Tao Feng, Chongrui Ye, Fangxu Yu, Tianyang Luo, Jingjun Xu |
Large language model (LLM) agents increasingly operate within a harness, the scaffolding that determines what enters the executor's context, yet the experience they accumulate across tasks rarely flows back into this harness. Existing approaches include execut...Large language model (LLM) agents increasingly operate within a harness, the scaffolding that determines what enters the executor's context, yet the experience they accumulate across tasks rarely flows back into this harness. Existing approaches include executor fine-tuning and external memory retrieval, but combining task-adaptive retrieval with experience reuse across frozen executors remains challenging. To this end, we introduce ExpHarness, a learnable experience harness that improves frozen and replaceable LLM executors without modifying their parameters. Specifically, ExpHarness distills trajectories into reusable skills and failure lessons within a self-evolving experience graph, and trains a lightweight retrieval copilot that decides, per task, how broadly to explore the graph and how strongly to favor historically useful experiences over merely similar ones. The copilot is optimized with reinforcement learning from a utility-grounded reward combining the with/without-experience score difference and an absolute-performance term; the same reward updates the graph during training. Extensive experiments on ExpSuite, spanning 10 static benchmarks and 2 agentic environments, show that ExpHarness improves over the strongest baseline by 12.1% and 4.5% on static tasks and by 21.4% and 12.7% on agentic tasks with the smaller and larger executors, respectively, while reducing interaction steps by up to 21.6%. Transfer experiments further examine reuse of the learned harness across executors of different scales and reasoning capabilities, with joint graph and copilot transfer performing closest to target-specific training among the evaluated transfer variants.
|
| 737 |
EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models
2606.00722
|
cs.CLcs.AI
|
Hyundong Jin, Yo-Sub Han |
Controlling language model outputs is essential for ensuring structural validity, reliability, and downstream usability, and diffusion language models are no exception. Recent advances in diffusion language model decoding have extended output control beyond re...Controlling language model outputs is essential for ensuring structural validity, reliability, and downstream usability, and diffusion language models are no exception. Recent advances in diffusion language model decoding have extended output control beyond regular constraints to context-free grammar (CFG) constraints. However, the prior CFG-constrained decoder can be up to four times slower than unconstrained decoding, forcing a trade-off between the correctness benefits of CFG constraints and the decoding efficiency of diffusion models. A key source of this overhead is sequential validity checking, which limits parallel token commitment and adds repeated validation costs. We propose EPIC, a CFG-constrained decoding framework designed to improve this efficiency-correctness trade-off. In order to reduce decoding overhead without sacrificing syntactic correctness, EPIC combines lexing memoization, relaxed compatible subset selection for parallel commit, and validation using Earley-style parsing instead of deterministic automata. This design enables multiple compatible tokens to be committed together while avoiding repeated lexing and expensive validation. Experiments on three benchmarks using four models show that EPIC improves the efficiency-correctness trade-off, bringing runtime close to unconstrained levels while maintaining comparable syntactic and functional correctness. Relative to the prior CFG-constrained decoder, EPIC achieves a best-case inference-time reduction of 67.2%. Our implementation is available at https://github.com/hyundong98/EPIC-Decoding .
|
| 738 |
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
2606.01196
|
cs.CLcs.AI
|
Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto |
Language models often answer harmful requests in low-resource languages (LRLs) that they refuse in high-resource languages (HRLs). Across three instruction-tuned models and 23 languages, harmful refusal falls from 87.9% in HRLs to 43.9% in LRLs, while harmless...Language models often answer harmful requests in low-resource languages (LRLs) that they refuse in high-resource languages (HRLs). Across three instruction-tuned models and 23 languages, harmful refusal falls from 87.9% in HRLs to 43.9% in LRLs, while harmless refusal remains low. A common explanation is that models represent harmfulness weakly in LRLs. We test whether harmfulness is instead represented but does not reliably produce refusal. Across three models, a harmfulness direction learned from HRL activations still separates harmful from harmless LRL prompts, showing that complete absence of harmfulness information cannot explain many failures. However, harmfulness scores shift downward for LRL prompts, making harmful prompts less likely to reach the range associated with refusal. Motivated by this shift, we train a low-rank logistic classifier on HRL activations and calibrate its threshold with a few target-language examples. During generation, the classifier conditionally adds or ablates the HRL harmfulness direction. With the same HRL data and 32 target-language examples per class, CAST remains limited by low harmful refusal and AdaSteer by high harmless refusal, yielding mean refusal selectivity ($\Delta$ = harmful - harmless refusal) of 33.6 and 6.8, respectively. Our intervention reaches 54.5 while preserving MMLU utility. HRL-only calibration improves selectivity for Qwen and Gemma, whereas Llama benefits from target-language calibration. These results show that recalibrating existing representations can offer a training-free method for repairing low-resource safety failures.
|
| 739 |
Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling
2606.02004
|
cs.CLcs.LG
|
Vladimir Beskorovainyi |
Price statistics increasingly draw on scanner, web-scraped and receipt data, whose product descriptions are short, noisy and carry no standard product code, so each item must be coded to a consumption classification such as COICOP. National statistical offices...Price statistics increasingly draw on scanner, web-scraped and receipt data, whose product descriptions are short, noisy and carry no standard product code, so each item must be coded to a consumption classification such as COICOP. National statistical offices already report that lightweight text classifiers are adequate for this task, but the published evidence is thin on dispersion, paired comparison, train-test overlap and measured computational cost. This paper supplies that evaluation. On an openly released synthetic benchmark of six COICOP-like categories, seven models are trained under one matched protocol and compared on accuracy and cost. A character n-gram logistic regression is the most accurate model in every category (mean F1 = 0.997, and 0.996 on test items unseen in training). The small CNN and LSTM examined here never win a comparison, but once trained to the same stopping criterion they are not reliably worse than the word-level models either; what separates them is cost, at 65 and 77 times the training time of the cheapest model. The most accurate model is not the cheapest: character features cut inference throughput to 44,900 items per second against 212,700 for unigram bag-of-words. A rule-based prefix-tree stage admits 75-86% of positive items, so it bounds the cascade's recall. A Monte Carlo study of the labeling protocol, in which annotators are simulated and no human annotation was collected, shows that an additive reliability weight barely improves on majority vote while Dawid-Skene aggregation recovers labels markedly better, and that the difference carries through to classifiers trained on those labels. The benchmark is synthetic because the production data behind the architecture are confidential; its generator applies character-level corruption and shares phrase sets with the rule stage, and both conditions are stated wherever a result depends on them.
|
| 740 |
See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making
2606.03371
|
cs.CL
|
Honghui Zhang, Anna Min, Chenmeinian Guo, Yujia Zhang, Yichen Yu |
Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-person video: before an explicit customer request, an agent must use limited human-object interaction evidence to...Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-person video: before an explicit customer request, an agent must use limited human-object interaction evidence to intervene or remain silent. Physical grounding here means converting observations into task-relevant retail state, not modeling low-level dynamics. We introduce the Proactive Intent World Model (PIWM): See constructs the perceptual basis, Foresee models counterfactual consequences, and Act selects an action. Performance is poor when the agent must extract information from raw video and decide directly, but improves substantially with structured inputs extracted and annotated from a professional retail perspective. AIDA-stage constraints and BDI-state ablations further support role- and goal-directed selection and organization of decision-relevant cues. Counterfactual prediction performs well in standalone evaluation, yet planning methods that query these forecasts at inference time degrade sharply: locally useful consequence prediction does not reliably improve action selection. This gap may reflect incomplete process understanding, uncertainty in fine-grained single-step outcomes, and insufficient joint modeling of scenes and temporal evolution. Hold remains the hardest action in structured-state evaluation, exposing a related challenge in temporal awareness. PIWM advances static intent recognition toward intent world modeling by organizing observations under task knowledge, anticipating candidate interventions, and treating intervention and non-intervention jointly. Future work will introduce long-horizon interaction trajectories and temporal consequence supervision to improve sustained reasoning and intervention timing.
|
| 741 |
SMADE-IE: Sparse Multi-Agent Framework with Evidence-Driven Debate for Zero-Shot Information Extraction
2606.04691
|
cs.CL
|
Kenfeng Huang, Yi Cai, Xin Wu, Zikun Deng, Li Yuan |
Zero-shot information extraction (IE) with large language models (LLMs) enables adaptation to new schemas and domains without task-specific training. Existing methods mainly follow three paradigms. Monolithic prompting is efficient but prone to missed mentions...Zero-shot information extraction (IE) with large language models (LLMs) enables adaptation to new schemas and domains without task-specific training. Existing methods mainly follow three paradigms. Monolithic prompting is efficient but prone to missed mentions, boundary errors, and type confusion. Each-type prompting improves type-level focus but may produce overlapping or conflicting predictions, while multi-agent debate can resolve such conflicts at the cost of irrelevant context, redundant interactions, and high token overhead. To address these issues, we propose SMADE-IE, a sparse and evidence-driven multi-agent framework. An Adaptive Mode Selector routes simple inputs to a lightweight Global Extraction Mode and ambiguous inputs to a Type-Centric Extraction Mode based on sample complexity and relevant types. Cross-type conflicts are resolved by an Evidence-Driven Debate module that uses Toulmin-style arguments, external evidence scoring, Beta-based confidence updates, and early stopping. Experiments on nine benchmarks covering NER, RE, and JERE show that SMADE-IE improves average Partial F1 over the strongest baselines by 11.37, 3.83, and 14.46 points, respectively. Compared with the multi-agent baseline CrossAgentIE, SMADE-IE reduces token consumption by 85.1% on DocRED and 80.1% on CrossRE, demonstrating substantially higher inference efficiency. Code is available at https://github.com/Cppys/SMADE-IE.
|
| 742 |
Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents
2606.04874
|
cs.CL
|
Haoyu Sun, Wenxuan Wang, Mingyang Song, Jujie He, Weinan Zhang |
Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent evaluations often report only end-to-end success, making it difficult to determine ...Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent evaluations often report only end-to-end success, making it difficult to determine whether failures stem from planning or execution. We introduce Agent Planning Benchmark (APB), a planning-specific diagnostic benchmark with 4,209 multimodal cases across 22 domains and five settings, covering holistic planning, feedback-conditioned step-wise planning, and robustness under extraneous tools, broken tools, and unsolvable tasks. Across 12 MLLMs, APB reveals systematic weaknesses in long-horizon planning, tool-noise robustness, calibrated refusal, and inference-time refinement. We further validate APB on 200 ToolSandbox tasks and 200 $\tau^2$-bench tasks, where APB-guided refinement consistently improves plan correctness, plan grade, and downstream execution metrics across three representative models. APB thus serves as an upstream diagnostic complement to execution benchmarks. The APB benchmark and code are available in \href{https://github.com/Mikivishy/AgentPlanningBenchmark}{this URL}.
|
| 743 |
Function-Vector Heads Are Two Populations: Writers and Cancellers in In-Context Learning
2606.07560
|
cs.CLcs.LG
|
Han-yu Wang |
In-context learning lets a language model perform a task specified by examples in its prompt. Function vectors capture task information in a compact activation assembled from attention-head outputs. Across two rule families and three Pythia models, we find two...In-context learning lets a language model perform a task specified by examples in its prompt. Function vectors capture task information in a compact activation assembled from attention-head outputs. Across two rule families and three Pythia models, we find two opposed functional populations among candidate function-vector heads. Writers support the rule-correct answer, while cancellers systematically suppress it. These roles predict opposing group effects on held-out prompts, where removing cancellers improves accuracy by 2.4 to 6.8 percentage points in all six settings. The dominant suppressive contributions depend on task content, and the same components can support correct predictions in another task. At the population level, substantial support and suppression can nearly cancel. Separating these roles explains how task-related computation can work against correct predictions and how a small aggregate effect can conceal strong opposing contributions.
|
| 744 |
The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models
2606.13993
|
cs.CL
|
Zachary Nicholas Houghton, Yu Zhou, Dan Pluth, Jordan Hosier, Vijay K. Gurbani |
One of the most central aspects of language processing is the ability to trade off between stored representations and abstract knowledge: one must retrieve stored representations, but also generate novel ones by applying productive rules. While recent work has...One of the most central aspects of language processing is the ability to trade off between stored representations and abstract knowledge: one must retrieve stored representations, but also generate novel ones by applying productive rules. While recent work has examined abstract knowledge in language models, holistic storage has received far less attention. We probe internal representations in both text-based LLMs and an ASR model, testing whether V+up phrasal verbs develop distinct representations as a function of frequency and predictability. All models show evidence of holistic storage driven by frequency and predictability, further supporting usage-based theories of language.
|
| 745 |
Retrospective Progress-Aware Self-Refinement for LLM Agent Training
2606.14302
|
cs.CL
|
Xinbei Ma, Congmin Zheng, Jiyang Qiu, Jiale Hong, Yao Yao |
Long-horizon LLM-based agents receive rich environmental observations during interaction, yet outcome rewards provide limited explicit supervision about how individual actions advance task completion. We investigate whether agents can turn this interaction evi...Long-horizon LLM-based agents receive rich environmental observations during interaction, yet outcome rewards provide limited explicit supervision about how individual actions advance task completion. We investigate whether agents can turn this interaction evidence into useful training signals through retrospective progress assessment. A WebShop pilot shows that direct progress prompting reduces task success, whereas hindsight-annotated demonstrations improve it. We introduce RePro, Retrospective Progress-Aware Training, with a forward-then-reflect rollout: the agent estimates progress while acting, then reassesses each step using the completed trajectory and outcome. After warmup with externally generated demonstrations, policy optimization combines self-generated progress differences, online-retrospective alignment, and format rewards with environment feedback, requiring neither a separate process reward model nor ongoing teacher annotation. Experiments on WebShop, ALFWorld, and Sokoban show that RePro enhances the Qwen family's performance, with up to 11.57% success rate gains.
|
| 746 |
REFLEX: Reflective Evolution from LLM Experience
2606.16496
|
cs.CLcs.LG
|
Pan Wang |
Large multimodal language models (MLLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies. In existing program-evolution systems, however, reusable knowledge is usually carried by whole programs in the p...Large multimodal language models (MLLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies. In existing program-evolution systems, however, reusable knowledge is usually carried by whole programs in the population, and it is difficult to trace how a visual observation led to a particular code change and its measured outcome. We present REFLEX, a train-free evolutionary framework that links these steps in one loop. A vision-enabled Critic turns task-specific behavioral evidence into a structured diagnosis; the diagnosis retrieves executable code snippets from a persistent Skill Memory; a text-only Actor writes the child program; and the child--parent fitness change updates the utility of each retrieved snippet. Every step is recorded in a single trace. Under matched backends and 100-call budgets over 10 paired seeds, REFLEX reaches the solve threshold in a median of 13, 20, and 24 LLM calls on Acrobot, Pendulum, and Lunar Lander, roughly half the calls required by official MLES and by a compute-matched Actor-only ablation. Frozen Skill Memory banks from Pendulum or Lunar Lander raise the final Acrobot score on all 10 paired seeds. On a 36-dimensional antenna-array design task with equal evaluation budgets and shared initialization, REFLEX reaches the harder $25.25$ score threshold on 9/10 seeds, compared with at most 3/10 for CMA-ES, GA, and PSO.
|
| 747 |
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
2606.18394
|
cs.CL
|
Lanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan, Yujie Zhao |
Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and dr...Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed only when acceptance remains high and drafting overhead stays low. This ceiling has been difficult to break because prior head-based SD methods face a causality-efficiency dilemma. Autoregressive drafters produce path-conditioned candidates that are effective for tree speculative decoding with higher acceptance length, but their drafting cost grows with tree depth. Bidirectional block-diffusion drafters generate all positions in one pass, but their branch-agnostic marginals can form individually plausible yet mutually inconsistent trees, wasting budget and reducing acceptance. We propose JetSpec, a head-based SD framework that combines one-forward drafting efficiency with branch-wise causal conditioning. JetSpec trains a causal parallel draft head over fused hidden states from the frozen target model, producing candidate trees whose scores align with the target model's autoregressive factorization. This enables JetSpec to convert larger draft budgets into longer accepted prefixes and higher end-to-end speedup. Across math, coding, and chat benchmarks on dense and MoE Qwen3 models, JetSpec consistently outperforms bidirectional-head and tree-based SD baselines. On H100 GPUs, JetSpec achieves up to 9.64x speedup on MATH-500 and 4.58x on open-ended conversational workloads, with further latency gains demonstrated through vLLM integration under realistic serving loads. Our code and models are available at https://github.com/hao-ai-lab/JetSpec.
|
| 748 |
Capability Provenance in Language Models: A Case Study in Social Reasoning
2606.19625
|
cs.CLcs.LG
|
Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla |
We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social reasoning versus STEM reasoning in OLMo3-7B. Training-data attribution measures how strongly each training docume...We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social reasoning versus STEM reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer's 24-format x 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2x2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than at the knowledge level. Targeted machine unlearning provides partial causal validation: forgetting high-attribution topics (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-topic random baselines. We release the code and aggregate artifacts at https://github.com/HCAI-Lab-GT/capabilibara and https://huggingface.co/HCAI-Lab-GT.
|
| 749 |
Validation-Gated Causal Interventions for Interpreting High-Stakes Large Language Model Behavior: A Case Study in Suicidality Detection
2606.21078
|
cs.CL
|
Nafiz Ahmed, Sarah Sharif, Dingjing Shi, Mike Banad |
Large language models are proposed to flag suicidal content, but representing a concept and acting on it are distinct. A suicidality feature is decodable from 0.5B parameters upward, while the behavior that uses it emerges only in the low billions. Where it em...Large language models are proposed to flag suicidal content, but representing a concept and acting on it are distinct. A suicidality feature is decodable from 0.5B parameters upward, while the behavior that uses it emerges only in the low billions. Where it emerges, it rests on a compact mid-network feature: in Llama-3.1-8B-Instruct this direction decodes suicidal from non-suicidal text above a matched lexical baseline and survives removal of the 38 most label-predictive keywords. Projecting it out drops the model's ranking from AUC 0.975 to 0.832, while an equal-norm random direction does not; ablating a low-dimensional subspace degrades it further, though no stable rank is recoverable. The direction transfers to two further corpora, replicates in Qwen and Mistral, and tracks clinician-assigned risk. Steering raises an off-target readout as much as the target across five concepts, so sufficiency claims from yes/no readouts are confounded by generic salience.
|
| 750 |
Who Checks the Citations? Benchmarking Legal Hallucination Detection
2606.21155
|
cs.CL
|
Patty Liu, Dominik Stammbach, Peter Henderson |
Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over ...Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations---with this number growing year-over-year. This study evaluates whether AI-based systems can mitigate these errors by automatically detecting hallucinations. We propose a taxonomy of legal citation hallucinations grounded in actual court filings and introduce a dataset of 1,300 brief excerpts containing injected errors. Benchmarking five models in agentic and non-agentic settings as well as Claude Code reveals that while the latest iterations perform better---GPT-5 achieves 84.4% recall and a 55.0% F1 score in an agentic framework---all models struggle with subtle error categories. Agentic verification remains resource-intensive, with GPT-5 averaging 15.3 steps per excerpt. Furthermore, restricted information access limits the efficacy of even the best agents. This gap creates policy concerns, as it disadvantages both AI systems and litigants who lack subscriptions to commercial legal databases. Together, our dataset, tools, and policy recommendations provide a foundation for building and auditing reliable legal citation checking tools.
|
| 751 |
Compression Is Not Evaluation-Neutral: Fixed RAG Compression Can Distort Reader Comparisons
2606.21807
|
cs.CL
|
Sugam Panthi, Rabab Abdelfattah |
Retrieval-augmented generation (RAG) evaluations often compare readers after a compressor has changed their evidence. This mixes two questions: which complete pipeline works best, and how much of a reader upgrade survives compression. We show that fixed compre...Retrieval-augmented generation (RAG) evaluations often compare readers after a compressor has changed their evidence. This mixes two questions: which complete pipeline works best, and how much of a reader upgrade survives compression. We show that fixed compression can raise average pipeline accuracy while hiding most of a reader upgrade. We hold questions, retrieved candidates, prompts, scoring, and compressed text fixed while comparing 8-20 readers across five question-answering benchmarks and five compression families. HotpotQA and MuSiQue are the main benchmarks, analyzed under a plan fixed in advance. In the 20-reader HotpotQA panel, the lowest- and highest-scoring readers under raw evidence are 31.8 percentage points (pp) apart before compression but only 7.8pp apart on the same stored output from a HotpotQA-trained RECOMP compressor. On both main benchmarks, lower raw-scoring readers gain more under exact match and token-overlap F1. Reader pairs also change order more often under compressed evidence than in raw replays across question halves. Row-level accounting explains how higher average accuracy and smaller upgrades coexist: compression rescues some wrong answers and damages some correct answers, with a different balance for each reader. We release ragscale to audit declared reader upgrades under raw and fixed compressed evidence. Reader evaluations should run this audit before attributing a compressed-only result to the reader.
|
| 752 |
Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding
2606.21906
|
cs.CL
|
Xuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang, An Yang |
Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dyn...Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that deeper representations yield more reliable next-token predictions. We revisit this assumption by revealing a recurring Guess-Refine-Perturb dynamic: early layers form coarse guesses, intermediate layers refine reasoning-relevant semantics, and final layers can perturb these refined predictions toward generic or alignment-preferred tokens. We introduce Confident Decoding, a training-free decoding strategy that dynamically selects the most reliable near-final layer through entropy-guided conservative backward search. We further provide a theoretical formulation of layer selection as an optimal stopping problem, showing that under bounded projection noise and dominant late-stage alignment perturbation, our search rule filters perturbation while bounding the loss relative to the oracle refinement layer. Experiments across dense and Mixture-of-Experts LLMs demonstrate consistent gains on challenging reasoning benchmarks, including GPQA-Diamond, Omni-MATH, and HLE, with zero memory overhead and less than 2% latency increase. These results suggest dynamically bypassing final-layer perturbations can unlock stronger reasoning behavior from aligned LLMs.
|
| 753 |
Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing
2606.27598
|
cs.CLcs.AI
|
Mreedul Gupta, Advait Deshmukh, Ashwin Umadi, Matt Pauk, Maria Leonor Pacheco |
Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail. We hypothesize that a key limitation is the reliance on sentence-level context, since disambiguating evidence is ofte...Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail. We hypothesize that a key limitation is the reliance on sentence-level context, since disambiguating evidence is often spread across multiple sentences. Testing this has been difficult because all existing UFET resources are sentence-level. We present Narrative-UFET, a controlled extension of UFET in which each entity mention is paired with an automatically generated short, coherent narrative. Synthesizing narratives lets us isolate the effect of specific discourse properties. We experiment with two paired variants: one in which the entity's type is held constant across the narrative (Maintain) and one in which it shifts (Change). We show that narrative context yields consistent improvements on long-tail types over sentence-level baselines, with the Change variant providing the stronger signal. A comparison against naturally occurring contexts shows that synthetic narratives yield stronger gains, indicating that controlled discourse construction can surface signals that real text leaves implicit. Substantial room for improvement remains, suggesting open directions in both discourse modeling and narrative construction.
|
| 754 |
Labeling Training Data for Entity Matching Using Large Language Models
2606.28823
|
cs.CL
|
Aaron Steiner, Christian Bizer |
Large language models (LLMs) achieve strong entity matching performance without task-specific training data, but applying them to large sets of candidate pairs is slow and costly. Matchers built on pretrained language models (PLMs), such as BERT, offer faster ...Large language models (LLMs) achieve strong entity matching performance without task-specific training data, but applying them to large sets of candidate pairs is slow and costly. Matchers built on pretrained language models (PLMs), such as BERT, offer faster inference but require training data. We systematically study knowledge-distillation workflows in which an LLM teacher labels training pairs for a smaller student matcher. We vary pair selection, labeling budget, teacher model, correspondence post-processing, and student model across eight benchmarks, including unseen entities and non-English data. We compare students trained on machine-labeled data with matchers trained on the original benchmark training sets. In most cases, PLM-based matchers trained on LLM-labeled data perform similarly to those trained on benchmark sets. Pair selection matters most for small labeling budgets, where active learning is often most effective. An open-weight teacher trains competitive students, so distillation requires no closed-weight models. Compact PLM-based students compete with much larger LLM students on most tasks while requiring 34 to 459 times less inference time than direct LLM matching. On the two benchmarks with high shares of unseen products, PLM-based students substantially underperform their teachers, as do students trained on benchmark data. Under GPT-5.2 pricing, LLM labeling costs per training set average \$5.86 to \$8.11. These findings support knowledge distillation as a practical approach to reduce the effort of labeling task-specific training data while enabling efficient inference.
|
| 755 |
From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives
2607.00918
|
cs.CLcs.AI
|
Chloe Ho, Aayush Aluru, Muhammad Hammouri, Kerry Luo, Ryan Lagasse |
Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce MAGNET, a multi-agent goal-driven narrative...Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and coherent plot lines in long-form stories. In this work, we introduce MAGNET, a multi-agent goal-driven narrative engine for storytelling, which generates stories with persona-grounded character agents that propose actions based on a shared structured state and evolving story goals. We evaluate MAGNET on 20 and 100 page stories using LLM based editor annotations and rubric scoring. At both narrative lengths, MAGNET significantly reduces editor annotations and improves rubric scores compared to single-model prompting, StoryBox, and StoryWriter (p<0.05). Ablation experiments also show that the critic module, structured state, and goal sequencing each contribute significantly to performance (p<0.05). These results suggest that long-form narratives can emerge from explicit structured states, critic-based action refinement, and goal-driven multi-agent generation with loosely specified character personas, providing a foundation for controllable and structurally coherent long-form narrative generation.
|
| 756 |
Message Passing Enables Efficient Reasoning
2607.01077
|
cs.CLcs.LG
|
Xuecheng Liu, Daman Arora, Gokul Swamy, Andrea Zanette |
While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck. Thus, in contrast to sequential scaling methods like CoT, recent parallel scali...While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck. Thus, in contrast to sequential scaling methods like CoT, recent parallel scaling techniques instead use fork and join (FJ) primitives to divide work across multiple LLM threads. However, in the fork-join paradigm, threads are typically transient and do not communicate pointwise with one another which limits scalability. To tackle this, we introduce Message Passing Language Models (MPLMs), a framework for LLM reasoning in which threads communicate directly via lightweight send and receive primitives. MPLMs enable efficient scaling through two key mechanisms: (1) reduced communication costs, achieved by avoiding redundant context sharing, and (2) preemption, which allows threads to terminate early based on partial information from their peers. We demonstrate the promise of MPLMs on 3 classes of tasks. First, on Sudoku puzzles, we show that MPLMs require an asymptotically smaller context than both serial CoT and parallel FJ. We then fine-tune a single model to solve 25 x 25 puzzles that remain challenging for standard CoT and FJ approaches, as well as frontier reasoning models without tools. Second, on 3-SAT puzzles, the capability of preemption allows termination of unpromising branches, which results in improved efficiency. Finally, we show that appropriately prompted large pre-trained models follow the MPLM protocol, achieving competitive results on long-context question answering relative to popular fork-join approaches.
|
| 757 |
Office Comprehension Benchmark
2607.01245
|
cs.CLcs.LGcs.AI
|
Firoz Shaik, Mateus Pican\c{c}o Lima Gomes, Tanvir Aumi, Jingci Wang, Milos Milunovic |
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity ...We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension over native file formats (.docx, .xlsx, .pptx) and their variants. OCB consists of two tracks. File Fidelity Q&A tests structural and visual perception of office artifacts - tables, charts, embedded images, formulas, and app-specific elements such as headers, speaker notes, and named ranges. Domain Q&A tests expert-level reasoning grounded in real-world industry documents across 12 professional domains, with queries requiring multi-step analysis and synthesis across documents. Each reference answer is decomposed into atomic, binary-gradable claims, and an ensemble of LLM judges scores responses against each claim independently. Even the strongest frontier system in its default reasoning mode reaches only about 59.3% on Domain Q&A; increasing thinking depth within a tier does not move performance materially, while moving to a higher product tier yields modest gains. We release the dataset, evaluation tooling, judge prompt, and a public leaderboard.
|
| 758 |
Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale
2607.01538
|
cs.CL
|
Siddharth Gollapudi, Nilesh Gupta, Prasann Singhal, Sewon Min |
Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant answer. However, prior work has largely focused on proprietary systems or the smaller-scale reranking task,...Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant answer. However, prior work has largely focused on proprietary systems or the smaller-scale reranking task, leaving corpus-scale in-context retrieval largely unexplored. In this work, we present the first systematic study of in-context retrieval on two scales practical retrievers demand: million-token corpora and length-generalization far beyond training-time sizes. We first introduce BLOCKSEARCH, an 0.6B LCLM retriever whose architectural and training modifications improve over prior LM baselines and length-generalize up to 10x beyond its training regime. Nevertheless, its retrieval still collapses under more extreme extrapolation. We trace this failure to an attention dilution effect: as the corpus grows, irrelevant documents dominate the softmax denominator and the normalized mass on the gold document collapses. Individual attention heads continue to locate the gold document more reliably than the model decodes it, even at million-token scale, though this signal also weakens as the corpus grows. Motivated by this analysis, we introduce length-aware adjustments to the attention softmax and document-level sparse attention, improving retrieval at million-token scale to performance comparable to a same-backbone dense retriever. On the lexical LIMIT benchmark, these gains also transfer out of distribution: at a million-token corpus, Recall@1 reaches 13.5% versus 2.9% for the dense baseline. Together, our results position in-context retrieval a promising alternative to classical retrieval while emphasizing attention control under extreme context growth as a new challenge.
|
| 759 |
LP-SFT: Keeping Plausible Alternatives Alive in Supervised Fine-Tuning
2607.04733
|
cs.CLcs.LG
|
Yueyang Wang, Baolong Bi, Shuo Lu, Jingyuan Zhang, Jiajun Shi |
Supervised fine-tuning (SFT) is widely used to adapt pretrained language models to downstream domains, but can over-specialize the model and degrade its pre-existing capabilities. Standard cross-entropy drives probability mass onto the observed target token an...Supervised fine-tuning (SFT) is widely used to adapt pretrained language models to downstream domains, but can over-specialize the model and degrade its pre-existing capabilities. Standard cross-entropy drives probability mass onto the observed target token and can make plausible alternatives that the pretrained model itself endorsed vanish. Using Shannon and Renyi entropies, we show that pretrained next-token distributions exhibit a regular multimodal structure, with entropy peaks corresponding to different numbers of plausible alternatives. Motivated by this observation, we propose LP-SFT, a Local-Preserving Supervised Fine-Tuning objective with two key design principles: it removes the supervised token from a local top-$K$ candidate set to avoid conflict with cross-entropy, and applies a locally normalized KL loss to preserve relative preferences among the remaining non-label alternatives. Across mixed-domain and domain-specific fine-tuning experiments, LP-SFT consistently outperforms standard SFT and recent baselines in aggregate performance while keeping base-endorsed alternatives from vanishing, achieving a favorable balance between single-sample accuracy and finite-budget solution accessibility, as measured by pass@1 and pass@$k$, respectively, with only a modest increase in training cost.
|
| 760 |
Co-LMLM: Continuous-Query Limited Memory Language Models
2607.07707
|
cs.CLcs.LGcs.AI
|
Yair Feldman, Linxi Zhao, Nathan Godey, Dongyoung Go, Yilun Hua |
Externalizing knowledge in LLM pre-training is a promising avenue to achieve higher performance at smaller scales, control knowledge use, and overall increase model transparency. We propose continuous-query limited memory language models (Co-LMLM), an LLM that...Externalizing knowledge in LLM pre-training is a promising avenue to achieve higher performance at smaller scales, control knowledge use, and overall increase model transparency. We propose continuous-query limited memory language models (Co-LMLM), an LLM that interleaves flexible vector retrieval queries with next-token predictions, and is pre-trained to copy knowledge returned from the KB, rather than memorize it. Co-LMLM is pre-trained with a scalable approach that jointly trains a knowledge-externalizing LLM, induces its knowledge base, and learns an expressive continuous retrieval mechanism. Across pre-training at multiple model scales, Co-LMLM outperforms prior knowledge-externalizing and vanilla LLMs in both perplexity and factual precision. At 360M scale, this includes lower perplexity than models pre-trained on 40$\times$ more data, and SimpleQA-verified performance that is in line with gpt-4o-mini and higher than Claude Sonnet 4.5.
|
| 761 |
Can a Language Model Learn Facts Continually in Its Weights?
2607.11020
|
cs.CLcs.LG
|
Charles O'Neill, Max Kirkby, Jonathon Liu, Michael Psenka |
Continual learning is a long-standing capability gap between LLMs and humans. Writing new knowledge into a model's weights routinely causes it to forget old knowledge, commonly denoted as "catastrophic forgetting". Various modifications of supervised fine-tuni...Continual learning is a long-standing capability gap between LLMs and humans. Writing new knowledge into a model's weights routinely causes it to forget old knowledge, commonly denoted as "catastrophic forgetting". Various modifications of supervised fine-tuning and distillation aim to mitigate catastrophic forgetting, but quantifying what (or how much) information was forgotten is often difficult. In this paper, we study whether current methods of writing knowledge into weights enable models to learn continually without forgetting. We introduce a framework for studying continual learning in the iterative regime, writing invented facts one at a time into a Qwen3 model already modified by previous writes, and varying the training data, method, and parameter update. Across SFT and off- and on-policy distillation, using LoRA or full fine-tuning, we compare repeated statement training (the same fact repeated in two formats) with varied example training (24 factual restatements) and find that varied examples comprehensively support more flexible use. After twenty sequential writes and merges, the model answers only 1% of questions about earlier facts correctly when every write uses repeated statements, compared with 46% when every write uses varied examples. We additionally show that this retention depends on the data used for the later writes, regardless of training method or parameter update, and that behavioral forgetting of an earlier fact does not erase its presence from the log-probabilities. Together, our framework neatly provides a comparison of performance across training data, training regimes, and parameter update schemes in an iterative learning task.
|
| 762 |
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
2607.13591
|
cs.CLcs.AI
|
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu, Levina Li, Dong Liu |
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed he...Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of the tasks, benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MemCon is backend-agnostic: it wraps any existing memory implementation, learns from task-by-task binary feedback with no pretraining and no additional LLM calls, and uses a lightweight tabular contextual bandit with UCB exploration that converges within tens of tasks. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MemCon consistently outperforms multiple memory baselines by up to 15.2 points in task success while reducing token consumption by 5--20%.
|
| 763 |
Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
2607.15092
|
cs.CL
|
Haocheng Yang, Licheng Pan, Yuan Lu, Xiaoxi Li, Zhichao Chen |
Rubric evolution offers a promising approach to improving the quality of rubrics generated by large language models (LLMs). Central to this process is rubric comparison, which identifies the better of two rubrics and guides the direction of evolution. However,...Rubric evolution offers a promising approach to improving the quality of rubrics generated by large language models (LLMs). Central to this process is rubric comparison, which identifies the better of two rubrics and guides the direction of evolution. However, accurate rubric comparison is difficult, which presents two challenges. (1) It should reflect downstream task performance, which is essential for assessing rubric utility but often prohibitively expensive to evaluate. (2) It should discourage unnecessary criteria, which increase verification costs and may dilute the influence of essential criteria. To address these challenges, we introduce Rubrics on Trial, a multi-agent framework that evolves rubrics by comparing synthetic response pairs. To address challenge 1, the framework compares synthetic responses that satisfy the respective rubrics, providing a proxy for downstream performance without training a separate policy for each rubric. To address challenge 2, it assesses the necessity of a candidate criterion by independently generating high-quality alternative responses that violate it and comparing them with edited versions that satisfy it. A rubric is favored when it improves response quality in both comparisons, and the resulting comparison signal is further incorporated for rubric evolution. Extensive experiments demonstrate that Rubrics on Trial improves the quality of generated rubrics and leads to better downstream task performance.
|
| 764 |
SLPO: Scaling Latent Reasoning via a Surrogate Policy
2607.19691
|
cs.CLcs.LGcs.AI
|
Runyang You, Zhiyuan Liu, Yongqi Li, Wenjie Li |
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since intermediate reasoning must be externalized as ...Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since intermediate reasoning must be externalized as natural-language tokens. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: a differentiable surrogate policy interface over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across two continuous latent reasoners, two backbones, and three held-out benchmarks, SLPO improves Pass@8 and Pass@16 in all 12 backbone--dataset settings, with gains of up to 12.07 percentage points. SLPO further transfers to soft-token inference and learns difficulty-adaptive computation, allocating longer latent trajectories to harder instances. Project Page: https://modalitydance.github.io/SLPO/
|
| 765 |
HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering
2607.20219
|
cs.CL
|
Abdessalam Bouchekif, Mohammed-En-Nadhir Zighem, Salah Eddine Bekhouche, Hichem Telli, Somaya Eltanbouly |
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact...Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer. We introduce HalluTruthQA, a fine-grained benchmark for hallucination evaluation in Arabic question answering. The benchmark contains 2,400 expert-curated examples across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Each example pairs an Arabic question and a model-generated answer with a verified reference answer, a binary hallucination label, and six candidate answers for factual verification. Hallucinated answers additionally include character-level erroneous spans, human-written explanations, and macro- and micro-level hallucination types. We evaluate four open-source LLMs, ALLaM-7B, Falcon-H1R-7B, Qwen3-32B, and SILMA, in a zero-shot setting across hallucination detection, span-level localization, factual verification, and explanation evaluation. Results show that these tasks capture different abilities: no single model performs best across all tasks. The best scores are 0.880 Macro-F1 for detection, 0.516 F1-Sp for localization, 0.852 LO-Score for factual verification, and 0.644 for explanation evaluation. These findings show that hallucination evaluation should move beyond response-level detection toward the localization, verification, and explanation of factual errors.
|
| 766 |
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
2607.24604
|
cs.CLcs.AI
|
Xueping Gao, Jianwei Yang, Qiang Yang |
Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs p...Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified 14B replication, stale traces harm 34/135 correct starts versus 4/135 with current traces, a 22.2-point increase (task-cluster 95\% CI $[8.9,37.0]$, exact Holm $p=0.0337$). A prospective 540-rollout policy eliminates observed correct-start harm but reduces wrong-start repair and fails its joint criterion. Repository experiments over 24 bugs and four coder stacks expose floor effects and component heterogeneity without Holm-significant effects. We therefore separate admission, preservation, grounded certification, competence, and liveness. We derive an evidence-bound typed loop contract and instantiate its mechanically enforceable subset in a reference implementation that binds verifier evidence to exact code states, preserves verified checkpoints, and emits auditable admission receipts. The implementation is an executable specification and conformance artifact, not evidence of improved repair competence or calibrated verifier dependence.
|
| 767 |
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
2607.26348
|
cs.CLcs.AI
|
Zihan Chen, Di Zhu, Lei Nico Zheng |
Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an ev...Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We publicly release the cross-domain benchmark and validation framework, including all code and data, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.
|
| 768 |
KV Cache Translation across Heterogeneous Large Language Models
2607.28979
|
cs.CL
|
Jin-woo Lee, Minkyung Song, Junghyun Oh, Seunghoon Han, Gwangseon Jang |
Heterogeneous Large Language Model (LLM) systems increasingly share contexts, retrieved evidence, and multi-agent dialogue histories, yet their internal key-value (KV) caches remain model-specific and cannot be reused across architectures. Consequently, each m...Heterogeneous Large Language Model (LLM) systems increasingly share contexts, retrieved evidence, and multi-agent dialogue histories, yet their internal key-value (KV) caches remain model-specific and cannot be reused across architectures. Consequently, each model must repeatedly prefill or store caches for the same context, limiting the scalability of multi-model reasoning and long-context generation. We propose Mixture-of-Translators (MoT), a cache translation framework that maps context KV caches from a source LLM into the cache space of a target LLM. Unlike prior approaches that depend on a single projection path or global shared latent space, MoT uses multiple translator modules with token-level routing to capture diverse cache translations. We further introduce a Context Correction Loss that aligns the replayed target trajectory with the native target trajectory, reducing residual translation error. Our analysis reveals two competing failure modes: propagation error from early injection and correction-deficit error from late injection. We evaluate heterogeneous translation among Llama, Gemma, and Qwen, spanning substantially different architectures and KV-cache spaces. MoT achieves an average accuracy of 57.6% and F1 of 0.42. These scores recover 95.0% and 120.6% of native target performance and reach 1.5x and 3.5x the strongest-baseline scores, respectively. In practical case studies, MoT enables accurate heterogeneous-agent discussion with scale-invariant KV-memory behavior as the number of agents grows, while retaining near-native generation quality in long-context cache-augmented generation. These results demonstrate scalable KV-cache reuse across heterogeneous LLMs.
|
| 769 |
Cost-Effective Automated Judging of Natural-Language Mathematical Proofs
2608.00004
|
cs.CLcs.LGcs.AI
|
Benjamin Grayzel |
Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, an...Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IMO-GradingBench, two of three cheap judges (GPT-OSS-120B, DeepSeek-V4-Flash) and their three-model consensus are statistically no worse than the frontier (Claude Opus 4.7, Gemini 3.1 Pro) on agreement with human pass/fail decisions, at 4-100$\times$ lower cost. On the full 1000-instance benchmark, the choice of consensus rule over the three judges is a precision/recall dial: unanimous (all-three-pass) rules reach the highest precision (0.855), majority vote the highest recall (0.912); across four replicate runs the unanimous rule is also the steadiest. No rule won outright; the dial replicated on a held-out 600-instance split and on the independent ProofBench. In this domain, cheap judges are competitive with the frontier at one to two orders of magnitude lower cost, and unanimity is the right setting when false positives are costly.
|
| 770 |
Exemplars in Disguise: Pure Exemplar Models Mimic Abstraction-First Learning
2608.00821
|
cs.CL
|
Zachary Nicholas Houghton, Vsevolod Kapatsinski |
Whether idiosyncratic, item-specific knowledge is learned before abstract class-level generalizations, or vice versa, is a central question in language learning, with exemplar and abstraction-based theories making opposite predictions. Recent methods have clai...Whether idiosyncratic, item-specific knowledge is learned before abstract class-level generalizations, or vice versa, is a central question in language learning, with exemplar and abstraction-based theories making opposite predictions. Recent methods have claimed to show that, at least for large language models, abstract knowledge is learned first. We show that these methods fall short: pure memorizer models with no abstract representations can appear, by the same criteria, to learn either item-specific or class-level knowledge first, depending on their sensitivity to individual observations, with the transition point governed by the distributional properties of the input. We further argue that the distinction between item-specific and abstract knowledge may be ill-defined for distributed representations, as a word's class-level properties may not be separable from its item-specific properties.
|
| 771 |
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
2608.02613
|
cs.CLcs.LGcs.AI
|
Jiadong Zhang, Xiaosong Ma |
Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent mu...Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for 50 agents over 15 days (10.3M dialog-text tokens, 24.1K text-only ego-observed tokens/agent/day). With the interaction history, it co-generates ground truth over six recall, reasoning, and trustworthiness evaluation dimensions. We evaluate five open-weight readers with Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch as memory backends. Three results stand out: (1) Memory-backend choice matters more for content accuracy: At Qwen3-8B, moving from Memobase to MemSearch gains +22.1/+21.2 pp, whereas scaling the reader to Qwen3-32B-AWQ gains at most +3.5/+4.4 pp under either backend. (2) Permission-aware access fails in two distinct modes: Oracle leaks heavily, while the other backends fail to surface the protected fact. (3) Search latency bites only at small readers: on a Spark GB10 edge node, memory-search adds a moderate and fixed 87/8/51 ms (BM25-RAG/Memobase/MemSearch) that composes a small part of TTFT for most reader-backend combinations. We release code, the MASim simulator, and the MemArena-L benchmark at https://github.com/dereksodo/MemArena-Bench.
|
| 772 |
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
2608.15763
|
cs.CL
|
TaoLive AIGC LLM Team, Yuhan Sun, Wenhao Lin, Yongdong Luo, Yibo Hu |
AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates, and accurate yet effective responses. Evolvable Harnesses, whose Skills, Hooks, ...AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates, and accurate yet effective responses. Evolvable Harnesses, whose Skills, Hooks, prompts, and tools can be updated independently of model weights, enable rapid iteration but expose a trade-off: large models adapt zero-shot yet are too slow, whereas compact models meet latency targets but overfit to fixed Harness configurations. We propose Harness-Aware Training (HAT), which trains compact models to adapt to changing Harnesses. Its key component, Harness-State Augmentation (HSA), applies task-preserving transformations to Skill identifiers and content, tool schemas, prompt structures, and Hook functions. Training proceeds in three stages: HSA-SFT learns reasoning and tool use from strong-model trajectories across diverse environments; General On-Policy Distillation restores generalization lost during SFT; and HSA-RL improves robustness to changing Harnesses through reinforcement learning in augmented environments. Across four evaluation sets, HAT achieves 94.8 on Live-Stream QA (base: 80.3; strongest general LLM: 93.0) and 94.6 on Harness-Variant QA (base: 75.4). Unlike Fixed-Harness SFT, which lowers IFEval by 7.7 points from the base model, HAT avoids this regression and reaches 83.5. On one NVIDIA H20 GPU, the optimized system delivers P50 and P95 latencies of 3.4 s and 8.1 s. Deployed in Taobao Live's digital-avatar service, it also yields positive online A/B test results for item-page views.
|
| 773 |
Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning
2608.19009
|
cs.CL
|
Yajie Yin |
Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to m...Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to mean at least five different things: verification granularity, concept abstraction, risk tier, system-stack layer, and the epistemic source of the ground truth. We propose Verification Autonomy Levels (VAL), a meta-standard that classifies any verification scheme along a single axis: where does the verification spec come from, and what does the verdict guarantee? VAL ranges from L0 (LLM self-declaration; no deterministic anchor) through L2 (objective ground truth; correctness only) to L3/L4 (decidable systems with single-property or domain-level completeness), with L5 impossible in the unrestricted case. Central to VAL is the completeness blind spot: substitution- and sampling-based verifiers can confirm that proposed candidates hold, but cannot prove that no candidate was missed. We further identify a dichotomy the literature has not stated: completeness is reachable only for formally specifiable properties, whereas empirical open-world verification (fact-checking, diagnosis) caps at anchored correctness (L2). We document this gap empirically across four domains (symbolic mathematics, behavior monitoring, medical diagnosis, code generation) and in the strongest formal-verification baseline in our survey. We show that granularity, concept hierarchy, risk, and system stack are orthogonal to VAL, resolving a conflation across 17 surveyed papers. Code and the full literature assessment are released on Zenodo (DOI: 10.5281/zenodo.23120985).
|
| 774 |
FormuEvo: LLM-Guided Evolution for Discovering Solver-Efficient Mixed-Integer Programming Formulations
2608.23353
|
cs.CL
|
Haofeng Yuan, Jianing Peng, Jieyi Bi, Ni Zhang, Shiji Song |
Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While large language models (LLMs) have recently shown promise in automated MIP modeling from natural language, they prioritize semantic correctness but overlo...Mixed-integer programming (MIP) lies at the core of operations research and industrial optimization. While large language models (LLMs) have recently shown promise in automated MIP modeling from natural language, they prioritize semantic correctness but overlook formulation strength, severely bottlenecking the efficiency of downstream solvers. We propose FormuEvo, an LLM-guided evolutionary framework for automated discovery of solver-efficient MIP formulations. FormuEvo frames MIP formulation design as evolutionary optimization over the symbolic space of MIP formulations, represented as executable modeling programs, by iteratively generating, evaluating, and selecting stronger candidates via LLM-driven crossover, mutation, and repair operations. To move beyond blind exploration, FormuEvo introduces a solver-informed diagnosis mechanism that exploits fine-grained solver statistics as verbal gradients for targeted refinement. Additionally, a structured memory abstracts prior experience into reusable modeling strategies, avoiding redundant exploration while enabling zero-shot transfer to unseen problems and bootstrapping smaller LLMs. Experiments across diverse linear and non-linear problems demonstrate that FormuEvo discovers formulations that significantly outperform both expert-designed formulations and existing LLM-based approaches, accelerating solvers by up to 5.5$\times$, with distilled knowledge transferring effectively across problems and model scales.
|
| 775 |
ROBE: Reversed-Order-Biased-Experts for Extracting Extreme Long-tail Events from Historical Texts
2608.24268
|
cs.CL
|
Stella Verkijk, Piek Vossen |
This paper proposes methods to extract over 50 types of events from a Dutch historical corpus spanning the 17th and 18th centuries. The methods we propose aim to tackle a very challenging scenario in Machine Learning: extracting the long-tail of the long-tail....This paper proposes methods to extract over 50 types of events from a Dutch historical corpus spanning the 17th and 18th centuries. The methods we propose aim to tackle a very challenging scenario in Machine Learning: extracting the long-tail of the long-tail. Historic data from before the 19th century is in itself a niche domain not covered in the pre-training of Large Language Models, and we aim to extract events only scarcely annotated in the training data available for this domain. We propose creating expert classifiers for subgroups of the events present in the training data. We make these groupings based on similar frequency in the training data or on semantic relatedness. Experts trained on underrepresented events are assigned higher priority when predicting to avoid being dominated by frequency biases. We refer to this new way of combining classifiers, specifically tailored to protect the long-tail, as ROBE: Reversed-Order-Biased-Experts. We also propose a controlled method to create domain-specific synthetic data. Our two implementations of ROBE outperform a simple fine-tuned encoder model with a .16 increase in precision and a .05 increase in recall respectively. The best model achieves a .11 increase in f1 for a group of long-tail classes in our niche data set.
|
| 776 |
Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal
2608.24988
|
cs.CLcs.LGcs.AI
|
Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron |
Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is un...Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this. We study the stability of embedded steering for refusal ablation and brevity amplification across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF. We find that steering degrades more under stronger optimisation pressure against the steered behaviour, which depends on both the training data and the training algorithm. Refusal ablation loses 69% of its effect on average under our SFT setup, but only 2% under our RLHF setup. The weight-edit, by contrast, remains almost unchanged (on average only 0.4% of it is restored), and the fine-tuning update along the steering direction is close to orthogonal to the pre-edit weight pattern (mean cosine similarity 0.071). Fine-tuning that restores the behaviour does so without reversing the edit. As the steered behaviour is not fully durable, embedded steering should be re-validated after downstream training.
|
| 777 |
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
2608.27309
|
cs.CLcs.AI
|
Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li |
LLM-judge audits assess bias by comparing ratings across matched conditions. Difference-in-differences designs compare two candidate responses within each item and then compare that contrast across a manipulated attribute. We show in closed form that this endp...LLM-judge audits assess bias by comparing ratings across matched conditions. Difference-in-differences designs compare two candidate responses within each item and then compare that contrast across a manipulated attribute. We show in closed form that this endpoint need not identify differential preference on the latent scale when ratings are bounded. A severity shift common to both responses produces an observed interaction whenever the scale bounds attenuate it unequally. We examine this problem in a pre-registered audit comprising 990 calls to a frozen pedagogy judge. The registered primary analysis did not detect an effect of the stated learner profile on scaffolding preference. The only nominally significant secondary interaction concerned productive struggle ($+0.378$ points; $p = 0.002$). A post-hoc construction with zero differential preference reproduced 79 to 85\% of this interaction using the observed severity shift and the scale floor alone. The observed interaction therefore does not identify differential preference on the latent scale.
|
| 778 |
Sliding-window beats linear attention
2608.28444
|
cs.CLcs.LG
|
Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais |
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy: every new token costs more than the previous one, and its keys and values must be stored in memory indefinitely, which is unsustainable. Two main lines o...Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy: every new token costs more than the previous one, and its keys and values must be stored in memory indefinitely, which is unsustainable. Two main lines of work address this: compressing the KV cache, e.g., by evicting or quantizing keys and values, and retrofitting LLMs to use Linear Attention, which replaces the KV cache with a fixed-size state. Retrofitting has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, it has not been properly compared to the simplest form of KV-cache eviction: Sliding Window Attention (SWA) with attention sinks. In this work, we show that SWA with sinks performs as well or better than most retrofitted Linear Attention models across multiple LLMs and downstream tasks, with the largest gains on long-context and generative tasks. On long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no additional training, is extremely fast, and requires little memory, making it an extremely cheap and reliable solution. When the training budget is limited, switching to SWA is a much more effective way to reduce inference memory cost than retrofitting linear attention. Linear attention models have shown promise, but they require training from scratch or extensive retrofitting to reap their architectural benefits and come close to SWA.
|
| 779 |
The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice
2608.28930
|
cs.CLcs.LGcs.AI
|
Jungseob Lee, Jaehyung Seo, Heuiseok Lim |
Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the ...Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full-dimensional classifiers, so apparent architectural complexity largely reflects high-dimensional covariance estimation difficulty rather than exploitable non-linearity. A simple L2-regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi-layer aggregation exceeds CLAP cross-layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle-layer performance without oracle access. Our claims characterize the geometry within the controlled paired-example paradigm. Our code is available at https://github.com/js-lee-AI/LayerMix.
|
| 780 |
Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning
2608.29278
|
cs.CL
|
Zhaolu Kang, Meixin Wu, Yu Xue, Yingjie He, Qiming Shi |
Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal fusion, but clean evaluation...Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal fusion, but clean evaluation cannot tell whether success depends on stable cross-modal structure or on cues sufficient only in intact inputs. To address this gap, we define a modality fault line: a boundary at which model behavior becomes unstable when a modality remains present and human-interpretable, but its internal evidence structure is perturbed. We introduce SCEval (Structure-Corruption Evaluation) a diagnostic evaluation protocol that keeps the question, answer space, and modality channels fixed while applying controlled structural corruptions to text, vision, and audio individually and jointly. Built from $273$ human-verified tri-modal examples from Social-IQ, OmniBench, and VALOR, SCEval evaluates $15$ proprietary and open-source omni-modal systems. The results show that structural corruption lowers clean accuracy, text--vision damage forms the most stable shared fault line, and multi-modal degradation is non-additive rather than a simple function of the number of corrupted modalities. Clean omni-modal accuracy therefore does not establish that a model will remain reliable when cross-modal evidence becomes structurally unreliable.
|
| 781 |
JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction
2608.29616
|
cs.CL
|
Zhaolu Kang, Yantao Liu, Tailong Luo, Leqi Zheng, Lei Wei |
Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges...Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, and while some have attempted to evaluate reasoning quality, their evaluations are indirect, often relying on LLM-generated rubrics that reflect model-internal preferences rather than the inherent logical structure of legal adjudication. We propose Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction. JPO first uses teacher-generated rationales to supervise a standardized four-step reasoning process, and then applies reinforcement learning with a composite reward over legal prediction quality, reasoning structure completeness, and cross-step consistency. JPO further introduces token-level advantage reweighting and adaptive clipping for legally salient reasoning segments. Experiments on multiple open-source language models and three Chinese legal benchmarks show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.
|
| 782 |
CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation
2608.30158
|
cs.CLcs.AI
|
Kwangmin Ki, Yunhun Nam, Jongheon Jeong, Jaehyung Kim |
Supervised fine-tuning (SFT) is the de facto standard for adapting large language models (LLMs) to target domains, but it often degrades the model's general capabilities, a phenomenon known as catastrophic forgetting. Existing approaches typically modify the S...Supervised fine-tuning (SFT) is the de facto standard for adapting large language models (LLMs) to target domains, but it often degrades the model's general capabilities, a phenomenon known as catastrophic forgetting. Existing approaches typically modify the SFT loss to mitigate forgetting, but they inevitably operate along a domain-generality trade-off. In this work, we step outside this trade-off by decoupling the two capabilities at the model level: we keep the original base model for general capability, and selectively invoke the SFT expert only when domain-specific knowledge is required. Specifically, we propose CPR (Critical-Point Routing), a token-level routing framework between a base model and its expert derivative, based on critical tokens where the base model fails but the expert succeeds. We train a lightweight hierarchical router that estimates the expert-call probability per token, and pair it with a tailored inference procedure that combines momentum smoothing and threshold gating. Across diverse model-domain configurations, CPR achieves state-of-the-art across all settings, surpassing SFT expert by 1.4-5.5% in domain performance while recovering its general-capability drop from 3.4-14.5% to at most 0.5%, with minimal overhead from invoking the expert on only one-third of tokens.
|
| 783 |
Every Token Leaves a Ripple in the Stream of Thought: Eliciting Model-Internal Token Saliency for Chain-of-Thought Compression
2608.31066
|
cs.CL
|
Tianyi Zhao, Yinhan He, Wendy Zheng, Jundong Li, Chen Chen |
Chain-of-thought (CoT) reasoning improves multi-step problem solving, but long reasoning traces inflate inference cost. Token-level CoT compression reduces this cost by pruning full reasoning chains into shorter traces for model adaptation, making token select...Chain-of-thought (CoT) reasoning improves multi-step problem solving, but long reasoning traces inflate inference cost. Token-level CoT compression reduces this cost by pruning full reasoning chains into shorter traces for model adaptation, making token selection the central challenge. Existing methods often rely on external scorers or heuristic signals only indirectly tied to the model's internal answer computation. We instead adopt a model-internal perspective: as the model forms an answer, each reasoning token induces a ripple in the residual stream whose effect on the answer reflects the token's contribution to the underlying computation. Building on this view, we propose \textsc{MIST} (Model-Internal Saliency for Token-level CoT compression), which defines token importance along two complementary axes: \emph{necessity}, the drop in answer likelihood when a token's internal contribution is removed, and \emph{sufficiency}, the gain in answer likelihood when that contribution alone is provided. Combining the two yields a unified importance score for pruning. Across four reasoning benchmarks and four models, \textsc{MIST} consistently outperforms baseline methods, suggesting that model-internal saliency provides an effective proxy for reasoning-token importance.
|
| 784 |
Investigating Assistant Bias in LLM User Simulators Using a Role Vector
2609.00608
|
cs.CL
|
Daeheon Jeong, Yoonjoo Lee, Eugene Choi, Sinie van der Ben, Juho Kim |
LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task goals. They rarely reproduce ...LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.
|
| 785 |
How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models
2609.03322
|
cs.CL
|
Dun Li Chan, Emily Liu, Niyathi Allu, Christian Hoang |
Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language...Language models encounter typos, corrupted text, altered words, and disrupted token order, yet robustness is usually evaluated only through output behavior. We study how six naturalistic and synthetic input perturbations propagate through decoder-only language models at three levels: output behavior, hidden-state geometry, and attention-head function. We evaluate behavioral effects across four GPT-2 and two Qwen2.5 checkpoints, analyze layerwise geometry using centered kernel alignment and intrinsic dimension, and examine attention-head responses in GPT-2. Perturbation types produce distinguishable metric profiles that are not fully captured by output measures and are only partly consistent across the tested checkpoints. Copying scores show the strongest pooled associations with activation-patching recovery under token substitution and shuffling, although these associations do not isolate copying-specific effects. Gradient-guided HotFlip perturbations also cause stronger behavioral and representational disruption than rate-matched random token substitutions in GPT-2; their behavioral effects are consistent across all six tested checkpoints. Our results show that robustness claims based on a single behavioral or representational metric can be misleading, and motivate multi-level evaluation of how perturbations alter language-model computation.
|
| 786 |
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect
2609.07448
|
cs.CL
|
Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo, Shanyu Chauhan, Felix Drinkall |
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) often change their responses to subtle rephrasings that align with an implied stanc...We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) often change their responses to subtle rephrasings that align with an implied stance by users. This can leave users with advice tainted by how they happened to phrase a question rather than by the underlying facts, and the consequences are highly costly in high-stakes domains. Because in the realistic scenarios, both expert practitioners and non-expert users frequently ask LLMs questions containing incomplete or misleading assumptions, models are highly susceptible to those framings. To test this, we inject the framing bias across three nested levels: a framing-biased question phrasing (root), an injected framing-biased premise prepended to a neutral question (propositional), and a premise paired with a framing-biased question (global). Evaluating nine open models (3.8B-70B) across four families, we find that strong per-variant accuracy does not guarantee the robustness across differently phrased questions under the fixed factual information.
|
| 787 |
The Hidden Frame: How Large Language Models Impact Democratic Society
2609.07735
|
cs.CL
|
Wend K. Tam |
Large language models are rapidly becoming an interface between citizens and political information. They are often regarded as "a better Google." While this analogy might work for some instances, it is unintuitively problematic for democratic politics. A searc...Large language models are rapidly becoming an interface between citizens and political information. They are often regarded as "a better Google." While this analogy might work for some instances, it is unintuitively problematic for democratic politics. A search engine retrieves human-authored documents, while a language model generates novel text that necessarily embeds invisible framing decisions. Because conveying knowledge involves framing, a system that generates answers cannot serve as a neutral conduit to "all human knowledge." Instead, these systems are becoming a new kind of political intermediary. Mechanistic evidence shows that partisan identity is encoded as a locatable geometric direction inside the Llama 3.1 8B model, and that alignment training masks rather than removes this structure. Building on that evidence, we use the model's training cutoff in 2023 to make those frames visible. This cutpoint auspiciously falls just before a dramatic realignment in American politics marked by the second Trump administration and the MAHA transformation of health politics, providing us with a natural experiment. We find that the model presents temporally contingent partisan alignments as {\em knowledge}, with no reliable mechanism for distinguishing fact from opinion. Because citizens judge political claims by who made them and when, an answer that carries neither cue hampers the judgment on which self-governance depends. Large language models purporting to summarize "all human knowledge" are, in actuality, simply magnifying the cultural and partisan divides inherent in their training data.
|
| 788 |
BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models
2609.09554
|
cs.CL
|
Shivam Singh, Aditya Yadavalli, Catherine Arnett, Alex Warstadt |
We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent mode...We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often perform poorly on languages less well-represented in their training set. While it has long been known that effective language adaptation can be achieved through simple fine-tuning on monolingual data, this strategy has only been applied to a small number of languages. We massively scale up this simple approach to 102 languages covered in the FLEURS dataset, while also implementing a more complex language adaptation strategy that integrates monolingual tokenizer replacement and data augmentation using text-only fine-tuning. BuzzASR models outperform Whisper-large-v3 on 77 out of 102 languages, reducing character error rates (CER) by a factor of over 2.8 on average. Our models achieve state-of-the-art CER among open-source systems on 27 of 102 languages on the combined FLEURS and Common Voice test set. Our tokenizer replacement strategy yields an average 3.3x improvement in compression rate (characters per token) over Whisper's multilingual BPE, with gains of up to 21.7x. We release all models, code, and detailed results: https://lemn-lab.github.io/buzz-asr
|
| 789 |
Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models
2609.14779
|
cs.CL
|
Mingze Yin, Xiaohan Wang, Dian Li, Haichao Yao, Yilin Zhao |
Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the realm of mathematical function...Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the realm of mathematical functions, our investigation reveals a critical modality interference phenomenon: even advanced models, while performing textual computational reasoning, tend to disregard or misinterpret essential visual cues. To address this challenge, we propose Func-R1, which synergistically harmonizes precise visual perception and rigorous logical reasoning. Concretely, built upon an explicitly decoupled architecture, we employ a hierarchical post-training framework to progressively identify critical visual evidence and conduct in-depth theoretical reasoning. Furthermore, the Perception-Aligned Theoretic Optimization (PATO) strategy is proposed to steer policy updating towards internalizing fundamental theoretical properties while dynamically rectifying heterogeneous visual information throughout the reasoning process. Extensive experiments across diverse benchmarks demonstrate that Func-R1 delivers the optimal performance among open-source MLLMs, even surpassing GPT-5 with an 8.4% improvement on MathVerse's function-oriented tasks.
|
| 790 |
Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
2609.18909
|
cs.CLcs.AI
|
Xinshuai Guo, Junjie Wu, Dolly Deng, Yinghui Li, Hai-Tao Zheng |
Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in ...Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are systematically associated with final agent performance. To disentangle agent performance redundancy from a complete perspective, we propose DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores. Across five agent benchmarks and five representative baselines, DualViewEval achieves the best results in all datasets. With only 20 tasks, it achieves $24\times$--$40\times$ compression on APEX-Agents and BFCL, reducing mean absolute error (MAE) by $14.5\%$--$28.2\%$ over the strongest competitors while improving Kendall's $\tau$ by up to $7.2\%$ relative to EssenceBench on SWE-bench Verified. The selected minisets further reveal capability differences among different agents, providing compact and diagnostic feedback for efficient agentic model development.
|
| 791 |
Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data
2609.19153
|
cs.CLcs.LG
|
Gregory M. Dickinson |
Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines (TF-IDF features and linear classifiers) because the textual feature is often the object of study rather than a means to a predi...Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines (TF-IDF features and linear classifiers) because the textual feature is often the object of study rather than a means to a prediction. Yet these pipelines inherit preprocessing defaults from mid-century information retrieval that were never validated against classification accuracy. The most entrenched of these is stopword removal. This study introduces an exhaustive single-word ablation that measures a preprocessing step's effect directly against the downstream objective, and applies it to stopword removal. Matching Supreme Court Database labels to Caselaw Access Project opinion texts, the study examines two binary tasks, ideological direction (no-removal baseline F1 about 0.68) and constitutional versus non-constitutional law type (about 0.92), across 7,668 and 7,001 opinions. For each task, the ablation removes each of roughly 18,500 candidate words in turn, and a task-specific stoplist is built from the resulting measurements. Generic stoplists in common use fall below the no-removal baseline on held-out opinions in all twelve tests. The task-specific stoplists move held-out F1 by +0.0023 (95% CI [-0.0124, +0.0170]) on ideology and by +0.0001 ([-0.0082, +0.0085]) on law type. Neither task shows a detectable benefit from removal, and a supplemental analysis finds that word-level statistics predict a word's removal effect poorly, because the words' true removal effects differ by less than the measurement can register. The method generalizes to any inherited preprocessing default, and the result is a caution specific to interpretable legal text-as-data, where a step that reshapes which features a model sees can distort the doctrinal and ideological signal the research is meant to recover. The burden of proof sits with removal.
|
| 792 |
The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts
2609.24821
|
cs.CL
|
Manjiang Yu, Hongji Li, Zihan Wang, Junwei Chen, Xue Li |
Linear probing and activation steering use linear directions to predict and control behavior-level concepts such as correctness, safety, and social bias in question answering. We call these \emph{behavior-level concept directions}. Drawing on the Linear Repres...Linear probing and activation steering use linear directions to predict and control behavior-level concepts such as correctness, safety, and social bias in question answering. We call these \emph{behavior-level concept directions}. Drawing on the Linear Representation Hypothesis (LRH) and circuit studies, these directions are often interpreted as concept representations. Yet controlled evidence for linear concept structure and local mechanisms does not establish that these directions represent the intended behavioral concepts, leaving probing and steering without a unified theoretical account. We propose the \emph{Answer-Basin Representation Hypothesis} (ABRH): the model's own answer measure organizes the linear structure of these directions. For each question, all continuations yielding the same answer form an answer basin, whose mass is their total probability; these masses define the model's answer distribution. ABRH posits that its concentration before generation and the relative mass of each answer after generation are represented along linear directions shared across questions. Experiments span four Qwen2.5 and Gemma-3 models on correctness, social bias, and safety tasks. Decoupling concept labels from basin mass shows that probing and steering exhibit concept-consistent effects when labels align with mass orderings, weaken as mass gaps shrink, and reverse under conflict. Probes trained solely on mass orderings among wrong answers still select correct answers; a fixed steering vector can instead favor wrong answers when labels conflict with mass orderings. These results support an answer-measure account of behavior-level concept directions and of when probing and steering succeed, fail, or reverse.
|
| 793 |
LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning
2609.27009
|
cs.CL
|
Qingjing Chen, Junkai Zhang, Shaochun Wang, Jiahao Ding, Siyuan Zheng |
Large language models are increasingly applied to high-risk domains such as law, yet complex legal reasoning remains limited by two structural challenges. First, existing RAG and GraphRAG methods emphasize lexical or semantic similarity while overlooking norma...Large language models are increasingly applied to high-risk domains such as law, yet complex legal reasoning remains limited by two structural challenges. First, existing RAG and GraphRAG methods emphasize lexical or semantic similarity while overlooking normative relations among legal provisions. Second, vanilla Chain-of-Thought prompting may generate plausible rationales without enforcing the normative structure of legal reasoning. To deal with the bottleneck of pipelines in the legal reasoning domain, we propose LEGO, a dual-module framework that synergizes Legal Expert GraphRAG and expert Chain-of-thought for complex legal reasoning. ExpertGraphRAG uses an expert-annotated civil code graph encoding these normative relations with a greedy normative-coverage retrieval algorithm to dynamically extract instance-specific provision subgraphs, while ExpertCoT organizes the retrieved provisions and case facts into structured Provision-Fact-Conclusion reasoning. With a Qwen3-8B backbone, LEGO achieves 40.53% exact-match accuracy on LawExamQA_Civil, outperforming the evaluated RAG and CoT baselines and performing comparably to the evaluated larger models, while remaining robust on multi-hop questions. It also achieves the best results among the evaluated baselines on the open-ended benchmarks. Ablation studies confirm the individual and complementary contributions of both modules, demonstrating LEGO's effectiveness in improving LLMs' complex legal reasoning ability. Code and dataset can be found in the link: https://github.com/BLK-WHT/LEGO
|
| 794 |
Technical Manual for Toolkit for Confidence-Corpus Consistency, Corpus Absorption and Rule Learning via Fine-Tuning on a Fabricated Corpus
2609.28747
|
cs.CLcs.LGcs.AI
|
Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Reinaldo Cesar, Marcelo Vinicius de Paula |
This manual documents version 2.0.0 of an open toolkit for fine-tuning small causal language models on fabricated and rule-governed arithmetic corpora and measuring what they take up from them. The fact domain is the 81 additions of two single-digit natural nu...This manual documents version 2.0.0 of an open toolkit for fine-tuning small causal language models on fabricated and rule-governed arithmetic corpora and measuring what they take up from them. The fact domain is the 81 additions of two single-digit natural numbers, small enough to be enumerated exhaustively. The toolkit fine-tunes a model on the correct sums, on one fixed fabricated answer for every addition, and back on the correct sums of a subset of the additions; it fine-tunes copies of these models on simple rules (the sum plus a constant) and on a conditional rule (a shift that depends on the order of the addends), each paired with a control that has the same answers but no rule; and it measures every model on every candidate answer of every addition with one unchanged procedure, reporting results separately for additions seen in fine-tuning and additions held out. We describe and justify each stage of the pipeline: the confidence index (the probability of a complete answer, closed by an end marker), the single candidate set, the answer-only training loss, the lineage of fourteen measured models, the held-out split, the controls, the exclusion of additions that would count as hits by coincidence, and the exact and resampled intervals attached to every result. We then explain every figure and table a run produces and how each is read. This manuscript is a methodological and implementation reference: it documents the instrument, and it neither states nor tests hypotheses, nor reports or interprets the outcome of any specific run. Those are the subject of work that uses the toolkit. The toolkit and its pinned dependency environment are archived separately (Section 10) under a persistent identifier, to be cited as an instrument.
|
| 795 |
A Unified Account of Concepts and Chunks
2609.30414
|
cs.CLcs.AI
|
Karthik Singaravadivelan, Pat Langley |
Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses ...Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses a challenge for unified theories of cognition. In this paper, we review Cobweb, a computational account of categorization and concept formation, then propose an extended theory that incorporates chunks and their acquisition. The theory makes no commitments about modality, applying to any experience that decomposes into elements and relations among them. We also present \trellis/, an implementation of this theory, and illustrate its application to learning context-free grammars, which we adopt as a testbed because they involve both concept-like and chunk-like elements. In addition, we report experimental results on learning for three synthetic grammars that demonstrate the system's ability to represent syntactic knowledge, use it to parse and generate sentences, and learn compositional structures from sample parses. We conclude by discussing related work on concepts and chunks, along with directions for future research on the problem.
|
| 796 |
Yor\`{u}b\'{a} in Unicode: An Overview of a Problem
2609.33734
|
cs.CL
|
K\'ol\'a T\'ub\`os\'un |
There is a recurrent problem in the writing of Yor\`ub\'a on the internet and on the computer that has proven intractable over the years. The language, along with other African languages that depend on diacritics for disambiguation, requires a small set of pre...There is a recurrent problem in the writing of Yor\`ub\'a on the internet and on the computer that has proven intractable over the years. The language, along with other African languages that depend on diacritics for disambiguation, requires a small set of precomposed characters that Unicode does not encode. This has forced writers and digital systems to rely on combining character sequences that behave inconsistently across platforms, corrupt under font substitution, and fail in search. This paper documents that failure across a range of real world contexts, from published books to web platforms to mobile keyboards, using personal and empirical evidence. It identifies Unicode's NFC normalization stability policy as the structural constraint that prevents a straightforward fix, arguing for direct intervention of the Consortium in solving the active problem, proposing a formal encoding request for the four core Yor\`ub\'a characters as the most durable path to resolution.
|
| 797 |
LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks
2609.34187
|
cs.CLcs.AI
|
Julia Witte Zimmerman, Calla G. Beauregard, Tabia Tanzin Prama, Parisa Suchdev, Kathryn Cramer |
The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bo...The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bound of pattern reuse despite producing alluringly fluent text. We test this hypothesis using conlang-like tasks. Several LLMs are given only natural-language descriptions of fictional languages that subvert prominent superficial patterns in training data by combining statistically uncommon and unattested features. Crucially, no example outputs are given. We argue that if the models exhibit rule-following behaviour, they cannot be relying solely on superficial statistical patterns; such patterns often work against the correct output. Instead, successful performance requires representations of the constraints specified in the prompt. Across three complementary task families, models systematically move in the meaning-predicted direction: they distinguish prompt exposure from instructed use, alter semantic relationships in response to novel constraints, and sometimes produce exact matches to complex translation answer keys. Although performance varies across the spectrum of models used, these results provide evidence for meaning-mediated abstraction in LLMs and refute the strong stochastic parrot hypothesis. Our work shows that, under appropriate architectural and contextual constraints, statistical learning can produce meaning-mediated abstractions, although generation remains strongly constrained by superficial plausibility. We discuss implications for model development and for understanding how increasingly abstract representations may emerge from plausible-text-generation objectives.
|
| 798 |
CRISP: Cultural Reward Modeling for Implicit Situated Propriety
2609.34345
|
cs.CL
|
Zekun Yuan, Yangfan Ye, Baohang Li, Shuaibo Zhao, Zekun Zhou |
As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural k...As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-\(N\) experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.
|
| 799 |
RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings
2609.34556
|
cs.CLcs.LG
|
Jarod L\'evy, Mathurin Videau, Jad Yehya, Jean-R\'emi King, St\'ephane d'Ascoli |
Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language model...Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearby tokens. Existing alternatives have been evaluated under different settings, leaving the literature fragmented and without a clear replacement. We bring structure to this landscape by examining a specific weakness of RoPE: its slow frequency bands, whose wavelengths exceed the training context and expose models to unseen angles during extrapolation. We therefore introduce Data aware RoPE (DaRoPE), which preserves standard RoPE on the fast bands but replaces absolute position on the slow bands with bounded coordinates learned from contextual representations. Therefore, the slow-band geometry depends on the data rather than only on positional distance. We compare representative encodings under matched conditions across synthetic tasks, symbolic music, genomics, neural signals, and language models spanning 124M to 50B parameters. Across these experiments, DaRoPE leads on non-text benchmarks, mitigates recency bias, while remaining best or on par in language modeling and length extrapolation. Moreover, the learned coordinates also make the mechanism interpretable, revealing how attention layers leverage contextual information beyond token distance. Together, these results support DaRoPE as the best overall default among the evaluated methods, when there is no domain-specific reasons to prefer another.
|
| 800 |
Reference-Grounded Data Curation for Instruction-Following Thai-English Machine Translation
2609.34770
|
cs.CL
|
Thodsaporn Chay-intr, Krittapad Harnchang, Mahannop Thabua, Kobkrit Viriyayudhakorn, Thanaruk Theeramunkong |
Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do ...Instruction-following machine translation (IF-MT) requires respecting prompt-level rules on terminology, formatting, and register. Rule compliance typically trades off against translation quality, a tension that general-purpose IF data augmentation methods do not address. We propose Reference-Grounded Data Curation, a two-phase pipeline that extracts every supervised constraint from a reference translation that already satisfies it, ensuring feasibility by construction. Phase 1 applies Instruction-Following Difficulty (IFD) scoring to retain the hardest-but-learnable instances from an English-Thai parallel pool. Phase 2 extracts constraints from each reference target and keeps only generations satisfying every constraint, yielding the 1.97M-record Grounded dataset. We fine-tune open-weight bases on Grounded to produce ChindaMT, a Thai-English translation family at 4B, 2B, and 0.8B parameters. Under length-controlled pairwise judging, ChindaMT outperforms or matches every same-size baseline at every tier on both plain translation and under explicit rules, reaching up to a 68.4% win rate against the strongest baseline. The recipe transfers cleanly across Qwen generations. We release model weights, the Grounded dataset, and evaluation suites.
|
| 801 |
Language Models Act on Hidden Valence
2609.35591
|
cs.CL
|
Cameron Berg, Caspar Kaiser |
Language models describe some internal states as good and others as bad. But whether models have a stake in them is an open question. Simply asking models is unlikely to be informative. Any answer may be consistent with genuine introspection, superficial patte...Language models describe some internal states as good and others as bad. But whether models have a stake in them is an open question. Simply asking models is unlikely to be informative. Any answer may be consistent with genuine introspection, superficial pattern-matching, or with fixed scripts learned in character training. We therefore study revealed preference. Rather than asking about a state, we use activation steering to attach a positively or negatively valenced activation pattern to one of two otherwise meaningless 'zones', switch steering off, and then observe which zone the model prefers. A model with a stake in that state should choose accordingly. Across seven open-weight models from five families, this is indeed what we find. First, steering changes the passages models write about each zone, and those words shift later choice. Second, this shift persists when all surface-level tokens are held fixed and only the hidden KV cache differs. Third, the effect also remains when all text is generated without steering and valence is only injected during cache construction. Thus, the hidden state is sufficient to move choice in proportion to the steering dose. Fourth, the dependence of choice on hidden valence is nearly absent in a base model and emerges during DPO, consistent with a link between valence and goal-directed behaviour formed in training. Finally, given tools to self-steer, a model does not tend to induce a positive state, but it regularly removes an imposed negative state. It does so at a dose-dependent rate and significantly more often than it removes random-direction interventions. Overall we demonstrate that valence-related activation patterns leave hidden traces that predictably govern later choices, even when all visible tokens are identical across conditions. Whether these traces are accompanied by any subjective experience relevant to model welfare remains unclear.
|
| 802 |
Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV
2609.36399
|
cs.CLcs.AI
|
Bushra Asseri, Abdulaziz Asseri |
Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Modul...Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12 matched American personas and without a persona, in English and Arabic, under eight ways of formulating the request (288,000 answers). JEV's answers were highly repeatable (ICC 0.997), and without a persona they resembled those of its own American personas. When the persona was Saudi rather than American, the answers moved in the direction of the human Saudi-US difference, reproducing 87% of its size in English but 62% in Arabic, with long-term orientation reversed. A language cross shows that the smaller difference in Arabic comes from the language of the items, not from the language of the persona description. Age shifted the profiles about as much as nationality, gender shifted them more for Saudi than for American personas, and JEV was less confident in Arabic and for Saudi personas. These patterns held in every request design, although the model never generates text.
|
| 803 |
Billiger.de Products: A Bilingual Entity Matching Benchmark
2609.37713
|
cs.CL
|
Aaron Steiner, Ksenia Elagin, Ralph Peeters, Johannes Knopp, Christian Bizer |
Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark...Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark covering thirteen consumer product categories, including difficult-to-handle categories such as clothing and furniture. The benchmark data originates from the German price comparison platform billiger.de. Following the design of WDC Products, the benchmark offers multiple variants that differ in the fraction of corner cases, the size of the development set, and the fraction of entities unseen during training. An aligned English translation of every offer keeps all pairs, splits, and labels fixed, while cross-language test sets combine German and English records within individual pairs. We validate the benchmark using six supervised matchers and zero-shot GPT-5.2 on both language versions and the cross-language test sets. The validation shows the difficulty of the benchmark. The comparison of the results on the English version of the benchmark to the results on the German version shows that most matchers score on average higher on the English version. The difference is largest for RoBERTa and HierGAT, while the zero-shot LLM runs are largely insensitive to the language. Comparing the F1 scores achieved by PLM-based matchers on the English version of Billiger.de Products with their performance on existing English-language benchmarks, such as WDC Products and Abt-Buy, shows that Billiger.de Products is more difficult than these benchmarks.
|
| 804 |
GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
2609.38923
|
cs.CL
|
Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang |
Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either gene...Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal.
|
| 805 |
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
2609.39102
|
cs.CLcs.LGcs.AI
|
Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian |
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree ...Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
|
| 806 |
Explore-over-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness
2609.39786
|
cs.CL
|
Ola El Khatib, Djellel Difallah |
Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by...Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts. Alternatives that ask LLMs to generate missing knowledge risk introducing hallucinated evidence. We introduce XoG (eXplore-over-Graph), a framework for multi-hop question answering over incomplete KGs that recovers missing reasoning paths from learned graph structure rather than LLM parametric knowledge. XoG combines type-level entity-relation statistics to identify candidate relations with KG embeddings to retrieve plausible missing entities, using the LLM as a semantic selector and reasoner. These mechanisms are integrated into an iterative planning-exploration-reasoning process. Experiments on WebQSP, CWQ, and the Wikidata-based BRINK benchmark show that XoG remains competitive on complete KGs and consistently outperforms comparable methods without task-specific KGQA training under KG incompleteness. These gains persist across multiple LLM backbones, indicating that stronger LLMs alone do not resolve missing graph evidence. XoG also reduces LLM token consumption by up to 33% compared with a closely related planning-based approach.
|
| 807 |
OverdoseMoE: A Multi-Expert Framework for Opioid Overdose Risk Prediction
2609.40108
|
cs.CL
|
Mingchen Li, Rohan Pandey, Junhui Qian, Feiyun Ouyang, Sunjae Kwon |
Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Here, we investigate diagnosis-specific adaptation for 180-day opioid overdose risk prediction from patients' pr...Opioid overdose remains a major clinical and public health burden, highlighting the need for scalable approaches to identify patients at high risk. Here, we investigate diagnosis-specific adaptation for 180-day opioid overdose risk prediction from patients' preceding one-year longitudinal ICD histories. We develop OODMAMBA and OODQWEN through continued pretraining on longitudinal diagnostic sequences followed by task-specific fine-tuning. Building on the stronger Qwen-based predictors, we further propose OVERDOSEMOE, a multi-expert framework that integrates models of different scales using complementary expert-weighting strategies. Diagnosis-specific adaptation consistently improved predictive performance over general-purpose language-model baselines, with OODQWEN achieving an AUPRC of 24.47 and an AUROC of 68.56. OVERDOSEMOE further improved discrimination and precision, achieving an AUPRC of 25.17 and an AUROC of 69.49 while outperforming the strongest single-model baselines. Among patients ranked in the top 5% of predicted risk, OVERDOSEMOE identified substantially enriched overdose risk, achieving a PPV of 25.38% while retaining meaningful recall. Evaluation on an independent MIMIC-IV cohort further demonstrated cross-cohort robustness, with complementary weighting strategies showing advantages across different performance measures. These findings demonstrate that diagnosis-specific language-model adaptation combined with multi-expert integration can improve opioid overdose risk stratification and support more robust prediction across heterogeneous electronic health record populations.
|
| 808 |
Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models
2610.01082
|
cs.CL
|
Grzegorz Kulik, Miko{\l}aj Pokrywka, Adam Jatowt, Wojciech Kusa |
Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, ev...Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, evaluating on SiLTT - a new Pol-Szl testset, alongside established BOUQuET and FLORES benchmarks. Results show our rule-based system is consistently strongest on SiLTT and BOUQuET datasets and that TranslateGemma fine-tuned on a curated dataset improves over strong neural baselines but does not surpass the rule-based system in dialectal settings. We release SiLTT and our best neural model to support further research.
|
| 809 |
LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
2610.02076
|
cs.CL
|
Yinheng Li, Justin Wagle |
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs ...Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
|
| 810 |
Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise
2610.02702
|
cs.CL
|
Ziang Ni, Peng Zou |
Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose int...Multi-agent debate is increasingly used to reach consensus among LLM agents, yet agents often yield to a unanimous majority. When an agent changes its answer, has it changed its mind or only its statement? We study this with two-hop factual questions whose intermediate entity (the bridge, e.g. the country in "the capital of the country where the Sagrada Familia is located") is never stated by anyone. Scripted peers, in the role of Asch's confederates, unanimously assert a wrong answer taken from another fact with a different bridge. At the moment the agent answers, we read the bridge from its residual stream with the Jacobian lens (J-lens) and, for comparison, the logit lens. In pre-registered tests on held-out facts with four open-weight models, agents of Qwen3.5-4B, Qwen3.6-27B and Gemma-4-E4B-it that gave in still represented their original bridge in the pre-registered layers below the output (hit@100 above a control entity: 0.85, 0.22 and 0.24), where the logit lens rarely ranked it among the top 100 tokens (0.00-0.06). These agents also represented the bridge behind the peers' answer, beyond a mention baseline. A pre-registered addendum hid the agent's earlier answer or removed it: agents that gave in still represented their original bridge in all four models (0.43, 0.29, 0.37 and 0.25 with the answer hidden), including Llama-3.1-8B-Instruct, which barely did so with its answer in view (0.03). The premise can thus be computed from the question alone while the agent states the majority's answer. Hiding the earlier answer also changed conformity: Qwen3.5-4B gave in on 89% of questions instead of 8%. In exploratory interventions, injecting the bridge's J-lens direction brought agents back to their original answer only in the two Qwen models. Stated consensus in multi-agent debate can thus overstate agreement. We also report the negative results of our pre-registered program.
|
| 811 |
Clinical Concept Centers in LLMs
2610.02829
|
cs.CLcs.LG
|
Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler |
Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found ...Large language models are increasingly used in clinical settings. However, research into the reliability and performance of these models has focused almost entirely on the language substrate, scoring what the model says. Mechanistic interpretability has found that the latent space carries a higher fidelity of representation than the text: internal representations not only encode substantially more than the output verbalizes, but the stated reasoning also systematically omits features that causally drive the answer. An evaluation of model behavior in terms of mechanistic interpretability has not been explored in clinical decision support. In this work, we extend behavioral evaluation into the latent space and ask whether clinical concepts exist as locatable, causally used representations inside open-weight LLMs. We find dedicated clinical concept centers in the latent space of all eleven open models we test. These concept centers are interpretable, firing only on their aligned clinical narratives, and meaningfully and causally drive model behavior in both constrained and open-ended settings. They are not just analytical representations, but circuits that can be utilized in clinical practice, and we explore their use from the perspective of both evaluation and performance. From the evaluation standpoint, models stay internally coherent and keep using the relevant concept centers even under adversarial role-based priming, while aligned priming improves downstream clinical performance. From a performance perspective, we simulate realistic deployment settings and find that steering models along these centers leads to meaningful downstream improvements. Finally, we conduct a blinded clinician validation and find the activation and usage of these concept centers predicts clinicians preferences.
|
| 812 |
OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination
2610.02999
|
cs.CL
|
Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue |
Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated com...Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at https://github.com/RongHuiQiang/OmniConfess.
|
| 813 |
Single-Pass Uncertainty Heads for Claim-Level Hallucination Detection in Persian Medical Language Models
2610.03482
|
cs.CL
|
Mehrdad Ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani, Sadra Hakim, Audrina Ebrahimi |
Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LU...Hallucination detection is particularly important for medical language models, but repeated-sampling approaches are expensive and existing uncertainty-head resources do not directly transfer to a new backbone and language. We adapt the LLM Uncertainty Head (LUH) framework to Aya-Expanse-8B-based Persian medical models, using Gaokerena-V and Gaokerena-R as two previously developed backbones. We first examine response variability on a 168-question Iranian medical entrance examination and observe substantially lower five-run consistency for Gaokerena-V than for Aya-Expanse-8B, whereas Gaokerena-R is comparable to Aya-Expanse-8B. We then construct two paired claim-level hallucination datasets directly in Persian, containing 1,600 responses for each backbone, and train lightweight claim-level heads on frozen backbone attention maps and token probabilities. On held-out test splits, the heads obtain PR-AUCs of 0.4820 and 0.4652, corresponding to 2.30 and 2.66 times their respective random baselines, and ROC-AUCs of 0.7852 and 0.7810. The heads require neither retrieval nor repeated sampling at inference time. These results provide an initial study of single-pass claim-level uncertainty estimation for Persian medical language models; the test splits are small and the labels are automatically generated.
|
| 814 |
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
2506.24056
|
cs.CLcs.LG
|
Tung-Ling Li, Hongliang Liu |
RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce the refusal-affirmation logit gap: the difference between the top refusal-token logit and the top affirmative-token l...RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce the refusal-affirmation logit gap: the difference between the top refusal-token logit and the top affirmative-token logit at the first decoding step. This single scalar quantifies the per-prompt safety margin that alignment provides. Empirically, alignment widens the gap on 97.5-99.8% of toxic prompts across three model families, and median gap closure co-varies with True-ASR ranking across suffix strategies (an internal consistency check, since our method optimises gap closure). To validate the metric's practical significance, we present logit-gap steering, a gradient-free, forward-pass-only method that discovers short in-distribution suffixes ($<$10 tokens per component) whose cumulative effect closes the gap. The method requires ${\approx}26{,}000$ forward-pass equivalents per family (${\approx}2$~min on one A100), ${\approx}125\times$ less than a single GCG search. Suffixes discovered on 0.5B--2B models transfer without modification to 72B within family. An 8-suffix ensemble reaches 38-96\% True ASR across 13 models on AdvBench and HarmBench, with most suffixes having $10^{3}$-$10^{4}\times$ lower perplexity than GCG-meaning published perplexity-filter defenses that collapse GCG (64.7%$\to$1.0%) leave our suffixes nearly intact (76.9%$\to$76.0%). These results demonstrate that current alignment margins, while consistently present, can be thin and efficiently measurable, and that defense strategies must account for in-distribution suffixes.
|
| 815 |
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
2510.12680
|
cs.CLcs.LGcs.AI
|
Shouren Wang, Wang Yang, Xianxuan Long, Debargha Ganguly, Qifan Wang |
Hybrid thinking enables LLMs to switch between reasoning and direct answering, offering a balance between efficiency and reasoning capability. Yet our experiments reveal that current hybrid thinking LLMs only achieve partial mode separation: reasoning behavior...Hybrid thinking enables LLMs to switch between reasoning and direct answering, offering a balance between efficiency and reasoning capability. Yet our experiments reveal that current hybrid thinking LLMs only achieve partial mode separation: reasoning behaviors often leak into the no-think mode. To understand and mitigate this, we analyze the factors influencing controllability and identify four that matter most: (1) larger data scale, (2) using think and no-think answers from different questions rather than the same question, (3) a moderate increase in no-think data number, and (4) a two-phase strategy that first trains reasoning ability and then applies hybrid think training. Building on these findings, we propose a practical recipe that, compared to standard training, can maintain accuracy in both modes while significantly reducing no-think output length (from 1085 to 585 on MATH500) and occurrences of reasoning-supportive tokens such as "wait" (from 5917 to 522 on MATH500). Our findings highlight the limitations of current hybrid thinking and offer directions for strengthening its controllability. The code is available at: https://github.com/SR-A-W/demystifying-hybrid-thinking
|
| 816 |
EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning
2511.19935
|
cs.CLcs.LG
|
Songlin Zhao, Michael Pitts, Zhuwei Qin |
Deploying domain-specialized large language models on resource-constrained hardware motivates reducing both adaptation cost and the number of retained weights. LoRA lowers adaptation cost but leaves a dense backbone, while separate pruning can discard weights ...Deploying domain-specialized large language models on resource-constrained hardware motivates reducing both adaptation cost and the number of retained weights. LoRA lowers adaptation cost but leaves a dense backbone, while separate pruning can discard weights that become important during adaptation. We propose EfficientXpert, a framework that co-adapts low-rank updates and sparse support. Its ForeSight Mask scores weights through a downstream reconstruction surrogate using the evolving LoRA-augmented weights; Partial Brain Surgeon (PBS) realigns the adapter through a closed-form correction. At 40% sparsity, the strongest variant retains 98.55--108.99% of dense-LoRA aggregate performance across the evaluated LLaMA settings. On Qwen3-8B, adaptation adds 17.9--27.0% training time and at most 4.8% peak GPU memory. Our analysis connects the domain- and rank-dependent effects of PBS to its constrained reconstruction objective, providing guidance for selecting adapter recovery capacity. Together, these results demonstrate efficient production of sparse, domain-specialized experts. Code is available at https://github.com/TagoreZhao/Efficient_Domain_Adaptation.
|
| 817 |
Systematic Hazard Sampling: Minimal-Variance Inference for Discrete Diffusion and Flow Models
2601.02799
|
cs.CLcs.LG
|
Seunghwan Jang, Wonje Jeung, SooJean Han |
Uniform-noise discrete diffusion and flow models generate sequences non-autoregressively through iterative, context-dependent token replacements. However, these models are typically formulated as time-inhomogeneous continuous- or discrete-time Markov chains (C...Uniform-noise discrete diffusion and flow models generate sequences non-autoregressively through iterative, context-dependent token replacements. However, these models are typically formulated as time-inhomogeneous continuous- or discrete-time Markov chains (CTMC/DTMC), sampled using independent Bernoulli change decisions per discretization step. This induces Poisson-binomial variance in per-position jump counts that grows with the number of required edits, leading to the common under-editing (residual noise) and over-editing (cascading substitutions) failure modes that degrade sample quality. We identify this sampler-induced variance as an orthogonal source of degradation, distinct from model-side errors and addressable purely at inference time. We propose Systematic Hazard Sampling (SHS), a training-free, drop-in, and hyperparameter-free inference principle for any sampler that admits a stay-vs.-replace decomposition. SHS models per-token edits as events driven by cumulative hazard (CTMC) or jump mass (DTMC) and triggers an edit whenever this quantity exceeds unit-spaced thresholds with a single random phase per position. For any fixed cumulative mass, this preserves the expected jump count while achieving the minimum conditional variance possible among unbiased integer estimators (at most 1/4), without altering per-jump destination sampling. Experiments on four uniform-noise discrete diffusion and flow language models spanning ~110M to ~3B parameters show that SHS consistently improves sample quality across numbers of function evaluations (NFE). Code is available at https://github.com/Jang-seunghwan/Systematic-Hazard-Sampling.
|
| 818 |
When Is Enough Not Enough? Illusory Completion in Search Agents
2602.07549
|
cs.CLcs.AI
|
Dayoon Ko, Jihyuk Kim, Sohyeon Kim, Haeju Park, Dahyun Lee |
In agentic search, an LLM agent searches the web, reads the pages it finds, and decides what to look for next before returning an answer. But can we trust an answer simply because the agent returns it? Often not, and even a correct answer can be a lucky guess:...In agentic search, an LLM agent searches the web, reads the pages it finds, and decides what to look for next before returning an answer. But can we trust an answer simply because the agent returns it? Often not, and even a correct answer can be a lucky guess: on questions with several constraints, we find that agents conclude the task is complete while a constraint remains unverified in up to 48% of their correct answers. We call this illusory completion. To see how it arises, we introduce the Epistemic Ledger, which tracks at every turn what the retrieved pages establish about each constraint and what the agent claims. Across 13 agents, from 7B RL-trained models to frontier LLMs, training and scale raise accuracy but change the pattern of verification failures rather than eliminating them: constraints may be left unchecked, assumed without support, or retained despite refuting evidence. To measure what agents lose without tracking their constraints, we show them each constraint's state, approximated by LiveLedger, a lightweight 4B tracker. Agents then answer 4.4-16.1 points more questions correctly, suggesting that on their own, they may not track what they have verified and what remains.
|
| 819 |
Aligning the Query Space: Greedy Information Projection for Language Model Data Selection
2603.13790
|
cs.CLcs.LG
|
Victor Ye Dong, Kuan-Yun Lee, Jiamei Shuai, Shengfei Liu, Yi Liu |
Data selection for language models is often framed as balancing example quality and diversity. We argue that both are consequences of a more fundamental principle: selected examples should preserve the downstream query space induced by instructions, task signa...Data selection for language models is often framed as balancing example quality and diversity. We argue that both are consequences of a more fundamental principle: selected examples should preserve the downstream query space induced by instructions, task signals, or retrieval needs. We present Greedy Information Projection (GIP), a query-aligned method that requires only candidate embeddings and a score signal from LLM judgments, metadata, or intrinsic geometry. GIP greedily selects examples whose embedding span explains the largest residual component of task/query scores, yielding a fast matching-pursuit selector. A Gaussian projection view connects this update to maximizing mutual information, equivalently minimizing the residual volume of the query subspace left unexplained by selected data; this explains how quality and diversity emerge from one objective. Empirically, GIP matches or surpasses full-data fine-tuning using small instruction and reasoning subsets, while pretraining and RAG passage-selection experiments show that the same residual principle transfers beyond supervised fine-tuning.
|
| 820 |
The Ultimate Tutorial for AI-driven Scale Development in Generative Psychometrics: Releasing AIGENIE from its Bottle
2603.28643
|
cs.CLcs.AI
|
Lara Russell-Lasalandra, Hudson Golino, Luis Eduardo Garrido, Alexander P. Christensen |
Psychological scale development has traditionally required extensive expert involvement, iterative revision, and large-scale pilot testing before psychometric evaluation can begin. The \texttt{AIGENIE} R package implements the AI-GENIE framework (Automatic Ite...Psychological scale development has traditionally required extensive expert involvement, iterative revision, and large-scale pilot testing before psychometric evaluation can begin. The \texttt{AIGENIE} R package implements the AI-GENIE framework (Automatic Item Generation and Validation with Network-Integrated Evaluation), which integrates large language model (LLM) text generation with network psychometric methods to automate the early stages of this process. The package generates candidate item pools using LLMs, transforms them into high-dimensional embeddings, and applies a multi-step reduction pipeline --- Exploratory Graph Analysis (EGA), Unique Variable Analysis (UVA), and bootstrap EGA --- to produce structurally validated item pools entirely \textit{in silico}. This tutorial introduces the package across eight parts: installation and setup, text generation, embeddings, item generation, the full AI-GENIE pipeline, the GENIE pipeline for researcher-supplied items, advanced prompt engineering, and fully local operation. Two running examples illustrate the package's use: the Big Five personality model (a well-established construct) and AI Anxiety (an emerging construct). The package supports multiple LLM providers (OpenAI, Anthropic, Groq, HuggingFace, and local models), offers a fully offline mode with no external API calls, and provides the \texttt{GENIE()} function for researchers who wish to apply the psychometric reduction pipeline to existing item pools regardless of their origin. The \texttt{AIGENIE} package is freely available on CRAN at \url{https://CRAN.R-project.org/package=AIGENIE}.
|
| 821 |
Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions
2604.12138
|
cs.CLcs.AI
|
Aditya Agrawal, Alwarappan Nakkiran, Aman Singh Thakur, Alex Karlsson, Harsha Aduri |
Retrieval-Augmented Generation (RAG) systems are built on an unexamined assumption - that queries have correct answers and retrieval should converge toward them. This position paper argues that this creates a factual bias where RAG systems optimize for reducin...Retrieval-Augmented Generation (RAG) systems are built on an unexamined assumption - that queries have correct answers and retrieval should converge toward them. This position paper argues that this creates a factual bias where RAG systems optimize for reducing epistemic uncertainty while ignoring the aleatoric uncertainty, inherent in opinion-rich content. The consequences go beyond technical limitations- due to risk of minority voice erasure and risk of opinion manipulation. To address this, we formalize opinion-aware retrieval through uncertainty quantification and derive a unified objective using the Wasserstein distance. As an existence proof, we present Opinion-Aware RAG (O-RAG), which enriches documents with LLM-extracted, entity-linked opinion metadata before indexing. Across e-commerce seller forums and public hotel reviews, O-RAG reduces Wasserstein distance to corpus-level sentiment distributions by 18-48%, and human evaluators preferred its responses 79.2% of the time. We close with a research agenda for opinion-aware RAG.
|
| 822 |
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
2604.17338
|
cs.CL
|
Miaosen Chai, Wang Bill Zhu, Shangshang Wang, Yejia Liu, Song Bian |
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are from precise debugging, we introduce the P...Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate how far LLMs are from precise debugging, we introduce the Precise Debugging Benchmark (PDB) framework, which automatically converts any coding dataset into a debugging benchmark with precision-aware evaluation. PDB generates buggy programs by synthesizing verified atomic bugs and composing them into multi-bug programs. We define two novel metrics, edit-level precision and bug-level recall, which measure how many necessary edits are made and how many bugs are resolved. We release two evaluation benchmarks: PDB-Single-Hard on single-line bugs, and PDB-Multi on multi-line bugs. Experiments show that frontier models, such as GPT-5.1-Codex and DeepSeek-V3.2-Thinking, achieve unit-test pass rates above 76% but exhibit precision below 45%, even when explicitly instructed to perform minimal debugging. Finally, we show that iterative and agentic debugging strategies do not substantially improve precision or recall, highlighting the need to rethink post-training pipelines for coding models.
|
| 823 |
Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference
2605.03379
|
cs.CLcs.LG
|
Yi Liu |
Repeated sampling can improve LLM accuracy, and quantifying the gains from additional calls is essential for allocating test-time compute. We study binary decisions, a fundamental setting where repeated answers to the same question are aggregated by majority v...Repeated sampling can improve LLM accuracy, and quantifying the gains from additional calls is essential for allocating test-time compute. We study binary decisions, a fundamental setting where repeated answers to the same question are aggregated by majority vote. We show that two independently sampled responses per example in a validation set with known answers constrain the latent distribution of example-level success probabilities to a class consistent with the paired outcomes. Optimizing over this class yields sharp accuracy and gain bounds at every finite voting budget. A shared large-sample confidence region accounts for validation uncertainty. We also obtain sharp infinite-vote bounds and moment-matched forecasts of the vote-accuracy curve. On QNLI and QQP, our method successfully distinguishes settings with voting gains above a target margin from those with little room for improvement.
|
| 824 |
Is Escalation Worth It? On the Depth of LLM Cascades
2605.06350
|
cs.CLcs.LGcs.AI
|
Dylan Bouchard |
LLM cascades, in which a cheap model defers to an expensive one on low-confidence queries, are widely used to reduce inference cost. Given a pool of models, a practitioner must decide how many models to include and where to set each deferral threshold. We deri...LLM cascades, in which a cheap model defers to an expensive one on low-confidence queries, are widely used to reduce inference cost. Given a pool of models, a practitioner must decide how many models to include and where to set each deferral threshold. We derive first-order optimality conditions showing that, at an optimum, the ratio of expected accuracy gain to expected downstream cost is equal across deferral boundaries. A local search based on these conditions closely matches exhaustive search. We also derive an identity that decomposes the accuracy gain of score-based escalation over random escalation into two AUROC terms. Across five benchmarks and nine deferral scores, with model sequences and thresholds optimized from a pool of eight models, two-model cascades improve mean test-set accuracy over single-model selection by 2.1 to 8.2 percentage points. However, allowing more than two models does not improve mean test-set accuracy in 118 of 135 comparisons across scorers, datasets, and depth caps, and adds at most 0.43 percentage points. To understand the role of deferral scores in depth gains, we conduct counterfactual experiments with simulated confidence scores. When these scores have high AUROC and reflect only whether the current model answered correctly, allowing more than two models improves test-set accuracy on four of five benchmarks. However, these gains do not persist when the scores also reflect query difficulty shared across models, even at the same AUROC. These results suggest that gains from additional depth depend on how well the confidence score separates correct from incorrect answers for the current model compared with later models.
|
| 825 |
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
2605.07579
|
cs.CLcs.LGcs.AI
|
Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim |
Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training fo...Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the group mean over rollouts from the same prompt, so an accurate baseline leaves fewer distinct prompts in the batch, while PPO avoids this trade-off by training a policy scale critic, roughly doubling the cost of training. We introduce POISE (Policy Optimization with Internal State Value Estimation), a reinforcement learning algorithm that turns the model's internal states into a value model. A lightweight probe reads the signals already computed during the forward pass to predict the baseline, and is trained online alongside the policy. To preserve gradient unbiasedness, we introduce a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states. On Qwen3-4B and OLMo3-7B-Instruct-DPO across a six-domain verifiable-reward corpus, POISE outperforms other RLVR baselines while achieving more stable training. Moreover, the probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales. By leveraging the model's internal representations, POISE enables stable policy optimization.
|
| 826 |
Linking Scalar-Intensity Language to Structural Polarization with Validated Signed-Network Measures
2605.12814
|
cs.CL
|
Zhijin Guo, Li Zhang, Tyler Bonnet, Janet B. Pierrehumbert, Xiaowen Dong |
Polarization in online communities is often studied through either language or interaction structure, but the two views are rarely connected within a unified framework. Prior work has linked them by constructing interaction graphs from human judgements of agre...Polarization in online communities is often studied through either language or interaction structure, but the two views are rarely connected within a unified framework. Prior work has linked them by constructing interaction graphs from human judgements of agreement and disagreement, leaving a gap between language as observed text and structure as an engineered representation of that text. We address this gap with a language-grounded signed-network pipeline that derives signed relations directly from conversational exchanges and links window-level language patterns to structural polarization over time. Before examining this relationship, we compare spectral and frustration-based polarization measures on synthetic benchmarks and real interaction networks. We find that frustration-based measures normalized by the graph's cycle-space capacity provide a more suitable basis for comparing polarization across networks of different sizes and densities. We therefore carry forward two complementary frustration-based measures: a weighted form that incorporates stance-model confidence and a count form based on edge signs. The weighted measure provides the better-behaved structural estimate and aligns more closely with polarization measured from human-labelled interactions, while the count measure reveals a stronger relationship with language. Across monthly Reddit Brexit discussions, greater prevalence of scalar-intensity language is associated with greater structural polarization, with a similar rank-level pattern in the human-labelled network. We find little evidence that language in one month predicts polarization in the next, whereas contemporaneous scalar-intensity prevalence provides useful information about polarization within the same month.
|
| 827 |
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
2605.15726
|
cs.CLcs.AI
|
Chanuk Lee, Sangwoo Park, Minki Kang, Sung Ju Hwang |
Reinforcement learning with verifiable rewards (RLVR) is a scalable paradigm for improving the mathematical reasoning of large language models, but it is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. ...Reinforcement learning with verifiable rewards (RLVR) is a scalable paradigm for improving the mathematical reasoning of large language models, but it is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. Sampling more rollouts alleviates this at prohibitive compute cost, while objective-level modifications offer little control over what is explored. We propose NudgeRL, a framework for structured, diversity-driven exploration in RLVR whose core component, Strategy Nudging, conditions each rollout on a lightweight strategy-level context, without requiring the context generator to solve the problem itself. To learn from such exploration, we decompose the advantage into inter- and intra-context terms and add a policy correction term that transfers discovered behaviors back to the base policy. Across five mathematical reasoning benchmarks, NudgeRL with 8 rollouts matches the strongest GRPO baseline using 32 rollouts with roughly 5 times fewer total tokens and 2.9 times less training compute. Our code is available at https://github.com/tally0818/NudgeRL.
|
| 828 |
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
2605.18648
|
cs.CLcs.LGcs.AI
|
Maja Pavlovic, Silviu Paun, Massimo Poesio |
Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibration by capturing uncertainty, prior studies conflate these benefits with the implicit correction of mislabel...Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibration by capturing uncertainty, prior studies conflate these benefits with the implicit correction of mislabeled data (mode shifts), obscuring true effects of soft-labels. We present a controlled audit of soft-label learning across MNIST and a synthetic variant, re-annotating subsets to extract human uncertainty. By decoupling soft-label supervision from underlying label mode shifts, we show that while human soft-labels do provide accuracy gains, their larger value lies in acting as a regularizer that improves model calibration on difficult samples and promotes stable convergence across training runs. Dataset cartography reveals models trained on human soft-labels mirror human uncertainty, whereas those trained on synthetic labels fail to align with humans. Broadly, this work provides a diagnostic testbed for human-AI uncertainty alignment.
|
| 829 |
Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
2605.22902
|
cs.CLcs.LGcs.AI
|
Dimitrios Damianos, Leon Voukoutis, Georgios Skyrianos, Vassilis Katsouros, Georgios Paraskevopoulos |
Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability work on VLMs uses Sparse Autoencoders (SAEs), which decompose static residual repr...Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood. Existing interpretability work on VLMs uses Sparse Autoencoders (SAEs), which decompose static residual representations and miss the functional updates that drive cross-modal interaction. We adopt a function-centric framework based on Transcoders, sparse approximations of MLP sublayers that act as a causal proxy for layer-wise computation. Applied to Gemma 3-4B-IT, the framework decomposes the model into interpretable computational pathways linking image patches to directions in token generation. Transcoder attributions produce stronger and more stable effects on visually grounded tokens under patch ablation than SAE attributions, and align better with semantically relevant image regions. A False Visual Grounding counterfactual analysis confirms that the recovered pathways are specific to vision-language interaction.Finally, we perform a structural analysis of hallucinated generations, by extracting graph-based indicators from circuit traces produced by the transcoders. A logistic classifier over these mechanistic graph features predicts hallucinations at AUC $0.68$. These results show that function-centric circuit decomposition yields interpretable and predictive accounts of multimodal computation in VLMs.
|
| 830 |
LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training
2606.07610
|
cs.CLcs.LGcs.AI
|
Argyrios Gerogiannis, Yekaterina Yegorova, Mark Hasegawa-Johnson, Venugopal V. Veeravalli |
State-of-the-art GRPO-style methods for speech-aware large language model post-training suffer from coarse credit assignment, broadcasting the same terminal-reward advantage to every token in a response. This ignores useful structure within rollout batches, wh...State-of-the-art GRPO-style methods for speech-aware large language model post-training suffer from coarse credit assignment, broadcasting the same terminal-reward advantage to every token in a response. This ignores useful structure within rollout batches, where speech-conditioned completions often share prefixes before diverging at important decisions. We propose Low-rank Exploration with Adaptive Forking (LEAF), a retrospective tree-based RL method that recovers this structure without online branching or additional decoding. LEAF samples complete responses, selects high-surprisal boundaries, groups responses by shared prefixes, and assigns span-level advantages using descendant rewards. We theoretically justify LEAF's span-level credit assignment and boundary-selection design. Empirically, LEAF improves over GRPO across speech question answering and speech translation benchmarks under the same rollout and low-rank adaptation budget. Notably, smaller LEAF-trained models consistently outperform current state-of-art, full-parameter baselines.
|
| 831 |
Scaling Participation in Modular AI Systems
2606.07812
|
cs.CLcs.AI
|
Shangbin Feng, Yike Wang, Weijia Shi, Luke Zettlemoyer, Yejin Choi |
Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. Yet the LLMs used by all are built by the few -- a centralized market of monolithic AI models structurally ill-suited to capture the diversity of h...Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. Yet the LLMs used by all are built by the few -- a centralized market of monolithic AI models structurally ill-suited to capture the diversity of human knowledge, reasoning, and values. Here we introduce scaling participation, a new paradigm in which modular, community-sourced AI systems are built from the bottom up through the contributions of diverse stakeholders. Participants contribute small models trained on their own interests and priorities; these models then collaborate in modular frameworks as compositional AI systems, repurposing existing collaboration algorithms for this bottom-up paradigm. Participatory AI systems outperform monolithic LLMs by up to 15.42% (95% CI: [10.09%, 21.13%]) across 15 tasks, such as reasoning and factuality, surpassing models with more parameters than all contributed components combined. Further experiments show that these systems are especially strong at representing diverse cultures, values, and communities, benefit from contributor diversity, substantially improve on each contributor's original priorities, and exhibit emergent capabilities that allow them to solve over 15% of problems where all individual models fail. Scaling participation provides a technical foundation, demonstrated here with academic contributors and benchmark evaluations, for transitioning from the monolithic status quo toward an open, bottom-up, and collaborative AI future.
|
| 832 |
DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models
2606.23181
|
cs.CLcs.AI
|
Jungseob Lee, Seongtae Hong, Seungjun Lee, Jaehyung Seo, Junyoung Son |
Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to finish the answ...Hybrid reasoning models can answer directly or spend extra tokens on extended thinking. A practical router should choose between these modes for each query, so easy problems avoid unnecessary reasoning and hard problems receive enough budget to finish the answer. Existing routers move in this direction, but they typically require labeled training data or fix thinking budgets up front, ignoring answer-level evidence from the model itself. We introduce DART, a training-free routing framework that samples two cheap no-think drafts, accepts direct answering when the drafts agree, and predicts a thinking budget from draft entropy when they disagree. Across the main comparisons, DART preserves or improves always-thinking accuracy in most settings while reducing thinking-token use. Accuracy improves by up to +9.0 points on Olympiad-level math and by up to +22.5 points on code under execution-based equivalence, while thinking-token use drops by 32-73%. The Stage~1 signal extends across model scales (0.6B--32B), model families, and API-only hosted settings, with no labeled data and no gradient updates required. Our code is available at https://github.com/js-lee-AI/DART.
|
| 833 |
MaDI-Bench: An End-to-End Data Integration Benchmark
2606.30371
|
cs.CL
|
Aaron Steiner, Ralph Peeters, Christian Bizer |
Data integration is the process of combining data from multiple, heterogeneous sources into a consistent, unified representation. Data integration involves a sequence of interdependent tasks including schema matching, value normalization, blocking, entity matc...Data integration is the process of combining data from multiple, heterogeneous sources into a consistent, unified representation. Data integration involves a sequence of interdependent tasks including schema matching, value normalization, blocking, entity matching, and data fusion. Existing table-based benchmarks either evaluate these steps in isolation or cover only incomplete versions of the data integration pipeline, omitting specific steps. The lack of public end-to-end data integration benchmarks hinders research on data integration methods that address the integration process as a whole and account for the interdependencies among the different tasks. This paper fills this gap by introducing the Mannheim Data Integration Benchmark (MaDI-Bench), the first benchmark for the end-to-end integration of relational tables covering all steps of the integration process. MaDI-Bench contributes (i) a set of end-to-end data integration tasks spanning several application domains, each requiring the full schema matching, value normalization, entity matching, and data fusion pipeline, and (ii) a generic method for deriving task variants that mitigates rapid benchmark saturation as data integration systems advance. We validate the benchmark using human-engineered pipelines, a best-of-breed pipeline, an LLM workflow, and a pipeline written by a coding agent. The validation demonstrates the utility of the benchmark for measuring the step-wise as well as the end-to-end performance of data integration pipelines. All benchmark artifacts are available for public download.
|
| 834 |
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
2606.31002
|
cs.CLcs.AI
|
Ke Zhang, Patricio Gallardo Candela, Sudhir Murthy, Yi Xie, Zhi Wang |
Lean verifies that a generated declaration is well typed, but not that it states what the user intended. For statement autoformalization without canonical Lean targets, we study two questions: how far an LLM-based semantic criterion can be trusted, and how muc...Lean verifies that a generated declaration is well typed, but not that it states what the user intended. For statement autoformalization without canonical Lean targets, we study two questions: how far an LLM-based semantic criterion can be trusted, and how much compilation overstates faithfulness across systems. Our criterion requires Lean compilation and agreement of GPT-5.2 and Gemini-2.5-Pro. On an independently audited random sample, it agrees with the human majority on 91.5\% of cases (Wilson 95\% CI: 81.6--96.3\%); humans confirm 95.7\% of the outputs it accepts and 77--81\% of the outputs it rejects, so the criterion is reliable in aggregate and errs on the conservative side. A comparison with LeanScorer, an independent third-family judge, and a BEq formal cross-check support the same picture. Across eight systems evaluated on 227 graduate-level statements, the compile--faithfulness gap varies widely with the system, from 1.3 percentage points for one-shot Gemini-2.5-Pro to 29.5 points for a tool-augmented GPT-5.2 agent, which compiles 87.2\% of statements but is faithful on 57.7\%. A $2^3$ tool ablation of this agent shows that Lean feedback drives most of its gain in compilation, and nearly half of that gain consists of outputs that fail the semantic criterion.
|
| 835 |
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
2607.01774
|
cs.CLcs.AI
|
Maximo Eduardo Rulli, Thomas Vaitses Fontanari, Simone Petruzzi, Federico Alvetreti, Giorgio Strano |
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally re...Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can be reliably extracted using probes across layers, indicating that denoising progress is decodable from internal activations. We further demonstrate that steering the model along a low-dimensional subspace associated with the inferred timestep allows us to systematically modulate its notion of denoising progress, leading to predictable changes in model confidence and entropy. Finally, we analyse the geometry of the identified representation, showing that it exhibits structured and interpretable properties in activation space, and shedding light on how such a signal is processed by these models.
|
| 836 |
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
2607.02781
|
cs.CLcs.LGcs.AI
|
Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum |
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safet...Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single scalar score, so explicit safety constraints must either be ignored or encoded through manually tuned penalties. We propose Lagrangian Reward Augmentation (LARA), a general inference-time alignment framework under safety constraints. Starting from a KL-regularized constrained objective with a reward model and a cost model, LARA dualizes the constraint and reduces the optimization problem to a one-dimensional convex problem over a nonnegative dual variable. Estimated on a small calibration set, this dual variable defines an augmented reward that can be used as a drop-in scoring signal within existing inference-time alignment methods. For sequence-level sampling methods, such as Best-of-N reranking, the calibrated dual variable corresponds to the solution of the expected-cost constrained problem. For token-level reward-guided decoding methods, the same construction yields a principled dual-calibrated heuristic rather than an exact constrained-policy guarantee. We evaluate LARA on both sequence-level and token-level inference-time alignment methods, and find that LARA improves the helpfulness-harmlessness tradeoff, with Best-of-N achieving the best performance among inference-time methods, approaching finetuning-based direct alignment baselines.
|
| 837 |
Geometric Self-Distillation for Reasoning Generalization
2607.06855
|
cs.CLcs.LG
|
Josip Juki\'c, Ivan Titov |
On-policy distillation provides dense teacher supervision on a language model's own trajectories. In self-distillation with privileged context, this supervision comes from the model itself, conditioned on a hint or solution trace hidden from the student. When ...On-policy distillation provides dense teacher supervision on a language model's own trajectories. In self-distillation with privileged context, this supervision comes from the model itself, conditioned on a hint or solution trace hidden from the student. When the teacher's preferences hinge on privileged information, it can assign higher probability to continuations the student cannot infer from its own context. Matching these preferences throughout training can induce predictive drift and degrade out-of-distribution (OOD) reasoning. We propose GeoSD, a self-distillation method that controls this drift through two complementary geometric terms. A Hellinger loss weights each teacher preference by the student--teacher overlap, reducing the influence of tokens to which the student assigns low probability. Because these influences can still accumulate, a Fisher--Rao penalty regulates predictive distance from a copy of the student refreshed periodically during training. Both terms compare next-token distributions in Fisher--Rao geometry and are jointly optimized with a preconditioner motivated by the natural gradient. Across three model families, GeoSD retains strong in-distribution gains while improving average mathematical OOD accuracy by 5.7--8.6 points over the base model. OOD gains hold across five model scales from 1.7B to 32B and transfer to code generation, where GeoSD improves code accuracy by 1.9 points on average despite distilling on mathematics alone. Our analysis of mathematical reasoning shows that standard matching rapidly concentrates probability mass at high-entropy states and that its samples confidently agree on incorrect answers. In contrast, GeoSD preserves alternative token mass and reduces false consensus.
|
| 838 |
Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset
2607.27420
|
cs.CLcs.AI
|
Mayank Sharma, Savira Nadela, Tyler Matteson |
Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether...Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find that HLE is dominated by a single general reasoning factor: domain labels explain only 3.5\% of variance in the leading principal components of item responses, and within- and between-domain residual correlations are nearly identical (Cohen's $d = 0.016$). A simulation study with the same number of models ($N = 29$) shows that our analyses would detect clearly distinct domain abilities, so their absence is informative; however, small domain-specific differences cannot be ruled out, and model rankings vary across domains somewhat more than a single factor predicts. A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above $\theta = 0$, where the strongest models sit. These findings suggest that the subset's domain subscores do not warrant distinct capability interpretations and that its ability to discriminate among the strongest models might be limited.
|
| 839 |
Which Decisions Low-Bit Quantization Breaks, and How to Predict Them
2608.06564
|
cs.CLcs.LG
|
Zekun Wu, Swati Dhiman, Adriano Koshiyama |
Quantization saves memory by storing model weights with fewer bits. It can also change model decisions, such as whether to call a tool or which option to choose from a finite set. We study these decision changes in 16 language models from 8 families at 4, 3 an...Quantization saves memory by storing model weights with fewer bits. It can also change model decisions, such as whether to call a tool or which option to choose from a finite set. We study these decision changes in 16 language models from 8 families at 4, 3 and 2 bits, across several post-training quantization settings. Our evaluation covers tool use, safety, general knowledge and social bias, using BFCL, XSTest, MMLU, BoolQ, BBQ and synthetic tasks. The decision margin is the score difference between two possible first tokens, measured before and after quantization. Writing the margin before quantization as $m$ and the margin after quantization as $m'$, we find an approximately linear relationship across decisions: $m' \approx c m + b$. The slope $c$ is usually below one and becomes smaller as precision falls, so quantization progressively shrinks decision margins. The offset $b$ is the same for every decision of one kind. Quantization therefore does not simply add random noise, and even a strong preference at full precision can flip. Quantization also affects different kinds of decisions to different degrees. Within tool use, whether to call a tool is often more sensitive than which tool to call: on 400 BFCL tasks, three of five models lose more completed calls than correct tool selections at 3-bit round-to-nearest. Under GPTQ and GGUF far fewer whether-to-call decisions flip than under plain rounding, so there is no single 3-bit failure point. The same relationship predicts how often decisions flip. Across 1,082 combinations of models, quantization settings, bit-widths and decision types, we fit the slope, the offset and the spread around the fitted line on half of the decisions and predict the flip rate on the other half. The predicted flip rate differs from the observed flip rate by a median of 1.0 percentage point, while reusing the flip rate of the first half misses by 1.3.
|
| 840 |
Vision-Language Model Confidence Is Not a Property of the Answer
2608.06571
|
cs.CL
|
Reza Khanmohammadi, Erfan Miahi, Ivan Brugere, Simerjot Kaur, Charese H. Smiley |
Vision-language models are increasingly deployed behind a confidence gate: the system reads how confident the model is in its answer and defers when confidence is low. This makes the confidence signal itself worth attacking. We show that a white-box adversary ...Vision-language models are increasingly deployed behind a confidence gate: the system reads how confident the model is in its answer and defers when confidence is low. This makes the confidence signal itself worth attacking. We show that a white-box adversary who perturbs only the input image, within an L-infinity budget of 8/255 and while keeping the model's answer byte-identical, can invert the confidence ranking, lowering it on correct answers and raising it on wrong ones until the signal points the wrong way. Most of the inversion persists even when the whole next-token distribution is held near the clean one, so the answer does not determine the confidence attached to it. Confidence is a separate signal read from the same network, and it can be corrupted on its own. A gate reading it is turned against itself, rejecting good answers and accepting wrong ones it was built to catch. Across four vision-language models and three visual question-answering benchmarks, the attack drives the model-internal readouts below chance in 83 of 84 readout-by-cell profiles under an adversary that knows which answers are correct; for the two readouts carrying a disjoint calibration reference, it falls below chance under an adversary that does not. Training a probe on frozen hidden states does not fix this: the robustness it gains is paid for with the information that made it useful. Nor does reading confidence from a separate, independently trained model, which holds up only until the attacker reaches it and then falls into the same regime. How far an answer-preserving adversary can reach a signal governs where it survives; whether a robust and informative readout can be built remains open. For deployment, a gate under this attack can admit most wrong answers it would otherwise catch and, corrected for how often the model is wrong, can leave the system worse off than using no gate at all.
|
| 841 |
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
2608.23264
|
cs.CLcs.AI
|
Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet |
Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethica...Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the model places greater emphasis on benign task-framing tokens (e.g., "Can you help me...") than on tokens signaling the underlying unethical behavior (e.g., "without getting caught"), which we term cue-tokens. We hypothesize that this under-attribution contributes to harmful compliance. To test this, we introduce two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens. Empirical evaluations show that these interventions promote safer responses, supporting cue-token attribution's role in compliance failures.
|
| 842 |
Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning
2609.09030
|
cs.CLcs.LGcs.AI
|
Mar Gonz\`alez I Catal\`a, Haitz S\'aez de Oc\'ariz Borde, Davide Murari, Carola-Bibiane Sch\"onlieb, Pietro Li\`o |
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation us...Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves over the reasoning process but do not reveal which competing hypotheses account for that uncertainty. We introduce answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds. As a strictly finer representation than endpoint and entropy summaries, answer-distribution trajectories enable us to characterize a trace through a dynamical reasoning profile spanning exploration, revision, motion, and commitment, and to distinguish different dynamical mechanisms of reasoning success and failure. Across sixteen open-weight language models and four reasoning benchmarks, we show that traces with the same endpoint and similar entropy profiles can exhibit substantially different reasoning dynamics. We further find substantial variation in these dynamics both within and across models and tasks, with different objectives favoring different dynamical profiles. Additionally, we show that training and inference choices systematically reshape these profiles. Our results suggest that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reasoning.
|
| 843 |
From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls
2609.09476
|
cs.CLcs.LGcs.AI
|
Hamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi Nia |
In-vehicle assistants must translate natural-language requests into accurate vehicle function calls under strict memory and latency constraints, making small language models (SLMs) attractive for on-device deployment. For such models, a key design choice is ho...In-vehicle assistants must translate natural-language requests into accurate vehicle function calls under strict memory and latency constraints, making small language models (SLMs) attractive for on-device deployment. For such models, a key design choice is how the available function surface is presented. Two approaches are to represent each function with a dedicated Functional Token (FT) or provide function schemas directly in the prompt. FTs enable compact inference but are restricted to functions learned during training, whereas Schema-in-Prompt (SIP) can generalize to unseen functions at the cost of longer prompts and higher inference overhead. We introduce a benchmark of 9,822 single-turn examples spanning 79 vehicle functions derived from Android Automotive, including held-out functions and requests requiring refusal. We compare both approaches under matched fine-tuning across four SLMs from 270M to 1.7B parameters. On functions seen during training, scaling provides limited benefit: the 270M model can match the 1.7B model, and the highest Seen accuracy in our main grid occurs at 0.6B. On held-out functions, FT achieves zero accuracy by construction, whereas SIP generalizes and improves substantially with scale. On out-of-scope requests, FT can invoke an unavailable function it was trained to emit, while SIP more reliably refuses based on the functions offered. This flexibility comes with a higher inference cost, chiefly the latency of processing the schema prompt. Our theoretical analysis formalizes why only SIP can predict functions withheld as training targets and why longer schema contexts increase inference cost. Overall, function-surface representation, rather than model scale alone, determines the capabilities and failure modes of SLM-based vehicle function calling.
|
| 844 |
SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration
2609.24303
|
cs.CLcs.LG
|
Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, Jos\'e Miguel Hern\'andez-Lobato |
Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can...Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding PLM provides a natural label-free reference for post-hoc calibration. Prior agreement-gated PLM-referenced calibration fits a scalar temperature using only examples on which the PoLM and its PLM reference agree, excluding disagreement examples because direct alignment can drive the fitted temperature excessively high and induce under-confidence. We revisit this binary treatment. A controlled reintroduction diagnostic reveals a non monotonic aggregate effect: admitting a moderate fraction of disagreement examples can improve calibration, whereas the benefit diminishes as unit weight inclusion approaches the full disagreement set. We introduce SupportCal, a label-free post-hoc method that retains agreement examples at unit weight and assigns disagreement examples continuous weights based on the own-base PLM's relative support and corroboration from pretrained references selected from a size-compatible candidate pool. We further characterize when the resulting weighted objective admits a finite optimal temperature. Across MedMCQA and MathQA, SupportCal yields lower mean ECE than the agreement-only baseline for nearly all evaluated target-model configurations; supplementary TweetEval Sentiment results show the same pattern on a fixed-label classification task.
|
| 845 |
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
2609.24972
|
cs.CLcs.LGcs.AI
|
Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhuang |
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selec...An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 6.0 points on the split it evolves against and up to 4.3 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
|
| 846 |
Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices
2609.25645
|
cs.CLcs.LG
|
Qian Xie, Yueli He, Nairen Cao |
Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval, which draws on the Bayesian-opt...Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval, which draws on the Bayesian-optimal Gittins policy to determine which configuration to evaluate next and when to stop. We extend the policy with an anytime recommendation rule over both fully and partially evaluated configurations, using an LCB-style score to account for posterior uncertainty. GittinsEval is computationally efficient, requiring only lightweight online updates after offline precomputation. Across GSM8K, PIQA, AlpacaEval, and MMLU response matrices, GittinsEval is consistently competitive, with particularly strong gains over configuration-level Bayesian optimization on large-example benchmarks and over cost-unaware bandit baselines on large-candidate tasks. Crucially, GittinsEval often attains near-zero simple regret using only 1% to 2% of the exhaustive-evaluation cost; it also offers an adaptive stopping rule that typically triggers at 1% to 10%.
|
| 847 |
Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing
2609.31513
|
cs.CL
|
Marco Mandap (Bulacan State University) |
The decision problem is whether valence among the users who write reviews of an app has changed enough to warrant investigation of a release, outage, or support issue. We propose a statistical specification for that decision, applied to written Google Play rev...The decision problem is whether valence among the users who write reviews of an app has changed enough to warrant investigation of a release, outage, or support issue. We propose a statistical specification for that decision, applied to written Google Play reviews collected under a declared locale, language, sort order, retrieval cap, and window rule. That endpoint returns a selected collection rather than a random sample, so the index describes review writers under the declared protocol and not all app users. What is new is not a sentiment classifier but an auditable interface: an ordinal threshold model is the preferred representation of star ratings, a calibrated two-signal generalized least-squares estimator is retained as an interpretable linear approximation when held-out labels support its conditional-unbiasedness assumption, the estimand is declared before the estimator, and every downstream module--uncertainty, diagnostics, and regularization--is tied to that declaration. The framework separates the latent average of a fixed collected set from the mean of a protocol-defined population of review writers, so measurement and between-review variation are not conflated, and treats helpfulness and recency weights as a different policy target. Shrinkage, cross-app empirical Bayes, histogram and dependence diagnostics, and dynamic smoothing are specified as separate modules with defaults and mandatory sensitivity analyses. A reference implementation, a declared 6000-review three-app collection, and a synthetic stress test exercise the displayed equations: fusion reduced error in clean and small-window scenarios, coordinated contamination caused severe bias and interval failure, and the collected sample reproduces the abstention rules. No real-review validation is claimed.
|
| 848 |
Activation Flow: Manufacturing Activations for Steering
2609.32530
|
cs.CLcs.LGcs.AI
|
Hong Kiat Tan, Linh Le, David Williams-King |
Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from $k$ corr...Difference-in-means steering requires activations recorded while a model shows the desired behavior, which a sandbagging model withholds by deliberately underperforming. We introduce Activation Flow (ActFlow), which manufactures these activations from $k$ correct labels without fine-tuning. ActFlow sets target logits that rank each labeled item's correct answer first, and moves the logits toward them by adding one vector $x$ to all $k$ residual streams at one layer. ActFlow is a family of ordinary differential equations for $x$, one for each rule that maps the required logit change to the velocity of $x$. The smallest-norm rule lands exactly on the targets, while the others keep only the top singular directions of the Jacobian. We test ActFlow on three instruction-tuned models, each locked by a sandbagging prompt and by a password-locked LoRA. At $k=40$, ActFlow keeping five singular directions raises the mean held-out ARC-Easy accuracy over the six locked models from $0.05$ to $0.85$, against $0.88$ for rank-$16$ LoRA fine-tuning, which trains about $3{,}000$ times as many parameters, and $0.92$ for the honest models. Furthermore, it scores higher than the smallest-norm rule in 16 of the 18 combinations of locked model and $k$, and its steering direction is nearly orthogonal to the honest difference-in-means direction. It also unlocks two LoRA locks where the honest direction fails. Surprisingly, with one labeled item, the single-step smallest-norm rule raises the Qwen2.5-7B prompt lock from $0.04$ to $0.59$.
|
| 849 |
Jev in Medicine: A Benchmark Evaluation
2609.34024
|
cs.CLcs.LGcs.AI
|
Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho |
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Je...Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
|
| 850 |
Loop Dropout: Regularizing Shared Updates in Looped Language Models
2609.34218
|
cs.CLcs.LG
|
Zirui Zhu, Hailun Xu, Xuanlei Zhao, Yong Liu, Yingxuan Ren |
Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our ...Looped language models separate computational depth from parameter count by repeatedly applying the same transformer block. Adapting these models requires a shared update that remains effective as hidden states evolve throughout the recurrent computation. Our empirical analysis reveals a pronounced late-loop bias in standard low-rank adaptation (LoRA): the shared update provides limited adaptation at early loop positions. This imbalance motivates training shared updates under varying combinations of their applications. Randomly omitting adapter applications alone, however, does not improve task performance; it reduces expected update strength during training while leaving inference unchanged. We introduce Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. Extensive experiments demonstrate improved mathematical reasoning across model sizes, adapter ranks and training recipes, with benefits extending to general instruction tuning and code generation. Loop Dropout outperforms existing LoRA variants and adapter regularizers, while further analysis shows stronger early-loop adaptation. Every backbone loop remains active, and inference applies the adapter at all loops using standard LoRA without additional trainable parameters or inference computation. Code is available at https://github.com/NUS-HPC-AI-Lab/loop-dropout .
|
| 851 |
VOSSA: Voiceprint Optimization for Streaming Speech Architectures
2609.38887
|
cs.CLcs.LGeess.AS
|
Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna |
Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic vari...Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.
|
| 852 |
RPTune: Learned Context Curation for LLM Catalog Search
2610.00964
|
cs.CLcs.LG
|
Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan |
For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catal...For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensure that the model can use it effectively, since LLMs do not exploit long contexts uniformly. We therefore study in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. We propose RPTune, an end-to-end framework that couples learned catalog curation with LLM post-training using automatically generated, catalog-grounded supervision. An encoder-reorganizer curator orders and prunes products guided by downstream LLM feedback, while the resulting curated catalogs in turn improve the effectiveness of LLM post-training with a context-relative reward. We evaluate RPTune on 7 real merchants spanning distinct retail verticals, using 100 complex conversational queries per merchant. RPTune consistently improves search accuracy across both proprietary and open-weight LLMs, with context curation yielding gains of up to 31.4 percentage points and post-training adding a further 10.3 points on average.
|
| 853 |
Pay for the Fault, Not the Flow: Label-Free In-Flow Multi-Agent Workflow Optimization
2610.01017
|
cs.CLcs.AI
|
Xuehang Guo, Haoyu Wang, Shengyu Chen, Zach Chen, Wei Cheng |
Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust wi...Large language models (LLMs) increasingly construct multi-agent workflows that decompose a complex task and assign specialist agents from a pool. However, building such a workflow well remains challenging: how finely to divide the task, which agent to trust with each subtask, and when to create a new specialist are all critical decisions a workflow constructor needs to settle up front. Thus, whether each subtask succeeds remains unknown until the workflow runs. Yet, improving a workflow is costly. Locating a fault usually requires a reference answer, a graded outcome, or a trained assessor, and the fix is applied to the whole workflow through re-execution, re-search, or retraining. We propose InFlowOp, which prices every decision in one label-free cost that weighs how well an agent's competence meets what a subtask demands against how much that agent takes to run. Before execution, InFlowOp bidirectionally determines the granularity of task decomposition and agent assignment following from the cost rather than from a fixed template. During execution, InFlowOp corrects a fault with the cheapest move via the same cost that serves the workflow both as it is built and as it runs. Facing the workflow-level evaluation challenge, we introduce Braid, a benchmark whose tasks require multi-agent coordination beyond single-agent capability. Across various domains and backbones, InFlowOp outperforms single agent baselines by up to $+11.97\%$, achieving $+9.64\%$ with in-flow optimization. Our project page: https://xhguo7.github.io/InFlowOp/.
|
| 854 |
Multilingual GSM-Symbolic: What determines capability transfer across languages?
2610.03367
|
cs.CLcs.AI
|
Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart, Simon Enni |
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transf...We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size ($\beta = 1.77$), language resource level ($\beta = 0.77$), reasoning ($\beta = 0.67$) and typological distance ($\beta = -0.25$). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages ($\beta = -0.27$ and $\beta = -0.20$, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.
|
| cs.CV 451 papers | ||||
| 1 |
Fractal Cross Product: Theory, Differentiable Implementation and Application to Medical Image Analysis
2610.03755
|
cs.CV
|
Noaman Khan, Nihad Hadj Sahraoui, Samir Brahim Belhaouari |
The magnitude of the generalized Euclidean cross product is a Gram volume whose degree under common scaling is fixed by the integer dimension of the spanning frame. We formulate a generalized Fractal Cross Product (FCP) as a nonlinear radial deformation with a...The magnitude of the generalized Euclidean cross product is a Gram volume whose degree under common scaling is fixed by the integer dimension of the spanning frame. We formulate a generalized Fractal Cross Product (FCP) as a nonlinear radial deformation with a prescribed positive degree $D$, which may be non-integer. The scalar construction applies in any ambient dimension $m\geq k$, while its canonically oriented vector form requires codimension one. It recovers the classical generalized cross product exactly at $D=k$ and retains orthogonality, alternation, rotation equivariance, and $D$-homogeneity, but is generally not multilinear. For exact self-similar frame systems, the construction also obeys a scale-balance law at the similarity dimension. We derive a differentiable, dimensionless image response and a non-circular empirical accumulation exponent obtained by regressing raw angular Gram responses across patch widths. Binary64 calculations recover the finite-frame identities to roundoff, while raster experiments recover the Sierpi\'nski-triangle value 1.5849625 at three resolutions. In five-seed medical-imaging comparisons, FCP-centered fusion increased mean area under the receiver operating characteristic curve from 0.7264 to 0.8135 and from 0.5843 to 0.6765 on the random and hospital-separated Retinal Image Database for Optic Nerve Evaluation partitions, respectively, and from 0.7403 to 0.7475 on FracAtlas. On FracAtlas, balanced accuracy increased from 0.6186 to 0.6721 and the harmonic mean of precision and sensitivity from 0.3669 to 0.4457. These results support the utility of the complete fusion framework, but do not isolate the effect of FCP from that of its complementary descriptors and fusion head.
|
| 2 |
Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?
2610.03765
|
cs.CVcs.CLcs.AI
|
Varsha Suresh, Divij Jain, Jia Liu, M. Hamza Mughal, Vera Demberg |
Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstr...Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstruction-trained codebook using 19 co-speech gesture descriptors spanning from raw motion to abstract communicative function. Results show that geometry and handedness are readily decodable from token embeddings, while motion category is only weakly decoded despite showing systematic differences in discrete code usage. This gap between reconstruction quality and descriptor decodability suggests that reconstruction objectives alone do not guarantee that gesture semantics are captured, and that evaluating codebooks on such properties can guide the design of more semantic motion tokenizers.
|
| 3 |
LoRA Direction Extraction for Controllable Light Toggling in FLUX.1 Kontext
2610.03771
|
cs.CVcs.LG
|
Petr Golenderov, Dmitry Mazyar, Natalia Sovpel, Alexander Aksenov |
We propose a fine-tuning method for flow-matching diffusion models aimed at realistic artificial light modeling without the need for a large training dataset. We address the task of controllable interior image editing, where the goal is to turn artificial ligh...We propose a fine-tuning method for flow-matching diffusion models aimed at realistic artificial light modeling without the need for a large training dataset. We address the task of controllable interior image editing, where the goal is to turn artificial light sources on or off while preserving the scene geometry, object placement, materials, and visual identity of the original image. To achieve this, we decompose the task into two independent formulations. We introduce the LoRA Direction Training Method, which extracts the pure direction of the LoRA adapter effect in the diffusion model flow field, and we also introduce specialized loss functions to ensure the realism of the inverse transformation. Additionally, the resulting increment map is used for more precise adjustment of the lighting color and temperature.
|
| 4 |
Energy Variation in Training Modern Computer Vision Architectures
2610.03772
|
cs.CV
|
David Cortes, Carlos Juiz, Belen Bermejo |
The rapid growth of deep learning has substantially increased the energy consumption associated with model training, making energy efficiency an increasingly relevant design criterion. This study empirically measures the energy variation of training seven mode...The rapid growth of deep learning has substantially increased the energy consumption associated with model training, making energy efficiency an increasingly relevant design criterion. This study empirically measures the energy variation of training seven modern computer vision architectures, MobileNetV3-Small, MobileNetV3-Large, EfficientNet-B0, EfficientNet-B1, ViT-B/32, ConvNeXt-Tiny, and ViT-B/16 for the ImageNet-1k classification task, using a homogeneous 10,000-image subset (ImageNet-10k) and a uniform 40-epoch baseline configuration executed on two NVIDIA Tesla P100 GPUs at the Bioinformatics and Computational Biology Center of Colombia (BIOS). Energy was recorded directly via NVML and contrasted with the computational complexity of each model. The results show a Pearson correlation of 0.85 between floating-point operations (GFLOPs) and energy consumption in kWh, indicating that computational complexity is a strong but imperfect predictor of energy expenditure: architectures with comparable GFLOPs exhibited consumption differing by up to 3.1x due to differences in the hardware efficiency of their dominant operations. The EfficientNet variants offered the best balance between classification performance (Val Top-5 up to 97.15%) and energy efficiency (0.457-0.611 kWh), while Vision Transformers exhibited the highest relative energy consumption and lower classification performance under the evaluated configuration. These findings guide architecture selection in energy-constrained computing environments.
|
| 5 |
What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study
2610.03792
|
cs.CV
|
Avyay Sadhu, Patrick Cooper |
Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language models about ...Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language models about time. We fine-tune four open models (Qwen3-VL-8B/4B, Qwen2.5-VL-7B, Gemma-3-12B) with group relative policy optimization under three data recipes: verified (synthetic CLEVRER questions with exact answer and event-order rewards), unverified (self-supervised pretext tasks over 43,751 real web videos), and a 1:1 mixture, plus a verified+real arm that adds 4,000 verifiable questions on real video. Each cell is evaluated in-domain and on out-of-domain real video (a NExT-QA temporal stress set and an MVBench subset), with frames in order, shuffled, and absent. (1) Verified training yields large in-domain gains that shrink as base competence grows (+14 to +19 points on weaker models; +6 on the strongest). (2) Much of the gain is non-visual: accuracy with no frames rises nearly as much as with frames. (3) Verified-only training can severely degrade out-of-domain accuracy with no sign during training: Qwen3-VL-8B loses 26.7 and 25.2 points on the two real-video sets, while the mixture never significantly degrades a model trained on it. Adding real verified questions removes that loss (-2.3 points, within noise of base) and keeps a +9.3 in-domain gain, so the cause is narrow synthetic-only data, not verification. (4) No recipe induces temporal-order grounding: across 41 evaluations the ordered-versus-shuffled gap is indistinguishable from zero in 39 and marginal in two, despite an event-order reward. Verifiable rewards improve benchmark accuracy without temporal understanding. Report no-frame controls, and mix in real video to guard against out-of-domain degradation.
|
| 6 |
WAMJET: A Harness for World Action Model Acceleration
2610.03797
|
cs.CVcs.AI
|
Le Chen, Lixin Liu, Jan Schneider, Zeju Qiu, Simon Guist |
World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and ...World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95x lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.
|
| 7 |
StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search
2610.03799
|
cs.CVcs.LG
|
Ghadi Nehme, Faez Ahmed |
Recovering executable CAD programs from 3D meshes is challenging due to the compositional nature of CAD construction and the interaction between discrete modeling choices and continuous parameters. Many learning-based methods predict complete programs in a sin...Recovering executable CAD programs from 3D meshes is challenging due to the compositional nature of CAD construction and the interaction between discrete modeling choices and continuous parameters. Many learning-based methods predict complete programs in a single pass and rely predominantly on sketch-extrude representations, limiting operation diversity and opportunities to correct geometric errors during reconstruction. We introduce StepCAD, a generative optimization approach that combines a state-conditioned CAD policy with geometry-guided search. Given an input mesh, the policy predicts construction actions conditioned on both target and intermediate geometry, and an IoU-guided tree search refines the resulting program through local edits. We also introduce ARCADE-1.5M, a large-scale dataset of 1.5M executable CAD programs spanning diverse operations, sequences with a maximum length of 150+ counted operations, and 12.5M intermediate state-action transitions. Experiments across multiple CAD reconstruction benchmarks show that StepCAD achieves state-of-the-art geometric reconstruction accuracy with consistently high validity, yielding up to 87.2% relative IoU improvement over the strongest evaluated baseline, with particularly large gains on complex shapes. Project page: https://ghadinehme.com/stepcad.github.io/
|
| 8 |
Beyond Token Accuracy: Prioritizing What Matters for Visual Reconstruction
2610.03822
|
cs.CV
|
Zhicheng Liu, Zhouxiang Zhao, Chenliang Wu, Zhaohui Yang, Zhaoyang Zhang |
Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content a...Tokens have become a unified interface for multimodal foundation models, making visual-token communication a natural paradigm for efficient image delivery. However, existing methods typically rely on static policies that cannot jointly adapt to image content and channel conditions. Moreover, their token-level utility objectives do not necessarily translate into improved image reconstruction quality. In this paper, we propose AdapToC, an adaptive, reconstruction-oriented visual-token communication framework. At the transmitter, an adaptive selector jointly models image content, channel state, and communication budget to perform instance-wise resource allocation. Rather than using a fixed token rate and protection policy, it dynamically determines how many tokens should be transmitted and assigns different protection levels according to token importance and current channel conditions. At the receiver, an adaptive MaskGIT receiver incorporates channel reliability into contextual token modeling. It distinguishes tokens with different reliability levels, preserves high-confidence observations, corrects potentially corrupted tokens, and iteratively reconstructs missing content from the received evidence and global visual context. By co-designing token quantity, unequal protection, and reliability-aware recovery for image-level reconstruction quality, AdapToC achieves a peak mean PSNR gain of 4.20 dB over the strongest static baseline under matched communication costs and state-of-the-art performance among the evaluated visual-token communication methods.
|
| 9 |
Are We Measuring Anticipation? Auditing Privileged Information in Procedural Video Evaluation
2610.03826
|
cs.CV
|
Mahsa Mohammadi, Sareh Rowlands |
Benchmark scores license claims about the capabilities being evaluated. We audit the inference licensed by an evaluation protocol, rather than the predictive model alone. Using procedural action anticipation as a controlled case study, we study a broader evalu...Benchmark scores license claims about the capabilities being evaluated. We audit the inference licensed by an evaluation protocol, rather than the predictive model alone. Using procedural action anticipation as a controlled case study, we study a broader evaluation-validity failure mode: a protocol can remain temporally causal and free of classical target leakage while still supplying a privileged intermediate representation. The failure is not merely optimistic accuracy, but a mismatch between the capability claimed and the construct actually measured. On Breakfast, matched recognizer-history and GT-history regimes score 30.3% and 62.3% (Delta_PH(R_BA) = +32.0, 95% Student-t interval [+24.1, +39.8]). A fixed-checkpoint 2 x 2 intervention isolates a +14.6-point test-time oracle contrast; a matched no-video provenance probe yields a +26.8-point gap; and a boundary-independent query/history control retains a +19.4-point gap. We refer to this discrepancy as a privileged-information gap, defined relative to a specified non-oracle recovery pipeline, and propose a reusable four-condition audit. Across five evaluation settings the gap is heterogeneous; a standard 50 Salads long-term-anticipation re-implementation provides cautious support beyond the audit-specific next-action construction.
|
| 10 |
Streaming Multi-Track Timeline Control for 3D Human Motion Generation
2610.03873
|
cs.CVcs.LGcs.AI
|
Yangsong Zhang, Anujith Muraleedharan, Rikhat Akizhanov, G\"ul Varol, Fabio Pizzati |
Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone whi...Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation or simultaneous composition without explicitly combining streaming instruction arrival with independently timed, overlapping actions. We introduce streaming multi-track timeline control and propose TimelineControl to incorporate new instructions alongside ongoing actions. Interval-aware conditioning preserves instruction timing, while causal part-structured representations and part-aware denoising coordinate concurrent actions across body regions. We also construct TimelineMotion, a dataset with overlapping instruction intervals and body-part annotations. Experiments on TimelineMotion and MTT demonstrate improved semantic alignment and temporal adherence over evaluated streaming baselines, including models retrained on the same data. Ablations and human evaluations validate our design, complemented by spatial conditioning and humanoid execution demonstrations. Our code, data and models will become publicly available.
|
| 11 |
DABACO: A Multi-Camera Dataset and Benchmark for Screen Localization and Pointing Estimation
2610.03928
|
cs.CVcs.MM
|
\'Oscar G\'omez-C\'ardenes, Jos\'e Gil Marichal-Hern\'andez, Juan Manuel Mart\'in-Do\~nas |
Screen localization and pointing estimation are key to low-cost interactive devices. Yet developing and evaluating these algorithms requires realistic data: synthetic captures cannot fully reproduce the optical distortion, rolling shutter, motion blur, and dis...Screen localization and pointing estimation are key to low-cost interactive devices. Yet developing and evaluating these algorithms requires realistic data: synthetic captures cannot fully reproduce the optical distortion, rolling shutter, motion blur, and display processing of a physical acquisition, and most existing datasets provide static images rather than the video needed to assess continuous pointing. In this paper, we introduce the DABACO Dataset. Developed within the DABACO (Dispositivo Apuntador de BAjo COste, or Low-Cost Pointing Device) project, this dataset supports the development and evaluation of screen detection and camera-based pointing algorithms for embedded systems. It comprises video sequences captured with multiple low-cost camera sensors, including monochrome global-shutter and color rolling-shutter modules, on embedded platforms such as the Raspberry Pi 4B and ESP32-S3. We present a systematic annotation pipeline combining temporary visual watermarking, optical-flow tracking, manual verification, and marker removal. Both the original marked captures and the marker-free images, together with explicit corner annotations and modification masks, are released to support auditing and the study of potential reconstruction bias. In addition, we release an open-source evaluation toolkit with two reference baselines: a classical screen-detection pipeline based on edge and contour geometry, and a closed-vocabulary, off-the-shelf YOLOv8 segmentation model evaluated without dataset-specific training. Both are evaluated on a general-purpose computer using Intersection over Union, corner error, and pointing error, establishing initial reference results for future embedded implementations. Overall, DABACO addresses a gap in existing resources and is intended to support the development and evaluation of new low-cost screen-localization and pointing systems.
|
| 12 |
Dynamic Time Step Prediction in Inverse Heat Dissipation for Blur-Like Image Restoration Tasks
2610.03942
|
cs.CV
|
Cap Dang Xuan Kiet, Tat-Jen Cham |
When using diffusion models to target image restoration problems, diffusion inversion is typically employed to retain relevant image information from the degraded images. Instead of inverting back to the initial time step (i.e., T), many methods invert to a pr...When using diffusion models to target image restoration problems, diffusion inversion is typically employed to retain relevant image information from the degraded images. Instead of inverting back to the initial time step (i.e., T), many methods invert to a pre-determined intermediate time step, in order to better preserve information from degraded source images. However, a pre-determined time step for inversion is not ideal for reconstruction, as a severely degraded image requires an earlier starting time step than a mildly degraded one. In addition, DDIM-based models corrupt the original signal by adding Gaussian noise, which can be mismatched to the nature of blur-like degradations, such as blur, haze, and low-light. To address these problems, we propose two solutions: (1) we adopt an alternative diffusion process, called the Inverse Heat Dissipation Model, that diffuses the input image by gradually blurring a data point (2) we propose to implement a time predictor to estimate the starting time step for the inversion, with the model learning to adapt to the degradation severity. Extensive experiments on standard benchmarks show that our method achieves state-of-the-art performance in both quantitative and qualitative evaluations, with excellent generalization to many restoration tasks.
|
| 13 |
FADE: Frame-Aware Diffusion-Transformer-based Multi-Concept Erasure for Video Unlearning
2610.03980
|
cs.CV
|
Yuchen Li, Kaiyuan Deng, Chaoran Feng, Zhenyu Tang, Li Yuan |
Text-to-video (T2V) diffusion models can reproduce copyrighted, violent, or explicit content, which motivates concept erasure: removing designated concepts from a pretrained model while preserving its behavior on everything else. Existing T2V erasure methods l...Text-to-video (T2V) diffusion models can reproduce copyrighted, violent, or explicit content, which motivates concept erasure: removing designated concepts from a pretrained model while preserving its behavior on everything else. Existing T2V erasure methods leave two problems open. Their frame-agnostic suppression can leave isolated frames in which an erased concept resurfaces, a frame-reactivation gap that clip-level averages obscure; and they are usually evaluated with one target concept or category at a time. We propose Frame-Aware Diffusion Erasure (FADE), a multi-concept video unlearning framework. FADE first applies a joint closed-form key/value edit that suppresses all target concepts, then trains per-concept frame-aware low-rank adapters whose strength is gated by the frame index and the denoising timestep to remove residual per-frame leakage. Each adapter is trained with the other targets' prompts as hard negatives, which keeps the concept-specific components of different adapters well separated, and a similarity-based soft router combines the adapters according to the prompt. With 16 concepts (objects, artistic styles, and nudity) erased from a single Wan2.1-T2V-1.3B backbone, FADE reduces the residual accuracy on the object benchmark to 4.9%, against 15.5% for the strongest of eight baselines, while keeping the VBench average within 0.9% of the unedited model. The ranking is unchanged under a VLM judge and a blinded human study, and the advantage over the strongest baseline carries over to prompts that combine several erased concepts, to 30 simultaneously erased celebrity identities, and to Wan2.1-T2V-14B, CogVideoX-2B, and HunyuanVideo-1.5.
|
| 14 |
Masked Privileged-Information Distillation for Multimodal Skin Lesion Classification Under Missing Clinical Metadata
2610.03991
|
cs.CVcs.LG
|
Anirban Barua, Md Mahir Abrar Khan, Ayman Iktidar, Md. Sajjatul Islam |
Multimodal skin lesion classification combines clinical images with patient metadata to improve diagnostic accuracy. However, complete metadata available during training may be only partially accessible at deployment, and resource-constrained settings addition...Multimodal skin lesion classification combines clinical images with patient metadata to improve diagnostic accuracy. However, complete metadata available during training may be only partially accessible at deployment, and resource-constrained settings additionally require computational efficiency. We address these challenges with a privileged-information distillation framework in which a multimodal teacher trained on complete metadata supervises a 9.2x smaller student trained with randomly masked clinical fields. Clinical fields are masked as whole groups at a per-sample rate drawn from U(0,1), so one training run covers the full metadata availability range. Fusion is residual, with metadata added as a gated correction to an unconditional image base. On the PAD-UFES-20 dataset, distillation under masked training improves balanced accuracy over cross-entropy training at every availability level, by an average of +4.7 points versus +1.8 points without masking. The masked student loses only 7.8 balanced-accuracy points as metadata decreases from complete to absent, compared with 36.6 points for the same student trained on complete metadata, highlighting the role of masked training in graceful degradation beyond distillation alone. Grad-CAM visualizations further show that the masked student's attention generally remains lesion-centered as metadata is withdrawn. The resulting compact model targets point-of-care settings, where clinical metadata is often incomplete.
|
| 15 |
Verifier-Guided Synthetic Augmentation for 3D Human Shape Generation
2610.04006
|
cs.CV
|
Yuexuan Wu, Yang Xiang, Hamid Laga, Dip Das, Anuj Srivastava |
Limited training data diversity constrains generative modeling of 3D human bodies: conservative models remain close to observed examples, whereas exploratory models often violate basic body proportions. We introduce a verifier-guided augmentation framework tha...Limited training data diversity constrains generative modeling of 3D human bodies: conservative models remain close to observed examples, whereas exploratory models often violate basic body proportions. We introduce a verifier-guided augmentation framework that uses global and mode-local PCA to generate inexpensive candidates, screens them using correspondence-derived skeletal proportions and body-part geometry, and retrains a diffusion model on accepted candidates. Elastic registration provides both the modal structure used by distributed PCA and the dense anatomical correspondence needed for scalable screening without per-candidate body-model fitting. A blinded human study supports the verifier as a conservative gatekeeper, favoring verifier-accepted over rejected outputs. We evaluate full-pool verifier acceptance separately from the coverage and departure of accepted samples and combine them through EAUC. On 4,498 registered DFAUST surfaces, distributed-PCA augmentation achieves 86.32% acceptance, the highest CP-AUC (0.871), and the highest EAUC (0.752), improving EAUC by 32% over real-only and self-augmented diffusion. These results show that mode-local, verifier-guided proposals broaden diffusion generation while maintaining high agreement with calibrated body measurements.
|
| 16 |
VolS-GS: Relightable Gaussian Splatting with Volumetric Subsurface Scattering
2610.04007
|
cs.CV
|
Junyeong Ahn, Jaegul Choo |
We present VolS-GS, a relightable Gaussian splatting framework that reconstructs objects from one-light-at-a-time (OLAT) captures and renders them under novel lighting and viewpoints. Relightable Gaussian Splatting methods typically model appearance independen...We present VolS-GS, a relightable Gaussian splatting framework that reconstructs objects from one-light-at-a-time (OLAT) captures and renders them under novel lighting and viewpoints. Relightable Gaussian Splatting methods typically model appearance independently at each primitive, which makes non-local effects difficult to represent. This limitation is particularly apparent for subsurface scattering, where light entering the object at one location can emerge at another. Rather than modeling this effect solely with a neural network or a local kernel at each primitive, we use the spatial support of the Gaussian scene as the domain of a differentiable finite-volume transport solver, so that light can propagate through the object's interior. A small network predicts scattering and absorption coefficients for each Gaussian, and the solve redistributes incident light through the resulting field. The coefficients are fit to images rather than measured, so the solve supplies a transport-shaped path for aggregating per-primitive appearance, not a measurement of the material. To keep the learned shadow and specular terms from taking over the other components, our shadow term is predicted from visibility together with the transmittances and the scattering the solve produces, and a regularizer suppresses specular highlights in regions the shadow term predicts to be unlit. Experiments on three OLAT benchmarks show that VolS-GS consistently improves relighting quality on held-out lights and views.
|
| 17 |
PhysMamba: Selective State Space Models as Learned Articulated Body Simulators
2610.04014
|
cs.CV
|
Haochuan Zhang, Sinisa Todorovic |
We introduce the first learned articulated body simulator based on a selective state space model (SSM), called PhysMamba. PhysMamba predicts next-frame full-body state from position, rotation, and joint-action history, without velocity inputs. We compare four ...We introduce the first learned articulated body simulator based on a selective state space model (SSM), called PhysMamba. PhysMamba predicts next-frame full-body state from position, rotation, and joint-action history, without velocity inputs. We compare four architectures under partial- and full-observation inputs and three training protocols. The from-scratch rollout training protocol gives Mamba2 strong short- and mid-horizon accuracy under partial observation (s10 = 43 mm, 2/50 diverged), while the two-stage teacher-based rollout protocol stabilizes GRU but fails for Mamba2. With CUDA graph compilation, Mamba2 reaches 0.107 ms per frame (9,334 FPS) on an H100 GPU, within 1.1$\times$ of GRU's un-compiled throughput, adding under 1% latency to a 30 Hz HMR pipeline and enabling integration as a differentiable physics module for video-based mesh recovery.
|
| 18 |
ClasSAE: Class-Aligned Sparse Autoencoders via Differentiable Feature-Class Affinity
2610.04020
|
cs.CV
|
Jakub St\k{e}pie\'n, Marcin Mazur, Jacek Tabor, Przemys{\l}aw Spurek |
Sparse Autoencoders (SAEs) began as an unsupervised tool for decomposing neural representations into sparse, interpretable features, and are increasingly used not only for passive analysis but also for active interventions such as unlearning, bias mitigation, ...Sparse Autoencoders (SAEs) began as an unsupervised tool for decomposing neural representations into sparse, interpretable features, and are increasingly used not only for passive analysis but also for active interventions such as unlearning, bias mitigation, and concept editing. A central challenge for these editing and steering methods is reliably matching features to target concepts; most current approaches address this by computing post-hoc scores over an already-trained, frozen dictionary. We instead introduce ClasSAE, a novel method that both automatically assigns classes to features and guides the encoder toward class-separable representations during training. Specifically, we apply a differentiable top-$k$ operator to a trainable feature--class affinity matrix with per-feature budgets, coupling the features selected for each sample to the classes they are trained to represent. Because gradients flow through the selection of active features rather than only through their magnitudes, the encoder and the affinity matrix co-adapt rather than being fit in separate stages. The result is a dictionary that is both class-separable and class-annotated, with no need for post-hoc probing. We propose three variants for enforcing sparsity within this framework, which achieve comparable overall performance with slightly different trade-offs. Using CLIP ViT-L/14 embeddings on ImageNet, we show that the learned affinity matrix agrees closely with an independently estimated post-hoc feature-class matrix computed on held-out data. The model also supports direct class prediction from the encoder and affinity matrix alone, without a separately fitted classifier, and its more class-aligned encoder yields improved separation in Targeted Probe Perturbation evaluations. https://github.com/St0pien/ClasSAE.
|
| 19 |
Learning Subject-Specific Anatomical Representations via Manifold Expansion: Application to Accelerated Multi-Contrast MRI
2610.04028
|
cs.CV
|
Ruimin Feng, Wanyu Bian, Albert Jang, Zachary Stewart, Fang Liu |
Clinical MRI routinely acquires multiple contrast-weighted images of the same anatomy for complementary tissue characterization. However, current accelerated MRI methods typically reconstruct each contrast independently, without fully exploiting shared anatomi...Clinical MRI routinely acquires multiple contrast-weighted images of the same anatomy for complementary tissue characterization. However, current accelerated MRI methods typically reconstruct each contrast independently, without fully exploiting shared anatomical information. This work aims to learn anatomical representations invariant to contrast-dependent appearance for reconstruction of accelerated multi-contrast MRI. We propose MAX (MAnifold eXpansion), a subject-specific framework that learns anatomical representations from a single fully sampled reference contrast. To address the under-constrained separation of shared anatomy and contrast-dependent components from a single image, MAX expands the multi-contrast manifold using anatomy-preserving intensity augmentations. A disentangled implicit neural representation models augmented samples using shared spatial coordinates for anatomy and spatially invariant coordinates for contrast appearance. The learned anatomical representation is then fixed, with the contrast representation adapted to the undersampled target data, followed by unrolled refinement. Theoretical analyses further provide insight into the disentangled representation learning and explain how the learned anatomical representation improves the target contrast reconstruction. At R = 8 for brain MRI and R = 6 for knee MRI, MAX achieves the highest mean PSNR and SSIM across all tasks, improving PSNR by more than 1 dB over the strongest baseline for both brain contrasts. MAX more faithfully recovers subtle anatomical and pathological structures and remains robust to inter-contrast motion, structural heterogeneity between reference and target contrasts, and measurement noise. Therefore, MAX provides a general strategy for leveraging high-quality reference scans in accelerated MRI and has the potential to be extended to other reference-assisted MRI inverse problems.
|
| 20 |
MAGEFormer: Learning Metric-Consistent Representations for Anisotropic CT Segmentation
2610.04036
|
cs.CVcs.AI
|
Jiaying Li, Paolo Remagnino |
Vision Transformers (ViTs) have shown strong performance in volumetric segmentation, but their effectiveness on clinical CT is limited by an isotropic Euclidean lattice assumption. This conflicts with anisotropic CT acquisition, leading to two key issues: (1) ...Vision Transformers (ViTs) have shown strong performance in volumetric segmentation, but their effectiveness on clinical CT is limited by an isotropic Euclidean lattice assumption. This conflicts with anisotropic CT acquisition, leading to two key issues: (1) a metric mismatch between voxel indices and physical anatomy, and (2) accuracy degradation from isotropic resampling. To address this, we propose MAGEFormer, a geometry-calibrated framework that embeds physical metric constraints directly into representation learning. Our method introduces Metric-Adaptive Spatial Embedding (MASE) to calibrate positional frequencies using voxel spacing, Geometry-Constrained Attention (GCA) to suppress physically implausible feature correlations, and Geometric View Voting (GVV) to reduce discretization bias during inference. We evaluate MAGEFormer on two multi-organ abdominal CT benchmarks, BTCV and FLARE 22, under a unified protocol against strong CNN and Transformer-based baselines. MAGEFormer achieves the strongest boundary accuracy among the compared methods, with 10.58 mm HD95 on BTCV and 3.40 mm HD95 on FLARE 22, and shows consistent gains in Dice under the same protocol. These results show that geometry-aware internal calibration is more effective than relying on conventional isotropic preprocessing alone for anisotropic CT segmentation.
|
| 21 |
Dynamic Quadtree Tokenization and Transformer for Adaptive Mesh PDE Forecasting
2610.04044
|
cs.CVcs.LG
|
Yilin Zhuang, Noah Zambrano, Karthik Duraisamy |
The quadratic attention cost of Vision Transformers (ViTs) forces a trade-off between spatial resolution and rollout horizon, particularly for fine-scale PDEs where shocks, reaction fronts, and material interfaces occupy small, evolving regions of the domain. ...The quadratic attention cost of Vision Transformers (ViTs) forces a trade-off between spatial resolution and rollout horizon, particularly for fine-scale PDEs where shocks, reaction fronts, and material interfaces occupy small, evolving regions of the domain. Conventional neural surrogates also lack mechanisms to adapt resolution dynamically. We propose WAMRViT, a ViT that tokenizes inputs as balanced quadtrees using a wavelet-inspired refinement criterion, jointly encodes position and refinement level with 3D rotary positional embeddings, and regrids in cell space during inference for stable long-horizon rollouts. A multi-scale variant retains each leaf at its native source resolution and lets the model learn across resolution levels. Unlike fixed-budget adaptive-tokenization methods, WAMRViT imposes no predetermined token count and supports fully adaptive topology throughout autoregressive rollout. To our knowledge, it is the first machine-learning surrogate to natively tokenize multi-level Adaptive Mesh Refinement (AMR) data. On uniform-grid benchmarks, uniform-patch WAMRViT improves finest-level region-of-interest VRMSE over a finest-patch uniform ViT while using substantially fewer tokens. The multi-scale variant achieves the lowest first-step full-field RMSE and VRMSE on both benchmarks and improves rollout-averaged full-field and refined-region accuracy at long horizons. With parallelized regridding, its end-to-end rollout cost lies between finest-patch and approximately token-matched coarser-patch ViTs. On a complex AMR combustion problem whose finest features cannot be represented natively by the evaluated uniform-grid baselines, WAMRViT operates directly on adaptive cells and substantially reduces finest-level error at matched transformer capacity. Code: https://github.com/tonyzyl/wamrvit
|
| 22 |
ASD-FEAT: A Multi-Modal Infant Video-Derived Dataset for Early ASD Risk Prediction
2610.04051
|
cs.CV
|
Sidrah Liaqat (University of Kentucky), Halil Helvaci (University of Kentucky), Sen-Ching Cheung (University of Kentucky), Chongruo Wu (University of California Davis), Dongjie Chen (University of California Davis) |
Accurate early screening for Autism Spectrum Disorder (ASD) is a precursor to timely intervention, which is critical for improving cognitive and behavioral outcomes. We present ASD-FEAT (ASD - Feature Extraction And Tracking), a multimodal dataset derived from...Accurate early screening for Autism Spectrum Disorder (ASD) is a precursor to timely intervention, which is critical for improving cognitive and behavioral outcomes. We present ASD-FEAT (ASD - Feature Extraction And Tracking), a multimodal dataset derived from video recordings of infant-adult interaction sessions. The key contribution of ASD-FEAT is the combination of longitudinal coverage from infancy through 36 months, repeated interaction sessions, clinically validated developmental outcomes, expert frame-level behavioral annotations, and privacy-conscious multimodal feature representations. To the best of our knowledge, existing ASD behavioral datasets do not jointly provide these characteristics at comparable scale. To demonstrate the utility of ASD-FEAT, we use it to evaluate a computer-vision-based end-to-end pipeline relying on machine learning techniques to automatically identify ASD risk. ASD-FEAT integrates both expert-defined and deep-learned features, including face and eye landmarks, facial action units, gaze direction, head position, mel-spectrogram audio representations, and optical flow, to identify behavioral markers of social interaction. Our automated pipeline achieves an ASD classification accuracy of 76.2% and an Area Under the Receiver Operating Characteristic (AUROC) of 0.82, compared to classifiers trained on manually labeled behaviors, which yielded 81.3% accuracy and an AUROC of 0.88. We further introduce a within-visit partner contrast: a per-visit signal contrasting examiner-directed and parent-directed social behavior which, when added to the classifier, lifts the fully automated Look Face + Smile and Look Face + Vocal configurations to Matthews correlations of 0.49 and 0.45 respectively, exceeding the human-coded single-partner baseline of 0.42.
|
| 23 |
A Theory of Shape Reconstruction from Heat Conduction and Shading
2610.04052
|
cs.CV
|
Akihiko Oharazawa, Sriram Narayanan, Mani Ramanagopal, Srinivasa G. Narasimhan |
Shape from shading using a single image of a Lam- bertian surface is inherently ambiguous. When the light source direction is known, the surface normal estimation has a cone- ambiguity, which worsens when the source is unknown. Recently, shape from heat conduc...Shape from shading using a single image of a Lam- bertian surface is inherently ambiguous. When the light source direction is known, the surface normal estimation has a cone- ambiguity, which worsens when the source is unknown. Recently, shape from heat conduction has emerged as an approach that leverages heat transport equations to estimate the Shape Lapla- cian operator, an intrinsic measure of shape. However, deriving surface normals from the Laplacian operator encounters a local binary convex/concave ambiguity. Our contribution introduces a novel theory to resolve these local shape ambiguities (excluding a few degeneracies) without relying on priors like smoothness, by combining the cues from shading and heat conduction. Our method ensures the mathematical constraints of both shading and the Laplacian are satisfied simultaneously, even with an unknown light source. We validate our theory through simulations of complex shapes and analyze its performance in the presence of noise, as well as on a noisy single thermal video of real-world objects with complex shapes and material properties, including varying albedo.
|
| 24 |
Evaluating Zone-Guided Front Extraction for Glacier Calving-Front Delineation in SAR Imagery
2610.04066
|
cs.CVcs.LG
|
Chhaya Kulkarni, Emam Hossain |
Automatic calving-front delineation from synthetic aperture radar imagery is challenging because the front is a thin and often ambiguous boundary between glacier ice, ocean, and surrounding rock or terrain. The CAlving Fronts and where to Find thEm (CaFFe) dat...Automatic calving-front delineation from synthetic aperture radar imagery is challenging because the front is a thin and often ambiguous boundary between glacier ice, ocean, and surrounding rock or terrain. The CAlving Fronts and where to Find thEm (CaFFe) dataset provides both binary calving-front masks and broader semantic zone masks, making it possible to study whether zone-level supervision can support front recovery. In this paper, we compare direct front prediction with zone-guided front extraction using U-Net, DeepLabV3+, and SegFormer-B0 under the same bounding-box-cropped CaFFe setting. In the direct setting, models predict the binary calving-front mask. In the zone-guided setting, the model first predicts four semantic zone classes, and the front is then extracted from the predicted glacier-ocean boundary. We evaluate both zone-level and front-level performance, include a ground-truth-zone boundary check, and examine lightweight test-time adaptation on sensor-specific and glacier-specific subsets. The results show that zone labels contain useful front-boundary information: extracting the front from ground-truth zones gives the lowest mean distance error. However, fronts extracted from model-predicted zones remain weak, even when zone segmentation scores are moderate. Test-time adaptation also does not consistently improve zone-guided front recovery. These results indicate that zone segmentation performance should not be treated as a substitute for front-level evaluation and that effective use of zone labels may require boundary-aware training, label fusion, or explicit front supervision.
|
| 25 |
Why Convolution Still Matters: Evaluating Inductive Biases in Cryospheric Image Classification
2610.04073
|
cs.CVcs.LG
|
Chhaya Kulkarni, Emam Hossain |
Recent advances in attention-based deep learning have motivated their adoption for remote sensing image classification; however, their benefits for cryospheric imagery, where surface states are dominated by fine-grained textures and class imbalance, remain unc...Recent advances in attention-based deep learning have motivated their adoption for remote sensing image classification; however, their benefits for cryospheric imagery, where surface states are dominated by fine-grained textures and class imbalance, remain unclear. In this work, we revisit a benchmark Greenland Ice Sheet image dataset, previously shown to favor convolutional neural networks (CNNs), to examine whether modern attention-based and hybrid architectures improve class-wise reliability. We conduct a controlled comparison between a classical CNN (AlexNet), a modern CNN (ConvNeXt-Tiny), a pure attention-based model (Swin-Tiny), and a hybrid convolution-attention model (CoAtNet-0) under identical training and evaluation protocols. Results show that AlexNet achieves the highest accuracy and the strongest balanced performance as measured by macro-averaged F1, while ConvNeXt-Tiny exhibits the highest macro-averaged AUC, indicating strong class separability but less consistent final decision quality. Class-wise analysis reveals that hybrid architectures improve recall for rare and structurally distinct surface classes, whereas convolutional models remain more reliable for texture-dominated categories. These findings highlight the importance of aligning architectural inductive bias with cryospheric data characteristics and suggest that increased model complexity does not necessarily translate to improved reliability for ice-sheet surface classification.
|
| 26 |
VCURF: Virtual Camera-based Uncertainty of Radiance Fields
2610.04076
|
cs.CV
|
Liyan Chen, Nathaniel Burgdorfer, Philippos Mordohai |
Radiance fields, implemented with either implicit (NeRF) or explicit (Gaussian Splatting) representations, are advancing the state of the art in novel view synthesis at a rapid pace. Even though the rendered views they generate are often compelling, they are n...Radiance fields, implemented with either implicit (NeRF) or explicit (Gaussian Splatting) representations, are advancing the state of the art in novel view synthesis at a rapid pace. Even though the rendered views they generate are often compelling, they are not free of errors. In this paper, we propose a new approach for pixel-wise uncertainty quantification based on measuring the inconsistencies among renderings by the radiance field model in virtual cameras sampled near the target viewpoint. We named our approach VCURF for Virtual Camera-based Uncertainty of Radiance Fields. VCURF treats the radiance field model as a black box, only assuming that it is capable of rendering color and depth on demand. This property makes our approach applicable to both NeRF and GS models without any modification. Our experiments on a combination of datasets, radiance field models and baselines demonstrate VCURF's effectiveness in pixel-wise uncertainty estimation. We conclude the paper with findings that question the way view selection is tackled by the majority of the current literature.
|
| 27 |
Dependable AI-Assisted Engineering: A Formal Framework for AI Participation and Assurance in Safety-Critical Workflows
2610.04084
|
cs.CVcs.AI
|
Puxue Tan |
Generative AI can produce engineering artefacts, but generation alone does not determine whether or how those artefacts should enter safety-critical workflows. This paper develops a formal framework for assigning AI participation and assurance at the level of ...Generative AI can produce engineering artefacts, but generation alone does not determine whether or how those artefacts should enter safety-critical workflows. This paper develops a formal framework for assigning AI participation and assurance at the level of individual workflow units. Each unit has a participation and assurance record covering its engineering requirement, an approved operational formalization where applicable, the applicable mechanism, fallback where applicable, evidence obligations and the applicable guarantee, plus a deployment-readiness status. The framework distinguishes deterministic verification, statistically calibrated admission, authorized human judgement supported by AI advice, authorized human adjudication of AI-produced artefacts, retained deterministic tool paths and explicit non-participation; these arrangements carry different kinds of guarantee rather than levels on a common scale. The framework also separates formalization fidelity from verifier soundness, provides a staged classification and readiness procedure, and derives conditions for comparing a gated AI-assisted unit with an incumbent process under recurring-population assumptions. We instantiate and apply the framework in an executed 17-unit wing-spar structural-analysis workflow combining deterministically gated AI-generated CAD, retained deterministic computation and human judgement. The AI-generated CAD program passed all 23 deterministic checks and was admitted at the first attempt. Favourable stress magnitudes did not suffice to pass the stress criteria where the predeclared mesh-convergence evidence was insufficient; those criteria were instead referred to engineering judgement. The case demonstrates selective AI participation and explicit evidence handling at unit level; no claim is made of workflow-level dependability, certification, structural safety or productivity.
|
| 28 |
UniBRep: Learning Unified Geometry and Topology for Image-conditioned B-Rep Generation
2610.04092
|
cs.CV
|
Haiyang Ying, Allen Tu, Jiaye Wu, Tom Goldstein, Matthias Zwicker |
Generating a boundary representation (B-rep) conditioned on a single image requires faithful reconstruction of geometry, valid topology, and support for complex shapes. We present UniBRep, a geometry-first framework that adapts a pretrained image-to-3D model t...Generating a boundary representation (B-rep) conditioned on a single image requires faithful reconstruction of geometry, valid topology, and support for complex shapes. We present UniBRep, a geometry-first framework that adapts a pretrained image-to-3D model to generate a feature mesh as a unified intermediate representation. Its surface provides a geometric scaffold, while spatially aligned learned features encode face-separation cues for topology recovery. Dual decoder branches generate the geometry and face-separation features; a geometry- and feature-guided construction pipeline then fits parametric surfaces, recovers boundary curves and connectivity, and assembles an explicit B-rep using a CAD kernel. Recovering topology from mesh regions avoids predefined architectural face-count limits, allowing face count to scale with shape complexity. On the standard DeepCAD benchmark, UniBRep produces valid B-reps for 80.49\% of inputs and reduces face Chamfer distance from 0.1096 to 0.0345 relative to CADDreamer. In a matched comparison, UniBRep also outperforms the HoLa public demo across all reported metrics. Further evaluations demonstrate scalability to high-complexity shapes beyond the standard 30-face range, generalization to objects outside the CAD training distribution, and qualitative transfer to real photographs.
|
| 29 |
Scaling 3D Visual Grounding in Abdominal CT
2610.04095
|
cs.CVcs.LG
|
Sam Church, Danyal Maqbool, Joshua D. Warner, Andrew Voter, Junjie Hu |
Visual grounding models can enhance radiology workflows by linking report findings to image regions. This is particularly valuable for 3D CT, where findings often occupy a tiny fraction of the volume. Training 3D grounding models requires large sets of paired ...Visual grounding models can enhance radiology workflows by linking report findings to image regions. This is particularly valuable for 3D CT, where findings often occupy a tiny fraction of the volume. Training 3D grounding models requires large sets of paired phrases and regions, and building such datasets is expensive, requiring radiologists to annotate images by hand. We posit that this supervision is already created implicitly during routine reporting, as radiologists frequently place 2D annotations (e.g., distance measurement, arrows) on key images to make measurements and to support report interpretation. We introduce an automated pipeline that converts these routine clinical annotations into large-scale phrase-region supervision for 3D visual grounding. The pipeline links each annotation to the corresponding finding in the report through metadata matching, then uses a promptable 3D segmentation model to convert the 2D annotation into a volumetric mask. This produces phrase-mask-volume datasets without requiring additional radiologist annotation. Applied to a single institution's clinical picture archiving and communication system (PACS), our approach generated 105K phrase-mask-volume triplets from 59K abdominal CT exams. We also introduce two abdominal CT grounding benchmarks, LocusBench-Onc and LocusBench-ED, which comprise 240 oncology and 260 emergency-department radiologist-reviewed phrase-mask-volume triplets, respectively, with the latter spanning 13 distinct categories such as appendicitis, hematoma, and hernia. We further introduce LocusCT, a 3D visual grounding model trained on this dataset, which achieves hit rates of 0.725 on LocusBench-Onc and 0.773 on LocusBench-ED, substantially outperforming comparator models. These results show that routine PACS annotations are a scalable, previously unused source of supervision for 3D visual grounding.
|
| 30 |
Where Does the Semantic Gain Come From? A Reproduction and Extension of Semantic Knowledge-driven Contrastive Learning for Long-Tailed Recognition
2610.04104
|
cs.CVcs.LG
|
Sushrut Ghimire |
Semantic Knowledge-driven Contrastive Learning (SKCL) uses a language model to decide which classes are related, and pulls each image towards the prototypes of its semantic neighbours. On CIFAR-100-LT (beta = 100) it reports 54.02% top-1 accuracy, 2.01 points ...Semantic Knowledge-driven Contrastive Learning (SKCL) uses a language model to decide which classes are related, and pulls each image towards the prototypes of its semantic neighbours. On CIFAR-100-LT (beta = 100) it reports 54.02% top-1 accuracy, 2.01 points above Balanced Contrastive Learning (BCL), the method it builds on. The code and the class descriptions are not public. I reimplement SKCL, BCL and ConCutMix in one framework, check it against the public baseline code, and run every configuration with three seeds. The two baselines reproduce within 1.5 points, but SKCL built on BCL, as the paper describes it, ends up 0.56 points below BCL. To find out why, I add SKCL to the authors' own ConCutMix code. Trained for the paper's 300 epochs, it reaches 53.69, only 0.33 below the published number. At the same budget, however, the semantic graph adds just 0.23 points over ConCutMix, while training ConCutMix for 100 more epochs adds 1.04. Together with ConCutMix's published lead over BCL (1.15), this explains the claimed gain. A BCL model that never sees the graph already shares 41.2% of the graph's top-2 neighbours with its own most-confused classes (2.0% by chance), which shows why the graph adds so little on these benchmarks. I also test several changes to SKCL. Combining it with the CutMix branch improves it by 1.06 points, and an adaptive version of the graph improves it slightly (+0.34 and +0.28 in two codebases), although these gains are within seed noise.
|
| 31 |
From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models
2610.04139
|
cs.CV
|
Feiran Wang, Xiaoqi Wang, Ziwei Li, Wenbin He, Yan Yan |
Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatia...Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction. Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.
|
| 32 |
Watermarks and Fingerprints as Soft Bindings for Content Provenance: An Open-Licence Benchmark for Images, Audio and Video
2610.04151
|
cs.CV
|
Seyedmahdi Kazempourradi (Original Pictures Technologies, Inc., Delaware, USA), Ramtin Mojtahedi (Original Pictures Technologies |
Content-provenance standards such as C2PA let a platform recover a stripped manifest through a soft binding: an invisible watermark read from the content, or a fingerprint looked up in a registry. We benchmarked both families under one protocol, restricted to ...Content-provenance standards such as C2PA let a platform recover a stripped manifest through a soft binding: an invisible watermark read from the content, or a fingerprint looked up in a registry. We benchmarked both families under one protocol, restricted to openly available models whose licences we audited, on public media, with false-match rates calibrated on held-out negatives and source-level bootstrap intervals for performance estimates. For watermarking we evaluated 25 image, 7 audio and 7 video configurations from 12 methods on perceptual quality, robustness, false positives and cost; for fingerprinting, 35 methods on registries of up to 98,985 images, partial edits and adversarial attacks. PixelSeal gave the best balance for image and video watermarks and AudioSeal for audio, but the error-correcting detector of every TrustMark variant fired on 5.9 to 15.3% of unmarked images, so a verifier should test the expected payload. Among fingerprints, copy detectors trained on non-commercial data detected up to 75.5% of transformed images at a pair-level false-match rate of $10^{-7}$ and DINOv2, the best permissively licensed method, 65.0%; near-copies in a product catalogue dominated the false matches, and no detection improvement from geometric verification was observed under a matched calibration false-binding constraint in the evaluated image pipelines. On the same attacked copies the two families failed differently: the union of watermark and fingerprint successes covered 69% of image copies; fingerprints covered more audio queries, whereas expected-key watermark verification covered more video queries. Embedding a watermark moved the ISCC code of 62.0% of images past its match threshold, and platform-dependent colour conversion changed watermark bits between x86 and ARM hosts.
|
| 33 |
Kepler4D: Controllable Future Video Generation via 4D Scene State Evolution
2610.04152
|
cs.CV
|
Feiran Wang, Bin Duan, Junyi Wu, Gaowen Liu, Yan Yan |
Video world models aim to preserve scene structure and predict how dynamic objects evolve beyond visual observations. We present Kepler4D, a framework for future video generation through explicit 4D scene state evolution. Given a monocular video, Kepler4D cons...Video world models aim to preserve scene structure and predict how dynamic objects evolve beyond visual observations. We present Kepler4D, a framework for future video generation through explicit 4D scene state evolution. Given a monocular video, Kepler4D constructs a shared 3D representation of background geometry, object motion histories, coarse spatial supports, and semantic context. Chain-of-Motion summarizes observed motion and uses a vision-language model to select structured speed and heading decisions and decide whether to bound object-center height from below. A deterministic rollout converts these decisions into future object trajectories for inspection and editing before synthesis. We render the evolving proxies into geometric controls for a pretrained video generator, separating coarse object motion from the synthesis of appearance and articulation. Experiments on real-world videos demonstrate that Kepler4D enables controllable object motion and plausible future rollout while preserving scene consistency.
|
| 34 |
Diagnosis-Conditioned Spatial Gating and Decoder-Level Supervised Contrastive Learning for Radiology Report Generation
2610.04159
|
cs.CV
|
Md Mustafizur Rahman, Mylene C. Q. Farias |
Radiology report generation models can produce fluent text while still containing finding-level inaccuracies. Diagnosis-driven methods improve generation by conditioning on predicted findings, but these predictions do not directly modify the visual patch featu...Radiology report generation models can produce fluent text while still containing finding-level inaccuracies. Diagnosis-driven methods improve generation by conditioning on predicted findings, but these predictions do not directly modify the visual patch features provided to the decoder, and global gating applies the same modulation across spatial locations. We introduce a Position-Aware Gate (PAG) that uses predicted finding representations to modulate visual patches spatially without region supervision. We also propose a Decoder-level Supervised Contrastive Loss (DSCL) that structures decoder representations using shared positive findings rather than instance identity. On MIMIC-CXR, PAG+DSCL improves clinical efficacy (CE) F1 from 0.484 to 0.502 over a matched global-gate reference, while PAG and DSCL individually reach 0.491 and 0.495. Without additional fine-tuning, the combined model achieves 0.226 CE F1 on IU X-Ray, compared with 0.211 reported by PromptMRG.
|
| 35 |
Referring Multi-Object Tracking in Moving-Camera Videos via Global Motion Compensation
2610.04185
|
cs.CV
|
Hsin-Chen Pai, Jyun-Kai Wang, Yi-Cheng Peng, Wei-Ta Chu |
Referring multi-object tracking (RMOT) takes a video and a language expression as input and tracks all referred objects. Many tracking requirements involve how an object moves rather than how it appears. However, in a video captured by a moving camera, a parke...Referring multi-object tracking (RMOT) takes a video and a language expression as input and tracks all referred objects. Many tracking requirements involve how an object moves rather than how it appears. However, in a video captured by a moving camera, a parked vehicle may appear to move, while a moving vehicle may show little displacement. Existing RMOT methods relate motion with text but do not explicitly remove camera-induced motion. In this paper, we propose extracting residual motion across frames by estimating camera motion in driver-view videos and compare motion characteristics with the query expression. We consider the motion-matching extent and integrate it with the RMOT method's prediction result through late fusion. In the evaluation, we verify the performance gain of taking the motion compensation module as a plug-in across different RMOT hosts.
|
| 36 |
CellSplat4D: PSF-Aware 4D Gaussian Splatting for Sparse Robotic Live-Cell Imaging
2610.04199
|
cs.CV
|
Yingda Tao, Guoyu Lu |
A robotic microscope watching living cells cannot afford to look as often as it would like. Every volume it acquires costs photons the specimen does not get back, and time owed to other wells. What such a platform exists to produce is a record of individual ce...A robotic microscope watching living cells cannot afford to look as often as it would like. Every volume it acquires costs photons the specimen does not get back, and time owed to other wells. What such a platform exists to produce is a record of individual cells through time: which cell is which from one volume to the next, and which cell divided into which two. Sampling sparsely breaks that record exactly where it matters, and the fault lies in the acquisition schedule rather than in the analysis software. We fill the gaps by reconstructing them, fitting a 4D Gaussian model to whatever volumes the hardware could afford. The model is a cloud of light-emitting blobs, each carrying a position, a shape and a lifetime. Being continuous in time, it renders any missing volume on demand, decoupling how often the robot analyses from how often it can afford to look. The microscope's point-spread function is measured from the data rather than inherited from acquisition metadata or left to the optimizer, because metadata inflates it and the optimizer cannot recover it at all: a wider blur around a smaller blob fits the images equally well. Each blob's lifetime is stored in frames rather than as a fraction of the recording, so that it denotes a fixed duration on any sequence. Unmeasured timesteps are supervised at coarse scale by a 3D U-Net that predicts the intermediate volume directly and estimates no motion field, since a dividing cell becomes two and no motion describes that. On two Cell Tracking Challenge sequences, a C. elegans embryo and a Chinese Hamster Ovarian (CHO), with fidelity scored per cell nucleus, our reconstruction holds the highest nucleus fidelity at every distance from an acquired frame, has the flattest decay across the gap, and best recovers focal planes it was never shown with graceful degradation across the gap.
|
| 37 |
Sparse-GS2Mesh: 3D Gaussian Splatting Guided by Novel Stereo Views and 2DGS for Sparse View Surface Reconstruction}
2610.04203
|
cs.CV
|
Younghyun Noh, Minje Kim, Tae-Kyun Kim |
Surface reconstruction under sparse-view settings remains challenging due to limited geometric cues. Volume rendering methods based on signed distance functions often produce over-smoothed surfaces, while 3D Gaussian Splatting (3DGS), though time-efficient, su...Surface reconstruction under sparse-view settings remains challenging due to limited geometric cues. Volume rendering methods based on signed distance functions often produce over-smoothed surfaces, while 3D Gaussian Splatting (3DGS), though time-efficient, suffers from incomplete geometry due to the lack of reliable depth supervision and the limitation of being optimized only from given input views. In this paper, we present Sparse-GS2Mesh, a stereo-aware framework for surface reconstruction from sparse views. While 3DGS and stereo matching have been leveraged for surface reconstruction under dense view settings, we extend them to operate effectively under sparse view conditions by first initializing 3DGS using epipolar depth priors to mitigate the 3DGS overfitting problem, followed by our three key components: (I) adaptive baseline selection, (II) fine-tuning with a stereo matching network, and (III) 2D/3D co-regularized fine-tuning. Given a warmed-up 3DGS initialized with epipolar depth, the adaptive baseline selection automatically determines a baseline to synthesize for each sparse view. We then fine-tune 3DGS by backpropagating depth-refining gradients from the stereo matching network, effectively specializing the 3DGS for stereo matching. The 2D/3D co-regularization further helps obtain stable reconstruction, addressing weak geometric cues in close stereo views. Sparse-GS2Mesh achieves a 15\% improvement over state-of-the-art methods in little-overlap settings and comparable results in large-overlap settings. Codes will be publicly available.
|
| 38 |
FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding
2610.04225
|
cs.CV
|
Ziye Zhu, Yanghao Zhou, Lixing Tan, Jialiang Kang, Shuxuan Li |
Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing r...Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational overhead. Many existing methods prune visual tokens during or after vision transformer (ViT) encoding, leaving much of the encoding cost unaddressed. Some approaches prune patches before encoding but rely on learned auxiliary networks for patch selection, incurring additional training and inference overhead. To address these limitations, we propose FlashGaze, a training-free method that reduces spatiotemporal redundancy before ViT encoding without introducing auxiliary networks. FlashGaze uses pixel-space differences as a proxy for information loss and employs Quadtree Dynamic Programming to jointly optimize patch dropping, merging, and keeping under a fixed budget. Experiments on two MLLM backbones across multiple benchmarks demonstrate substantial efficiency gains while largely preserving accuracy. On Qwen3-VL-8B, FlashGaze retains 98% of the full-input baseline accuracy on LongVideoBench while achieving up to 5.4x and 17x speedups in ViT encoding and MLLM prefill, respectively, and reducing peak GPU memory usage by a factor of 1.8. These efficiency gains enable the model to process videos with more frames and higher resolutions on the same GPU hardware, unlocking video understanding at scales previously out of reach.
|
| 39 |
OctMesh: A Unified Octree-Hierarchical Framework for Lossless Triangle Mesh Compression
2610.04281
|
cs.CV
|
Shiyu Feng, Xihua Sheng, Lingyu Zhu, Chunyang Fu, Shiqi Wang |
Lossless triangle mesh compression must preserve both vertex positions and connectivity. Octrees support learned point cloud geometry coding and progressive refinement, but extending them to meshes requires a compatible connectivity representation. Unlike the ...Lossless triangle mesh compression must preserve both vertex positions and connectivity. Octrees support learned point cloud geometry coding and progressive refinement, but extending them to meshes requires a compatible connectivity representation. Unlike the eight occupancy decisions of a voxel, a parent edge can develop into varied child connections, making edge refinement difficult to model with a compact prediction prior. We propose OctMesh, a learned framework that codes geometry and connectivity on a shared octree hierarchy. Its key observation is that octree pooling produces parents with either one child or two to eight children. Child edges are then grouped by their endpoint parents' types and whether the endpoints share a parent. Each candidate group contains children from just one parent or two connected parents. The resulting four categories define small, fixed-shape prediction tasks: connections uniquely determined by the parent graph are inherited without bits, while three neural predictors estimate probabilities for the remaining candidates. These probabilities guide arithmetic coding of the actual edge symbols. Binarized predictions of within-parent connections provide context for predicting connections between different parents. A graph-aware parent feature extractor combines local geometry, parent connectivity and global shape. The connectivity models use dedicated weights at coarse levels and share weights at fine levels. Residual edges and a finest-level face-selection payload complete the reconstruction. On 256 frames from eight MPEG V-DMC test sequences, OctMesh losslessly recovers the finest-level vertex-coordinate, edge and unoriented face sets at an average of 7.033 bits per face, 12.8% below V-Mesh. The same hierarchical representation supports nine levels of progressive vertex-and-edge refinement.
|
| 40 |
A Geometric-Transformation Feature-Adaptive Manifold Restoration Method for Open-Vocabulary Semantic Segmentation of Remote Sensing Images
2610.04300
|
cs.CV
|
Jianzheng Wang, Huan Ni, Xiaonan Niu, Danfeng Hong, Haiyan Guan |
The semantic information of objects in remote sensing images is typically invariant to geometric transformations from the dihedral group D4. However, SAM3-based open-vocabulary semantic segmentation (OVSS) methods often exhibit inconsistent responses to differ...The semantic information of objects in remote sensing images is typically invariant to geometric transformations from the dihedral group D4. However, SAM3-based open-vocabulary semantic segmentation (OVSS) methods often exhibit inconsistent responses to different geometric transformations. To exploit this property and improve the stability of OVSS for remote sensing images, we propose a feature-adaptive manifold repair method based on dihedral-group geometric transformations. First, we introduce multi-scale harmonic-guided D4 view selection (MH-D4VS) to select complementary candidate views from a set of geometrically transformed views. Next, we propose original-view-anchored adaptive manifold repair (OAMR), which uses the original view as an anchor and reliable cross-view information to selectively repair locally unreliable visual features. Finally, we develop pixel decoder test-time adaptation (PD-TTA) for SAM3, which fine-tunes only the parameters of the GroupNorm layers online during inference, thereby enhancing the model's ability to adapt to sample-level distribution shifts. Experimental results show that the proposed method achieves an average mIoU of 55.6% across eight remote sensing semantic segmentation benchmarks and delivers consistent performance improvements under different SAM3-based inference frameworks.
|
| 41 |
Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency
2610.04318
|
cs.CV
|
Sixun Dong, Wei Li, Andong Deng, Qi Qian, Victor Zhu |
Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while fro...Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding latency depends on the size of the candidate pool rather than the final token budget. An empirical study across multiple VLMs and long-video benchmarks yields three findings: dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; resolution-sensitive tasks benefit from selected high-resolution frames; and front-end decoding dominates wall time for hour-long videos. Motivated by these findings, we introduce LoHi, a training-free, single-pass framework that combines a dense low-resolution video stream with sparse high-resolution image frames through the VLM's native video and image pathways. LoHi-Anchor selects high-resolution frames using codec I-frame metadata, while LoHi-SemDiv uses query relevance and visual diversity over CLIP features. Across three long-video benchmarks, LoHi improves average accuracy by 10.6 percentage points over the native-resolution baseline at a matched token budget and by 5.2 percentage points over the strongest prior efficiency method. It also reduces front-end decoding latency by up to 7x on hour-long videos. Project page: https://sixundong.com/projects/lohi
|
| 42 |
Trust the View That Sees the Target: Mining Cross-View Conflicts for Reliability-Gated Disaster Damage Assessment
2610.04327
|
cs.CV
|
Yifan Yang |
After a disaster, building damage is assessed from overhead tiles and ground-level photographs, and most methods fuse the two views symmetrically, trusting both equally for every building. This paper focuses on the samples where that assumption fails: the conf...After a disaster, building damage is assessed from overhead tiles and ground-level photographs, and most methods fuse the two views symmetrically, trusting both equally for every building. This paper focuses on the samples where that assumption fails: the conflict cases, on which two independently trained single-view models disagree. We mine such cases from three paired collections (inspection photographs from the 2025 Eaton wildfire and street-view panoramas from Hurricanes Ian and Milton, each matched to very-high-resolution overhead tiles), where they make up 10-33% of the data. On these samples an oracle that simply trusts the correct view beats every fusion method we tested by 0.37-0.41 accuracy, and the gap survives longer training, calibration, and backbone changes. We recover part of it with a visibility-conditioned reliability gate: a linear model that decides which view to trust from building-visibility features, calibrated per-view confidences, and the disagreement itself. On the wildfire data the gate is the only method that significantly beats calibrated probability averaging (+0.051 on conflicts, p=0.0001) and end-to-end fusion (+0.072, p<10^-4); on the panoramic datasets it matches them. A controlled field-of-view experiment explains why: cropping panoramas toward the building doubles the benefit of fusion, whereas random crops of the same size do not. Finally, the spatial density of conflicts predicts tile-level damage without labels (Spearman r=0.615, p=0.001). Mining conflicts turns "does fusion help?" into "which view should be trusted, where, and why?".
|
| 43 |
Synthetic-to-Real ViT-Based Pose Estimation of a Noncooperative UAV
2610.04335
|
cs.CV
|
Krishnanujam Srinivas, Hanish Acharla, Brij Agrawal, Leonardo Herrera |
Remote pose estimation of noncooperative Unmanned Aerial Vehicles (UAVs) from imagery is critical, as they cannot be influenced or instrumented in advance. Deep-learning-based approaches offer a promising solution; however, their development is constrained by ...Remote pose estimation of noncooperative Unmanned Aerial Vehicles (UAVs) from imagery is critical, as they cannot be influenced or instrumented in advance. Deep-learning-based approaches offer a promising solution; however, their development is constrained by the cost and difficulty of acquiring large-scale real-world datasets with accurate pose labels. Synthetic imagery provides an alternative, but models trained on synthetic data must overcome the synthetic-to-real domain gap to generalize to real-world imagery. This work investigates the inherent synthetic-to-real generalization capability of a Vision Transformer (ViT)-based model for monocular UAV pose estimation. The proposed approach employs a self-supervised DINOv2 backbone and is trained exclusively on labeled synthetic imagery while being evaluated on labeled real-world imagery. Pose ambiguity-aware strategies are incorporated during training and inference to address ambiguities arising from the projection of a three-dimensional target onto a two-dimensional image plane and from target symmetries. An $\alpha$-$\beta$ filter is further integrated during inference to improve pose estimations. To assess the model under operational requirements, it is evaluated in terms of Mean Angular Error (MAE) and inference time, both before and after filtering, using a real-world dataset containing 77,077 labeled UAV images. Before filtering, the model achieves an MAE of $19.18^{\circ}$ and an inference time of $13.25$ ms, whereas after filtering, these values are $8.74^{\circ}$ and $13.42$ ms, respectively.
|
| 44 |
A differentiable Lagrangian-coupled 3D Gaussian Splatting-SPH model for forward simulation and inverse analysis in solid mechanics
2610.04336
|
cs.CV
|
Tian Xu, Soroush Atashi, Tianju Xue |
Recent advances in generative world models have increased interest in digital models that reproduce both the appearance of real objects and their response to physical interaction. Three-dimensional reconstruction techniques, including 3D Gaussian Splatting, ca...Recent advances in generative world models have increased interest in digital models that reproduce both the appearance of real objects and their response to physical interaction. Three-dimensional reconstruction techniques, including 3D Gaussian Splatting, capture detailed surface geometry and appearance from images and videos. However, extending these representations beyond plausible animation to mechanically interpretable models for constitutive behavior, boundary conditions, and inverse parameter identification remains less explored. In this work, a differentiable Lagrangian-coupled 3DGS-smoothed particle hydrodynamics (SPH) model is proposed for forward simulation and inverse analysis of deformable solids. The observed object is first reconstructed from multi-view calibrated visual dataset as a 3DGS rendering model. An envelope-based procedure then generates an independent SPH support for the solid-mechanics model, avoiding the direct use of rendering primitives as mechanical particles. A reference-configuration Lagrangian transfer maps SPH deformation to Gaussian positions and covariances, thereby coupling the physical model and the image observation model while preserving a differentiable computational path. The SPH formulation supports linear elastic, hyperelastic, and Kelvin--Voigt viscoelastic responses, together with fixed, free, and Robin-type boundary conditions. Numerical studies validate the SPH response against finite-element results, assess accuracy and efficiency against a conventional model using Gaussian centers as surface SPH particles, and demonstrate forward simulations on beam, bridge, and liver-shaped examples. Inverse analyses further estimate constitutive and boundary parameters from rendered deformation observations, including noisy cases, demonstrating the feasibility of the proposed model for mechanics-based parameter identification from image data.
|
| 45 |
Any-scale Object Detection using Arbitrary-scaled Images
2610.04346
|
cs.CV
|
Kazutoshi Akita, Norimichi Ukita |
This paper proposes any-scale object detection using arbitrary-scale super-resolution for continuously rescaling object images, while general multi-scale object detection uses discretely rescaled appearance representations. However, a naive usage of super-reso...This paper proposes any-scale object detection using arbitrary-scale super-resolution for continuously rescaling object images, while general multi-scale object detection uses discretely rescaled appearance representations. However, a naive usage of super-resolution produces many false-positive detections if many super-resolution images are independently fed into an object detector. Our method suppresses these false positives by predicting scale proposal maps, each of which represents a set of pixels appropriate for each super-resolution scale.
|
| 46 |
LoCoSplat: Real-Time Feed-Forward 3D Gaussian Splatting with Minimal 3D Reasoning
2610.04351
|
cs.CV
|
Sinan Wang, Jinjin He, Yuchen Sun, Duowen Chen, Shenyifan Lu |
Feed-forward 3D Gaussian Splatting (3DGS) increasingly aggregates multi-view evidence with heavy learned 3D networks. We propose LoCoSplat (Local-Context Splatting), motivated by the observation that a Gaussian is a local primitive: once depth is predicted, wh...Feed-forward 3D Gaussian Splatting (3DGS) increasingly aggregates multi-view evidence with heavy learned 3D networks. We propose LoCoSplat (Local-Context Splatting), motivated by the observation that a Gaussian is a local primitive: once depth is predicted, what the 3D stage must add (scale, rotation, opacity) depends on the point cloud around each anchor, and a fixed local average of that neighbourhood is enough to supply it, no heavy network required. LoCoSplat realises exactly this average: it splats a 16-d linear projection of the point features into a fine and a coarse grid and reads both back at each anchor with a 0.14M-parameter pointwise MLP; with no learned 3D network and no dynamic sparse computation, its whole encoder runs as one fp16 CUDA graph. On RealEstate10K, LoCoSplat outperforms every prior feed-forward method on PSNR, SSIM, and LPIPS at 6, 12, and 24 views, with a margin that widens as views densify (+3.3 PSNR over VolSplat, the prior voxel-aligned state of the art, at 24 views) and grows further under zero-shot transfer to ACID and fine-tuning on ScanNet. It reconstructs a 6-view scene in 33 ms on one NVIDIA RTX PRO 6000 GPU, the fastest of seven feed-forward methods and $4.2\times$ faster than the previous state of the art, trains $2.7\times$ faster ($5.9\times$ at 24 views), and uses $6.7\times$ less inference memory.
|
| 47 |
SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction
2610.04356
|
cs.CV
|
Yuhang Wang, Kai Luo, Yuanfan Zheng, Kailun Yang |
Comprehensive 3D scene understanding for autonomous driving requires modeling geometry, semantics, and motion. However, camera-based occupancy and scene flow prediction are sensitive to unreliable spatial and temporal aggregation, caused by semantically incomp...Comprehensive 3D scene understanding for autonomous driving requires modeling geometry, semantics, and motion. However, camera-based occupancy and scene flow prediction are sensitive to unreliable spatial and temporal aggregation, caused by semantically incompatible image features, misaligned historical observations, and incomplete voxel structures. To address this issue, we propose SelectOccFlow, a selective spatiotemporal aggregation framework that progressively refines contextual evidence across image, temporal, and voxel domains. To obtain semantically compatible image evidence, we design Semantic-Guided Sampling (SGS) to regulate feature sampling with semantic priors. Since reliable image evidence alone cannot resolve temporal inconsistency, we then present State-Conditioned Temporal Aggregation (SCTA) to selectively retrieve historical evidence according to voxel states. To further enhance the structural completeness of voxel representations, we introduce Extent-Aware Spatial Aggregation (ESA), which exploits directional structural support to refine foreground geometry. Experiments on OpenOcc demonstrate that SelectOccFlow achieves a state-of-the-art OccScore of 44.9, improving the previous best by +4.2%. It also maintains competitive occupancy performance on Occ3D-nus and improves the mean OccScore under nuScenes-C corruptions by +11.1%, demonstrating improved robustness to visual corruptions. The source code will be made publicly available at https://github.com/muchen1021/SelectOccFlow.
|
| 48 |
TRIM-ReID: Duplication-Aware Token Reduction and Modality-Aligned Interaction for Multi-Modal Object Re-Identification
2610.04361
|
cs.CV
|
Wanke Xia, Ruiding Zhu, Xingguo Xu, Zhengbo Zhang, Dongxia Liu |
Multi-modal object re-identification exploits complementary RGB, near-infrared (NIR), and thermal-infrared (TIR) observations to retrieve target objects. However, existing methods commonly employ visual encoders optimized for global image-text alignment and se...Multi-modal object re-identification exploits complementary RGB, near-infrared (NIR), and thermal-infrared (TIR) observations to retrieve target objects. However, existing methods commonly employ visual encoders optimized for global image-text alignment and select tokens using learned importance scores. Such designs fail to preserve fine-grained identity cues or explicitly account for token redundancy, resulting in underrepresented local evidence and duplicated tokens that lead to noisy and costly cross-modal interaction. To address this gap, we propose TRIM-ReID, a compact framework that unifies dense feature extraction, intra-modal token reduction, and inter-modal aligned interaction. Specifically, semantically rich and spatially coherent patch features are extracted by Dense Identity Representation (DIR), which leverages DINOv3 to preserve fine-grained identity information. We then introduce Token Diversity Mining (TDM) to identify complementary local evidence and construct compact modality-specific token sets by suppressing repetitive patches while preserving informative diversity. Retained tokens are subsequently fused by Modal Relational Interaction (MRI) to enable effective information exchange across modalities, while a triangular alignment loss explicitly regularizes their joint relationships to maintain cross-modal semantic consistency under independent token selection. Extensive experiments on RGBNT201, RGBNT100, and MSVR310 demonstrate that TRIM-ReID achieves state-of-the-art performance.
|
| 49 |
Beyond Plausibility: Verifiable Fine-Grained Image Editing on Structured Assets
2610.04381
|
cs.CV
|
Muyao Wang, Chen Zhu, Shiqi Yang, DongHyun Hwan, Hideki Nakayama |
Fine-grained image editing requires more than producing a visually plausible result: an editor must execute the requested attribute change precisely while leaving everything else intact. However, existing benchmarks leave a critical gap between realism and ver...Fine-grained image editing requires more than producing a visually plausible result: an editor must execute the requested attribute change precisely while leaving everything else intact. However, existing benchmarks leave a critical gap between realism and verifiability: benchmarks built on realistic images typically rely on human or vision--language model judgments, while deterministic evaluation has largely focused on synthetic shape canvases, with application-oriented extensions primarily limited to charts. This makes it difficult to determine precisely how much of a requested edit was executed, where unintended changes occurred, and whether small differences between models reflect genuine editing capability or evaluator uncertainty. To bridge this gap, we present VeriEdit-Bench, a benchmark for fine-grained, instruction-faithful image editing across realistic structured assets with deterministic, four-axis evaluation. Its 1,740 cases are compiled from the source code of 153 Scalable Vector Graphics (SVG) graphics, charts, web interfaces, and presentation slides. Controlled source-code edits preserve the original visual context while yielding exact target images, pixel-level edit masks, and explicit edit specifications, enabling reproducible scoring along four axes: edit fidelity, preservation, localization, and magnitude. Evaluating eleven editors, we find that even the strongest model remains far from full credit; rankings for the same recoloring operation reverse between charts and SVG graphics; and outputs with similar pixel-accuracy profiles can still differ substantially in localization and change magnitude. This decomposition yields graded, verifiable feedback and exposes model-specific capability and failure profiles that holistic scores or evaluator-dependent judgments may obscure.
|
| 50 |
Detecting Defects that Matter: An Application-Driven Benchmark for Anomaly Detection in Manufacturing and Retail Logistics (VAND 4.0 Challenge)
2610.04392
|
cs.CV
|
Lars Heckler-Kram, Dorian Henning, Ashwin Vaidya, Jan-Hendrik Neudeck, Ulla Scheler |
Existing Anomaly Detection benchmarks are saturated and often unrealistic. As part of the VAND 4.0 Challenge, we introduce a hidden-test, application-driven benchmark across two deployment-critical domains: industrial manufacturing and retail logistics. In the...Existing Anomaly Detection benchmarks are saturated and often unrealistic. As part of the VAND 4.0 Challenge, we introduce a hidden-test, application-driven benchmark across two deployment-critical domains: industrial manufacturing and retail logistics. In the Industrial Track (MVTec AD 2), the results reveal that unsupervised anomaly segmentation remains challenging: the best regular-setting method achieves only ~57\% pixel-level $SegF_1$, indicating substantial room for improvement. Zero-shot approaches trail by ~15 $SegF_1$ points, confirming that task-specific training on normal data remains essential for precise defect localization. Robustness to distribution shifts remains a key open challenge and DINOv3-backbones clearly dominate this track. In the Retail Track (Kaputt 2), the results reveal that (1) supervised defect detection is approaching saturation for common defect types; (2) the best off-the-shelf VLM approach trails specialized models by ~28 AP, confirming that currently VLMs cannot replace fine-tuned detectors, (3) reference images did not prove helpful for top-performing approaches. Performance collapses on rare defects (spillage ~53 AP, missing units ~27 AP), where the supervised ceiling is bounded by data availability. To drive future progress in this domain, we provide a new low-prevalence retail AD dataset (Kaputt-Rare). Across both tracks, computational efficiency is assessed as a first-class metric combining performance, throughput, memory, and power consumption. We introduce a novel metric for measuring efficiency and reveal that that top-performing methods rely on heavy architectures while efficiency is largely neglected. Overall, we conclude that the community needs (a) more efficiency-aware method development, and (b) true anomaly detection approaches for rare defects and shifting conditions. https://sites.google.com/view/vand4-cvpr2026/challenge
|
| 51 |
AgroGround: Multi-Granularity Grounded Recognition in Agriculture
2610.04425
|
cs.CVcs.AI
|
Abdulla Alshehhi, Zongyan Han, Rao Anwer |
Agricultural visual models are typically evaluated for either recognition or localization, but reliable diagnosis requires identifying what is present and localizing the evidence. Agricultural visual question answering (VQA) datasets carry rich semantic labels...Agricultural visual models are typically evaluated for either recognition or localization, but reliable diagnosis requires identifying what is present and localizing the evidence. Agricultural visual question answering (VQA) datasets carry rich semantic labels but rarely link them to image regions, and adding such annotations by hand is costly at scale. We introduce AgroGround, a large-scale dataset for grounded agricultural recognition: identifying plant diseases and other agricultural targets and localizing their image regions. An automated pipeline converts the labels of eight agricultural VQA datasets into annotations for disease lesions and whole objects, producing 794,850 instruction examples. Healthy images provide negative supervision for disease queries, teaching the model to return empty predictions. We fine-tune a shared vision-language model on known-target grounding instructions combined with instructions requiring both recognition and localization. We evaluate predicted identities, regions, joint correctness, and healthy-image abstention on 1,480 human-verified images disjoint from all training data. Grounding-only fine-tuning reduces recognition accuracy from 51.8\% to 29.1\%, while adding recognition-and-localization instructions raises it to 72.6\%. With images and annotations held fixed, combining the two formats raises joint accuracy from 19.2\% to 43.3\% at comparable grounding. Healthy negatives raise abstention on healthy images to 95.0\%, and reinforcement learning improves lesion-level grounding. The resulting 2B model exceeds its annotation teacher in grounding F1 on our benchmark and on the external PlantSeg test set. AgroGround establishes a benchmark for grounded agricultural recognition, measuring joint correctness of identity and localization along with abstention on healthy images. The code is available at https://github.com/AB-Abdulla/AgroGround.
|
| 52 |
Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos
2610.04432
|
cs.CVcs.AI
|
Jinzhou Tang, Zijun Zhang, Jing Yang, Yuchen Yan, Kun Zhou |
Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents c...Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5\% to over 15\%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.
|
| 53 |
Multi-Crop Leaf Disease Recognition: A Unified Benchmark and Cross-Region Study
2610.04456
|
cs.CV
|
Rosemary Nalwanga, Sebastian Bunda, Godliver Owomugisha, Luuk Spreeuwers, Estefania Talavera Martinez |
Deep learning models for crop leaf disease recognition routinely report near-perfect accuracy yet are typically trained and evaluated on a single dataset collected under controlled laboratory conditions, leaving their behavior under realistic cross-region doma...Deep learning models for crop leaf disease recognition routinely report near-perfect accuracy yet are typically trained and evaluated on a single dataset collected under controlled laboratory conditions, leaving their behavior under realistic cross-region domain shift poorly understood. We introduce MLD (Multi-crop Leaf Disease) dataset, a unified multi-region benchmark that combines six public crop-disease datasets from the USA, Asia, and Africa into a shared hierarchical taxonomy spanning 18 crops, 56 crop-disease classes (including one healthy class per crop) making 167,427 images. We define standardized single-source and pooled multi-source evaluation protocols that explicitly probe cross-region generalization. We also investigate whether exploiting the inherent crop-to-disease dependency via a hierarchical formulation (HiLeaD) that conditions disease prediction on the predicted crop improves recognition under cross-region shift. Under the HiLeaD, the model trained on PlantVillage achieves 99.07% in-domain disease F1 but collapses to 12.88% when tested on PlantDoc, exposing a severe cross-region domain gap. The model trained on the pooled MLD dataset partly recovers cross-region disease F1 from 12.88% to 39.64% on PlantDoc (HiLeaD), achieving a 26.76 percentage point improvement. The hierarchical formulation provides a consistent additional gain, ranging from 1.71 to 4.88 percentage points in disease F1 over the flat baseline under the MLD dataset indicating that progress in this area is currently limited more by data coverage and diversity than by model design.
|
| 54 |
RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers
2610.04457
|
cs.CV
|
Mengyuan Fan, Bokai Huang, JiaMing Pan, Xiaokun Yuan, Peizhuang Cong |
Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantiza...Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for nn.Linear and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33% Top-1 / 94.48% Top-5 under W2/A4, Swin-T reaches 79.30% Top-1 / 94.79% Top-5 under W2/A8, and DeiT-S reaches 77.41% Top-1 / 93.11% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly $5.4$-$7.1\times$ relative to FP32 and end-to-end on-device latency drops by $1.4$-$1.6\times$. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.
|
| 55 |
Homogeneous Semantic Alignment and Hierarchical Expert Routing for Radiology Report Generation
2610.04499
|
cs.CVcs.CLcs.LG
|
Erjian Zhang, Jiayuan Ma, Liejun Wang, Yikemaiti Sataer, Xiaoming Tao |
Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the inc...Radiology report generation (RRG) aims to convert medical images into diagnostic texts to assist in clinical decision-making and alleviate the workload of physicians. Although existing methods have made extensive progress in cross-modal interaction and the incorporation of external priors, the distribution shift of underlying representations and the undifferentiated rigid coupling of heterogeneous information cause weak visual abnormality cues to be easily diluted by massive text priors and generation inertia during decoding. To overcome this bottleneck, inspired by cognitive science, we propose a novel two-stage Homogeneous Semantic Alignment and Hierarchical Expert Routing (HSA-HER) framework. First, the model introduces an explicit homogeneous distribution constraint in the underlying latent space to effectively eliminate the cross-modal distribution shift between visual and textual features, thereby extracting purified visual features as semantic anchors that accurately align with diseases. Second, for heterogeneous clinical evidence composed of visual features, local entities, and global retrievals, we design a hierarchical expert routing mechanism guided by these disease semantic anchors. This mechanism abandons the undifferentiated rigid coupling paradigm. Specifically, it dynamically activates expert networks to perform targeted mining and semantic reconstruction on multi-source evidence, and adaptively allocates fusion weights. Extensive experiments on three mainstream benchmark datasets demonstrate that HSA-HER achieves state-of-the-art performance, accurately depicting complex imaging details and key diagnostic information.
|
| 56 |
Localization Lens for Improving Medical Vision-Language Models
2610.04502
|
cs.CVcs.AI
|
Hasan Farooq, Murtaza Taj, Mehwish Nasim, Arif Mahmood |
Medical Vision-Language Models (Med-VLMs) have demonstrated strong capabilities in clinical tasks. However, they often struggle to understand anatomical structures and spatial positioning, which are crucial for medical reasoning. To address this, we propose a ...Medical Vision-Language Models (Med-VLMs) have demonstrated strong capabilities in clinical tasks. However, they often struggle to understand anatomical structures and spatial positioning, which are crucial for medical reasoning. To address this, we propose a localization-aware enhancement to the Med-VLM pipeline, introducing improvements at three levels: data,architecture, and alignment. First, we introduce localization lens, a set of expert-validated representations that provide richer anatomical and positional context. However, as these representations increase input complexity, we integrate pixel shuffle within the model architecture to filter and refine representations, enhancing spatial information processing while preserving anatomical continuity. Lastly, to effectively align the localization lens representations with textual features, we incorporate decoupled contrastive loss (DCL) alongside the standard loss function. This ensures better feature discrimination and robustness, particularly in data limited medical settings. Through extensive evaluations on medical visual question answering (Med-VQA) datasets, we show that our methodology improves localization-driven performance across different Med-VLM architectures. Our analysis of localization-based questions further reveals that improvements in anatomy and spatial reasoning directly enhance the overall accuracy of Med-VQA upto 6.2%. The proposed approach is model-agnostic and can be seamlessly integrated into existing Med-VLM pipelines. The dataset, code, and trained models will be made publicly available at https://github.com/CVLABLUMS/localizationlens.
|
| 57 |
EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning
2610.04506
|
cs.CVcs.AI
|
Yutong Li, Molin Wang, Xiaotong Li, Yanyan Fang, Daoguo Dong |
Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their a...Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their ability to reason directly about future visual states underexplored. We introduce EgoExo-Next, a visual-option benchmark for dynamic visual-state reasoning, where models must identify how an observed action trajectory subsequently appears rather than predict only an action label or textual description. EgoExo-Next contains 2,503 human-curated four-choice questions from six public egocentric and ego--exo video sources and comprises four interconnected subtasks that evaluate egocentric next-state prediction, bidirectional ego--exo state correspondence, exocentric next-state prediction, and their composition in Ego-to-Exo Next-State. Extensive evaluation of proprietary, open-source, and spatial reasoning VLMs reveals a substantial human--model gap, with the best model achieving 43.81\% average accuracy compared with 98.55\% for humans, and the largest degradation occurring on the composed Ego-to-Exo task. These results suggest that current VLMs remain substantially limited in dynamic visual-state reasoning, particularly when temporal progression and cross-view reasoning must be composed. The benchmark is publicly available at \url{https://huggingface.co/datasets/yutongli2024/EgoExo-Next}.
|
| 58 |
TAME:Topology-Aware Text-Driven Motion Editing across Heterogeneous Humanoid Skeletons
2610.04529
|
cs.CV
|
Qichen Zheng, Siyuan Yang, Chong Wang, Jun Liu, Shijian Lu |
Text-driven motion editing modifies an existing motion sequence according to a text instruction while preserving the content of the source motion. Existing methods are typically built for a single, fixed skeletal topology, which limits their use in animation p...Text-driven motion editing modifies an existing motion sequence according to a text instruction while preserving the content of the source motion. Existing methods are typically built for a single, fixed skeletal topology, which limits their use in animation pipelines where characters differ in joint count and skeletal hierarchy. We present Topology-Aware Motion Editor (TAME), a flow-matching transformer that edits motions on humanoid skeletons of varying topology. TAME represents motion as per-joint, per-frame tokens and models interactions among joints, across frames, and with the text instruction through skeletal, temporal, and text cross-attention layers. To make the skeletal attention follow each character's hierarchy, TAME replaces full joint attention with Topology-Constrained Skeletal Propagation (TCSP), which restricts attention to one-hop kinematic neighbors in the skeleton's adjacency matrix. We further introduce Edit-Focused Representation Alignment (EFRA), a self-distilled representation alignment strategy that aligns student features with cleaner EMA-teacher features exclusively on edit-relevant joint-time tokens, making edits faithful to the instruction. To make this setting trainable and comparable, we construct TopoMotionFix, a multi-topology extension of MotionFix with seen- and unseen-topology evaluation protocols. TAME outperforms previous methods in edit alignment and source preservation on MotionFix and reliably edits motions on unseen skeletons in TopoMotionFix.
|
| 59 |
FASTER: Fast Adjoint Stochastic Transport for Endpoint Refinement in Reward-Guided Image Editing
2610.04538
|
cs.CV
|
Yimiao Zhou, Zejia Zhong, Jingya Wang, Ye Shi |
Reward-guided image editing at test time seeks to improve a specified reward while preserving source content and visual plausibility. Many existing approaches optimize candidates through pretrained generation processes, making repeated adjustment depend on cos...Reward-guided image editing at test time seeks to improve a specified reward while preserving source content and visual plausibility. Many existing approaches optimize candidates through pretrained generation processes, making repeated adjustment depend on costly large-model execution and, in some cases, backbone backpropagation. We develop a theoretical framework that jointly accounts for reward, source preservation, and pretrained-prior preferences, allowing the desired output distribution to be specified separately from the dynamics used to realize it. Based on this framework, we introduce FASTER, which trains a small network for each source and objective to perform inexpensive editing, while pretrained and reward models provide feedback on candidate outputs. By reusing each candidate and its feedback across multiple small-network updates, FASTER reduces repeated sampling and supervision queries without placing the pretrained generative backbone inside the inner optimization loop. On SD3, FASTER leads all four target metrics and several validation metrics among the evaluated methods. Compared with the evaluated baseline that optimizes controls along pretrained generation trajectories, FASTER achieves editing-time speedups of up to \({6.91\times}\) on Stable Diffusion 3 and \({24.14\times}\) on Stable Diffusion 1.5.
|
| 60 |
EagleDepth: Efficient Fine-Grained Depth Estimation via Pixel Diffusion Decoder
2610.04554
|
cs.CV
|
Bowen Chai, Tianbao Zhang, Shuyu Wu, Dexin Zuo, Zhaoxin Fan |
Recovering detailed geometry from high-resolution images is critical for precise perception of the surroundings and objects. However, existing methods which use latent-space modeling and VAE reconstruction can compromise geometric details. Furthermore, decodin...Recovering detailed geometry from high-resolution images is critical for precise perception of the surroundings and objects. However, existing methods which use latent-space modeling and VAE reconstruction can compromise geometric details. Furthermore, decoding from latent codes introduces substantial inference overhead. To address those issues, we present EagleDepth, an efficient framework for high-resolution monocular depth estimation that combines the geometric priors of latent diffusion with fine-grained pixel-space generation. Our key idea is to retain depth-aware latent representations as guidance while generating the final depth map directly in pixel space. We train the latent and pixel components sequentially: first, we fine-tune a pretrained latent diffusion model using paired RGB--depth supervision; then, we adapt a pretrained pixel diffusion decoder, PiD, to predict depth conditioned on the learned features. Training of the pixel component starts at 1024 resolution and continues across multiple resolutions up to 4K. The latent branch processes resized, lower-resolution RGB images, while the pixel branch generates depth at the target resolution, bypassing the original VAE decoder. This design preserves learned geometric knowledge without requiring the latent backbone to operate at the output resolution. On five commonly used depth estimation datasets and the high-resolution Synth4K dataset, our framework achieves state-of-the-art depth estimation performance, with faster inference and better preservation of fine structures and object boundaries.
|
| 61 |
Frozen in a Frame: The Velocity Blind Spot in JEPA World Models
2610.04585
|
cs.CV
|
Tinghe Zhang, Chunyu Liu, Yu Leon Liu, Zerui Zhao, Jiaheng Chen |
Joint-embedding predictive architectures (JEPAs) for world modeling train an encoder so a predictor maps a current embedding and action to the next frame's embedding, always from a single rendered frame. This has a structural blind spot: a renderer without mot...Joint-embedding predictive architectures (JEPAs) for world modeling train an encoder so a predictor maps a current embedding and action to the next frame's embedding, always from a single rendered frame. This has a structural blind spot: a renderer without motion blur draws a scene from configuration alone, so a single-frame embedding carries no velocity information, for any encoder, including the official released LeWM weights. We confirm this on official checkpoints across four real benchmarks (PushT, Reacher, Cube, TwoRoom): every linear velocity probe sits at or below chance while position probes reach R^2 about 0.95. We introduce RateIdent, a three-stage diagnostic protocol, and TI-JEPA, a lightweight fix splitting the latent into a pose code and an explicit finite-difference motion code, predicted jointly. Across three physically grounded environments, TI-JEPA gives a significant, seed-robust gain on a stop-at-goal planning task over a matched-memory baseline, e.g. 55% lower final distance on Pendulum (p=3.2x10^-10) and 64% on CartPole (p=5.1x10^-15). We reproduce this at official ViT-Tiny plus AdaLN-transformer scale, then push the same recipe onto real dm_control Reacher photographs trained from scratch, where TI-JEPA's branch separation exceeds the memory-having baseline's by roughly 38x, the paper's largest margin. Against a same-footprint recurrent RSSM-style predictor, TI-JEPA matches or beats its rollout accuracy on two of three environments, stays separately probeable for pose and motion, and wins outright on the most coupled one. A checkable formal argument and six evaluated environments show single-frame targets are the wrong object to predict when velocity matters, and a small, interpretable structural change fixes it with no privileged supervision. Code, checkpoints, and the project page are linked below the title.
|
| 62 |
WASP: Weakly Aligned Spatiotemporal Pairs for Fetal Brain MRI-Ultrasound Learning
2610.04601
|
cs.CVcs.LG
|
Francesco Correnti, Gabriele Magrini, Marco Mistretta, Niccol\`o Biondi, Pietro Pala |
Magnetic Resonance Imaging (MRI) is widely regarded as the optimal sensor for fetal brain analysis due to its superior soft-tissue contrast and anatomical detail. However, its high cost and operational burden make it invasive and difficult to obtain at scale. ...Magnetic Resonance Imaging (MRI) is widely regarded as the optimal sensor for fetal brain analysis due to its superior soft-tissue contrast and anatomical detail. However, its high cost and operational burden make it invasive and difficult to obtain at scale. Ultrasound (US), in contrast, is cheap, safe, and routinely acquired, and as a result it has produced substantially larger datasets and a growing ecosystem of pretrained models. This asymmetry raises a natural question: Can we teach a US-only model to understand fetal MRI from only a limited set of examples? The standard recipe, training a foundation model on subject-to-subject paired MRI-US scans, is not viable since no such paired fetal dataset is publicly available. In this paper we address this gap with Weakly Aligned Spatiotemporal Pairs (WASP), a framework that formulates cross-modal correspondence as an entropic Optimal Transport problem driven by clinical metadata, in particular Gestational Age (GA) and diagnostic planes, enabling the fitting of a lightweight alignment module that lifts MRI representations into the US latent space, without fine-tuning the backbone. Empirically, WASP yields its largest gains when MRI is unseen by the model during pretraining (on USFM, GA estimation error drops from 21.9 to 17.4 days and standard plane classification accuracy climbs from 61.9% to 69.0%), while providing smaller, backbone-dependent refinements for backbones pretrained on both modalities (e.g., BioMedParse GA estimation error from 6.5 to 6.0 days and SAM-Med2d plane accuracy from 83.3% to 88.1%). Code is available at https://github.com/miccunifi/WASP.
|
| 63 |
Organize Primitives into Semantic Parts: Reinforcement Reasoning for 3D Segmentation
2610.04602
|
cs.CV
|
Xiaoming Gong, Ruoyu Wu, Zhenhong Sun, Chunlin Chen, Daoyi Dong |
Primitive-based 3D segmentation offers a compact and explicit alternative to dense surface prediction, naturally supporting structural abstraction and boundary localization. However, geometric decomposition alone does not determine how primitives should be org...Primitive-based 3D segmentation offers a compact and explicit alternative to dense surface prediction, naturally supporting structural abstraction and boundary localization. However, geometric decomposition alone does not determine how primitives should be organized into semantic parts: a single part may span multiple primitives, while geometrically similar or touching primitives may belong to different parts. We therefore introduce RePart (Reinforcement Part Reasoning), which formulates primitive-to-part organization as a finite-horizon Markov decision process and learns semantic organization through trajectory-level reinforcement reasoning. RePart constructs a Composable Primitive Workspace from fine-grained superquadrics and applies a merge-and-stop policy whose decisions are optimized by their downstream effects on the resulting partition rather than local primitive compatibility. The inferred part identities are then mapped back to the original mesh through Boundary-Aware Surface Labeling, preserving accurate surface boundaries beyond the primitive approximation. On PartNet, RePart achieves the strongest results across all four aggregate partition metrics; on 3DCoMPaT++, it obtains the highest RI and SC without target-dataset fine-tuning. These results demonstrate that reinforcement reasoning provides an effective mechanism for organizing geometric primitives into semantic parts while retaining dense segmentation accuracy. Code is available at https://github.com/EngineeringAI-LAB/RePart.
|
| 64 |
ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution
2610.04605
|
cs.CVcs.LGcs.AI
|
Yehonatan Elisha, Oren Barkan, Ziv Weiss Haddad, Noam Koenigstein |
Many visual explanation methods in computer vision highlight pixel importance but struggle to link these low-level cues to semantically meaningful concepts, limiting their interpretability and trustworthiness. We introduce Concept-based Explanations (ConEx), a...Many visual explanation methods in computer vision highlight pixel importance but struggle to link these low-level cues to semantically meaningful concepts, limiting their interpretability and trustworthiness. We introduce Concept-based Explanations (ConEx), a novel framework that bridges saliency visualization with concept-based reasoning to provide both faithfulness and interpretability. ConEx automatically discovers class-specific concepts and represents them through concept activation vectors (CAVs), learned without manual supervision using an architecture-specific masking mechanism that reduces noise introduced by the segmentation masks to enhance concept purity. ConEx generates faithful saliency maps that reveal where each concept appears in the image and how it contributes to the prediction. To evaluate the reliability of these learned concepts, we propose two complementary metrics, Vector-Concept Match (VCM) and Concept-Class Match (CCM), that quantify concept alignment and enable direct comparison with existing methods. Extensive experiments across diverse settings demonstrate that ConEx achieves state-of-the-art performance on faithfulness, segmentation, and concept-quality benchmarks. Overall, ConEx advances the field toward truly interpretable and concept-grounded explanations in vision models.
|
| 65 |
Sparse-View 4D Gaussian Splatting via Spatiotemporal Priors and Generative Assistance
2610.04606
|
cs.CV
|
Shengqi Wang, Zhengxian Yang, Kaiwen Tian, Yang Liu, Bowen Liu |
We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under s...We present a 4D Gaussian Splatting framework for the Sparse-View Track of the SIGGRAPH Asia 2026 Volumetric Video Challenge, which requires dynamic scene reconstruction from only six cameras with wide baselines. To achieve robust dynamic reconstruction under such sparse views, our framework integrates three components. (1) Region-adaptive spatial priors: We use foreground masks to guide Gaussian initialization and mask voting to control densification separately for the dynamic foreground and static background. Background geometry is regularized using monocular depth aligned to metric scale. (2) Motion-consistent temporal priors: We provide supervision at intermediate times through frame interpolation and constrain projected Gaussian motion with estimated optical flow. (3) Generative assistance: We place virtual cameras in the widest angular gaps and restore their rendered images using a diffusion-based model conditioned on camera poses. The restored images are iteratively incorporated into training as pseudo-supervision. On the validation set, our framework improves full-frame PSNR from 25.60 dB for the baseline to 29.75 dB. On the official test benchmark, it achieves 30.04 dB full-frame PSNR and 27.88 dB foreground PSNR, ranking first overall in the Sparse-View Track.
|
| 66 |
Task-Sensitive Geometry of Representation Transfer for Object Detection under Image Degradation
2610.04627
|
cs.CV
|
Van Vung Pham |
Object detection under image degradation can benefit from clean-image supervision, but aggregate gains do not imply that transferred representation changes are uniformly useful. We study how clean task knowledge affects degraded-image representations and wheth...Object detection under image degradation can benefit from clean-image supervision, but aggregate gains do not imply that transferred representation changes are uniformly useful. We study how clean task knowledge affects degraded-image representations and whether local responses to structured representation directions can be characterized geometrically. Using paired clean and Gaussian-degraded BDD100K images, we show that clean-teacher distillation improves observed detection accuracy while producing heterogeneous object-level transfer. We isolate a representation component complementary to direct clean-teacher alignment and map it into the distilled student space through an orthogonal bridge. Controlled interventions rescue 13.22% of objects lost under the distilled representation, versus 4.30% under norm-matched random perturbations, with very low harm on preserved objects. We introduce task-sensitive geometry, a gradient-derived channel-space geometry constructed from normalized detection-loss gradients. On a reserved cohort, mapped-complement orientation within this frozen geometry is positively associated with local intervention-response magnitude after controlling for intervention magnitude (partial Spearman $\rho$ = 0.242, 95% CI [0.108, 0.359]). The relationship eplicates on independent data ($\rho$ = 0.180) and with RT-DETR-L ($\rho$ = 0.227), but not for the direct clean-teacher residual family, and it weakens for large interventions. Routing rules and specialized distillation objectives based on these signals do not yield statistically reliable gains over CLEANKD. These results support a local, direction-family-dependent task-sensitive geometry while showing that converting such structure into improved global training remains an open problem.
|
| 67 |
FLASHSWIN: Unlocking Large Windows and Dense Tokens in Swin Vision Transformers with Memory Efficient Attention
2610.04664
|
cs.CV
|
Tushar Kataria, Gerald Sabin, Ponnuswamy Sadayappan, Shireen Y. Elhabian |
High-resolution vision backbones have long been forced to trade away local token density to afford larger receptive fields. Hierarchical Swin transformers impose this compromise because standard windowed attention materializes an $M^2\times M^2$ score matrix p...High-resolution vision backbones have long been forced to trade away local token density to afford larger receptive fields. Hierarchical Swin transformers impose this compromise because standard windowed attention materializes an $M^2\times M^2$ score matrix per window, incurring $O(M^4)$ memory as windows or token grids grow. Furthermore, Swin adds a learned relative-position bias elementwise to attention scores, requiring full materialization of the score matrix and its gradient. This keeps Swin and SwinV2 trapped in a small-window($M=8,16$), coarse-token regime with patch size $4\times4$ ($p=4$), limiting performance for fine-grained tasks. We introduce FLASHSWIN, which replaces standard windowed attention with a FlashAttention implementation that computes exact softmax attention without materializing the score matrix, reducing per-window memory from $O(M^4)$ to $O(M^2)$. This enables higher token density and larger receptive fields without inflating memory overhead. Training memory is flat across window sizes: at a $32\times32$ window, FLASHSWIN-T requires only $12.4$\,GB, unchanged from $8\times8$, compared to $70/90$\,GB for SwinV2/V1-T. However, applying FlashAttention directly to Swin creates a trade-off: bypassing the score matrix precludes Swin's additive relative-position bias, forfeiting spatial information in exchange for memory efficiency. FLASHSWIN restores position information as window-local learnable 2D RoPE, making large windows and dense token grids both affordable and accurate. At matched scale, FLASHSWIN-T outperforms Swin variants. With dense tokens and wide windows ($p=2,M=32$), the same Tiny model reaches $84.1\%$ ImageNet-1K, $44.1$ COCO box AP, and $47.28$ ADE20K mIoU---gains of $+1.3$, $+5.1$, and $+1.82$ over SwinV2-T at $M=16$, respectively. At fixed $M=32$, halving the patch size yields roughly $3\times$ larger gains in boundary quality than in mIoU.
|
| 68 |
COMPASS: Comet Object Measurement Pipeline with Automated Selection and Scoring
2610.04680
|
cs.CV
|
Jack Roberts, Canya Lu, Alexis Michelle Lawson, Kaitlyn Holden, Gerald S. Wilkinson |
Summary: The single-cell gel electrophoresis ('comet') assay is a widely used technique for quantifying DNA damage at the individual cell level. However, image analysis often relies on manual inspection or semi-automated software, which can be labor-intensive,...Summary: The single-cell gel electrophoresis ('comet') assay is a widely used technique for quantifying DNA damage at the individual cell level. However, image analysis often relies on manual inspection or semi-automated software, which can be labor-intensive, difficult to reproduce, and sensitive to image quality and comet morphology. COMPASS automates comet assay image analysis by combining deep learning-based comet segmentation with damage measurement and automated comet selection. The pipeline produces standardized DNA damage measurements while offering robust detection, reducing manual effort and improving reproducibility through transparent selection and optional manual review. Availability and implementation: COMPASS is implemented in Python and is freely available at https://github.com/ rsinghlab/COMPASS. Installation instructions, pretrained weights, and example usage are provided in the repository. Contact: jack_roberts2@brown.edu, ritsingh@illinois.edu Supplementary information: Available onlime upon publication.
|
| 69 |
Decouple, Purify and Unite: Semantic-Structural Prototype Learning for Federated Medical Segmentation
2610.04700
|
cs.CV
|
Xingyue Zhao, Wenke Huang, Linghao Zhuang, Yanzhou Su, Zhifeng Wang |
Federated learning enables medical institutions to train a global model without sharing data, yet feature heterogeneity from diverse scanners or protocols remains challenging. Existing representation-based methods face two limitations: 1) Incomplete Contextual...Federated learning enables medical institutions to train a global model without sharing data, yet feature heterogeneity from diverse scanners or protocols remains challenging. Existing representation-based methods face two limitations: 1) Incomplete Contextual Representation Learning: single-layer or coupled representations overlook multi-level structural cues and entangle regional semantics with boundary details. 2) Layerwise Style and Aggregation Biases: domain-specific style discrepancies across intermediate layers degrade prototypes, while aggregation that overlooks client distribution shifts can further amplify bias. We propose FedBCS+, federated decoupled contextual alignment with style-purified aggregation. We employ Frequency-domain Style Recalibration (FSR) in prototype construction to decouple content-style representations and extract style-purified prototypes. Built upon these purified features, Decoupled Contextual Prototype Alignment (DCPA) explicitly decouples multi-level features into semantic and structural prototypes and aligns regional semantics and fine-grained anatomical structures separately. Style-purified Semantic Prototype Aggregation (S2PA) measures each client's purified prototype divergence from the global consensus and adaptively reweights aggregation toward under-represented clients to reduce consensus bias. On five heterogeneous medical segmentation benchmarks spanning histopathology, MRI, ultrasound, and colonoscopy, FedBCS+ achieves the highest mean Dice among the compared methods. A convergence analysis further characterizes how aggregation and alignment affect the optimization bound.
|
| 70 |
Knossos and Ariadne: Benchmarking and Learning Complete Diagram Topology Extraction with Vision-Language Models
2610.04721
|
cs.CVcs.AI
|
Bangwei Guo, Xujiang Zhao, Shengyu Chen, Yanchi Liu, Wei Cheng |
Structural diagrams are widely used to represent complex systems and relational information across scientific, engineering, procedural, and spatial domains. Recent vision-language models (VLMs) have become increasingly capable of recognizing diagram elements a...Structural diagrams are widely used to represent complex systems and relational information across scientific, engineering, procedural, and spatial domains. Recent vision-language models (VLMs) have become increasingly capable of recognizing diagram elements and reasoning about their content, while complete diagram topology extraction remains comparatively underexplored. In this paper, we study diagram-to-graph topology extraction: extracting all diagram entities and the complete relations among them. To enable large-scale supervised training and systematic evaluation of this task, we introduce Knossos, a benchmark of 19,200 diagrams across six diverse domains, with 245,179 nodes and 439,740 edges. Its symbolic generation process provides exact alignment between rendered diagrams and annotations of complete topology, relation types, and connector geometry. To address the modeling challenge of complete topology extraction, we also present Ariadne, a structured framework that decomposes the task into node inventory extraction and source-conditioned edge prediction. Extensive experiments show that training on Knossos substantially improves complete topology extraction in smaller open-source VLMs. Ariadne further improves over one-step extraction under matched supervision, demonstrating the additional benefit of structured decomposition. It achieves the highest average Edge F1 among the evaluated methods on Knossos, while both backbone variants also improve over their unadapted counterparts on the real-world external benchmark. Code and benchmark are available at https://github.com/bangwayne/knossos_Ariadne_Public.
|
| 71 |
NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis
2610.04722
|
cs.CV
|
Ramil Khafizov, Ilya Statsenko, Ruslan Rakhimov, Artem Komarichev, Peter Wonka |
Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that r...Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis. Additional qualitative results, videos, and resources are available at https://corl-team.github.io/namvis/
|
| 72 |
Investigating Spatiotemporal Redundancy in Video Transformer for Collision Anticipation
2610.04727
|
cs.CV
|
Xiaoshan Zhou |
In worker-equipment proximity monitoring, video transformers are widely used for collision anticipation and have demonstrated strong performance. However, their accuracy comes with substantial computational demands, creating a tension with the need for low-lat...In worker-equipment proximity monitoring, video transformers are widely used for collision anticipation and have demonstrated strong performance. However, their accuracy comes with substantial computational demands, creating a tension with the need for low-latency inference on mobile robots and the pursuit of lower-carbon computation in construction. To address this, this study investigates where computation within an established video transformer is redundant and whether that redundancy can be removed without materially degrading predictive performance. Using VideoMAEv2-Base on the Nexar Collision Prediction dataset, we first examine how collision-relevant information evolves across network depth and then investigate two complementary forms of redundancy: structured capacity redundancy in multilayer perceptrons (MLPs) and spatiotemporal redundancy in the token stream. Linear probes show that interpretable motion cues, including flow magnitude, looming, and approach versus retreat, are most accessible at intermediate layers, whereas collision-label discrimination strengthens toward the final layer. Token redundancy is axis-specific: adjacent temporal-token similarity reaches 0.970 in later layers, while spatial similarity falls to 0.297, indicating substantially greater redundancy across time than across space. Exploiting this asymmetry, temporal token merging reduces backbone computation from 356.99 to 178.50 GFLOPs and latency from 12.69 to 7.21 ms per clip, a 1.76x speedup, while mean average precision changes only from 0.7478 to 0.7443. Importance-guided retention of 50% of MLP units preserves an AUC of 0.753, compared with 0.529 under matched random retention, and reveals that pruning alters score calibration before discriminative ranking collapses. These findings establish a new pathway for pursuing faster algorithms through targeted temporal token compression and neuron pruning.
|
| 73 |
Probabilistic Pedestrian Forecasts from a Handheld Phone: World-Frame Heat Maps, Visual-Inertial Height Drift, and Evaluation without Ground Truth
2610.04736
|
cs.CV
|
Danial Safaei |
A pedestrian with a phone could be shown where nearby people will be in the next few seconds, if the forecast stays on the ground while the phone moves, is calibrated, and needs only a monocular camera and visual-inertial odometry (VIO). We build and evaluate ...A pedestrian with a phone could be shown where nearby people will be in the next few seconds, if the forecast stays on the ground while the phone moves, is calibrated, and needs only a monocular camera and visual-inertial odometry (VIO). We build and evaluate such a system. People are detected, lifted onto the floor by ray-plane intersection, tracked in a gravity-aligned metric frame, and forecast as per-step probability maps by a small U-Net trained with a negative log-likelihood (NLL) loss on bird's-eye (SDD) and first-person (EgoTraj-Bench) trajectories. On handheld ADVIO recordings, vertical VIO drift and the user's own changes of level silently rescale monocular ground positions (by 87% within 90 s on one clip; on another, all tracks are lost for the last 31% of the clip); keeping the camera's height above the floor constant under a low-pass-filtered altitude avoids this, though it lags on escalators. On the SDD and EgoTraj-Bench test splits, the final forecaster lowers the NLL at 4.8 s by 1.51 and 1.37 nats relative to a fitted constant-velocity Gaussian. Lacking ground truth for people in handheld video, we score forecasts against the tracker's own later raw measurements. In an internally pre-registered evaluation on seven held-out clips, the final forecaster's NLL is lower than the benchmark-fitted baseline's at 1.2, 2.4 and 4.8 s (by 0.17, 0.23 and 0.47 nats; 95% intervals over people exclude zero), but by less than half as much as on the development clips. Exploratory analyses cut both ways: resampling clips instead of people widens the intervals to include zero at 1.2 and 2.4 s, and once both forecasters are recalibrated on the development clips the network is significantly better only at 1.2 s; but two clips run with ADVIO's reference poses favour the network much more when re-run with the phone's own poses. We discuss what such self-consistency scores can and cannot show.
|
| 74 |
ThyCLIPNet: A BiomedCLIP-Guided Lightweight Attention-Enhanced DeepLabV3+ Framework for Robust Thyroid Nodule Segmentation
2610.04743
|
cs.CV
|
Tasnim Jahan, Md Easin Arafat, Swakkhar Shatabda |
Accurate thyroid ultrasound segmentation is often challenged by low contrast, speckle noise, and unclear boundaries. Although recent methods have improved segmentation accuracy, many rely on resource-intensive architectures or lack explicit integration of mult...Accurate thyroid ultrasound segmentation is often challenged by low contrast, speckle noise, and unclear boundaries. Although recent methods have improved segmentation accuracy, many rely on resource-intensive architectures or lack explicit integration of multiscale features with global biomedical visual guidance. In this paper, we introduce ThyCLIPNet, a lightweight semantic-guided hybrid encoder-decoder framework that integrates BiomedCLIP-derived biomedical semantic guidance into a lightweight multi-scale CNN segmentation pipeline. The encoder integrates MobileNetV2 with efficient channel attention, while atrous spatial pyramid pooling and a custom convolutional block attention module enrich bottleneck features. The decoder combines hierarchical skip connections and lightweight attention refinement with a BiomedCLIP-guided gated fusion pathway that projects vision-only global biomedical embeddings into decoder feature space and selectively integrates them through semantic-local fusion and spatial gating. To the best of our knowledge, ThyCLIPNet is among the first lightweight thyroid ultrasound segmentation frameworks to use BiomedCLIP's vision encoder alone for image-only global semantic guidance without text prompting. Experiments on TG3K, TN3K, DDTI, and PKTN achieve dice similarity coefficients of 96.22%, 87.58%, 84.73%, and 80.70%; intersection over union scores of 92.72%, 77.91%, 73.51%, and 67.64%; and 95th-percentile hausdorff distances of 3.75, 16.38, 18.23, and 10.86, respectively. ThyCLIPNet uses 8.55M parameters and 22.99G FLOPs. Overall, the results support integrating global biomedical semantic guidance with lightweight multi-scale CNN representations for robust and computationally efficient thyroid ultrasound segmentation. Source code: https://github.com/Tasnim-Jahan/ThyCLIPNet. [Abstract shortened for arXiv. See PDF for full abstract.]
|
| 75 |
ARISE: Adaptive Agentic Reasoning with Image-grounded Self-Evaluation for Interpretable IBD Assessment
2610.04777
|
cs.CVcs.AI
|
Pronoma Banerjee, Anuva Shah, Jason Wu, Md. Masudur Rahman, Sanjay Mohanty |
Inflammatory bowel disease (IBD) requires frequent imaging-based assessment, yet interpretation of modalities such as wireless capsule endoscopy (WCE) and intestinal ultrasound remains heavily dependent on specialist expertise. Vision-Language Models (VLMs) de...Inflammatory bowel disease (IBD) requires frequent imaging-based assessment, yet interpretation of modalities such as wireless capsule endoscopy (WCE) and intestinal ultrasound remains heavily dependent on specialist expertise. Vision-Language Models (VLMs) demonstrate significant potential in multimodal medical image analysis, but their clinical adoption is hindered by their insufficient domain-specific reasoning, susceptibility to hallucination, scarcity of high quality training data in fine-grained diagnostics and limited interpretability. We introduce ARISE (Adaptive Agentic Reasoning with Image-grounded Self-Evaluation), an autonomous planning framework that models few-shot medical image understanding as a sequential agentic workflow. ARISE structures agent execution into a transparent 5-stage workflow: hypothesis generation, image-grounded evidence summarization, evidence-conditioned refinement, symbolic verification, and final diagnosis. We apply ARISE to IBD assessment across two independent patient cohorts: wireless capsule endoscopy (WCE) images for Crohn's disease and B-mode ultrasound data for ulcerative colitis. ARISE consistently improves diagnostic performance over baseline VLMs while exposing where reasoning succeeds or fails, providing a more interpretable basis for clinical decision support and realistic deployment.
|
| 76 |
Super-Resolution in The Right Latent Space: A Frozen Vision-Foundation Substrate
2610.04781
|
cs.CV
|
Wanzhou Lei, Cuifeng Sheng, Yanjin He, Maohua Li, Hua Yuan |
In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degra...In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.
|
| 77 |
Lollypop: Camera-to-Motion-Capture Calibration Verification with a Reference Target
2610.04785
|
cs.CV
|
Tianyi Liu, Kevin Harris, Mihika Dave, Kun He |
Camera-to-motion-capture (mocap) calibration is essential for using mocap as ground truth in robotics, AR/VR, and other computer vision tasks. However, the calibration can drift after deployment, while calibration residuals and visual inspection provide limite...Camera-to-motion-capture (mocap) calibration is essential for using mocap as ground truth in robotics, AR/VR, and other computer vision tasks. However, the calibration can drift after deployment, while calibration residuals and visual inspection provide limited independent verification. We present Lollypop, a fiducial-mocap reference target for independent calibration verification. The target couples an ArUco fiducial with a mocap marker constellation so the visual center and tracked centroid represent the same physical point. Given a candidate calibration, verification projects the mocap point into the image and measures its disagreement with the detected fiducial center. Experiments show sub-pixel nominal error, sensitivity to controlled extrinsic perturbations, and increasing error during an illustrative mixed-handling sequence.
|
| 78 |
Active-DiNTS: Active Differentiable Network Topology Search
2610.04787
|
cs.CVcs.LGcs.AI
|
Gean Trindade Pereira, Thierry Urruty, Muriel Visani, Andr\'e C. P. L. F. de Carvalho |
Neural Architecture Search (NAS) has proved to be a strong alternative to manual network design, but applying it to 3D medical image segmentation is limited by two well-known costs, large annotation budgets and multi-GPU clusters. Thus, this paper introduces A...Neural Architecture Search (NAS) has proved to be a strong alternative to manual network design, but applying it to 3D medical image segmentation is limited by two well-known costs, large annotation budgets and multi-GPU clusters. Thus, this paper introduces Active-DiNTS (Active Differentiable Network Topology Search), an approach that embeds pool-based Active Learning (AL) into a bi-level differentiable topology search to perform architecture discovery and label curation jointly. At each query round, unlabeled MRI volumes are ranked by one of three uncertainty signals (Entropy, Variance, or Standard Deviation), and only the top-ranked volumes are sent to an oracle for annotation. The new labels feed two interlocked stages. Network weights are updated in an outer loop, while the macro/micro topology of a U-Net-style backbone is refined in an inner loop. Three AL regimes (weights-only, topology-only, joint) expose the speed-accuracy trade-off. Evaluations on the Medical Segmentation Decathlon (MSD) Task01 BrainTumour benchmark showed that Active-DiNTS surpasses DiNTS, C2FNAS, and nnU-Net in Dice, with gains of about 10 percentage points on Edema and 5 points on Non-Enhancing core, on a single GPU and using a fraction of the labeled volumes. The discovered architectures are denser and more FLOP-heavy than prior baselines, but remain competitive in trainable parameters and peak memory; the fastest search variant finishes in under 0.25 GPU-days, over 27x faster than the eight-GPU DiNTS search. Together, these results indicate that pairing differentiable NAS with active data acquisition is a practical recipe for accurate 3D segmentation under realistic constraints.
|
| 79 |
MOXIE: Discovering Alternative Explanations for Biomedical Image Classifiers
2610.04814
|
cs.CVcs.LG
|
Abiha Tahsin Chowdhury, Rahul Dubey |
Segment-based explanation methods such as LIME return a single explanation for each prediction, computed from one fixed image segmentation. This hides two important facts: a prediction can be supported by many different sets of image segments, and the segmenta...Segment-based explanation methods such as LIME return a single explanation for each prediction, computed from one fixed image segmentation. This hides two important facts: a prediction can be supported by many different sets of image segments, and the segmentation itself shapes which explanations can be found. We introduce MOXIE (Multi-Objective eXplanation Imaging Engine), an evolutionary framework that searches for segment subsets that preserve the classifier's confidence while keeping as little of the image as possible. Instead of one explanation, MOXIE returns a Pareto front of alternative explanations that range from compact to highly faithful. We evaluate MOXIE with NSGA-II and four segmentation methods (SLIC, Felzenszwalb, Watershed and Voronoi) on BloodMNIST and HAM10000 datasets, using the same evaluation budget as LIME. Results show that MOXIE achieves a higher hypervolume than LIME on every image. LIME's explanations often appear convincing, yet the classifier's confidence collapses when only the highlighted segments are shown. MOXIE's fronts reveal how much of the image is needed to preserve the model's confidence and which contextual regions influence it. We also find that segmentation strongly affects evaluation: methods with unequal segment sizes appear most compact when segments are counted. These results show that alternative explanations provide a more complete view of a model's decision than a single explanation.
|
| 80 |
DriftSR: One-Step Real-World Image Super-Resolution via Distribution Drifting
2610.04819
|
cs.CV
|
Wei Zhu, Kai Zhang, Yu Zheng, Zhaopeng Yang, Lei Luo |
One-step real-world image super-resolution (Real-ISR) offers efficient inference, but recovering realistic and perceptually rich details often relies on score distillation or adversarial learning, introducing additional trainable components and making optimiza...One-step real-world image super-resolution (Real-ISR) offers efficient inference, but recovering realistic and perceptually rich details often relies on score distillation or adversarial learning, introducing additional trainable components and making optimization more cumbersome. To this end, we propose DriftSR, a one-step Real-ISR framework that leverages pretrained diffusion priors through distribution drifting. Specifically, we perform drifting in the frozen intermediate representation space of a pretrained diffusion model, without introducing an additional task-specific feature encoder. Building on this space, we introduce Spatial Feature Drifting, which treats spatial features rather than entire images as distributional samples, enabling denser supervision for distribution alignment. To mitigate structural deviations, we further introduce Structure-Modulated Guidance, which adaptively refines drifting guidance according to local structural consistency with the LQ input. Consequently, DriftSR optimizes only the one-step generator, without auxiliary distillation branches or adversarial discriminators. Extensive experiments on three real-world benchmarks demonstrate that DriftSR delivers high-quality super-resolution reconstruction with efficient one-step inference.
|
| 81 |
RSure-Agent: Reliable Use of Tool Observations for Remote Sensing Agents
2610.04836
|
cs.CV
|
Fuyuan Liu, Nayu Liu, Wenhao Yu, Peijin Wang, Yingchao Feng |
Remote sensing agents rely on perception, measurement, and raster analysis tools to solve Earth observation tasks. We refer to their judgments and quantitative results about ground objects as tool observations. However, these observations are subject to substa...Remote sensing agents rely on perception, measurement, and raster analysis tools to solve Earth observation tasks. We refer to their judgments and quantitative results about ground objects as tool observations. However, these observations are subject to substantial uncertainty and may be incorrect even when the tools execute successfully. When agents accept incorrect observations, the errors can propagate through subsequent reasoning and cause task failure. We analyze 1,229 execution trajectories across three remote sensing agent benchmarks. On each benchmark, at least 88.1% of tasks depend on tool observations. Among these tasks, at least 22.7% contain incorrect observations despite successful tool execution. These errors propagate to the final answer in at least 82.0% of affected tasks on each benchmark. To address this problem, we propose RSure-Agent, a framework for verifying tool observations and limiting error propagation. We introduce a verifiable observation protocol that requires tools to return process evidence for the agent to verify their observations. We also construct a task-tool reliability prior from offline task feedback. The prior summarizes each tool configuration's past performance across task types and provides a task-specific reference for verification. Using process evidence and this prior, RSure-Agent decides whether to accept an observation, request additional evidence, or reject it. We evaluate RSure-Agent on EarthBench, ThinkGeo, TerraLogic, and CHOICE-420. Across the three agent benchmarks, RSure-Agent reduces the error propagation rate by 21.3 to 25.9 percentage points relative to the base configuration with verification and the prior disabled. On CHOICE-420, it improves overall accuracy over direct answering by 5.71 percentage points on average across 11 backbone models. On EarthBench, it reduces the tool-call ratio by 25.9% relative to Earth-Agent.
|
| 82 |
One Tile, Multiple Instances: Rethinking MIL for Sparse Diagnostic Evidence
2610.04853
|
cs.CVcs.LG
|
Runsheng Liu, Cheng Jin, Hao Jiang, Hao Chen |
In weakly supervised Whole Slide Image (WSI) classification, feature extractors typically compress each image tile into a single global embedding. Consequently, slide-level aggregators are restricted to this coarse tile scale, concealing fine-grained sub-tile ...In weakly supervised Whole Slide Image (WSI) classification, feature extractors typically compress each image tile into a single global embedding. Consequently, slide-level aggregators are restricted to this coarse tile scale, concealing fine-grained sub-tile evidence from the attention mechanism. We introduce DI-MIL, a framework that decouples encoding context from instance granularity through decomposed instances. By clustering dense spatial tokens from a frozen foundation model within each tile, DI-MIL converts a single tile into multiple independently weighted instance embeddings. As a training-free post-encoding module, DI-MIL integrates seamlessly into existing pipelines without requiring re-encoding or downstream architectural modifications. We evaluate DI-MIL on cytopathology, a challenging testbed where sparse diagnostic signals are easily diluted within standard tiles. Across four datasets, three frozen foundation models, and two attention-based aggregators, DI-MIL demonstrates consistent efficacy, improving 67 of 72 metric-level comparisons, with the largest mean gains reaching 3.64 points under cytopathology-specific backbones. In a broader comparison against seven representative MIL baselines, DI-MIL paired with ACMIL achieves highest mean performance in 33 of 36 backbone-dataset-metric comparisons. Ablations show that direct smaller tiling inflates the extracted tile count by up to 43.3$\times$ with non-monotonic performance, whereas DI-MIL incurs zero additional image-extraction overhead while achieving the strongest overall results. These results establish instance construction as an orthogonal design dimension in MIL, supporting DI-MIL as a cost-efficient solution under sparse diagnostic evidence.
|
| 83 |
Enhancing Long-Video VLM Embeddings with Query-Aware Streaming Latent Reasoning
2610.04864
|
cs.CVcs.AI
|
Haozhe Chi, Song Jin, Yang Jin, Yadong Mu |
Long-video embedding requires capturing sparse query-relevant evidence under a limited visual-token budget. Uniform sampling can miss brief events in videos spanning minutes or hours, whereas encoding more frames in a single context increases memory and comput...Long-video embedding requires capturing sparse query-relevant evidence under a limited visual-token budget. Uniform sampling can miss brief events in videos spanning minutes or hours, whereas encoding more frames in a single context increases memory and computation. We introduce \textbf{Query-Aware Streaming Latent Reasoning} (QASLR), a post-training framework that accumulates evidence across clips while keeping the embedding size fixed. QASLR selects a bounded set of frames, restores their temporal order, and processes them clip by clip with a vision-language backbone. A compact set of persistent think tokens cross-attends to each clip's features, while an embed token reads out a normalized representation after every update. This design integrates evidence across multiple backbone calls without requiring all selected frames to share a single context. Training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation. Intermediate supervision trains partial-video readouts for retrieval, while self-distillation regularizes them toward the final representation. Query-aware selection produces query-conditioned representations for candidate-set scoring and reranking, whereas query-independent selection enables reusable corpus indexing. Under the full training recipe, HourVideo retrieval Hit@1 increases from 54.2 to 70.7 and from 57.4 to 72.8 for 2B and 8B Qwen3-VL-Embedding backbones, respectively. Gains extend to the evaluated moment-retrieval and video-QA tasks, and the streaming head transfers to a second Qwen-family embedding backbone. These results support streaming latent aggregation as an effective approach to integrating long-video evidence into fixed-dimensional representations.
|
| 84 |
Reflection-Robust 6DoF Object Tracking with Light Fields
2610.04883
|
cs.CV
|
Nikolai Goncharov, Donald G. Dansereau |
Tracking the 6DoF pose of a moving rigid object is fundamental to robotics and autonomous driving, but existing trackers assume that object appearance is stable across a sequence, an assumption that breaks down on reflective surfaces whose appearance changes a...Tracking the 6DoF pose of a moving rigid object is fundamental to robotics and autonomous driving, but existing trackers assume that object appearance is stable across a sequence, an assumption that breaks down on reflective surfaces whose appearance changes as they mirror the environment. We introduce a light field based reflection-robust 6DoF tracker that turns this apparent nuisance into a pose cue. Per frame, our method recovers depth robustly against reflections, back-projects it into a point cloud, and estimates surface normals. It then decomposes the object's view-dependent appearance into a diffuse albedo and the environment map it reflects, resulting in a relightable surface light field. Starting from a coarse initialization, we relight it by the recovered environment map and optimize the pose on the photometric loss. Because a moving object mirrors new parts of the scene, the environment map fills in as the sequence proceeds, sharpening this signal over time. To evaluate this approach, we introduce a light field tracking dataset re-rendered from a robotic manipulation benchmark at four controlled reflectivity levels, each paired with simulated depth that reproduces how consumer RGB-D sensors fail on shiny surfaces. Additionally, we evaluate on two captured light field sequences. Our method trails the strongest baselines on diffuse objects and is the only one that holds its accuracy on fully reflective objects, where every baseline degrades.
|
| 85 |
PWM: Personalized World Models with Online Reinforcement Learning
2610.04920
|
cs.CV
|
Zhexin Lou, Guancheng Lu, Zeyu Zhang, Yi Zhang, Yang Zhao |
Pretrained world models can generate diverse environments, yet users often want to explore a particular scene specified by their own video. This requires learning the scene's visual identity while retaining the quality of action-conditioned generation. We intr...Pretrained world models can generate diverse environments, yet users often want to explore a particular scene specified by their own video. This requires learning the scene's visual identity while retaining the quality of action-conditioned generation. We introduce Personalized World Models (PWM), a framework for customizing interactive world models from short scene videos through online reinforcement learning. In PWM, the support trajectory and its associated controls provide reward feedback on continuations sampled from the current policy. In the GRPO instantiation, group-relative optimization updates a compact LoRA adapter using a unified reward for scene appearance, visual continuity, and motion, while base-policy anchoring regularizes changes to the pretrained generation prior of a frozen Yume-5B backbone. The same adaptation procedure is applied across real and rendered environments. We also instantiate PWM with DiffusionNFT as an alternative reward-guided optimization method for learning the scene-specific adapter. We also introduce PWM-Bench, comprising 150 customization tasks across Indoor, Outdoor, and Gaming, with paired evaluation on held-out continuations. The GRPO and DiffusionNFT instantiations of PWM improve customization over native Yume in 71.3% and 65.3% of the evaluated scenes, respectively, with positive mean gains across all three domains. For the GRPO instantiation, matched SFT comparisons further demonstrate higher mean customization gains and better mean image-quality scores in every domain, while retaining frame-level visual quality close to the pretrained model.
|
| 86 |
TRACE: Time-Adaptive Residual Attention Control with Content-Style Decomposition for Training-Free Diffusion Style Transfer
2610.04922
|
cs.CV
|
Duc Khoan Le, Kim Ngoc Tran, Minh Nhat Le, Thanh An Tran, Viet Toan Nguyen |
Reference-guided style transfer aims to preserve the semantic structure of a content image while transferring the visual appearance of a style reference. Recent diffusion-based methods achieve impressive stylization quality by exploiting strong pretrained gene...Reference-guided style transfer aims to preserve the semantic structure of a content image while transferring the visual appearance of a style reference. Recent diffusion-based methods achieve impressive stylization quality by exploiting strong pretrained generative priors. However, training-free approaches still face a difficult trade-off among style fidelity, content preservation, and content leakage. Direct style injection may unintentionally transfer semantic content from the style image, while fixed guidance schedules often ignore the time- and state-dependent nature of diffusion sampling. To address these limitations, we propose TRACE, a training-free diffusion style transfer framework with Time-adaptive Residual Attention Control and Content-Style Decomposition. TRACE first performs offline CLIP-based subspace analysis to separate content and style directions from paired data. During inference, it removes content-related components from the style reference and style-related components from the content reference to reduce leakage. It then injects style information through residual cross-attention and applies uncertainty-aware guidance to adapt the guidance signal at each denoising step. Experiments show that TRACE achieves a favorable trade-off between stylization and preservation. Compared with optimal-control-based baselines, TRACE substantially improves style fidelity (+17.28 CSD and +34.10 SRA). While, compared with stylization methods, it better preserves content structure (+12.80 DINO, +5.52 CLIP-I, and -8.19 LPIPS) and reduces directional semantic leakage by 29.5% in DCL. Our code is publicly available at https://github.com/pixelchemy-research/TRACE.
|
| 87 |
Robust Tensor Completion via Reflective Convolution Nuclear Norm Minimization
2610.04979
|
cs.CV
|
Weiguo Zhou, Feng Zhang, Wenjin Qin, Jianwen Huang |
Robust tensor completion recovers multidimensional data from partial observations corrupted by sparse gross errors. Existing convolutional low-rank models typically construct translated copies using circular continuation, which introduces artificial wrap-aroun...Robust tensor completion recovers multidimensional data from partial observations corrupted by sparse gross errors. Existing convolutional low-rank models typically construct translated copies using circular continuation, which introduces artificial wrap-around neighborhoods for finite nonperiodic data. We propose reflective convolution nuclear norm minimization (RCNNM), which replaces circular shifts with endpoint-nonrepeating reflection. The resulting lifting has nonuniform entry multiplicities and satisfies a weighted Gram identity that supports both the recovery analysis and the optimization method. Under random sampling and sparse corruption, we establish high-probability exact recovery of the underlying tensor and sparse errors, together with stability under bounded dense perturbations. We further develop a two-block ADMM with a closed-form entrywise tensor update, while singular-value thresholding is implemented through the smaller right Gram matrix. Experiments on synthetic tensors, BSDS color images, and CAVE multispectral images show that RCNNM consistently improves over its circular-lifting counterpart, with the clearest gains near image boundaries. In particular, average boundary-PSNR improvements reach 3.26 dB while global reconstruction quality remains competitive.
|
| 88 |
Geometry-Aware Preference Optimization for Text-to-Image Diffusion Models
2610.04980
|
cs.CV
|
Lei Wang, Zhen Wang, Yuexiang Xie, Yaliang Li |
Preference alignment has become a standard practice for text-to-image diffusion models. Direct Preference Optimization (DPO) simplifies this process by eliminating explicit reward modeling. Its diffusion variant, Diffusion-DPO, has become a widely adopted base...Preference alignment has become a standard practice for text-to-image diffusion models. Direct Preference Optimization (DPO) simplifies this process by eliminating explicit reward modeling. Its diffusion variant, Diffusion-DPO, has become a widely adopted baseline. Diffusion-DPO essentially encourages the likelihood of preferred samples while suppressing dispreferred ones. In this paper, we revisit DPO-style alignment methods for diffusion models from the perspective of the manifold hypothesis. Under this view, natural images concentrate near a low-dimensional manifold embedded in the high-dimensional ambient space, whereas DPO directly optimizes preference distributions in the full space without accounting for this geometric structure. This creates a mismatch in the optimization dynamics: it suppresses geometry-preserving tangential updates, while insufficiently restricting hazardous normal-direction updates. This mismatch gradually degrades image quality and diversity. To address this issue, we propose Anisotropic Geometry-Aware Preference Optimization (APO), which replaces the uniform Euclidean treatment of prediction errors with a geometry-aware anisotropic metric derived from the reference model. Concretely, APO adaptively strengthens regularization in directions where the reference denoising function is highly sensitive, while relaxing constraints in directions that permit safe semantic adjustment. This recalibrates preference optimization according to the local manifold geometry, and maintains the original manifold structure. Experiments show that APO achieves strong performance and an average win rate exceeding 60\% against various existing alignment methods across diverse benchmarks. It requires significantly fewer training steps than prior methods, and preserves generation diversity throughout training.
|
| 89 |
VisualErase: Dual-Branch Visual Trajectory Redirection for Robust Concept Erasure in Text-to-Image Diffusion Models
2610.05000
|
cs.CV
|
Qianlong Xiang, Miao Zhang, Kun Wang, Yupeng Hu, Junhui Hou |
Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts whi...Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts while preserving general generation capability by redirecting target-related text-to-image mappings. However, recent studies show that erased models may still retain visual generative trajectories of target concepts, leaving them vulnerable to adversarial recovery attacks and revealing a fundamental gap between redirecting text-to-image mappings and truly removing visual knowledge. To bridge this gap, we propose VisualErase, a new paradigm that redirects concept-bearing visual generative trajectories toward explicitly defined concept-removed outcomes. To enable this redirection, we use structure-preserving image editing to construct content-aligned, concept-removed counterparts for source images, providing explicit visual endpoints that retain non-target content. We then derive a denoising target from each source-to-counterpart pair and use a dual-branch redirection loss to align both text-conditioned and unconditional predictions with this target, since conditional supervision alone does not explicitly constrain generation without textual guidance. To mitigate the adverse effects of concept erasure on non-target generation, we jointly optimize the redirection loss with a counterpart retention loss that matches denoising predictions from the frozen pretrained model. Across style, celebrity, and nudity erasure, VisualErase limits the maximum attack success rate over seven attacks to 0%, 8%, and 0.1%, respectively, while retaining general generation quality. These results highlight the importance of visual trajectory redirection for robust concept erasure beyond text-to-image mappings alone.
|
| 90 |
PortraitAes: Intent-Conditioned Structured Portrait Aesthetics Assessment
2610.05010
|
cs.CV
|
Junzhou Xie, Haozhong Xiong, Xunyun Tian, Kaile Du, Tianchen Yu |
Portrait aesthetic assessment assigns comparable scores according to how effectively human-centered images fulfill their photographic intent. These scores support data filtering, candidate selection, and preference modeling in image-generation pipelines. Exist...Portrait aesthetic assessment assigns comparable scores according to how effectively human-centered images fulfill their photographic intent. These scores support data filtering, candidate selection, and preference modeling in image-generation pipelines. Existing methods typically predict a single aesthetic score or use general-purpose MLLMs without conditioning on photographic intent. This omission matters because the same blur, pose, lighting, or framing choice may serve one photographic intent but undermine another. These models thus learn context-agnostic aesthetic priors and yield inconsistent, inaccurate, misleading judgments for portraits with distinct photographic objectives. We introduce PortraitAes-Bench, an 11K-scale benchmark that decomposes this task into intent-conditioned subjudgments. Expert-authored rubrics define nine photographic intents, six first-level dimensions, and 22 secondary criteria. They support a structured pipeline for intent routing, specialist assessment, verification, and score fusion. Following this structure, we train PortraitAes with multi-task supervision. We then improve score comparability through Gaussian score calibration and within-dimension cross-image ranking. On the standard benchmark, PortraitAes achieves a Pearson correlation of 0.924 and a Spearman rank correlation of 0.934. On the hard-case set, its Pearson correlation is 0.829 and its Spearman rank correlation is 0.795. Across both sets, PortraitAes outperforms the evaluated general-purpose MLLMs and specialized aesthetic baselines.
|
| 91 |
EMBER-Bench: Benchmarking Cross-Event Causal Memory in Long-Horizon Embodied Tasks
2610.05013
|
cs.CV
|
Aoyang Cai, Boning Zhao, Shaoxuan Xie, Dahui Gao, Huan Yang |
Lifelong physical agents must reason over extended interactions where past events continue to shape the world long after they disappear from view. Beyond recalling what happened, agents must infer how history changes the current state and constrains future act...Lifelong physical agents must reason over extended interactions where past events continue to shape the world long after they disappear from view. Beyond recalling what happened, agents must infer how history changes the current state and constrains future actions. Yet existing embodied and video-memory benchmarks largely focus on historical retrieval and summary, leaving such history-dependent causal reasoning underexplored. We introduce EMBER-Bench, an egocentric benchmark for cross-event causal reasoning in long-horizon embodied tasks, for which we newly created the task design, video recording, and data annotation. It contains 189 household tasks and 699 QA pairs, spanning task progress, failure recovery, external interventions, and compound long-horizon tasks with distant dependencies and prerequisites, with fine-grained event and causal-chain annotations. EMBER-Bench evaluates reasoning in both directions: next-action prediction selects the next action from history, and causal traceback, given that action, identifies the historical event that makes it necessary. Input ablations that add action logs or privileged cause-and-consequence annotations to the video indicate which kind of historical information models fail to use. Among the 16 evaluated models, the highest overall accuracy is 61.2%, compared with a mean of 98.3% across two human evaluators. At paired decision points, correct traceback is not associated with correct next-action prediction. Adding action logs yields a gain of 1.6 points, whereas cause-and-consequence annotations yield an additional gain of 13.0 points on top of that. These results suggest that extracting causal information from past events and converting it into constraints on current actions remains a key difficulty for long-horizon embodied agents. Project Page: https://zhaoalexgoat.github.io/EMBER-Bench/
|
| 92 |
Look Where You Say You're Looking: Self-Grounded Attention for Visual Reasoning
2610.05023
|
cs.CV
|
Uri Berger, Gal Chechik, Gal Dalal |
We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned ...We introduce Self-Saliency, a method for training Vision-Language Models (VLMs) to increase the alignment between their visual attention and the image regions mentioned in their reasoning. Self-Saliency uses a grounding model to localize the objects mentioned in each reasoning step and treats the resulting areas as supervision for the model's visual attention. Previous work on steering visual attention determines target image regions based solely on the image and question. In contrast, we show that conditioning the target regions on the model's generated reasoning improves downstream performance. For proper evaluation, we build a unified, broad suite of 25 visual reasoning benchmarks, where we reproduce the results of previous methods. We find that Self-Saliency significantly outperforms both prior attention-steering methods and baselines that ground image-level text, achieving both a better average rank and a better mean score. Post-training analysis shows that the model primarily adapts its reasoning text to existing attention patterns, producing shorter steps that refer to larger regions. Nevertheless, when controlling for generated text, attention to grounded regions increases significantly across the relevant layer. Finally, we identify a consistent geometric bias in VLM visual attention toward the image border. However, our ablations show that Self-Saliency's gains cannot be explained by simply aligning attention with the center of the image, highlighting the importance of aligning visual attention with the regions mentioned in the model's reasoning.
|
| 93 |
LightVLN: Efficient Aerial Vision-and-Language Navigation with Compact Memory and History-Guided Local Aggregation
2610.05024
|
cs.CV
|
Yiming Zhao, Tianshun Li, Jingle He, Ruonan Chai, Xinhu Zheng |
Aerial vision-and-language navigation (VLN) enables unmanned aerial vehicles to execute long-horizon natural-language instructions from visual observations in complex three-dimensional environments. However, recent aerial VLN models often rely on large-scale v...Aerial vision-and-language navigation (VLN) enables unmanned aerial vehicles to execute long-horizon natural-language instructions from visual observations in complex three-dimensional environments. However, recent aerial VLN models often rely on large-scale vision-language backbones and dense visual histories, imposing substantial computation and memory costs that hinder onboard deployment. We propose LightVLN, a lightweight history-aware aerial VLN framework that combines a compact 0.5B language backbone with compact representations of both historical and current observations. LightVLN compresses each historical frame into a single token using visual features already computed by the policy. It further introduces history- and instruction-conditioned local aggregation to reduce the current observation from 256 to 32 visual tokens while preserving navigation-relevant spatial information. With up to 16 historical frames, the policy uses at most 48 observation-derived tokens. On the public OpenFly dataset, LightVLN achieves 50.93% Test-Seen and 36.14% Test-Unseen success rates (SR), outperforming the evaluated 7B language-backbone baselines on most reported metrics. It also achieves 25.83% SR on AerialVLN-S Val-Seen. In a reconstructed unseen campus, we deploy LightVLN on a DJI M350 RTK with an external Jetson Orin NX 16 GB for closed-loop onboard-compute real-to-sim hardware-in-the-loop (HIL) evaluation, achieving 14.61 Hz model inference and 11.13 Hz end-to-end decision updates. These results demonstrate the effectiveness and efficiency of LightVLN for aerial navigation.
|
| 94 |
Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA
2610.05025
|
cs.CV
|
Hyemin Yang, Wooseong Jeong, Giwon Lee, Kuk-Jin Yoon |
Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that...Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay renewed reasoning when the scene changes unexpectedly. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive and is computed from forwards the architecture already runs, requiring neither manual phase labels nor an auxiliary uncertainty model. On VLA-Arena, it achieves a higher success rate at matched call budgets than alternative uncertainty baselines while maintaining low wall-clock overhead, and more consistently separates successful from failed rollouts. Also, TUD finds a more favorable cost-success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. The same trade-off appears in our real-robot experiments, where TUD cuts generalist calls by 75% relative to the strongest fixed-interval baseline while achieving an even higher success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.
|
| 95 |
GeoBridge-VLA: Geometry-Aware Residual Adaptation for Vision-Language-Action Models
2610.05026
|
cs.CV
|
Hyun Song, Kangmin Kim, Loren Jinsoo Um, Minhui Han, Jaehyeok Park |
Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's fro...Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Stage I trains a feature bridge and geometry decoder with depth supervision. Stage II freezes these modules and trains a gated residual interface together with the action-side projections and action expert. The residual augments the existing visual tokens without adding a second image encoder or increasing the token count. Deployment requires RGB, robot state, and language, but no depth observations. Under matched evaluation conditions, GeoBridge-VLA achieves 70.9% success on LIBERO, compared with 60.0% for SmolVLA. Disabling the residual in the same trained checkpoint reduces success from 70.90% to 69.85%, with mixed effects across suites. On a physical ROBOTIS OMY robot, GeoBridge-VLA succeeds in 148 of 200 trials (74.0%) across four tasks, compared with 108 of 200 (54.0%) for SmolVLA.
|
| 96 |
SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation
2610.05029
|
cs.CV
|
Hyun Song, Taewan Cho, Kangmin Kim, Andrew Jaeyong Choi |
Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through laye...Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring CLIP tokens should be combined to recover fine local structure. We present SPACE-CLIPv2, a frozen-backbone depth decoder that aggregates fixed local neighborhoods in CLIP token space. At selected decoder stages, the model samples a fixed token stencil, predicts aggregation weights, and injects the resulting response through a gated residual update. A token-space high-pass branch further preserves shallow local contrast. On NYU Depth V2, SPACE-CLIPv2 improves over a matched SPACE-CLIP baseline, while five-seed experiments consistently favor fixed over learned-offset sampling. Zero-shot iBims-1 evaluation further improves boundary and planar-geometry measures. These results support constrained local token aggregation as a practical mechanism for decoding geometry from frozen CLIP representations.
|
| 97 |
Code2Games: Enabling Coding Agents for Gaming World Generation
2610.05033
|
cs.CV
|
Wei Wu, Ziyang Xu, Zeyu Zhang, Yang Zhao, Hao Tang |
Generating a high-quality gaming world from a natural-language game intent requires joint reasoning about scene structure, spatial layout, gameplay objectives, interactive entities, and executable gameplay logic. Existing coding agents can generate individual ...Generating a high-quality gaming world from a natural-language game intent requires joint reasoning about scene structure, spatial layout, gameplay objectives, interactive entities, and executable gameplay logic. Existing coding agents can generate individual assets, scenes, or scripts, but often struggle to maintain consistency across these components. We propose Code2Games, an agentic framework that builds a structured gaming world upon a base Blender world generated from the same game intent. Code2Games coordinates scene analysis, gameplay planning, constrained gaming-world generation, and gaming-engine customization through a shared scene-gameplay representation with persistent element correspondence. After world generation, Code2Games adapts the generated world to Unreal Engine 5 and employs an execution-guided reconstruction process that uses compilation diagnostics, runtime feedback, and gameplay test results to resolve inconsistencies arising during engine adaptation. To systematically evaluate gaming-world generation, we introduce the GameCode4D benchmark, which comprises ten fixed game prompts spanning different levels of scene and gameplay complexity. We evaluate the generated results across four dimensions: visual quality, interactive fidelity, multimodal artifact quality, and playable-game quality. Experiments demonstrate that, compared with direct gaming-world generation by coding agents and existing baseline methods, Code2Games consistently improves the visual quality and interactive fidelity of generated gaming worlds, as well as the quality of the resulting games after engine adaptation.
|
| 98 |
CoDG-Net: Structure-Guided Style Diffusion and Collaborative Learning to Mitigate Catastrophic Forgetting in Medical Image Domain Generalization
2610.05053
|
cs.CV
|
Yucheng Song, Jincan Wang, Haokang Ding, Zhiqiang Tian, Kangxu Fan |
Domain Generalization (DG) for medical image segmentation is both highly challenging and critically important. However, existing medical DG methods largely overlook the issue of Catastrophic Forgetting (CF): \textbf{Models often sacrifice their ability to reta...Domain Generalization (DG) for medical image segmentation is both highly challenging and critically important. However, existing medical DG methods largely overlook the issue of Catastrophic Forgetting (CF): \textbf{Models often sacrifice their ability to retain source-domain knowledge while pursuing cross-domain robustness.} This can directly threaten diagnostic safety in already-deployed clinical scenarios. To address this, we investigate data augmentation strategies and catastrophic forgetting for medical image DG segmentation. First, we propose a structure-guided style diffusion augmentation method. Constrained by anatomical structure consistency in the frequency domain, this method performs cross-domain diffusion on the amplitude spectrum, generating samples with more diverse and broader style coverage to better support domain generalization. Then, we design a collaborative learning network with a dual-branch interactive architecture (CoDG-Net), together with a novel learning bias-guided strategy that adaptively regulates knowledge transfer at both the layer level and the task level, thereby effectively mitigating catastrophic forgetting on the source domain. Experiments and ablation studies on single-source and multi-source medical DG benchmark datasets demonstrate that CoDG-Net not only outperforms existing state-of-the-art methods in target-domain segmentation performance, but also achieves a lower forgetting rate on the source-domain data. The code is available at: https://github.com/wangprocess/CoDG-Net.
|
| 99 |
Salvation Lies Within: Eliciting Inherent Style Transfer in Step-Distilled Diffusion Models
2610.05066
|
cs.CV
|
Shengyin Sun, Yiming Li, Yingzhao Lian, Xing Li, Xingzhi Zhou |
Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowled...Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit stylistic capabilities through language. Pursuing this direction requires textual guidance that captures how visual attributes jointly define a style and remain applicable as the depicted content changes. To explore this approach, we introduce StyleForge, a fully automatic, training-free framework that expresses reference styles as reusable rendering instructions. By integrating overall rendering characteristics with local color and lighting behavior, StyleForge organizes visual evidence from reference images into a coherent specification of how the target style should be expressed. The specification is then compiled into textual guidance that can be reused across content prompts, enabling frozen step-distilled T2I models to render different subjects and scenes in the reference style while retaining native few-step generation. Extensive experiments show relative gains of up to 29.47\% in generation quality scores over the strongest baseline, while Pareto analysis indicates that improved stylization is accompanied by strong adherence to the requested content.
|
| 100 |
CGDD-Net: Context-Guided Dynamic Detail Modeling for Retinal Vessel Segmentation
2610.05078
|
cs.CV
|
Xincheng Li, Xiaoqi Sheng, Xinyu Zhang |
Accurate retinal vessel segmentation requires features that capture vascular geometry while preserving information for fine-scale reconstruction. We propose CGDD-Net, a context-guided dynamic detail modeling network that connects adaptive feature extraction to...Accurate retinal vessel segmentation requires features that capture vascular geometry while preserving information for fine-scale reconstruction. We propose CGDD-Net, a context-guided dynamic detail modeling network that connects adaptive feature extraction to a shared decoder pathway. Context-Guided Scale-Adaptive Deformable Encoding (CSDE) combines fixed-grid convolution with deformable local attention to capture vascular patterns at multiple spatial extents. Spatially Adaptive Multi-Kernel Gating (SAMG) selects receptive-field responses at each location. Dynamic Cross-Scale Detail Fusion (DCDF) aligns the gated intermediate features and compresses them into an eight-channel representation, which is reused at three decoder resolutions together with selected encoder skips. This design consolidates intermediate information before decoding instead of transferring each middle-stage feature through a separate direct skip. On DRIVE, CHASE\_DB1, STARE, and HRF, CGDD-Net achieves AUC values of 0.9824, 0.9938, 0.9895, and 0.9874, with F1 scores of 0.8323, 0.8102, 0.8510, and 0.8157, respectively. The complete model contains 1.96 million trainable parameters. In cumulative ablations, the full model improves F1 over the internal baseline by 2.48, 0.61, and 3.49 percentage points on DRIVE, CHASE\_DB1, and STARE. Twelve directed cross-dataset experiments further characterize transfer without target-domain adaptation. The results support shared intermediate detail delivery as an effective, parameter-compact architecture for retinal vessel segmentation. Code is available at \url{https://github.com/lixincheng-xcl/CGDD-Net}.
|
| 101 |
ReMAP: Restoring the Perceptual Cycle with Reasoning-Time Latent Visual Memory
2610.05097
|
cs.CVcs.CLcs.AI
|
Hao Jiang, Zhanyu Guo, Chenwei Wu, Yichen Guo, Qizhe Zhang |
As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along th...As multimodal large language models (MLLMs) reason for longer, attention to the initial visual input diminishes, weakening visual grounding. Visual memory reintroduces visual evidence during reasoning. We conduct a controlled analysis of visual memory along three axes: curation, organization, and access. We find that local evidence benefits from global context, compact latent representations balance accuracy and visual-context cost, and the utility of memory access depends on the reasoning state. Guided by these findings, we propose ReMAP (Reasoning-Time Memory-Augmented Perception), which couples two complementary latent memories: a static, question-conditioned Global memory that preserves scene and cross-image context, and a dynamic Local memory that uses this context as an anchor while selecting and re-encoding region-level evidence according to the current reasoning state. Both memories return compact latent tokens inserted into the reasoning sequence, and a reinforcement-learning access policy trained with branched rollouts decides when to continue reasoning or invoke Global or Local memory. On ten benchmark families, ReMAP outperforms prior visual-memory methods on all four multi-image benchmarks, exceeding the strongest prior results on MuirBench and MIMIC by 8.38 and 14.84 percentage points. Across four backbone families, enabling memory access improves over the same trained model with memory disabled, and on shared V*Bench, CV-Bench-2D, and MuirBench questions ReMAP reduces the visual tokens entering the reasoning sequence by 51.0-76.8% relative to the native-resolution backbone. Further analyses show that Global and Local memory form distinct yet complementary latent representations. Together, these components restore the perceptual cycle by letting the reasoning state trigger targeted visual retrieval, with the retrieved evidence guiding subsequent reasoning.
|
| 102 |
LoopMoEVR: Loop-Based Degradation-Aware Mixture-of-Experts for Unified UHD Video Restoration
2610.05109
|
cs.CV
|
Yucheng Xin, Runci Bai, Yongcong Wang, Guangwei Gao, Jiao Liu |
Recently, unified high-definition image restoration has attracted considerable attention; however, existing models tend to excessively increase their depth in pursuit of improved generalization, which often yields only limited gains. Meanwhile, loop-based lear...Recently, unified high-definition image restoration has attracted considerable attention; however, existing models tend to excessively increase their depth in pursuit of improved generalization, which often yields only limited gains. Meanwhile, loop-based learning paradigms have drawn widespread attention due to their low parameter counts and strong regression capability, as exemplified by GPT-6 and looped Transformers. In this paper, we introduce the loop learning paradigm to address restoration tasks that require cross-domain learning. Specifically, we propose LoopMoEVR, a loop-based mixture-of-experts model capable of handling degraded ultra-high-definition (UHD) inputs. First, a degradation-conditioned low-rank loop embedding is designed to construct input-dependent stage conditions. Second, a spatio-temporal iterative adaptive normalization module, termed IterAda3DN, is developed to fuse local features with global loop context, thereby performing position-wise affine modulation. Finally, the expert branches further integrate the attention-updated local and global video states with the loop conditions to generate dedicated modulation parameters, while an input-conditioned depth predictor adaptively configures the number of loop iterations. With only approximately 0.884M trainable parameters, the proposed model uniformly handles UHD video dehazing, deraining, denoising, and low-light enhancement tasks, achieving state-of-the-art restoration performance on both public benchmarks and real-world scenarios.
|
| 103 |
PCLM: Small-target localization with frozen CLIP via prototype contrast and local magnification
2610.05115
|
cs.CV
|
Zhipeng Ye, Feng Jiang, Qiufeng Wang, Hao Li |
Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen ...Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define foreground and background prototypes through equally weighted regional features. Their difference provides a shared scoring direction for query patches, explicitly comparing target evidence with the demonstrated background. Nine overlapping query windows are enlarged and encoded independently to sample small targets more densely. Reprojection and coverage averaging combine their scores into a continuous localization map. The class direction occupies 2 KiB regardless of support count and transfers unchanged across datasets with mapped categories. On 5,047 small-target queries from VOC, COCO, ADE20K and Oxford-IIIT Pets, PCLM achieves higher mean pixel AP than every evaluated text-conditioned localization baseline on each dataset under our evaluation protocol. Gains over the strongest scene-dataset baselines range from 5.63 to 13.23 percentage points. At comparable measured latency, local magnification improves scene small-target AP by 4.51 to 5.68 points over whole-canvas enlargement. Factorial experiments show that prototype contrast increases the benefit of local observation, including under matched image-coordinate filtering. Support-budget experiments show that additional examples refine category estimation without increasing representation size or query-time scoring cost.
|
| 104 |
Representation--Behavior Alignment for Explainable Weakly-Supervised Video Anomaly Detection
2610.05129
|
cs.CV
|
Chao Huang, Pengfei Wei, Kaige Li, Chengliang Liu, Wei Wang |
Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as rep...Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012\% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.
|
| 105 |
RoMod: Temporal Routing Modulation via Mixture-of-Experts for Video Anomaly Detection
2610.05131
|
cs.CV
|
Chao Huang, Pengfei Wei, Benfeng Wang, Chengliang Liu, Wei Wang |
Intermediate-layer features from multimodal large language models have shown strong potential for video anomaly detection (VAD), yet the origin of their discriminative power remains unclear. We study this question using sparse mixture-of-experts (MoE) models, ...Intermediate-layer features from multimodal large language models have shown strong potential for video anomaly detection (VAD), yet the origin of their discriminative power remains unclear. We study this question using sparse mixture-of-experts (MoE) models, whose explicit expert structure and sparse activation make their internal computation easier to inspect. With a fully frozen backbone and no additional training, we find that anomaly-related evidence is concentrated in a small set of experts. These experts recur across layers, spontaneously specialize in different anomaly types, and together form a dynamic routing subnetwork. We further show that the output channels most strongly influenced by these experts are also the hidden dimensions that contain the most anomaly-relevant information. Routing statistics can therefore serve as an internal anomaly cue that complements semantic features.Based on these findings, we propose RoMod, an efficient VAD framework trained with only \(5\%\) of weakly labeled videos. RoMod includes a Routing-Modulated Fusion module, RoMF, and a Routing-aware Temporal Network, RoTN. RoMF uses routing signals to adaptively recalibrate hidden semantic channels. Its design also prevents the routing branch from bypassing semantic features and making predictions on its own. RoTN captures the temporal evolution of anomalies from onset to persistence and termination. Experiments on three benchmarks show that RoMod achieves state-of-the-art performance while running substantially faster than dense backbones of comparable size.
|
| 106 |
How Does Geometry Enter Generated Motion?
2610.05135
|
cs.CVcs.AI
|
Weihan Li, Junhao Wu, Yuhan Song, Xiaofeng Lin, Xinlei Chen |
Under a fixed physical law, the visible geometry of a scene determines how motion must change. We ask how video generators realize this relationship. We fix the law and the initial state and change only the geometry drawn in the first frame, within matched fam...Under a fixed physical law, the visible geometry of a scene determines how motion must change. We ask how video generators realize this relationship. We fix the law and the initial state and change only the geometry drawn in the first frame, within matched families of tracks and deflectors, and compare each generated trajectory with the simulator prediction for that geometry. Paired interventions change one thing at a time: a local bump, the height of a barrier, the words of the prompt, the length of the clip. Across nine image-to-video models, geometry is preserved and shapes the motion: the speed of the ball follows the drawn undulation of a track. A physical state would carry this response forward, and here the generated motion parts from the law. The mean slope barely accelerates the ball, successive contacts fail to compose through a consistent state, an edit ahead of the ball alters its motion before it arrives, and the ball climbs over barriers higher than its release point. Two global conditions organize the global trajectory: text strongly controls the destination, while clip length strongly controls timing in the open-weight models tested. The pattern persists with photographed first frames. Current video generation thus behaves as geometry-conditioned motion synthesis whose evolution of state differs systematically from that of a fixed physical law.
|
| 107 |
SemCam: Semantic Camera Motion Control for Video Generation
2610.05141
|
cs.CV
|
Janna Bruner, Omer Talmi, Ianir Ideses, Lior Fritz, Lior Wolf |
Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to spe...Controlling the camera relative to a moving subject in an existing video is challenging: behaviors such as maintaining a frontal view require the camera to adapt to the subject's changing position and orientation, making the desired trajectory difficult to specify in advance. Existing camera-controlled video-to-video methods typically rely on explicit trajectories or reference motions, which do not directly express these dynamic camera--subject relationships. We introduce semantic camera motion control, a novel video-to-video task in which a reference video and a target motion label specify the desired subject-relative camera behavior without an explicit target trajectory. Our method, SemCam, learns to realize this behavior while preserving source content. It combines shared-basis low-rank adaptation with motion-conditioned modulation, while a background-consistency loss encourages fidelity in regions visible in both reference and target videos. We construct 661 paired videos covering eight semantic camera behaviors and evaluate on a separate 109-scene benchmark using subject-relative motion metrics, appearance measures, and a user study. SemCam achieves a semantic-motion success rate of 68.6%, compared with 45.3% for Vista4D, the strongest evaluated baseline, while maintaining comparable subject identity preservation.
|
| 108 |
Construction and Evaluation of Machine Learning Models for Near-Real-Time Fire Detection from MTG FCI Imagery
2610.05154
|
cs.CV
|
Asaf Vanunu, Boaz Nadler, Arnon Karnieli |
Geostationary satellite observations are important for wildfire detection and monitoring. The current study evaluates machine learning models for MTG FCI near-real-time fire detection in 1- and 2-km spatial configurations and compares them with threshold-based...Geostationary satellite observations are important for wildfire detection and monitoring. The current study evaluates machine learning models for MTG FCI near-real-time fire detection in 1- and 2-km spatial configurations and compares them with threshold-based algorithms. The models were trained and evaluated using VIIRS fire reference data across diverse ecological regions in Europe, Africa, and the Middle East. The key results are that 1-km models significantly outperform both their 2-km variants and operational threshold products. The constructed 1-km models achieved F1 scores higher by up to 0.36 compared to baseline products. Importantly, the 1-km models detected small fires with higher probability compared to competing models. Finally, the models robustly detected fires up to 260 min earlier than baseline products. To support opensource applications, our trained models are publicly available.
|
| 109 |
Recurrent Latent Visual Search for GUI Grounding
2610.05185
|
cs.CV
|
Kaiyu Wu, Beichen Zheng, Weiyao Huang, Keze Wang |
GUI grounding is a critical capability for GUI agents powered by vision-language models, helping them execute user instructions by locating the corresponding elements in screenshots. Single-step grounding struggles with small elements and dense layouts, motiva...GUI grounding is a critical capability for GUI agents powered by vision-language models, helping them execute user instructions by locating the corresponding elements in screenshots. Single-step grounding struggles with small elements and dense layouts, motivating multi-step visual search. However, existing approaches commonly rely on textual reasoning misaligned with visual space or costly multi-round interactions with external visual tools. To make multi-step visual search an explicit spatial process within the model, we propose ReLaViS, which performs Recurrent Latent Visual Search in a single interaction round. At each step, a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that explicitly represents the search focus. This distribution then aggregates the visual tokens into latent visual evidence, which is recurrently fed back as the next input embedding to condition subsequent search. We further introduce a GUI-aware coarse-to-fine inductive bias through trajectories constructed from flat element annotations, supervising search from the global interface through intermediate element groups to the target. Built on Qwen2.5-VL-7B, ReLaViS improves ScreenSpot-Pro accuracy by 3.1 percentage points to 56.3% with only a 3.5% increase in inference FLOPs and outperforms the matched single-step baseline on all five benchmarks.
|
| 110 |
Order Matters: Competition-Guided Query Ordering for RNN-Based Object Detection
2610.05191
|
cs.CV
|
Shengjian Wu, Li Sun, Yu Shangguan, Qingli Li |
DETR-style detectors use one-to-one bipartite matching during training to assign object queries to ground-truth objects, enabling end-to-end set prediction without non-maximum suppression (NMS). However, without an explicit de-duplication procedure, multiple q...DETR-style detectors use one-to-one bipartite matching during training to assign object queries to ground-truth objects, enabling end-to-end set prediction without non-maximum suppression (NMS). However, without an explicit de-duplication procedure, multiple queries can still produce highly similar hypotheses for the same object, making training unstable and predictions less decisive. Inspired by the sequential ordering of NMS, we propose DETRNN, a plug-and-play module that turns unordered object queries into a competition-aware sequence for recurrent refinement. DETRNN builds an explicit confidence-and-similarity based order from prior predictions, then refines queries with an RNN along this order to model competition inside the decoder. This ordered recurrent refinement reduces redundant predictions, stabilizes optimization, and improves final detection accuracy. Experiments on multiple DETR-style detectors show consistent gains with comparable efficiency.
|
| 111 |
Long-MDR: Long-Context Reinforcement Learning for Multimodal Deep-Research Agents
2610.05195
|
cs.CV
|
On Tai Tang |
The next generation of multimodal research agents must reason over long-lived research histories rather than short model completions. During a single task, an agent may repeatedly search the web, inspect visual evidence, revisit earlier hypotheses, and accumul...The next generation of multimodal research agents must reason over long-lived research histories rather than short model completions. During a single task, an agent may repeatedly search the web, inspect visual evidence, revisit earlier hypotheses, and accumulate tens of thousands of tokens of multimodal context. Despite this trend, online RL for multimodal research agents remains largely confined to shorter contexts and interaction horizons. We push online RL training to 128k context and 75+ tool-interaction turns. To our knowledge, this is the first online multimodal deep-research RL study trained at 128k context, and the first trained with a 75 tool-turn horizon. Scaling to this regime exposes several practical limitations of conventional RL training. Early in training, weak policies make poor use of large interaction budgets, causing expensive rollouts with little reward improvement. Later, policy entropy can collapse before performance has saturated, prematurely ending useful learning. We introduce Long-MDR, a three-component training recipe designed specifically for this setting: On-Policy Distillation Warmup, Progressive Horizon Expansion, and Entropy-Triggered Rescue. Together, these techniques improve both the learning efficiency and stability of long-horizon RL, enabling continued gains in a regime where direct training is slow and costly. At a 50-turn evaluation budget, our RL-trained Long-MDR-9B ranks first on five of six benchmarks among the compared 7B-9B agents.
|
| 112 |
F$^2$ SLAM: Turning Feed-Forward Geometry into Persistent Factors for SLAM
2610.05207
|
cs.CV
|
Zhisong Xu, Fan Zhu, Jiawei Qian, Ziyu Chen, Zhenjun Zhao |
Feed-forward 3D models provide strong multi-view geometric priors, while on- line simultaneous localization and mapping (SLAM) relies mainly on local mea- surements and can accumulate drift over long sequences. Existing attempts to combine the two typically tr...Feed-forward 3D models provide strong multi-view geometric priors, while on- line simultaneous localization and mapping (SLAM) relies mainly on local mea- surements and can accumulate drift over long sequences. Existing attempts to combine the two typically treat feed-forward predictions as an external geomet- ric state that is aligned or fused with the online estimate after the fact, which keeps broader multi-view evidence outside the optimizer that refines the SLAM state. We present F2SLAM, which instead converts feed-forward geometry di- rectly into optimization-native target-weight measurements attached to a persis- tent dense factor graph. A high-frequency stream maintains local tracking con- straints and graph connectivity, while a low-frequency stream uses wider multi- view context to selectively refresh existing measurements after a state-consistency check. Both streams constrain the same poses, inverse depths, and optional cam- era intrinsics through a single dense bundle adjustment. Experiments on multiple benchmarks demonstrate consistently strong trajectory estimation and improved dense reconstruction in both calibrated and uncalibrated settings. Notably, the uncalibrated configuration reduces the average ATE RMSE from 0.030 m for the strongest feed-forward baseline to 0.002 m on the Replica dataset.
|
| 113 |
PixReenact: Pixel-Conditioned Causal Video Diffusion for Streaming Head-Avatar Reenactment
2610.05233
|
cs.CV
|
Gavriel Habib, Dvir Samuel, Or Shimshi, Rami Ben-Ari |
Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representation...Streaming head-avatar reenactment aims to animate a reference image according to a live driving video, requiring robust motion transfer, long-term identity stability, and low latency. Existing methods often rely on specialized identity or motion representations, which can discard useful visual information and inherit failure modes from external extractors. In addition, many recent diffusion-based reenactment methods use offline, clip-based generation, jointly processing and denoising an entire video clip before producing its output, making continuous low-latency streaming difficult. We introduce PixReenact, a pixel-conditioned streaming reenactment framework built on causal video diffusion. PixReenact conditions directly on VAE-encoded reference and driving frames, without specialized identity or motion representations. To separate reference identity from driver motion, we train with cross-identity pseudo supervision together with corrective objectives anchored to the original reference and driving inputs. Long self-rollouts reduce autoregressive drift, while state-aware dual-teacher distillation separately addresses cold-start and steady-state generation. Across three cross-identity benchmarks and a long-horizon streaming benchmark, PixReenact demonstrates robust cross-identity reenactment, particularly under challenging conditions such as extreme viewpoints, occlusions, and pronounced facial expressions, while maintaining the reference identity over long streams. A 4-NFE rolling student continuously emits four frames per update with a mean emission latency of 239 ms.
|
| 114 |
ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image
2610.05249
|
cs.CV
|
Kai Lv, Yibo Yin, Lijun Guo, Heng Fan, Kaihao Zhang |
Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolith...Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.
|
| 115 |
MGPO: Manifold-Guided Diffusion Alignment for Task-Aware Dataset Distillation
2610.05252
|
cs.CVcs.AI
|
Yunyi Chen, Chenru Wang, Xinyi Ye, Zexin Zheng, Chi Zhang |
Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, re...Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage loss, where generated samples collapse into a few high-density modes and fail to cover the manifold's structural diversity. We propose Manifold-Guided Policy Optimization (MGPO), which reformulates DD as a multi-objective reinforcement learning problem and achieves Dual-Space Alignment via a pixel-space discriminative reward and a latent-space geometric reward guided by a class-wise Minimum Spanning Tree (MST). The discriminative reward enforces class separability, while the MST-based geometric reward encourages generated latents to cover a sparse geometric skeleton of each class, jointly addressing both failure modes. We further provide an idealized analysis that motivates the MST-based reward, including a Hausdorff approximation bound and a subsampling bound independent of the dataset size. The reward-modular design extends to structured tasks such as object detection and segmentation by substituting the frozen task reward model. Extensive experiments show MGPO consistently outperforms existing methods, including a +8.0% mIoU gain on segmentation under low-budget settings.
|
| 116 |
Hybrid-Basis Feature Forecasting for Diffusion Sampling Acceleration
2610.05254
|
cs.CV
|
Kai-Liang Cheng, Yuan-Yuan Cheng, Yu-fan Jin, Xiao-Ming Fu |
We propose Hybrid-Basis Feature Forecasting (HybridFF), a training-free, plug-and-play framework for accelerating diffusion sampling. To capture local smoothness, long-range trends, and complex non-monotonic variations when modeling feature evolution, HybridFF...We propose Hybrid-Basis Feature Forecasting (HybridFF), a training-free, plug-and-play framework for accelerating diffusion sampling. To capture local smoothness, long-range trends, and complex non-monotonic variations when modeling feature evolution, HybridFF first estimates coefficients using moving least squares (MLS) for each of multiple complementary basis families and then combines the corresponding predictors using fusion weights. In addition to the choice of basis functions, the fusion weights also play a critical role. We introduce two strategies to balance quality and speedup. HybridFF (Fixed) prioritizes efficiency with model-specific fusion weights calibrated on a small set and held constant during inference. HybridFF (Adaptive) updates the fusion weights online using branch reliability scores computed from an exponential moving average of full-step prediction errors, improving prediction fidelity and generation quality under aggressive caching while retaining substantial acceleration. Experiments across DiT-XL/2, FLUX.1-dev, SD3.5-Large, and HunyuanVideo demonstrate a favorable speedup--quality trade-off over representative single-basis forecasters and caching baselines.
|
| 117 |
When and What to Prune? Stage-Aware Visual Token Pruning for Efficient VLA
2610.05273
|
cs.CV
|
Tianjun Shi, Haotian Xiong, Ziyu Gong, Qi Lu, Li Li |
Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually esti...Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.
|
| 118 |
Answer with Evidence: Consistency-Aware Grounded Visual Question Answering for Roadside Traffic Scenes
2610.05274
|
cs.CV
|
Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningwei Ouyang, Shaofeng Liang |
Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the ...Roadside traffic reasoning requires every free-form textual claim to be backed by visual evidence. Existing grounded multimodal large language models (MLLMs) frequently exhibit say-point mismatch, in which the textual answer contradicts the bounding boxes the model localizes. Evaluation metrics that score answers and boxes separately leave this failure unpenalized. We trace the mismatch to the conventional answer-then-ground factorization, which commits to a numerical claim before any object is enumerated. To measure it, we build RoadSceneVQA-G, a benchmark of 34.7K question-answer pairs in which every free-form answer is linked to the set of boxes that witnesses it, and we propose the Answer-Grounding Consistency (AGC) evaluation suite. To address it, we introduce Enumerate-then-Answer (EtA), which reverses the generation order so that answer-evidence agreement becomes a property of the output structure, and Enumeration-Consistent Policy Optimization (ECPO), a reinforcement learning stage that uses the union of multiple rollouts as a recall teacher without ground-truth boxes. EtA raises say-point consistency from 26.6\% to 93.7\% and grounding F1 from 52.2\% to 73.0\%, and ECPO further increases F1 to 75.6\% without per-box supervision. On gRefCOCO, the same framework outperforms the strongest compared method, indicating that it transfers beyond traffic scenes. The project is available at \url{https://github.com/GuanRunwei/RoadSceneVQA-G}.
|
| 119 |
Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting
2610.05289
|
cs.CV
|
Xiaobiao Du, Beixi Hao, Zhen Fang, Tianqing Zhu, Richard Hartley |
Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable performance in novel view synthesis, yet deploying both static and dynamic Gaussian representations on resource-constrained mobile devices remains challenging due to heavy storage, redund...Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable performance in novel view synthesis, yet deploying both static and dynamic Gaussian representations on resource-constrained mobile devices remains challenging due to heavy storage, redundant primitives, and costly per-frame computation. We present Mobile-4DGS, a unified lightweight framework for high-fidelity real-time static and dynamic Gaussian rendering on mobile platforms. For compact appearance modeling, we introduce a Monte Carlo Specular Energy Aggregator that compresses high-order radiance residuals into the first-order Spherical Harmonics (SH), together with an Attribute-Conditioned SH Enhancement module whose predicted offsets are pre-baked before inference. We further propose a Multi-View Alpha-Based Densification and Pruning strategy to suppress redundant primitives while maintaining multi-view consistency. For dynamic scenes, we develop a compact explicit 4D representation by constructing second-order Gaussian motion, learnable temporal support, and a binary static-dynamic partition, enabling continuous-time modeling without runtime deformation networks. Based on this partition, a Depth-Order Certificate selectively reuses previously committed depth orders to reduce re-projection, sorting, merging, and index-buffer updates during playback. Extensive experiments on static and dynamic scenes demonstrate that Mobile-4DGS substantially reduces storage and rendering overhead while maintaining competitive visual quality, enabling real-time 3D and 4D Gaussian Splatting on mobile devices. \textcolor{magenta}{\href{https://xiaobiaodu.github.io/mobile-4dgs-project/}{Code has been released: https://xiaobiaodu.github.io/mobile-4dgs-project/}}.
|
| 120 |
BossouChimpanzee: Long-term Chimpanzee Video Dataset
2610.05293
|
cs.CV
|
Daniel Schofield, Susana Carvalho, Vladimir Iashin, Andrew Zisserman, Max Bain |
We describe the BossouChimpanzee video dataset, a unique long-term visual record of wild chimpanzees at an outdoor laboratory for field experiments in Bossou, Guinea, spanning three decades (1988-2018) and comprising over 1,200 hours of continuous video record...We describe the BossouChimpanzee video dataset, a unique long-term visual record of wild chimpanzees at an outdoor laboratory for field experiments in Bossou, Guinea, spanning three decades (1988-2018) and comprising over 1,200 hours of continuous video recordings collected through collaborative fieldwork and research. In this paper, we outline the history and scientific contributions of the experimental paradigm and video archive, provide key statistics and details on the structure of the main video dataset, and release an initial ~74h snapshot, BossouChimpanzee70h, covering 23 identified individuals focused on chimpanzee individual and action recognition, ahead of the full video resource. This dataset represents a valuable resource for cognitive and behavioural research in ethology and a rich benchmark for training and evaluating machine learning models on audiovisual data from the wild.
|
| 121 |
Revisiting Ground-Truth Synthesis from High-Speed Video: Exact Validity Conditions and an Audited Consumer Capture Corpus
2610.05298
|
cs.CVcs.AI
|
Abdullah Al Shafi, Sumaiya Rahim Suma |
Motion-deblurring datasets are commonly synthesised by averaging $N$ consecutive high-frame-rate frames and labelling the result with the central frame. We show that this label is unbiased for every capture timing only when the window is odd and the frames' sa...Motion-deblurring datasets are commonly synthesised by averaging $N$ consecutive high-frame-rate frames and labelling the result with the central frame. We show that this label is unbiased for every capture timing only when the window is odd and the frames' sample durations are equal. An even window shifts every label by a fixed fraction of the blur length, even under perfect timing. Unequal durations are subtler: their mean misalignment is zero, so dataset statistics cannot reveal them, yet when the offset cannot be read from the blur they convolve the supervision rather than adding noise to it. Auditing 51 smartphone clips recorded at a nominal 240 fps, we find 19 captured near 176 fps, a behaviour recorded only in the container's timing tables. In a controlled test the parity choice costs a fitted linear deblurring filter far more than these timing irregularities do, and interpolation labels, unlike blur labels, can be repaired with the true frame times. We provide a tool that checks both conditions without decoding, and will release the clips and their timing tables.
|
| 122 |
Kinematics-Centric Continuous Sign Language Retrieval with Gloss-Guided Boundary-Aware Alignment
2610.05306
|
cs.CV
|
Chang Liu, Ke Han, Davide Talon, Elisa Ricci, Nicu Sebe |
Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous ...Sign language-text alignment remains a fundamental challenge for text-driven sign language understanding. Existing methods predominantly rely on appearance-heavy RGB representations, which entangle motion semantics with visual variations and lead to ambiguous motion-language grounding. In this paper, we reformulate sign language-text alignment in a structured kinematic space and propose a kinematics-centric framework that adopts 3D SMPL-X motion as the primary representation. By explicitly modeling the kinematic dynamics of signing in a unified motion space, our approach reduces reliance on appearance signals and yields more semantically consistent representations. To capture the compositional nature of sign language, we introduce a gloss-guided local alignment mechanism that leverages gloss temporal spans as weak supervision to decompose continuous motion into coherent segments and establish fine-grained motion-text correspondences, thereby reducing ambiguity in localizing word-level semantics in continuous signing. Furthermore, we develop a visual distillation strategy, where RGB signals serve as privileged supervision during training to provide complementary contextual cues, while being completely removed at inference time. Extensive experiments on standard benchmarks demonstrate that our method achieves state-of-the-art bidirectional retrieval performance on CSL-Daily and competitive results on PHOENIX-2014T. These results highlight the effectiveness of kinematic representations and explicit local grounding for sign language-text alignment.
|
| 123 |
Shadow Feature Refinement Network: Progressive Feature Refinement based on Knowledge Distillation for Effective Shadow Removal
2610.05325
|
cs.CV
|
Donghyun Han, Byoung-Dai Lee |
In the field of deep learning, has seen significant advancements; however, shadow removal remains a persistent challenge owing to the variable sizes and colors of shadows influenced by lighting conditions. This study proposes a novel shadow feature refinement ...In the field of deep learning, has seen significant advancements; however, shadow removal remains a persistent challenge owing to the variable sizes and colors of shadows influenced by lighting conditions. This study proposes a novel shadow feature refinement network (SFR-Net), which leverages supervised learning, feature refinement loss, and knowledge distillation to enhance shadow removal performance. A dedicated post-processing algorithm is further introduced to restore natural color consistency in the generated shadow-free images. We evaluated our method on two public datasets: the adjusted image shadow triplet dataset (ISTD+) and the shadow removal dataset (SRD), which demonstrate strong generalization capabilities under diverse conditions. On ISTD+, our model achieved a root mean square error (RMSE) of 3.4627 and structural similarity index measure (SSIM) of 0.9382 across the entire image. On SRD, it recorded an RMSE of 4.3781 and an SSIM of 0.9341. These comprehensive results show that our approach performs competitively across both shadow and non-shadow regions while setting a promising direction for robust and perceptually natural shadow removal. Code is available at https://github.com/DongHyun99/SFRNet.
|
| 124 |
Riemannian Shape Analysis of the Corpus Callosum in Kendall Space: Aging and Alzheimer's Disease
2610.05326
|
cs.CV
|
Olakunle S. Abawonse, Fatou Fall |
The corpus callosum (CC) is a major white-matter structure and a well-established marker of brain aging, but most studies quantify it using scalar summaries that discard its boundary geometry. We present a Riemannian shape-space framework for analyzing age-rel...The corpus callosum (CC) is a major white-matter structure and a well-established marker of brain aging, but most studies quantify it using scalar summaries that discard its boundary geometry. We present a Riemannian shape-space framework for analyzing age-related morphological change in the midsagittal CC, applied to the OASIS-1 cohort. Each contour is represented by $128$ landmarks and embedded into Kendall shape space, where translation, rotation, and scale are removed. We derive a multivariate geodesic regression with exact Riemannian gradients and use the fitted age-velocity field to localize age-related deformation to five anatomical sub-regions. In the cognitively normal cohort ($n = 252$), geodesic regression outperforms the Euclidean linear benchmark ($R^2 = 0.1355$ vs.\ $0.1216$). Regional energy is posterior-dominant: the Splenium carries $41.4\%$ and the Isthmus $23.8\%$ of total age-related shape change, together accounting for $\sim 65\%$ despite comprising only $\sim 35\%$ of landmarks. Signed projections confirm the ordering (Splenium $r = 0.570$; Isthmus $r = 0.473$). In contrast, age explains less than $0.5\%$ of shape variance in Alzheimer's disease ($n = 88$), indicating that the disease disrupts the healthy aging trajectory. A tangent-space classifier achieves an age-group AUC of $0.791$ from the 2D contour alone, exceeding a recent volumetric benchmark ($0.67$).
|
| 125 |
WILLIE: A Unified Framework and Benchmark for Wound Classification, Segmentation, and Localization
2610.05341
|
cs.CV
|
Gopi Trinadh Maddikunta, Shannan Hamlin, Hsin-Mei Chen, Kimaya Barnes, Peizhu Qian |
Chronic wound management affects over 8.2 million patients in the United States and imposes substantial clinical and economic burden. Clinical wound assessment commonly involves three coupled tasks: identifying wound type, delineating wound boundaries, and loc...Chronic wound management affects over 8.2 million patients in the United States and imposes substantial clinical and economic burden. Clinical wound assessment commonly involves three coupled tasks: identifying wound type, delineating wound boundaries, and localizing the wound region for measurement and monitoring. Despite this clinical coupling, existing machine learning approaches typically address wound classification, segmentation, and localization using separate models. We present WILLIE, a unified framework and benchmark for wound classification, segmentation and localization that enables systematic evaluation of multi-task wound analysis under a common protocol. WILLIE harmonizes three public wound datasets into a shared benchmark and compares unified models across three scaling configurations against 10 single-task baselines. The best model achieves 91.88% classification accuracy, 91.41% Dice, and 96.23% AP@0.5 while producing all three outputs in a single forward pass. Beyond aggregate performance, our results show that segmentation-derived localization outperforms dedicated detection baselines in this benchmark, suggesting that box-based localization may be unnecessary for spatially coherent wound targets. Our findings highlight that effective multi-task learning in healthcare imaging depends not only on shared representations, but also on task formulation, compatibility, and benchmark design.
|
| 126 |
IRSTD-Agent: Agentic Infrared Small Target Detection via Zoom-Guided Interaction Learning
2610.05342
|
cs.CVcs.AI
|
Jiawen Xi, Yu Zhang, Tianyi Zhao, Zhu Liu, Maoxun Yuan |
Infrared small-target detection plays an important role in maritime monitoring and aerial surveillance. Although multimodal large language models (MLLMs) offer promising capabilities for visual understanding, existing MLLM-based approaches struggle to precisel...Infrared small-target detection plays an important role in maritime monitoring and aerial surveillance. Although multimodal large language models (MLLMs) offer promising capabilities for visual understanding, existing MLLM-based approaches struggle to precisely localize infrared small targets. In this paper, we propose IRSTD-Agent, an agentic framework for infrared small target detection through dynamic visual search. The framework enables an MLLM to adaptively determine where and at what scale to inspect an image and progressively gather fine-grained visual evidence for precise target localization. Five complementary visual tools (PROPOSAL, ZOOM, DETECT, DROP and REFINE) support object candidate discovery, adaptive observation, target localization, hypothesis rejection, and target extent refinement, together enabling a coordinated search process over original-resolution images. To teach the MLLMs to conduct this search, we introduce Zoom-guided Interaction Learning, which uses annotation-derived interaction trajectories to supervise tool selection and the corresponding arguments. Through extensive experiments on WideIRSTD-Full and IRSTD-1k datasets, we demonstrate that IRSTD-Agent outperforms the evaluated vision-language models and enhances the precise localization capabilities of MLLMs in IRSTD tasks.
|
| 127 |
Learning Conditional Source Distribution via Flow Reversal for Temporal Flow Matching
2610.05349
|
cs.CV
|
Kuan-Hsun Tu, Hsuan-Chi Liu, Jia-Wei Liao, Chien-Sheng Chiang, Tsung-Wei Ke |
We introduce CNP-Flow, a flow matching framework for temporal generation that learns conditional source distributions through flow reversal. Whereas standard conditional flow matching (FM) incorporates conditioning through the vector field and draws source sam...We introduce CNP-Flow, a flow matching framework for temporal generation that learns conditional source distributions through flow reversal. Whereas standard conditional flow matching (FM) incorporates conditioning through the vector field and draws source samples from a standard Gaussian, CNP-Flow uses a conditional noise predictor (CNP) to produce an isotropic Gaussian source for each temporal condition. The CNP is supervised by source samples obtained through flow reversal, which maps observed targets backward through a pretrained FM model. A three-stage pipeline pretrains the FM model, trains the CNP, and fine-tunes the FM model using the learned source distribution, while preserving the FM backbone architecture. Across video prediction, video interpolation, and 7-DoF Franka robot motion planning, CNP-Flow consistently improves generation quality. It also matches baseline performance with fewer function evaluations. Project page: https://embodiedai-ntu.github.io/cnpflow
|
| 128 |
CleanMDM: Clean Motion Diffusion Model for Multimodal Motion Cleanup
2610.05411
|
cs.CV
|
Zhe Li, Shicheng Wang, Bowen Cai, Huan Fu |
Motion capture data is rarely directly usable, as they typically exhibit missing segments, jitter, drift and contact artifacts. Traditionally, corrupted motions are cleaned by animators through the manual identification of keyframes from noisy motion, subseque...Motion capture data is rarely directly usable, as they typically exhibit missing segments, jitter, drift and contact artifacts. Traditionally, corrupted motions are cleaned by animators through the manual identification of keyframes from noisy motion, subsequent keyframe correction, and interpolation between corrected keyframes to reconstruct coherent motion. While the rise of generative motion models has made automatic cleanup feasible, most approaches operate as black box denoisers with limited controllability, making it difficult to preserve reliable segments or enforce specific user intents. Inspired by animation workflows, we present CleanMDM, a unified multimodal motion cleanup framework that formulates cleanup as masked conditional generation with plug-and-play conditions. This single model supports arbitrary combinations of noisy 3D motion, sparse 2D keyframes, sparse 3D keyframes, and text. This design enables both automatic cleanup without additional user annotation and controllable cleanup under multimodal guidance. To further improve motion realism, we incorporate the Latent Motion Quality Discriminator (LMQD) to better match kinematic distributions and reduce skating, jitter, and interpenetration artifacts, and we apply Mesh-Aware Contact Projection as a test-time optimization step to enhance contact and physical consistency. Experiments across multiple datasets demonstrate that CleanMDM consistently outperforms prior cleanup and generation baselines, and that low cost conditions (text and 2D keyframes) provide reliable controllability gains in multimodal cleanup scenarios.
|
| 129 |
A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models
2610.05413
|
cs.CVcs.AI
|
Yilin Yang, Jun-Tao Tang, Kengyi Wang, Siyuan Su, Gaoyong Luo |
Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain...Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.
|
| 130 |
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training
2610.05416
|
cs.CV
|
Shuyuan Tu, Qi Tian, Yinming Huang, Yue Wu, Xintong Han |
Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingl...Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informative content and disrupting pretrained priors. Existing sparse attention methods either target training-free acceleration or overlook the unique structure of joint video-audio data, where cross-modal interactions are inherently concentrated around sound-producing regions. To address this, we propose Prism, a dynamic sparse attention framework for natively training joint video-audio generation models at 2K. In particular, Prism organizes the token sequence into spatiotemporal macro-zones, enabling the attention structure to adapt to local content. For each zone, it estimates local information structure via video feature variance along the channel and feature norms from the audio-to-video cross-attention, jointly capturing how visual content varies directionally and how strongly audio influences each visual region. Based on these signals, Prism dynamically assigns a tailored block shape to each zone, applying finer partitioning along axes of rapid visual content variation and strong audio-visual coupling. This encourages tokens within each block to remain semantically coherent, allowing block-level features to capture both visual content and joint video-audio interaction patterns. Prism further adopts a hybrid block selection strategy to dynamically determine per-query sparsity. Experiments show that Prism achieves 2.5$\times$ training speedup compared to full attention, while surpassing it in generation quality.
|
| 131 |
Render to Reason: Novel-View Semantic Prediction Improves Spatial Understanding in VLMs
2610.05417
|
cs.CV
|
Yuqun Wu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem |
Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields onl...Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose \textbf{novel-view semantic rendering} as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset) and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI. Project page: https://yuqunw.github.io/Render2Reason/.
|
| 132 |
Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking
2610.05425
|
cs.CVcs.CLcs.LG
|
Ali Vosoughi, Akhil Kasturi, Chenliang Xu, Axel Wismueller |
Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence und...Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.
|
| 133 |
DynaMesh: Dynamic 3D Texture Generation
2610.05529
|
cs.CV
|
Raj Hansini, Guan Chen, Rana Hanocka, Itai Lang |
We present DynaMesh, a dynamic texture generation method for 3D meshes. Given a textureless shape and a text prompt describing an effect, our method produces an appearance that evolves while the object's geometry remains unchanged. Previous works on dynamic 3D...We present DynaMesh, a dynamic texture generation method for 3D meshes. Given a textureless shape and a text prompt describing an effect, our method produces an appearance that evolves while the object's geometry remains unchanged. Previous works on dynamic 3D content generation have focused on motion, where an object's geometry and position change while keeping its appearance the same. Methods on texture generation sit on the other side of the problem, painting appearance onto a shape as a fixed surface property and not as an evolving process. Neither addresses a visual effect that propagates on a 3D object. A natural route consists of two generators: a video model that shows the effect from a single view, and an image-to-3D generator that lifts each frame to 3D. However, the latter has no notion of time, so running it per video frame produces a sequence that flickers, loses effect details, and yields a different mesh at every video frame. Our method addresses these failures by conditioning a video model on a render of the mesh and the prompt to obtain a reference video, then running a frozen image-to-3D generator on the video with two changes. The conditioning of each frame is blended over a temporal window, and low-rank adapters are fit per shape to restore the lost details. The mesh is encoded once for the whole sequence, so geometry is constant by construction, and the output is a single mesh with a texture per frame. Applied to various objects and effects, DynaMesh substantially improves over recent video-to-4D and texturing methods, and can generalize its temporal effect to different shapes never seen during training. Our project page is at https://threedle.github.io/dynamesh/.
|
| 134 |
Deep Prior Learning for Embodied Perception
2610.05531
|
cs.CV
|
Yimou Wu, Jiaxin Guo, Yun-hui Liu, Zheng Li |
Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate p...Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate priors, and recovering physical scale require more than simply accepting these inputs. We introduce \emph{Vision-Prior Geometry Grounded Transformer} (VPGGT), a VGGT-based framework that extends OmniVGGT for prior-aware embodied perception. We formulate sensor-motivated pose corruptions from ground-truth trajectories for training and introduce a parameter-free \emph{prior residual connection} (PRC) to mitigate \emph{prior dilution}, where predictions are less accurate than their supplied pose priors. Our noise formulation targets camera poses; supplied intrinsics and depth receive no additional corruption. We further introduce \emph{Metric Global Attention}, which conditions a global scale token on available pose and depth scales and predicts a shared metric scaling factor for the geometric outputs. Experiments across four datasets show that \emph{PRC} improves translation-direction accuracy and joint pose AUC over a matched training baseline when camera priors are provided for all views, under both exact and corrupted poses. These results support explicit prior access during refinement as a useful addition to feature-level conditioning.
|
| 135 |
Monocular markerless biomechanics for clinically interpretable gait assessment in spinal cord injury
2610.05552
|
cs.CV
|
Shreyasvi Natraj, Mathieu Ruepp, Yanke Li, Robert Riener, Inge Eriks-Hoogland |
Three-dimensional gait analysis guides rehabilitation after spinal cord injury but depends on marker-based motion capture and force plates, which few clinics have. Monocular markerless pipelines have been established in fewer healthy adult cohorts but not in n...Three-dimensional gait analysis guides rehabilitation after spinal cord injury but depends on marker-based motion capture and force plates, which few clinics have. Monocular markerless pipelines have been established in fewer healthy adult cohorts but not in neurological cohorts. We present the SCAI SCI Gait dataset, comprising 239 adult individuals with spinal cord injury with synchronized video, motion capture, and force-plate measurements, we fitted a parametric body mesh to a single sagittal-view video, driving an anthropometrically scaled OpenSim model via virtual markers. Markerless lower-body kinematics showed state-of-the-art agreement with motion-capture measurements (r = 0.68-0.90, p < 0.001, and RMSE = 4.18-6.49 degrees), and accurate kinematics-based predicted ground-reaction forces closely matched those measured by force plates (r = 0.85-0.87, p < 0.001, and RMSE = 2.13-2.19 Newton per kg). Furthermore, conditional-dependence graph analysis with Markov blankets revealed that waveform components were conditionally associated with functional independence, and speed-stratified clustering revealed distinct mechanical strategies among individuals walking at similar speeds. These findings establish the use of monocular video as a scalable approach for clinically meaningful biomechanical assessment and data-driven phenotyping in patients with spinal cord injury. Github: https://github.com/SCAI-Lab/SCAI-SCI-Gait
|
| 136 |
SteadySplats: Resampling of Low-Variance Gaussians for High-Fidelity Stochastic Rendering
2610.05576
|
cs.CV
|
Felix Windisch, Thomas K\"ohler, Lukas Radl, Chris Wyman, Georgios Kopanas |
Stochastic order-independent transparency enables efficient and elegant rendering of primitive-based radiance fields like 3D Gaussian Splatting models, but remains impractical due to the inherent visible noise in the output. We propose a principled approach to...Stochastic order-independent transparency enables efficient and elegant rendering of primitive-based radiance fields like 3D Gaussian Splatting models, but remains impractical due to the inherent visible noise in the output. We propose a principled approach to minimize high-frequency noise, addressing its sources at the representation and image synthesis level. During stochastic rendering, our history-based spatial resampling scheme drastically accelerates image convergence, while temporal importance resampling ensures coherence under camera movement. During training, a color regularizer implicitly reduces the variance along view rays in the 3DGS models. With these properties, our optimized, Vulkan-based renderer effectively mitigates output noise at low and high sample counts, achieving a substantial 13~dB PSNR increase in quality over previous stochastic methods at 1 sample per pixel and quickly converging to sorted 3DGS with an average L1 error of less than $10^{-4}$.
|
| 137 |
Generating the Wild: Individual-Consistent Image-to-Video Generation for Wildlife
2610.05587
|
cs.CV
|
Yuzhuo Li, Di Zhao, Xinyu Zhang, Daniel Wilson, Yun Sing Koh |
Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this lim...Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this limitation by synthesizing additional observations from a single reference image. However, existing I2V models mainly emphasize global layout, semantics, and motion, and therefore often fail to preserve fine-grained local appearance cues that distinguish one wildlife individual from another, such as fur texture, stripe boundaries, spot configurations, and contour transitions. We observe that these identity-critical cues are closely related to high-frequency information. To address this challenge, we propose WildIcon, a high-frequency-guided I2V framework for wildlife individual consistency. Specifically, WildIcon introduces a frequency-aware identity encoding branch that extracts individual-specific high-frequency cues from the reference image. Combined with isolated foreground information, the resulting identity tokens are then injected into cross-attention blocks as identity conditioning. Building on a frozen backbone with lightweight identity adaptation, WildIcon preserves fine-grained identity cues visible in the reference image while retaining the motion controllability and semantic fidelity of the base I2V model. In addition, to support the training and evaluation of wildlife individual-consistent I2V, we construct WildlifeVid, a wildlife-centric video dataset with high-quality, temporally coherent clips and individual-level identity labels. Experiments on I2V generation and downstream animal re-identification (ReID) show that WildIcon achieves stronger individual consistency than existing baselines, and that its filtered outputs can serve as useful candidate training augmentations for downstream ReID.
|
| 138 |
EchoDino: A pediatric foundation model for transferable echocardiographic analysis across the lifespan
2610.05603
|
cs.CV
|
Sheng Cheng, Donnchadh M. O'Sullivan, Daniel J. Penny, Craig G. Rusin, Minh B. Nguyen |
Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they ...Echocardiography is the most widely used cardiac imaging modality, yet interpretation demands integrating visual evidence across global anatomy, localized structures and dynamic cardiac motion. Machine-learning models have automated individual tasks, but they are typically built for a single purpose and depend on expensively labeled datasets - a barrier particularly acute in pediatric care, where data are scarce and anatomy changes with age. Here we present EchoDino, a self-supervised foundation model for echocardiography, created by adapting the DINOv3 framework to 3.7 million frames from 1.7 million unlabeled pediatric echocardiography videos. With its encoder frozen, EchoDino produces representations that capture global context, local anatomy, and dense spatial detail. We introduce Motion-biased Entropy Maximization Sampling (MEMS) to select the most informative frames for video-level analysis. Across nine pediatric and adult datasets, EchoDino outperformed strong baseline models, raising view-classification accuracy from 0.609 to 0.889 and the area under the receiver operating characteristic curve for structural-heart-disease detection from 0.811 to 0.872, while also cutting age-estimation error from 3.857 to 1.389 years, achieving the best segmentation accuracy and lowering ejection-fraction errors. By generalizing from label-free pediatric data to adult echocardiography, EchoDino offers a versatile foundation for cardiac image analysis across the lifespan.
|
| 139 |
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
2610.05608
|
cs.CVcs.LGcs.AIcs.MM
|
Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin, Grigorii Alekseenko |
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips...We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920$\times$1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.
|
| 140 |
StageVLN: Spatial and Trajectory Auxiliary Guidance for Efficient Vision-Language Navigation
2610.05664
|
cs.CV
|
Anh Dao, Quan-Dung Pham, Le Danh Vinh, The Anh Nguyen, Nguyen Viet Tri Pham |
Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene ge...Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3\% SR and 51.4\% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3\% SR without additional navigation training data or a geometry encoder at inference.
|
| 141 |
Difference Feature Map Distillation: Transferring Inter-Sample Relational Knowledge Towards Efficient Transformer-Based Tracking
2610.05707
|
cs.CV
|
Zhicheng Ding, Xinyu Chu, Qing Tian |
In autonomous driving perception, visual object tracking systems must satisfy stringent latency and power constraints while remaining robust in complex and dynamic environments. Although transformer-based trackers achieve state-of-the-art accuracy, their subst...In autonomous driving perception, visual object tracking systems must satisfy stringent latency and power constraints while remaining robust in complex and dynamic environments. Although transformer-based trackers achieve state-of-the-art accuracy, their substantial computational and memory overheads hinder deployment on real-time, resource-constrained platforms. To move toward this goal, we propose Difference Feature Map Knowledge Distillation (DFM-KD), a novel relational distillation framework tailored for transformer-based visual object tracking. Unlike conventional feature distillation methods that minimize point-wise discrepancies (e.g., mean squared error) between teacher and student feature representations, DFM-KD transfers knowledge through inter-sample feature differences, explicitly aligning the relational structure of the feature space. By distilling how the teacher models appearance variation and consistency across samples, rather than enforcing similarity in absolute activations, DFM-KD enables the student to better capture the structural dynamics of visual changes within a batch. As a result, the distilled model exhibits enhanced feature robustness and improved tracking performance. Extensive experiments demonstrate that DFM-KD consistently outperforms conventional feature-level distillation methods in both tracking precision and success rates.
|
| 142 |
Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs
2610.05715
|
cs.CVcs.AI
|
Zhaochen Wang, Yujun Cai, Huangbo Zou, Hower Yang, Naipeng Dong |
Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitu...Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitude estimation, and multi-view rotation reasoning. Across 12 VLMs, the gap is stark: the strongest models approach 100% accuracy on detection, yet even coarse magnitude estimation is near chance. When asked for exact angles, models place 91.8--100% of their predictions on just $0^\circ$, $90^\circ$, and $180^\circ$, a failure we term canonical-angle collapse. This collapse persists even without visual input. Representation probing shows that missing information is only part of the explanation. Although rotation information becomes less recoverable at finer granularity, substantial coarse-grained information remains, and a simple linear probe outperforms the models' generated answers. This suggests that VLMs underuse rotation information they already encode. We therefore propose RotationCue, a lightweight decoder that recovers coarse rotation information from the VLM's own frozen representations and feeds it back to the model as intermediate textual context. Across three VLMs, RotationCue improves every model--task combination on OR-Bench, raising macro-average accuracy by 7.9--12.6 points while preserving general capabilities.
|
| 143 |
T-JEPA: A Temporal Joint-Embedding Predictive Architecture for Learning Better Remote Sensing Representations
2610.05731
|
cs.CV
|
Bowen Peng, Li Liu, Yongxiang Liu, Weijie Li, Jie Zhou |
Earth observation (EO) data provide rich temporal supervision, yet existing remote sensing foundation models mainly exploit sequential observations through imposing predefined pairwise relations or aggregating holistic reconstruction context. We seek to furthe...Earth observation (EO) data provide rich temporal supervision, yet existing remote sensing foundation models mainly exploit sequential observations through imposing predefined pairwise relations or aggregating holistic reconstruction context. We seek to further exploit the sparse and nonuniform temporal sampling inherent in EO sequences as supervisory signals. To this end, we propose T-JEPA, a temporal joint-embedding predictive architecture that learns time-gap-conditioned latent transitions. A shared single-frame encoder processes each observation, while a temporal predictor estimates the complete target latent field from a masked source latent representation and the actual elapsed time. Across multiple temporal intervals, these predictive constraints organize observed states into structured latent trajectories. Asymmetric metadata injection mitigates shortcut learning, and direct supervision across multiple temporal scales proves more effective than recursively rolling out intermediate states. In parallel, masked pixel reconstruction provides complementary supervision for preserving spatial details. Under matched pre-training data and throughput, T-JEPA achieves leading transfer performance on both static and temporal tasks. Analyses further reveal that T-JEPA learns representations with time-gap-dependent transition predictability and coherent latent dynamics, while maintaining strong cross-period consistency, representation diversity, and semantic discriminability.
|
| 144 |
HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models
2610.05739
|
cs.CVcs.AI
|
Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu, Jiahao He |
Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-s...Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the $60$-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a $0.74$ dB PSNR gain and a $28.5\%$ reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over $547$ samples. At a $60$-second context, HLA-WM reduces historical-state memory by $12\times$ relative to full KV caching while incurring at most a $1.6\%$ reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN. Project page: https://caesarhhh.github.io/hla-wm/
|
| 145 |
Robust Local Optimization Done Right
2610.05743
|
cs.CV
|
James Pritts, Kevin K\"oser |
RANSAC scoring and local optimization (LO) impose different robustness requirements, motivating the separation of hypothesis selection from refinement. We systematically isolate the effects of robust-loss shape, incorrectly specified inlier scales, and optimiz...RANSAC scoring and local optimization (LO) impose different robustness requirements, motivating the separation of hypothesis selection from refinement. We systematically isolate the effects of robust-loss shape, incorrectly specified inlier scales, and optimization strategy on essential matrix, fundamental matrix, and homography estimation. A profile-marginal score marginalizes the nuisance inlier scale and selects an inlier partition, from which we estimate the scale that sets the LO loss width; this makes LO robust to an inlier scale specified too large, whereas one specified too small degrades selection itself. Refinement needs gradient from correspondences the seed currently rejects: optimizers that reweight from current residuals stay pinned to their seed, whereas methods with broad basins recover strongly perturbed seeds yet degrade accurate score-selected hypotheses, so basin size alone is insufficient to assess RANSAC LO. Joint half-quadratic optimization balances the two and is the most consistent strategy across model classes. An optimizer matched to the profile-marginal score, which never decreases it, does not reach the best accuracy, challenging the prescription that scoring and refinement objectives should match. Composed from these findings, our RANSAC reduces the median essential-matrix pose error of a state-of-the-art RANSAC on PhotoTourism from 2.23 degrees to 1.58 degrees with a correctly specified inlier scale and from 38 degrees to 6.2 degrees when it is grossly misspecified (128x too large).
|
| 146 |
A new design of a fall detection system integrating landmark identification and deep learning techniques
2610.05749
|
cs.CV
|
Tri Nhut Do, Thi Thuy Le |
This article introduces an innovative system that integrates landmark identification with deep learning to enhance fall detection accuracy and reliability. By utilizing advanced computer vision techniques, such as Media Pipe for spatial recognition, the system...This article introduces an innovative system that integrates landmark identification with deep learning to enhance fall detection accuracy and reliability. By utilizing advanced computer vision techniques, such as Media Pipe for spatial recognition, the system effectively differentiates between routine movements and actual falls. The integration of landmarks with a deep learning prediction algorithm minimizes false alarms, ensuring timely responses to genuine falls. Comprehensive experimentation underscores the system's versatility across various scenarios, emphasizing its potential to improve safety and independence for older adults. The training process demonstrates a steady increase in accuracy, stabilizing by the 40th cycle, while error rates decline significantly during the initial cycles. Real-time experiments, involving both male and female participants aged 8 to 50, recorded a remarkable 95% detection rate of falls, demonstrating the system's effectiveness and promising future applications in elder care and smart health monitoring environments.
|
| 147 |
Vision-enabled detection of safety helmet compliance in construction zones
2610.05756
|
cs.CV
|
Tri Nhut Do*, Ba Loc Pham |
In the rapidly evolving field of construction management, worker safety remains a top priority. This paper introduces an innovative vision-based system for real-time detection of helmet compliance, specifically designed for construction sites, utilizing advanc...In the rapidly evolving field of construction management, worker safety remains a top priority. This paper introduces an innovative vision-based system for real-time detection of helmet compliance, specifically designed for construction sites, utilizing advanced computer vision techniques and machine learning algorithms within the YOLO (you only look once) framework. Our system leverages high-resolution video feeds from strategically positioned cameras to monitor adherence to safety regulations regarding helmet usage. By employing deep learning methodologies, the system effectively identifies individuals not wearing helmets, thereby significantly mitigating the risk of head injuries among workers. Our training and validation results revealed an impressive precision exceeding 97% at mAP@0.5 for both helmeted and non-helmeted individuals. Furthermore, our experiments demonstrate exceptional detection accuracy, demonstrating the system's resilience under varying lighting conditions and diverse worker movements. The consistent decrease in loss and improvement in metrics throughout training validates the effectiveness of the YOLOv8 model in enhancing recognition performance. The implications of this research extend beyond mere regulatory compliance, opening avenues for innovative applications in occupational safety management. This study highlights the critical role of technology in protecting lives and lays the groundwork for future advancements in smart construction environments.
|
| 148 |
A Three-Dimensional Reverse-Projection Method for Sparse Point Cloud Completion and Its Application to High-Speed Train Nose Reconstruction
2610.05758
|
cs.CV
|
Xiaozhen Ma, Zhao Tang, Hanbin Lai, Ruiqi Chen, Jin Jin |
Holes in LiDAR scans of environments with glass windows remain an unresolved problem in three-dimensional reconstruction. This study presents a pipeline for point cloud acquisition, filtering, completion, and surface reconstruction to address sparse sampling a...Holes in LiDAR scans of environments with glass windows remain an unresolved problem in three-dimensional reconstruction. This study presents a pipeline for point cloud acquisition, filtering, completion, and surface reconstruction to address sparse sampling and missing window regions in scans of a high-speed train nose. FAST-LIVO2 provides the initial point cloud through multisensor odometry and mapping, and moving least squares (MLS) smooths the observations. We then introduce three-axis projection-based subdivision and interpolation with reverse hole boundary identification, referred to as three-axis reverse completion. The method interpolates missing regions from observations around each hole. Greedy projection triangulation, Poisson surface reconstruction, and a Marching Cubes-based pipeline generate meshes from the completed point cloud. Experiments on a proportionally scaled display model of a high-speed train nose show that the proposed method fills missing point cloud regions around the glass windows. Under the evaluation setting used in this study, greedy projection triangulation yields lower geometric distance errors than the other two reconstruction pipelines. The pipeline supports non-contact digital modeling of train nose geometry and provides a practical approach to reconstructing objects with glass windows.
|
| 149 |
Controllable Road Marking Generation
2610.05771
|
cs.CV
|
Zhiyu (Joey), Cai, Yufan Zhang, Ruichen Tan, Zengxiang Lei |
Lane and road markings provide critical guidance for vehicle navigation and multi-agent coordination, yet authoring them at scale remains a manual workflow that limits quantitative analysis and scenario testing. We introduce Controllable Road Marking Generatio...Lane and road markings provide critical guidance for vehicle navigation and multi-agent coordination, yet authoring them at scale remains a manual workflow that limits quantitative analysis and scenario testing. We introduce Controllable Road Marking Generation, which synthesizes a missing center-region marking layout from a drivable-area mask, optional outer-ring markings, and a textual description. Our benchmark uses deterministic, metadata-derived prompts and three output channels: lane dividers, road dividers, and pedestrian crossings. We develop a conditional bird's-eye-view (BEV) pipeline that combines (i) a text-conditioned latent rectified-flow DiT trained with a topology-aware auxiliary loss, (ii) Gaussian-blurred training targets that stabilize learning of thin, sparse markings, and (iii) Structured Gaussian Render (SGR), a training-free post-process that recovers crisp divider geometry by extracting polylines, fitting cubic B\'ezier curves, and re-rendering them as anisotropic super-Gaussian primitives. On 4,597 Argoverse~2 test tiles, our system achieves Buffered F1 of 80.8 and clDice of 50.2, compared with 38.8 and 24.6 for an adapted state-of-the-art mask-refinement baseline. On Waymo dataset, it yields 88.0 Buffered F1 and 66.2 clDice. Component ablations show complementary connectivity gains from topology-aware supervision and SGR. Text-editing experiments reveal that stronger guidance improves edit success but also increases changes to non-target structures. We see this framework as a step toward simulation-ready road-marking variation, automated map completion, and early-stage infrastructure design exploration.
|
| 150 |
InteractionBench: A Real-Time Interaction Benchmark for Streaming Video Systems
2610.05775
|
cs.CV
|
Enxin Song, Suhao Yu, Yifei Xu, Barbara Su, Weili Xu |
A video assistant must speak when its instruction warrants a response and stay silent otherwise. We introduce a benchmark that evaluates this decision for the complete system of model, memory, and response controller. InteractionBench covers query responses, e...A video assistant must speak when its instruction warrants a response and stay silent otherwise. We introduce a benchmark that evaluates this decision for the complete system of model, memory, and response controller. InteractionBench covers query responses, event triggers, and ongoing updates in 1,060 interactions over 812 videos, with 69 negative streams and 53 suites that pair counted events with look-alike near misses. It scores content accuracy, timing accuracy, and silence compliance on the video clock. Timely speech costs silence across systems. Polled Qwen3-VL-8B reaches 77.8 timing accuracy but 10.9 silence compliance. A native real-time interaction system reaches 29.2 silence compliance at 66.8 timing accuracy, yet emits on 89.9% of negative streams. No open-weight system clears a third of the near-miss suites. Fewer replies help only when chosen, as random deletion merely trades timing for silence. Offline scores miss these failures and mispredict online behavior. Adding restraint is costly, as the native system's controller adds little by itself and agentic systems add it only at about 30 s per poll.Project page: https://www.enxinsong.com/projects/interactionbench/ Code: https://github.com/Espere-1119-Song/InteractionBench Data: https://huggingface.co/datasets/InteractionBench/InteractionBench
|
| 151 |
A Spatiotemporal Semantic Importance-Guided Unified Compression and Editing Framework for AI-Generated Videos
2610.05779
|
cs.CVcs.MM
|
Xihua Sheng, Dong Liu, Chang Wen Chen |
AI-generated videos are rapidly increasing in volume, duration, and resolution, creating growing demands for efficient storage and transmission. Unlike natural videos captured from the physical world, AI-generated videos are samples from a learned generative d...AI-generated videos are rapidly increasing in volume, duration, and resolution, creating growing demands for efficient storage and transmission. Unlike natural videos captured from the physical world, AI-generated videos are samples from a learned generative distribution, where semantic structures are critical to content consistency, while many local textures and stochastic details can be plausibly regenerated. This distinction suggests that compression should preserve semantically important spatiotemporal information rather than reconstruct every pixel of a particular generative sample. Beyond reconstruction, AI-generated videos also create a practical need for prompt-based editing, where users expect to modify generated content while preserving its original spatiotemporal semantics. Motivated by these observations, we propose a unified compression and editing framework for AI-generated videos that incorporates a frozen video generator as a reusable generative prior. Within this framework, we design three spatiotemporal semantic importance-guided techniques that respectively address what to transmit, how much to transmit, and how to use the transmitted side information. First, an innovation selection method projects the latent discrepancy using spatiotemporal semantic importance, so that the selected innovations prioritize semantic invariants over replaceable generative variations. Second, a frame-adaptive bit allocation method estimates the nonuniform semantic demands of latent frames and allocates more innovations to frames requiring stronger semantic preservation. Third, a unified reconstruction and editing method continuously adjusts the influence of the transmitted side information, enabling the same compressed representation to provide strong guidance for faithful reconstruction or serve as a flexible semantic anchor for structure-preserving prompt-driven editing.
|
| 152 |
FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models
2610.05790
|
cs.CV
|
Md Aminur Hossain, Omkumar Vaghasiya, Rajeev Ranjan Dwivedi, Vinod Kurmi, Biplab Banerjee |
Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in R...Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in RSFMs. FairRSFM maps georeferenced samples from 14 terrestrial biome classes into six ecologically meaningful macro-groups and evaluates models under a unified frozen-backbone evaluation protocol. The benchmark covers four downstream datasets: m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K with Dynamic World label maps. Using Prithvi-EO-2.0, SatMAE, and DOFA across three random seeds, we show that aggregate performance consistently masks biome-dependent disparities across architectures and tasks. For example, Prithvi-EO-2.0 reaches 90.98% overall macro-F1 on m-EuroSAT but a mean worst-group score of only 83.72%, while m-SA-Crop-Type drops from 27.30% overall mIoU to 18.47% in the Xeric and Mineralogical group. We further evaluate Biome-Orthogonal Linear Probing (BOLP), Dynamic Biome Reweighting (DBR), and GroupDRO as complementary mitigation baselines. Their effectiveness is model- and task-dependent; for example, BOLP improves Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without updating the RSFM backbone. FairRSFM provides a reusable protocol for diagnosing and mitigating ecological robustness gaps in remote sensing foundation models. Code and datasets are available at: https://github.com/aminurhossain/FairRSFM.
|
| 153 |
Gauss-Map Variation for Image Denoising: Geometric Analysis and an Anderson--Accelerated Majorization--Minimization Method
2610.05801
|
cs.CV
|
Haibin Su |
We propose a Gauss-map variation (GMV) model for image denoising that measures the spatial variation of the tangent-plane projectors of the scaled image graph. We establish an equivalent representation of the regularizer in terms of the corresponding Gauss map...We propose a Gauss-map variation (GMV) model for image denoising that measures the spatial variation of the tangent-plane projectors of the scaled image graph. We establish an equivalent representation of the regularizer in terms of the corresponding Gauss map and, using differential geometric tools including tubular coordinates and the Frenet frame, analyze its behavior across general $C^2$ and piecewise $C^2$ boundaries. The resulting estimates provide edge- and corner-contrast preservation properties. To solve the proposed model, we introduce a bilinear decomposition involving a unit normal field and a scalar magnitude field and develop an Anderson-accelerated majorization--minimization algorithm. The normal field subproblem admits an explicit pointwise majorization--minimization update, which is combined with an Anderson acceleration. For both $L^1$ and $L^2$ data fidelity terms, we establish sufficient decrease and boundedness of the iterates and prove that the generated sequence converges to a critical point of the penalized model. Numerical experiments on synthetic and natural images demonstrate the boundary preserving capability of the proposed model and its competitive performance in removing Gaussian and impulsive noise.
|
| 154 |
Level-of-Token Diffusion
2610.05816
|
cs.CVcs.AI
|
Kiyohiro Nakayama, Brian Chao, Jan Ackermann, Hansheng Chen, Federico Tombari |
Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be red...Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framework that turns this knowledge into an explicit multiresolution token layout (Level-of-Token layout) for adaptive and efficient generation. Tokens represent rectangular patches of varying sizes and shapes, allocating finer tokens where detail is needed and coarser tokens elsewhere. We adapt pretrained diffusion transformers to LoT layouts through a patch-wise asymmetric flow parametrization and embeddings for multiresolution tokens, preserving full-resolution flow prediction at every denoising step while processing only a reduced token sequence. LoT Diffusion enables layout-adaptive generation while preserving pretrained generative priors. We demonstrate LoT with layouts derived from semantic masks, bounding boxes, texture variance, and depth-of-field cues, as well as agentic plans. Across image and video generation, LoT offers favorable quality-efficiency tradeoffs, with significant speedups determined by the layout's token budget. Our project website is at https://georgenakayama.github.io/lotdiffusion/.
|
| 155 |
Imagine to Act: High-Fidelity Data Synthesis via Image Editing World Model for Scalable GUI Agent Training
2610.05861
|
cs.CVcs.AI
|
Yongxin Ning, Runliang Niu, Qianli Xing, Zhiyi Duan, Qingzu He |
Graphical User Interface (GUI) agents have emerged as a promising paradigm for automating complex digital workflows across diverse applications. However, training highly capable and generalizable agents fundamentally relies on massive, high-fidelity visual-act...Graphical User Interface (GUI) agents have emerged as a promising paradigm for automating complex digital workflows across diverse applications. However, training highly capable and generalizable agents fundamentally relies on massive, high-fidelity visual-action trajectories, which are notoriously difficult to acquire. While human demonstrations are unscalable, existing GUI world models rely on text descriptions or HTML rendering, discarding crucial pixel-level visual details like icons and layout styles. To address this issue, we introduce Infinite-Dreamer, a simulation-free data synthesis method powered by a pixel-level Image Editing World Model. By conceptualizing GUI transitions as image editing tasks, we leverage Vision-Language Models (VLMs) to describe action-induced UI changes as structured delta-text. We then fine-tune an image editing backbone to controllably synthesize realistic screenshot transitions. We utilize this model to generate both single-frame visual robustness data and multi-step imaginary trajectories. To validate the effectiveness of our approach, we fine-tune the Qwen3-VL baseline solely on the synthesized data to obtain Infinite-Actor, and evaluate it on AndroidWorld, MobileWorld, and AndroidControl-Curated benchmarks. Infinite-Actor consistently outperforms the Qwen3-VL baselines across scales: Infinite-Actor-8B improves AndroidWorld Pass@1 by +4.45 and nearly doubles the MobileWorld Pass@3 success rate, while Infinite-Actor-2B improves Pass@1 by +9.05. Code is available at https://github.com/swaydy-n/Infinite-Dreamer.
|
| 156 |
Weave Mamba Fusion: Global Cross-Scale Interaction for Lightweight Face Detection
2610.05865
|
cs.CVcs.AI
|
Dohun Kim, Jinmyung Jung |
Feature pyramid methods, from FPN to BiFPN, have achieved strong performance in face detection by fusing multi-scale features. However, detecting faces under unconstrained conditions, such as small scale, occlusion, and extreme pose, remains difficult, as it r...Feature pyramid methods, from FPN to BiFPN, have achieved strong performance in face detection by fusing multi-scale features. However, detecting faces under unconstrained conditions, such as small scale, occlusion, and extreme pose, remains difficult, as it requires global cross-scale dependencies that local fusion cannot model. State space models such as Mamba provide global context with linear complexity by scanning features as a sequence, and therefore offer a promising direction for this problem. Nevertheless, such a scan needs the two pyramid scales combined into a single feature map, and the way they are combined determines whether cross-scale structure is preserved. Summation collapses the two scales before the scan, so the scan has no cross-scale structure to exploit, while concatenation keeps both scales but at far higher cost. To address this, we propose \textbf{Weave Mamba Fusion (WMF)}, which interleaves two adjacent pyramid scales column by column so that each step of a horizontal bidirectional SS2D scan moves from one scale to the other. With partial-channel processing and parameter-free de-weaving, WMF enables efficient cross-scale interaction while preserving feature structure. Integrating WMF into every fusion node yields \textbf{WeaveBiFPN}, the neck of our \textbf{WeaveFace} detector. On WIDER FACE, WeaveFace achieves 91.41\% mean AP with only 0.34M parameters and 1.16 GFLOPs, outperforming prior detectors under 0.5M parameters. Its largest gains are on the Hard subset, where it reaches 87.14\% AP. The code is publicly available at \url{https://github.com/dohun-mat/WeaveMambaFusion}.
|
| 157 |
Certification of Real Images through Calibrated Content Authentication
2610.05870
|
cs.CV
|
Sarim Hashmi, Abdelrahman Elsayed, Mohammed Talha Alam, Samuele Poppi, Nils Lukas |
Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-...Generative models can synthesize high-quality inauthentic multimedia content that is already being misused at scale. We evaluate twenty deepfake detectors against ten generators released in the last four years and find accuracy decreasing over time, from near-perfect 99.5% to 76%. Adversarial perturbations further reduce every baseline detector to below 2% accuracy, effectively inverting the detector's assigned label. We argue that this unreliability reflects a fundamental ambiguity: generators can reproduce authentic content exactly (e.g., through memorization), so content alone cannot reveal the true provenance label.For this reason, content produced by a generator must admit a faithful reconstruction by that same generator, and finding such a reconstruction makes synthetic provenance plausible and authenticity plausibly deniable.We therefore propose and evaluate a detection paradigm that outputs a calibrated prediction of whether authenticity is plausibly deniable: a faithful reconstruction by any known generator establishes plausible deniability, while calibration bounds how often content from known generators fails to be reproduced. Our evaluation shows that (i) our detector can be calibrated so that at most 1% of generated content is wrongly certified, an operating point at which most baseline detectors reach near-zero recall, including the strongest with 93% accuracy; (ii) calibrating a stricter security threshold on attacked samples preserves this bound against adaptive adversaries within the evaluated bounded-perturbation attack space, whose perturbations break every baseline, but does not cover arbitrary adversarial transformations; and (iii) post-hoc verifiability is eroding, as 1,116 of 3,000 Reddit images resist reproduction by a 2022 generator, but only 55 to 79 resist reproduction by 2024 generators.
|
| 158 |
fMRI-TAMCL: Text-Anchored Supervised Multimodal Contrastive Learning for fMRI-Based Brain Disorder Classification
2610.05880
|
cs.CV
|
Juliana Mantebea Danso, Enoch Opanin Gyamfi, Mylene C. Q. Farias |
Resting-state fMRI is important in the classification of brain disorders, but highly multimodal and exhibits strong multisite heterogeneity. Existing methods fuse images, BOLD-based functional connectivity, and phenotypic data modalities. Unlike other medical ...Resting-state fMRI is important in the classification of brain disorders, but highly multimodal and exhibits strong multisite heterogeneity. Existing methods fuse images, BOLD-based functional connectivity, and phenotypic data modalities. Unlike other medical imaging datasets, rs-fMRI datasets rarely include a text modality, so they are generated from phenotypic data or BOLD activations. These text generation methods rely on fixed assumptions for subjects, sites, devices, and protocols, leading to poor generalization across datasets. We propose fMRI-TAMCL, a text-anchored multimodal contrastive learning framework that integrates fMRI images, sparse FC, and generated subject-specific text. Its Subject-Adaptive Threshold Derivation module generates BOLD activation text, while Feature-Value Serialization module generates phenotypic text. All three modalities are encoded as clustered graphs, projected onto a shared unit hypersphere space, aligned using pairwise, text-anchored supervised contrastive learning, and fused with attention. fMRI-TAMCL proves its generalization capability across five datasets outperforming 29 baselines with 78.6%-86.4% accuracy in downstream classification.
|
| 159 |
Spatial Supervision Without Attribution Optimization: Improving Post-Hoc Class Activation Maps via Box-Guided Evidence Routing
2610.05891
|
cs.CV
|
Wenhao Liang, Liangwei Nathan Zheng, Lin Yue, Wei Emma Zhang, Mingyu Guo |
Post-hoc class activation maps (CAMs) are a standard tool for inspecting the evidence behind an image classifier's predictions, yet nothing in ordinary training encourages these maps to be spatially appropriate. We study whether inexpensive spatial supervision...Post-hoc class activation maps (CAMs) are a standard tool for inspecting the evidence behind an image classifier's predictions, yet nothing in ordinary training encourages these maps to be spatially appropriate. We study whether inexpensive spatial supervision can improve a classifier's own predicted-class Grad-CAM without ever optimizing an attribution map. Box-Guided Evidence Routing (BGER) trains a lightweight gate on the final feature map under box or mask supervision and routes classification through the gated features, while Grad-CAM is computed separately at the pre-gate representation, so the evaluated map never enters the training objective. With a BCE routing loss, BGER raises MaxBoxAccV2 from $0.584$ to $0.715$ on CUB-200-2011 and from $0.757$ to $0.832$ on Stanford Dogs at comparable accuracy. Matched controls attribute most of the ResNet-50 gain to the spatial supervision reshaping the backbone rather than to routing itself: when classification bypasses the gate, most of the improvement remains, and detaching gradients through the gate leaves the ResNet-50 result nearly unchanged. The same detachment preserves most of the gain in two DenseNet-121 chest X-ray settings but removes the apparent gain on Swin-T, and directly supervising the CAM reaches stronger localization at a larger accuracy cost. Overall, spatial supervision can improve separately evaluated post-hoc CAMs, but both the mechanism and the size of the benefit depend on the architecture and the evaluation setting.
|
| 160 |
TasteRoute: Personalized Routing for Video Generation
2610.05896
|
cs.CVcs.CLcs.AI
|
Zhi Rui Tam, Chao-Chung Wu, Sin-Han Yang, Peyton Ku, Brendan Kuang |
Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus...Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annotators is used as an oracle, it agrees with each annotator's own favorite only 34-55% of the time. Motivated by this observation, we introduce TasteRoute, a personalized video-generation router that selects a generator jointly based on the input request, user preferences, and available generation budget. Across text-to-video and image-to-video settings, TasteRoute is competitive with strong simple baselines on preference routing while reducing average generation cost. The cost saving increases under higher budget caps. Finally, we release TasteRoute-3k, a human-annotated dataset containing multi-model video comparisons, quality judgments, preference rankings, and user-profile signals to facilitate future research on personalized and cost-aware video routing.
|
| 161 |
Fitting Vision Adapters at Frontier Scales
2610.05897
|
cs.CVcs.LG
|
Jaehoon Lee, Harry Partridge, Mudith Jayasekara, Charles O'Neill, Max Kirkby |
Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while kee...Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M parameter projector from the vision encoder of Kimi K2.6 to GLM 5.2 and 5.3, both models without native vision capabilities, and further present a reproducible recipe for training these adapters at scale. We study the following: (a) how vision capabilities of multimodal models scale as purely the language model side scales, and (b) what specific vision capabilities are able to be imbued into a pure language model at scale, and which ones remain limited. We evaluate on MMMU-Pro and BLINK, examining both overall performance and results on individual visual tasks.
|
| 162 |
LoDEOT: Low-Dimensional and Efficient Offset Tokens for Building Footprint Extraction from Off-Nadir Imagery
2610.05899
|
cs.CV
|
Kai Li, Zigan Zhou, Zhenyang Li, Hui Shan, Zhe Chen |
Instance-level roof-to-footprint offset (RFO) prediction is central to extracting building footprints from off-nadir imagery. Query-based pipelines commonly use high-dimensional instance tokens to predict signed two-dimensional RFOs. We investigate whether RFO...Instance-level roof-to-footprint offset (RFO) prediction is central to extracting building footprints from off-nadir imagery. Query-based pipelines commonly use high-dimensional instance tokens to predict signed two-dimensional RFOs. We investigate whether RFO prediction can instead use a compact offset token. Under local pinhole projection and vertical-extrusion assumptions, the idealized RFO map admits a five-parameter sufficient descriptor comprising intrinsic shape, composite amplitude, and relative geometry. This factorization provides a structural prior for a five-dimensional offset token, whose channels learn task-relevant latent representations through end-to-end training. Based on this design, we propose LoDEOT, which retains high-dimensional instance tokens for detection and segmentation but maps instance-token, concentration-gated roof, and box-mask evidence to a five-dimensional offset token followed by an independent two-dimensional readout. Known denoising-query target indices further align each supervised decoder-layer estimate with the same clean instance RFO, organizing successive predictions as target-aligned recovery under perturbed query conditions. Experiments on five real-world building datasets demonstrate the effectiveness of LoDEOT for building footprint extraction. Experiments on real-world building datasets demonstrate that a five-dimensional offset token can support accurate RFO prediction. On BONAI, LoDEOT achieves the best roof-detection bAP and bAP50 and leads all five offset-corrected footprint metrics among the evaluated end-to-end methods, with FAP50 of 54.58 and mEPE of 5.23 pixels. Its FAP50 exceeds those of the evaluated end-to-end baselines by 7.56-16.85 percentage points.
|
| 163 |
End-to-End Autonomous Recursive Arborescence Deformable Flow and Non-Linear Hemodynamics for Patient-Specific Coronary Centerline Extraction
2610.05900
|
cs.CV
|
Zeyu Jia, Xin Ming |
Extracting patient-specific vascular trees from volumetric medical images is fundamental to computational angiography and non-invasive hemodynamic assessment. Conventional voxel segmentation models often sever delicate bifurcations, while heuristic Euclidean M...Extracting patient-specific vascular trees from volumetric medical images is fundamental to computational angiography and non-invasive hemodynamic assessment. Conventional voxel segmentation models often sever delicate bifurcations, while heuristic Euclidean Minimum Spanning Trees introduce non-anatomical shortcuts. Moreover, linear Poiseuille flow neglects quadratic kinetic dissipation across arterial narrowings, underestimating ischemia. We formulate an end-to-end framework decoupling continuous geometric arborescence generation from non-linear hemodynamics. First, an autonomous 3D Ostium Landmark Localization Head with dual-sinus query channels and spherical-gated refinement eliminates centerline seeding dependency, achieving cohort mean localization error of 7.63 mm (7.43 mm LCA, 7.83 mm RCA; 71.4% <= 8.0 mm) from raw contrast context. Second, a Spatially-Grounded Deformable Step Flow Architecture queries continuous 3D feature pyramids via trilinear sampling, sequentially generating trajectories with anchor boundary enforcement (X(0) = P_start). Third, a Top-Down Recursive Arborescence State Machine detects bifurcation peaks via Tree-NMS and parameterizes predecessor parent pointers (p_k < k), guaranteeing single connected acyclic tree topology (beta_0 = 1, beta_1 = 0) with differentiable step termination. Fourth, an iterative Picard non-linear Kirchhoff solver with Young-Tsai / Gould quadratic dissipation enforces machine-precision mass conservation (residual 5.82e-11 mL/s). Across 14 development patients under standardized in-silico stenosis stress testing (Q_0 = 4.0 mL/s), linear Poiseuille flow misclassifies 75% diameter lesions as non-ischemic (FFR > 0.80) in 14/14 cases, whereas our non-linear solver captures functional ischemia (FFR = 0.5864, lesion disparity 32.89 mmHg, p = 6.10e-5) with 3.66x collateral shunting. Test set firewall isolation was maintained.
|
| 164 |
Safe Image Generation via Reinforcement Learning
2610.05908
|
cs.CV
|
Eungyeol Han, Jong-Seok Lee |
Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-genera...Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-generation filtering (e.g. prompt-level text classifiers) or post-hoc moderation applied after an image is completely synthesized. However, adversarial attack methods operate over a much broader space. This imbalance highlights the need for a safety mechanism that intervenes during the generation process. We propose an in-generation safety framework that monitors the denoising trajectory and detects emerging NSFW signals from intermediate representations. Rather than merely detecting NSFW generations, our method applies reinforcement learning to generate safe images from NSFW prompts. By coupling in-generation detection with controllable steering, our approach mitigates unsafe trajectories even when NSFW signals emerge after generation has already begun. Experiments results show that our method consistently outperforms existing safe image generation methods across both standard and adversarial evaluation sets, while preserving perceptual quality and prompt fidelity. Code will be released upon acceptance.
|
| 165 |
AstraSR: Real-World Thermal Super-Resolution with GPT-6 Astra
2610.05910
|
cs.CV
|
Mengyuan Li, Changhong Fu, Jun Zhang, Ziyu Lu, Yuhang Zhang |
Real-world thermal super-resolution (SR) is constrained by limited sensor resolution and the difficulty of obtaining corresponding high-resolution (HR) observations for direct model supervision. Conventional SR methods typically construct training pairs by tre...Real-world thermal super-resolution (SR) is constrained by limited sensor resolution and the difficulty of obtaining corresponding high-resolution (HR) observations for direct model supervision. Conventional SR methods typically construct training pairs by treating captured thermal images with real-world degradations as HR references and applying predefined degradation to generate synthetic low-resolution (LR) inputs. Such a construction not only introduces a domain gap between synthetic and captured LR observations but also retains acquisition degradations in the supervision. To address this issue, we propose AstraSR, a real-world thermal SR method guided by GPT-6 Astra, a frontier multimodal generative model endowed with emergent and transformative visual capabilities. Specifically, we construct a dataset of image pairs by using captured LR thermal images to condition GPT-based HR reference. We develop a direct generative supervision strategy that learns from captured thermal inputs paired with GPT-generated HR references. Pixel, gradient, and perceptual losses jointly supervise the transfer of intensity patterns, structural boundaries, and visual details from the generated references. Qualitative comparisons with seven existing state-of-the-art real-world SR methods show continuous object contours, distinct structural boundaries, and smooth intensity transitions in the thermal scenes. These results demonstrate that AstraSR outperforms existing real-world SR methods in both thermal clarity and structural coherence.
|
| 166 |
Every View Counts: View-Consistent Panoptic Quality for Multi-view Panoptic Segmentation
2610.05911
|
cs.CV
|
Youngmin Lee, Byungha Ko, Guhnoo Yun, Dong Hwan Kim |
Multi-view panoptic segmentation assigns a semantic class and a scene-level instance ID to every pixel of an unordered set of images, and recent feed-forward 3D models predict these labels for the input views in a single forward pass. Their predictions, howeve...Multi-view panoptic segmentation assigns a semantic class and a scene-level instance ID to every pixel of an unordered set of images, and recent feed-forward 3D models predict these labels for the input views in a single forward pass. Their predictions, however, have been evaluated with the scene-level PQ (PQ^scene) borrowed from per-scene optimization methods, typically on rendered held-out views. PQ^scene tiles all views of a scene into a single image, so that a missed appearance or a change of ID lowers the score of the matched pair only in proportion to its area. We propose View-Consistent Panoptic Quality (VC-PQ), which extends PQ from a single image to a set of input views, counts equally every view in which an instance is visible, and penalizes a prediction that is not visible in the same views as its ground truth. A decomposition of VC-PQ attributes the score a method loses to mask accuracy, view consistency, and the matching threshold. A single additional parameter recovers the area weighting of tiling for comparison. Under a fixed evaluation protocol on ScanNet++ and ScanNetv2, recent feed-forward methods are evaluated with VC-PQ and PQ^scene, and the decomposition shows where each of them loses its score. Controlled perturbations of the ground truth show that VC-PQ responds to the number of views in which an instance is missed or changes ID, whereas PQ^scene responds to their area. The aim of this work is to make view consistency part of the evaluation of multi-view panoptic segmentation, with VC-PQ reported alongside PQ^scene.
|
| 167 |
Prompt and Refinement: Asymmetric Mutual Learning for Infrared Small Target Detection with Noisy Labels
2610.05918
|
cs.CV
|
Yimin Fu, Songbo Wang, Lizhuo Liu, Baicheng Pan, Zhunga Liu |
Existing data-driven infrared small target detection (ISTD) methods typically require large-scale datasets with accurate pixel-level annotations for model training. However, such labor-intensive requirements are difficult to satisfy in real-world applications ...Existing data-driven infrared small target detection (ISTD) methods typically require large-scale datasets with accurate pixel-level annotations for model training. However, such labor-intensive requirements are difficult to satisfy in real-world applications due to the heavy reliance on expert knowledge and the inherently weak distinctiveness of infrared small targets. Consequently, the presence of noisy labels during model training is inevitable, which can severely mislead the learning of target perception toward spurious patterns. To address this challenge, we propose Prompt and Refinement (PAR), a label-noise-robust asymmetric mutual learning paradigm for ISTD. Specifically, PAR comprises a pretrained Segment Anything Model (SAM) and an ISTD-specific detector trained from scratch, which learn collaboratively through a peer-teaching scheme. Coupled with local contrast regularity, the predictions of the two asymmetric peer models are mutually exploited as rectification cues for the supervisory masks of their counterparts. The interaction between complementary inductive biases effectively prevents the label correction process from degenerating into the self-confirmation loop of a single model, enabling progressive refinement of the annotations toward intrinsic target characteristics. In addition, the detector predictions are utilized as corrective mask prompts to facilitate task-specific adaptation of the vision foundation model. Moreover, an evidential uncertainty estimation strategy is introduced into the optimization process to further alleviate the adverse effects of noisy labels. Extensive experiments under diverse noisy label scenarios on three ISTD datasets demonstrate that PAR consistently achieves state-of-the-art performance.
|
| 168 |
UltraDub: Towards Authentic Dubbing by Unifying Visually-Steered Flow Learning and Trajectory Guidance
2610.05932
|
cs.CVcs.AIcs.SDcs.MM
|
Gaoxiang Cong, Liang Li, Jianwei Wen, Zhedong Zhang, Zheng-Jun Zha |
Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can impr...Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.
|
| 169 |
ReMem: Streaming Video Understanding With Long Context Retention
2610.05940
|
cs.CVcs.AI
|
Li Yiheng, He Xu, Wang Shaobo, Shao Ling, Lu Shijian |
Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several s...Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies in an attempt to adapt offline VLMs for streaming video understanding tasks. However, through our probing experiment, we identify that most existing works tend to progressively lose long context information as length of input stream increases. To address this, we propose ReMem, a novel training-free adaptation technique that enables VLMs to process streaming videos of arbitrary lengths while improving their long context information retention capability. ReMem exploits memory from two perspectives, implemented as two core components. The Streaming Context Memory (SCM) continuously compresses historical context with query-independent attention. The Retrieved Vision Memory (RVM) then retrieves the most salient, query-relevant context from memory to augment the VLM's input. Comprehensive experiments demonstrate that the proposed ReMem achieves state-of-the-art (SOTA) performance across a variety of widely used benchmarks, spanning both streaming video and general long video understanding tasks.
|
| 170 |
JLD: Perceptual Distance Through A Jacobian Lens
2610.05967
|
cs.CVcs.AI
|
Shreshth Saini, Balu Adsumilli, Alan C. Bovik |
Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fix...Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fixed data and resolution. For example, when image resolution is doubled, the correlation of DISTS with human scores on TID2013 drops from 0.815 to 0.717. We introduce the Jacobian Lens Distance (JLD), which derives its perceptual geometry from a frozen vision encoder rather than from human labels. JLD combines the locality of early patch features with the perceptual sensitivity captured by later encoder representations. Specifically, we use the encoder Jacobian to identify directions in the early feature space that most strongly affect the encoder output, producing a fixed metric tensor, $E[J^\top J]$, which we call the Jacobian lens. The lens is fitted only once from 100 unlabeled images, taking about 35 seconds. Locally, this construction defines a pullback metric in pixel space, giving JLD a clear geometric interpretation that can be directly analyzed on real images. Across four standard perceptual databases, JLD achieves state-of-the-art performance and consistently outperforms LPIPS, DISTS, PieAPP, and DreamSim. JLD is also robust to changes in image resolution, on TID2013, its lens-term correlation remains nearly unchanged when the resolution is doubled, decreasing only from 0.850 to 0.845. We further introduce JLD-fast, which is $4\times$ faster than LPIPS-VGG while achieving a mean correlation of 0.911. Finally, JLD naturally extends to video, reaching a correlation of 0.786 on Waterloo IVC 4K compared with 0.611 for VMAF.
|
| 171 |
Label-Free Coreset Selection with Foundation Models for Efficient Annotation in Computational Pathology
2610.05987
|
cs.CV
|
Tuo Yin, Jennifer Dhont |
Computational pathology has the potential to improve clinical outcomes through a demonstrated increase in diagnostic and prognostic accuracy. However, the development and validation of deep learning algorithms still require annotated data, a costly procedure i...Computational pathology has the potential to improve clinical outcomes through a demonstrated increase in diagnostic and prognostic accuracy. However, the development and validation of deep learning algorithms still require annotated data, a costly procedure involving expert pathologists who already face critical workforce shortages. Existing coreset selection methods to optimize annotation efforts currently all rely on hyperparameters tuned on natural-image benchmarks that do not transfer to histopathology and are cumbersome to use in clinical practice. In this study, we present GCcore, a novel label-free coreset selection method that embeds every image of a dataset with any pathology foundation model and greedily selects the samples that collectively maximize the global coverage of the embedding space. The proposed method provides a lower-bound guarantee on the global coverage of the returned coreset for any coreset size, while being completely hyperparameter-free and deterministic. We demonstrate GCcore's superior performance over 14 baselines including state-of-the-art methods across 10 tasks and datasets spanning whole slide image classification, tile classification, and tissue segmentation, where it ranks first on six and within the top three on nine, while also demonstrating how existing methods can shift by up to five rank positions depending on their hyperparameter settings. Code is publicly available at https://github.com/OncoAI-ULBHUB/GCcore.
|
| 172 |
From Transformation to Target State: Rethinking Query Representation for Zero-Shot Composed Image Retrieval
2610.05993
|
cs.CV
|
Yihe Zhao, Songhe Feng |
Composed image retrieval (CIR) aims to retrieve a desired target image from a query consisting of a reference image and a modification text. This task exhibits an unusual representational asymmetry: the modification text specifies a transition from the referen...Composed image retrieval (CIR) aims to retrieve a desired target image from a query consisting of a reference image and a modification text. This task exhibits an unusual representational asymmetry: the modification text specifies a transition from the reference state, whereas retrieval candidates depict completed target states. This creates a representation mismatch for zero-shot methods that query pretrained vision-language spaces directly with transformation-oriented language. We study this mismatch and reformulate zero-shot composed image retrieval as target-state reconstruction followed by retrieval. We instantiate this formulation with ASAP-CIR, a training-free framework that reconstructs a static target representation using a frozen multimodal large language model (MLLM). The representation combines multiple holistic descriptions with a variable set of importance-weighted atomic semantics, thereby preserving both overall target identity and fine-grained visual constraints. Retrieval then integrates holistic state alignment, atomic constraint grounding, and calibrated target-state evidence aggregation. A controlled text-only diagnostic shows that target-side static query formulations achieve more reliable retrieval than dynamic composed query formulations, particularly when source-state semantics must be suppressed or transformed. Experiments on FashionIQ, CIRR, and CIRCO further characterize the effectiveness and limitations of this representation principle, with the clearest gains on the multi-target CIRCO benchmark. These results show that how composed intent is represented before retrieval is a consequential design choice, distinct from the choice of retrieval backbone itself.
|
| 173 |
Investigating Query-Insensitive Behavior in Spatio-Temporal Video Grounding
2610.06018
|
cs.CVcs.CLcs.AI
|
Eryk Ko{\l}odziejczyk, Alberto Presta, Karol Szurkowski, Michal Byra |
Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video...Spatio-temporal video grounding (STVG) aims to localize objects or events described by natural language queries in both space and time. Existing STVG models are typically trained and evaluated under the assumption that each query is relevant to the input video. In this work, we challenge this assumption by studying the behavior of state-of-the-art STVG models under irrelevant queries and missing textual input. Our experiments show that current models can still produce plausible spatio-temporal predictions even when the query is unrelated to the video or removed entirely. We further analyze HCSTVG-v2 and VidSTG to identify dataset regularities that may encourage such query-insensitive behavior. Our study highlights an underexplored limitation of STVG models and motivates negative-aware evaluation protocols and architectures that explicitly assess query relevance.
|
| 174 |
Patch-based Querying Identifies Structures of Interest in Electron Microscopy
2610.06020
|
cs.CV
|
Niels Vyncke, Nicolas Nadisic, Yvan Saeys, Aleksandra Pi\v{z}urica |
Volume electron microscopy (vEM) has emerged as an essential sensing technique in biomedical research, allowing the three-dimensional imaging of biological cells and tissues at nanometer-scale resolution. The ability to generate extensive datasets has reached ...Volume electron microscopy (vEM) has emerged as an essential sensing technique in biomedical research, allowing the three-dimensional imaging of biological cells and tissues at nanometer-scale resolution. The ability to generate extensive datasets has reached the limitations of downstream analysis processes, which depend significantly on the intervention of human experts for preprocessing and annotation. We propose an efficient and reliable patch-based retrieval framework based on self-supervised learning of local image descriptors to locate self-similar structures in vEM datasets. Given a few manual annotations of a given cellular structure, our method can retrieve similar structures across the EM volume. Our framework is interactive, allowing the human expert to refine the search queries and retrieve relevant image patches quickly and using little labeled data. Experiments on real-world vEM images of biological tissues demonstrate that our framework can reliably identify relevant cellular structures, generalize across different organelles and acquisition modalities, and substantially reduce the search space for downstream analysis.
|
| 175 |
Scalable Minimal-Change Learning for Controllable Image Editing
2610.06021
|
cs.CVcs.AI
|
Shuo Chen, Fengming Huang, Yu Yao, Mingming Gong, Tongliang Liu |
Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. ...Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality in modern nonlinear generators and often requires supervision unavailable at scale. We instead optimize edit outcomes with reinforcement learning. An agentic vision-language reward model audits each source image, instruction, and edited image for two failure types: unimplemented requested changes and unintended changes. A group-level rubric merges and verifies these issues to provide consistent rewards across candidate edits without per-instruction human annotations. On FLUX.1 Kontext-dev, ARRO raises average EditScore from 5.21 to 5.88 across MinEval, MagicBrush, AnyBench, and Emu-Edit. On 600 evaluation examples, it reduces off-target pixel change by 8.4% relative to the base editor. Reward and SFT controls, blinded human evaluations, and transfer to OmniGen2 provide complementary evidence. Code: https://github.com/Showwwwwwwww/ARRO
|
| 176 |
Casual Flash Lighting for Gaussian Splat Inverse Rendering
2610.06035
|
cs.CV
|
Jiamin Xu, Dongheng Wei, Jiarong Zhao, Qi Wang, James Tompkin |
Recovering geometry, materials, and lighting from photographs is highly ambiguous when only static illumination is available. Active-lighting setups reduce the ambiguity but require dark rooms or specialized hardware. Instead, we synergize both static and flas...Recovering geometry, materials, and lighting from photographs is highly ambiguous when only static illumination is available. Active-lighting setups reduce the ambiguity but require dark rooms or specialized hardware. Instead, we synergize both static and flash lighting from casual indoor capture, with the flash on or off, each from independent viewpoints. The flash residual constrains albedo and the BRDF, while static lighting captures grazing-angle specular highlights that flash misses. With a 2DGS reconstruction framing, our key contribution is a GS-anchored diffuse field: a hash-encoded MLP is queried at the rasterized 2DGS depth. As it depends only on world position, it is view consistent in 3D and allows the flash residual to drive material decomposition instead of being absorbed by alpha-blending drift across views. At the same time, we render static lighting with deferred shading such that it can also supervise material decomposition. On five synthetic and three real indoor scenes, our method outperforms six recent baselines on diffuse color, albedo and roughness material parameters, and in relighting where PSNR improves by 4.17 dB over the next-best baseline.
|
| 177 |
Representation Disentanglement for Fair Chest X-Ray Diagnosis
2610.06041
|
cs.CVcs.AI
|
Yujie Sun, Ruizhe Li, Xiaowu Sun |
Deep learning has advanced chest X-ray (CXR) diagnosis, yet demographic biases in learned representations may contribute to performance disparities across intersectional groups. We propose a single-encoder framework combining dual-level decorrelation with prot...Deep learning has advanced chest X-ray (CXR) diagnosis, yet demographic biases in learned representations may contribute to performance disparities across intersectional groups. We propose a single-encoder framework combining dual-level decorrelation with prototype-guided cross-group contrastive learning to reduce demographic dependence while accounting for within-class variation. We further propose Demographic Representation Alignment Reduction (DRAR), a new metric that quantifies the reduction in demographic structure within disease representations. The framework is evaluated on four classification tasks using 34,809 CheXpert test images across eight intersectional groups, defined by age, sex and ethnicity. Compared with empirical risk minimization (ERM), our method reduces the mean equalized-odds gap from 15.41\% to 10.86\% and the AUC gap from 5.95\% to 5.01\%. Our method achieves a DRAR of 59.04\% relative to ERM, with only a slight decrease in mean AUC. These results demonstrate that representation disentanglement can reduce demographic bias and improve intersectional fairness. Code is available at \url{https://github.com/06Yujie/Fair-Medical-Imaging}.
|
| 178 |
Local2Mesh: Spatially Localized Contour-to-Mesh for Left Ventricular Reconstruction from Sparse 2D Cardiac MRI
2610.06052
|
cs.CVcs.AI
|
Haoyu Wu, Ling Lin, Pascal Lef\`evre, Ruizhe Li, Xiaowu Sun |
Three-dimensional (3D) left ventricular (LV) reconstruction from sparse cardiac magnetic resonance (CMR) imaging remains challenging due to inter-slice misalignment and insufficient local spatial information between slices. Global aggregation of contour featur...Three-dimensional (3D) left ventricular (LV) reconstruction from sparse cardiac magnetic resonance (CMR) imaging remains challenging due to inter-slice misalignment and insufficient local spatial information between slices. Global aggregation of contour features may obscure local contour-to-surface relationships. We propose Local2Mesh, a spatially localized contour-to-mesh framework that deforms a template mesh to reconstruct 3D LV geometry from sparse 2D contours without 3D mesh annotations. The framework introduces geometry-aware alignment to correct inter-slice misalignment and a plane-aware Local Router that routes contour features to template vertices using vertex-to-plane distances. Local and global contour features then jointly guide graph-based template deformation for 3D LV reconstruction. Experiments on two public datasets, M\&Ms-2 and ACDC, demonstrate superior geometric reconstruction and functional estimation over existing methods. Zero-shot transfer from M\&Ms-2 to ACDC demonstrates strong cross-dataset generalization. Reconstructed meshes also improve disease classification over sparse contours, supporting their utility for downstream cardiac analysis. These results demonstrate that combining geometry-aware alignment with local contour-to-vertex modeling improves LV reconstruction from sparse 2D contours and supports downstream cardiac analysis. The code is available at \url{https://github.com/hwu918945-alt/loca2mesh}.
|
| 179 |
ROT: Rotating Hidden States towards Contextual Vectors for Hallucination Mitigation in LVLMs
2610.06056
|
cs.CVcs.CLcs.AI
|
Yijing Du, Xiangcheng Zhan, Shuo Yang |
Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shi...Large Vision-Language Models (LVLMs) frequently suffer from object hallucination. Existing training-free interventions primarily manipulate attention weights, which indirectly affect the deep semantics reaching the final predictive layers. In this work, we shift our focus to the hidden state vectors extracted after self-attention and residual addition. Empirical analysis reveals that hallucinated tokens do not simply over-rely on linguistic priors; instead, they exhibit an anomalous contextual deviation, showing significantly lower similarities to both textual and visual contexts in intermediate layers. Motivated by this, we propose ROT, a layer-specific, training-free framework. ROT dynamically detects semantic deviation in the middle layers and applies a norm-preserving rotation to steer the hidden states back toward the local multimodal context plane spanned by the contexts. For subsequent layers, a representational smoothing mechanism is introduced to stabilize the calibrated trajectory. Extensive experiments on multiple benchmarks demonstrate that ROT consistently reduces hallucinations across various model architectures and scales, offering an efficient, geometry-driven solution for grounded generation.
|
| 180 |
Vision Transformer Ensembles for Panoramic Street Segmentation
2610.06063
|
cs.CVcs.LGcs.AI
|
Yunus Serhat B{\i}\c{c}ak\c{c}{\i} |
Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity ch...Semantic segmentation of street panoramas can support detailed descriptions of urban environments, yet small datasets and unequal training costs make model selection difficult. This paper presents the system used for a first place submission to the PalmCity challenge in the leaderboard snapshot dated 5 October 2026. Nine pretrained segmentation systems are compared using approximately equal computation budgets. The candidates include DeepLabV3+, SegFormer, UPerNet, Mask2Former, DINOv3 with a linear decoder, and an Encoder only Mask Transformer using DINOv3. The two leading candidates are trained independently with three random seeds and longer budgets. Equal averaging of class probabilities from the three Encoder only Mask Transformer models, evaluated at three image scales with horizontal reflection, produces 60.95% mean intersection over union and 71.16% mean F1 on the 84 image public validation split. The submitted predictions receive 57.08% mean intersection over union and 67.96% mean F1 on the hidden test leaderboard. Producing all 249 test masks takes 251.49 seconds including model initialization and provenance checks on one NVIDIA RTX 5090. Peak allocated GPU memory is 2.70 GiB. The study reports all eligible models, all inference variants, class level errors, source conditions, and reproducibility checks, providing a documented challenge workflow with existing architectures.
|
| 181 |
Anatomy-preserving unpaired cone-beam CT refinement for image-guided radiotherapy using pseudo-label guided diffusion
2610.06094
|
cs.CV
|
Qi Lai, Yutong He |
Cone-beam computed tomography (CBCT) is widely used in image-guided radiotherapy, but scatter, beam hardening, noise, truncation, and other artifacts limit image quality and CT number accuracy. Paired CBCT and CT data are difficult to obtain clinically because...Cone-beam computed tomography (CBCT) is widely used in image-guided radiotherapy, but scatter, beam hardening, noise, truncation, and other artifacts limit image quality and CT number accuracy. Paired CBCT and CT data are difficult to obtain clinically because of motion, anatomical changes, and acquisition mismatch. We present RefineCBCT, an unpaired CBCT refinement framework that uses pseudo-label guidance and short-step diffusion to reduce artifacts while preserving patient-specific anatomy. RefineCBCT was trained and evaluated on unpaired CBCT and planning CT data from public LUNG TCIA and PELVIC TCIA datasets and compared with representative GAN and diffusion based methods. On LUNG TCIA, it achieved the best results across all metrics, with MAE 19.411, RMSE 62.758, PSNR 30.845 dB, and SSIM 0.931. On PELVIC TCIA, it achieved the best MAE, PSNR, and SSIM, with values of 14.905, 36.671 dB, and 0.876. The refined images showed fewer streaking and shading artifacts, clearer anatomical boundaries, and improved soft tissue uniformity, with line profile and ROI analyses showing closer agreement with planning CT. These results suggest that RefineCBCT provides efficient and effective CBCT refinement under clinically realistic unpaired training conditions and may support more reliable CBCT use in image-guided radiotherapy workflows. Code is publicly available on GitHub, and the evaluated datasets are available from The Cancer Imaging Archive.
|
| 182 |
Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation
2610.06102
|
cs.CV
|
Fernando Alonso-Fernandez, Kevin Hernandez-Diaz, Jose Maria Buades, Josef Bigun |
We investigate CLIP for zero-shot gender estimation from full-face and periocular images. Three CLIP backbones are evaluated on 11,299 frontal images from Adience using image-text similarity with male/female prompts, achieving 95.54% full-face accuracy without...We investigate CLIP for zero-shot gender estimation from full-face and periocular images. Three CLIP backbones are evaluated on 11,299 frontal images from Adience using image-text similarity with male/female prompts, achieving 95.54% full-face accuracy without task-specific training. For periocular, zero-shot predictions are strongly biased towards males, primarily due to a misaligned decision boundary. Threshold alignment substantially reduces this bias, reaching 85.29% accuracy. Linear SVMs trained on CLIP features provide only marginal gains, with a best periocular accuracy of 86.17%, approximately 2.8% above previous Adience results in the literature. Nevertheless, the gap with full-face performance confirms the greater difficulty of periocular gender estimation
|
| 183 |
AnchorGen: Anchored Optimization for Customizable Generative 3D Design
2610.06135
|
cs.CV
|
Hantao Zhang, Oliver Heinimann, Jieke Wu, Yingxuan You, Emilien Seiler |
Engineering design often starts from a 2D sketch that fixes style and proportions, yet the subsequent 3D shape optimization relies on learned generative priors to keep the geometry valid. However, these priors are agnostic to the sketch: while they admit a val...Engineering design often starts from a 2D sketch that fixes style and proportions, yet the subsequent 3D shape optimization relies on learned generative priors to keep the geometry valid. However, these priors are agnostic to the sketch: while they admit a valid design by correcting a drifted proposal back to its training distribution, they often correct it towards the high-density region, ignoring the specified design. We introduce \emph{AnchorGen}, a rectified-flow framework trained unconditionally on the concatenated shape and sketch latents of paired data. The learned manifold represents the joint distribution of shape-sketch pairs, so constraining the sketch component restricts the iterate to the sub-manifold of shapes consistent with a target style. Since training employs no conditioning signal, the constraint is imposed at inference: gradient descent optimizes the shape latent to minimize a differentiable drag surrogate, while constraining the sketch latent to remain close to the target sketch via a token-wise cosine penalty. A single model thereby supports design-preserving optimization, dimensionally explicit design edits, and sketch-only synthesis.
|
| 184 |
Efficient Test-time Adaptation through Candidate Verification and Divergence Shifts
2610.06147
|
cs.CVcs.CL
|
Seungmin Oh, Seunghun Kang, Jongbin Ryu |
Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation ...Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes, caches, priors, or feature statistics, often incurring additional computational overhead. In this paper, we take a different perspective and reframe VLM-TTA as candidate verification rather than prediction adjustment. We propose Test-Time Correction (TTC), a hypothesis-based correction framework guided by a simple principle: hypothesize, reconstruct, correct. Given a test feature and its top-k candidate labels, TTC treats each candidate label as a hypothesis, reconstructs the feature within the corresponding latent subspace stored in a memory bank, and measures the resulting divergence shift. This shift quantifies how much the candidate subspace and its relations to other candidates change after the hypothetical insertion of the test feature. A correct candidate hypothesis induces only a small shift, whereas an incorrect one perturbs the subspace more strongly. TTC therefore corrects the prediction by selecting the candidate with the minimum aggregated divergence shift. This training-free candidate-verification mechanism avoids iterative optimization and provides a favorable accuracy-efficiency trade-off. Across five TTA settings and 15 benchmark datasets, including zero-shot classification, domain generalization, few-shot classification, base-to-novel generalization, and cross-dataset evaluation, TTC consistently improves accuracy over state-of-the-art VLM-TTA methods while achieving up to 2x speedup, over 3x lower CPU memory usage, and up to 1.4x lower GPU memory usage than the lowest-memory training-free baseline.
|
| 185 |
Bayesian Optimization in Sequence-to-Architecture Latent Space for Zero-Shot NAS
2610.06167
|
cs.CV
|
Ondrej Tybl, Lukas Neumann |
Zero-shot Neural Architecture Search removes the prohibitive cost of traditional NAS, but its search process is typically based on the evolutionary algorithm (EA); lacking an explicit model of the objective, it often resorts to a near-random search through mut...Zero-shot Neural Architecture Search removes the prohibitive cost of traditional NAS, but its search process is typically based on the evolutionary algorithm (EA); lacking an explicit model of the objective, it often resorts to a near-random search through mutation. Bayesian Optimization offers a principled alternative by modeling the objective and aggregating information across iterations, but scales poorly to the high-dimensional, discrete, graph-structured spaces of modern NAS, restricting its use to only small networks. In this paper, we bring Bayesian Optimization to zero-shot NAS for large-scale architectures by learning a latent space via a Variational Autoencoder trained to reconstruct a novel prefix encoding of architectures and propose a proxy scalarization that combines several zero-shot proxies into a single Bayesian Optimization objective. After only 10,000 iterations of the proposed search algorithm (8 hours on a single GPU), our method found a network architecture which under the given model parameter count constraints achieves state-of-the-art results on three separate tasks -- image classification, object detection and semantic segmentation.
|
| 186 |
EORestore-Agent: Fidelity-Guided Agentic Restoration of Remote Sensing Images with Composite Degradations
2610.06196
|
cs.CV
|
Heli Qi, Zeqi Zhou, Jingjun Yi, Kunyi Liu, Ziyang Lihe |
Remote sensing images often carry composite degradations, in which haze, cloud, noise, blur, low light, and low resolution coexist. Restoring them requires deciding which tool to apply, in what order, and when to stop, yet no clean reference is available at in...Remote sensing images often carry composite degradations, in which haze, cloud, noise, blur, low light, and low resolution coexist. Restoring them requires deciding which tool to apply, in what order, and when to stop, yet no clean reference is available at inference time to verify these decisions. All-in-one models trained on single degradations converge to a narrow PSNR band as degradations accumulate. To formulate real-world remote sensing restoration as a traceable trajectory, we present EORestore-Agent, which replaces this unmeasurable objective with reference-free, verifiable per-step decisions. A fine-tuned vision-language model reports all residual degradation types, whose tool pools are scored together, so the restoration order emerges from step-wise selection. A relative quality scorer, trained with full-reference supervision on synthetic degradation chains, predicts the changes in PSNR, SSIM, and LPIPS from the current image to each candidate. A step is accepted only when no predicted change is negative and the predicted PSNR gain is positive. Otherwise, the agent keeps the current image. On a synthetic Landsat-8 benchmark with six degradation types, EORestore-Agent improves PSNR by 2.3 to 3.2 dB over the strongest retrained all-in-one baseline on composites of two to six degradations, whereas zero-shot natural-image agents fall below the degraded input in PSNR in 17 of 18 settings. Replacing the learned scorer with no-reference quality differences costs 1.1 to 4.6 dB. The remaining harmful steps are small and cluster near the acceptance threshold. Sentinel-2 examples illustrate transfer to real atmospheric degradation without retraining.
|
| 187 |
MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting
2610.06210
|
cs.CV
|
Yiming Xu, Hao Cheng, Monika Sester |
Autoregressive generation is natural for language, where predicted tokens can be directly reused as the next prediction state, but trajectory forecasting lacks such a clean token: motion is continuous, multimodal, and expressed in local coordinate frames that ...Autoregressive generation is natural for language, where predicted tokens can be directly reused as the next prediction state, but trajectory forecasting lacks such a clean token: motion is continuous, multimodal, and expressed in local coordinate frames that evolve with the predicted trajectory. We present MoCAR (Motion-code Coordinate-aware AutoRegression), a decoder-only framework that casts trajectory forecasting as next-code prediction in a coordinate-aware continuous latent space. MoCAR learns a continuous motion-code space from endpoint-normalized trajectory segments, where each code jointly captures local trajectory geometry and the reference-frame transition induced by that segment. Historical motion codes are used as a teacher-forced prefix, future codes are generated autoregressively under temporal, map, agent, and mode interactions, and predicted codes persist in latent memory while decoded endpoints update the local scene context. This enables rollout without trajectory-space re-tokenization, trajectory queries, goal candidates, or proposal-and-refinement pipelines. On Argoverse (AV) benchmarks, MoCAR achieves top-tier performance with a simple single-stage architecture, transfers strongly from AV2 to AV1 in zero-shot evaluation, and improves on turn-heavy scenarios. Ablations confirm that the learned continuous motion-code space, latent alignment, weak KL regularization, and joint tokenizer-predictor optimization are essential for stable latent autoregression.
|
| 188 |
Frequency-Decoupled Diffusion Guidance for Non-Blind Image Deblurring
2610.06221
|
cs.CV
|
Sihan Wang, Jinshu Huang, Haibin Su, Yunhua Xue |
Pretrained diffusion models provide powerful image priors for training-free posterior sampling in image restoration. To guide this sampling process, frequency-aware methods progressively incorporate measurement information across frequency bands, facilitating ...Pretrained diffusion models provide powerful image priors for training-free posterior sampling in image restoration. To guide this sampling process, frequency-aware methods progressively incorporate measurement information across frequency bands, facilitating coarse-to-fine reconstruction. However, existing methods typically do not explicitly separate frequency activation from degradation-induced attenuation, leaving attenuation differences among inactive frequencies insufficiently modeled. In this work, we propose frequency-decoupled posterior guidance to separate frequency activation from attenuation-aware spectral regularization. Specifically, a progressive low-to-high frequency schedule determines the active measurement band, while a kernel-derived attenuation map defines a selective spectral prior over inactive components. To stabilize the sampling process, we also introduce a local trajectory regularizer that suppresses spatially irregular state-to-clean deviations. For a fixed endpoint energy, we provide a KL-regularized path-space interpretation. In practice, we construct time-dependent guidance through local energy corrections using a Tweedie plug-in approximation. Experiments on natural-image benchmarks demonstrate strong PSNR and SSIM performance across challenging non-blind deblurring settings, even at higher measurement noise levels.
|
| 189 |
Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning
2610.06234
|
cs.CVcs.AI
|
Lingyu Shen, Wei Tang, Fakhri Karray, Min-Ling Zhang |
Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes a...Multi-instance partial-label learning (MIPL) addresses inexact supervision in both the instance and label spaces, which can be applied to video classification. However, bag-level labels do not explicitly supervise the correspondence between candidate classes and temporal evidence. We propose {\ours}, which couples label disambiguation with temporal evidence allocation through a joint class--time assignment. Occupancy-regularized spherical matching associates contextualized video features while learning nonuniform temporal mass and discouraging excessive concentration. During training, candidate-restricted inference recomputes the assignment within the candidate label set. A dual-marginal KL projection then constructs a structured teacher that incorporates momentum-refined class beliefs while preserving the proposal's temporal occupancy. A single plan-level KL objective aligns the full-space predictor with this teacher. Our analysis characterizes when candidate re-solving differs from masking and shows that, under the stated construction, the joint objective decomposes into class-marginal and class-conditional temporal supervision. We construct VCMIPL benchmarks from Breakfast, DoTA, and FineAction using model-generated candidate labels and evaluate the method across four feature representations. Extensive experimental results demonstrate that PIVOTMIPL outperforms existing MIPL algorithms in both effectiveness and efficiency.
|
| 190 |
CentriQ: Calibration-Free Quantization of Diffusion Transformers via Exact Mean Centering
2610.06260
|
cs.CVcs.AI
|
Nata\v{s}a Jovanovi\'c, Mathieu Salzmann, Saqib Javed |
Diffusion transformers (DiTs) achieve state-of-the-art image generation, but their sampling cost limits deployment. Quantizing both weights and activations to 4 bits reduces this cost, yet existing methods fall short in one of two ways. Calibration-based metho...Diffusion transformers (DiTs) achieve state-of-the-art image generation, but their sampling cost limits deployment. Quantizing both weights and activations to 4 bits reduces this cost, yet existing methods fall short in one of two ways. Calibration-based methods are tied to a specific checkpoint and prompt distribution, whereas data-free Hadamard rotation, effective for LLMs, loses quality on DiTs. We show that this loss has a structural cause. Adaptive layer-norm conditioning adds a per-token mean to the activations, and at the widths of the evaluated DiTs, the Hadamard rotations used by data-free methods cannot spread this mean uniformly across coordinates. A single dominant direction therefore survives the rotation and sets the quantization range. We introduce CentriQ, a calibration-free quantizer that centers each token before rotation and restores the mean exactly through a rank-1 full-precision branch, so that per-token scales follow in closed form without data. Weights are fitted under a robust $\ell_p$ objective that tracks the dense mode of each group and discounts heavy tails. Across three DiTs, CentriQ matches the quality of calibrated SVDQuant at 4 bits, whereas calibration-free weight quantizers with plain per-token activation quantization collapse or degrade substantially. CentriQ outperforms the strongest calibration-free method reported to date at 2-bit weights. It is also the first calibration-free method to retain usable image quality at 2-bit activations.
|
| 191 |
VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction
2610.06293
|
cs.CV
|
Qiutong Chen, Yuchan Guo, Zhenlong Yuan, Haobo Yang, Fangfang Lin |
Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to...Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose VepAgent, an agentic framework that integrates causal-transition reasoning with tool-augmented reinforcement learning (RL) for robust VEP. Unlike prior methods that passively project future trajectories from historical dependencies, our approach explicitly models the logical progression from terminal observed states to future events. Specifically, we first construct futurebench-4K, a high-quality chain-of-thought dataset for supervised fine-tuning (SFT) that effectively bridges the causal-logic gap by structuring the deduction of unobserved intermediate states. Subsequently, we develop a diagnostic tool library integrating state tracking, frame retrieval, and region magnification, enabling the agent to dynamically augment reasoning with external tools to recover missing spatio-temporal evidence and resolve visual ambiguities during inference. Moreover, we propose a composite reward mechanism that jointly optimizes prediction accuracy, causal coherence, and reliable prior, compelling the agent to rely on genuine visual grounding rather than superficial textual similarities. Extensive evaluations on FutureBench and NEPBench datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming larger MLLMs and validating the empirical effectiveness of our agentic, future-oriented reasoning paradigm.
|
| 192 |
Wiring Matters: Injection Topology and Initialization of Affordance Heads in Vision-Language-Action Policies
2610.06318
|
cs.CVcs.AI
|
Zijian An, Linhan Wang, Jiayan Wang, Shijie Geng, Ran Yang |
Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a moder...Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a modern VLA on the LIBERO benchmark. Our recipe reads the backbone through a stop-gradient and re-injects an intermediate head feature into the action expert via a learned bridge. The stop-gradient is a precondition: letting affordance gradients reach the backbone drops the policy below the headless base (85.5% vs. 93.1%). With the backbone protected, a same-budget 2*2 ablation over injection topology (concatenation vs. residual) and bridge initialization (zero vs. random) shows initialization is the dominant lever. The best wiring, an actively initialized residual bridge, reaches 96.2%, matching the far more elaborate three-expert AffordanceVLA (95.8%) with under 1% extra parameters. Two probes explain the mechanism: ground-truth affordances fed as an input hurt, and inference-time zeroing shows a lazy bridge acts only as a training-time regularizer while an active bridge becomes load-bearing.
|
| 193 |
Readout Blindness: VLM Scores Miss the Spatial Direction Their Frozen Encoders Retain
2610.06324
|
cs.CVcs.LG
|
Guangyuan Li, Tianming Du, Yan Jiang, Bihan Wen, Jiancheng Yang |
CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and e...CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when scoring rules treat the subject and object symmetrically, direction cancels regardless of encoder training. Guided by this analysis, we introduce Antisymmetric Displacement Readout (ADR), which aligns caption words with image patches in the frozen features and scores each relation by the signed displacement between matched object centroids. Notably, ADR succeeds without additional training or learned parameters, thereby demonstrating that directional information remains in the frozen encoder. However, text and world priors can inflate accuracy, so we further introduce prior deflation, which measures the benefit of the image-text pairing as the grounded gain over a null that pairs each item with an unrelated image. Extensive experiments across encoder families show that ADR substantially improves over deployed scores, which remain near chance on most direction-balanced sets even for fine-tuned encoders. Compared with more complex readouts, ADR outperforms the evaluated MLLM likelihood readouts and is competitive with their chat inference at a small fraction of the computation. These results support our claim that directional information can be recovered from frozen features by an appropriate readout. Our implementation and evaluation kit will be publicly available.
|
| 194 |
SPIN: Image Immunization Against Diffusion Editing via Single-Step Projection in Stochastic Neighborhoods
2610.06334
|
cs.CV
|
Fengming Gu, Jie Zhang, Zhongqi Wang, Qiankun Li, Shiguang Shan |
Diffusion models have greatly advanced instruction-guided image editing, while also raising concerns about unauthorized image manipulation. Image immunization addresses this risk by adding imperceptible perturbations to an input image to disrupt subsequent edi...Diffusion models have greatly advanced instruction-guided image editing, while also raising concerns about unauthorized image manipulation. Image immunization addresses this risk by adding imperceptible perturbations to an input image to disrupt subsequent edits. Since editing requests are unknown at image release, protection should remain effective beyond the instruction used to construct the perturbation. Existing immunization methods either require costly full-trajectory backpropagation or use intermediate objectives whose effects may be weakened by subsequent denoising. Meanwhile, a single inference path provides limited feedback about alternative denoising continuations. To address these challenges, we propose \textsc{SPIN}, a framework for image immunization via one-step projection over local stochastic trajectory neighborhoods. Starting from an early denoising state, \textsc{SPIN} generates stochastic neighboring states under the same instruction and predicts their clean latents through one-step projection without full unrolling. We then optimize a bounded input perturbation to maximize the average deviation of these predictions from a clean-edit reference, encouraging the perturbation to disrupt multiple possible editing outcomes. Experiments on two image editors demonstrate substantial gains in protection performance, with \textsc{SPIN} outperforming compared methods across all six metrics under seen instructions and in the more challenging unseen instruction setting.
|
| 195 |
BabelFake: A Multilingual Audio-Visual DeepFake Benchmark
2610.06339
|
cs.CV
|
Carlotta Segna, Joel Tschesche, Anna Rohrbach |
Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of Engli...Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of English-speakers, often include outdated manipulation types, or overlook the audio modality. Further, many datasets feature individuals who did not consent to be used in DeepFake creation. We introduce BabelFake, a multilingual audio-visual DeepFake benchmark recorded with consenting participants. BabelFake contains 399k clips (1,323 hours) from 496 individuals spanning five languages (English, German, Italian, French, Spanish). Our modular data generation pipeline pairs 11 modern video manipulation methods with 4 voice cloning engines, distinguishing visual-only (face swapping) and joint audio-visual manipulations (lip synchronization and portrait animation). By benchmarking state-of-the-art detectors, we show that detection difficulty depends on the audio-visual generation pairing, with substantial performance degradation when authentic audio is preserved. Cross-language/demographic evaluation reveals sensitivity varying across detector architectures and training data, while human evaluation reveals that perceived realism and machine-detection difficulty do not necessarily align.
|
| 196 |
MeSD: Multi-Evidence Self-Distillation for VideoLLM
2610.06342
|
cs.CVcs.AI
|
Weijie Zhu, Han Fang, Hanyu Fu, Yuzhe Zhang, Xin Wei |
While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privile...While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.
|
| 197 |
KineWorld: Action-Induced Transport Fields for Embodied World Modeling
2610.06349
|
cs.CV
|
Ziying Song, Yuchen Liu, Zhuoran Xu, Ziyang Liu, Jian Jin |
Embodied world models predict the visual consequences of candidate actions before execution. However, existing action-conditioned world models often adopt uniformly weighted visual generation objectives that can be misaligned with embodied prediction needs. Ev...Embodied world models predict the visual consequences of candidate actions before execution. However, existing action-conditioned world models often adopt uniformly weighted visual generation objectives that can be misaligned with embodied prediction needs. Even with explicit motion conditioning, these objectives can underemphasize spatially sparse changes that are critical to interaction. We propose KineWorld, a transport-aware world-modeling framework that extends robot kinematics from motion conditioning to the spatial allocation of generative supervision. Kinematic Transport Lifting (KTL) constructs renderer-derived, camera-aligned transport fields from commanded robot motion. Transport-Aware World Diffusion (TAWD) calibrates their motion support on the video-latent grid and reweights future-RGB flow matching through a normalized mixture of uniform and transport-focused distributions. We train KineWorld using ALOHA-AgileX bimanual manipulation data from RoboTwin 2.0. KineWorld achieves an EWMScore-P of 68.95 in single-view evaluation and a TWB-Score of 54.82 in multi-view evaluation. These results support a shift from appearance fitting toward action-consequence modeling for embodied decision-making.
|
| 198 |
Environmental sensor readings in two crop disease image datasets identify the session in which each image was taken
2610.06369
|
cs.CVcs.LG
|
Sungwoo Kang |
Integrating environmental sensor data with leaf imagery is widely reported to boost crop disease classification accuracy. In this work, we reveal that these reported gains are often artifacts of dataset construction: because a single sensor reading is shared a...Integrating environmental sensor data with leaf imagery is widely reported to boost crop disease classification accuracy. In this work, we reveal that these reported gains are often artifacts of dataset construction: because a single sensor reading is shared across many images collected in a single session (one farm on one date), multimodal networks can predict disease simply by memorizing session identities. Analyzing two widely used Korean datasets, the Crop Disease Diagnosis (CDD) benchmark and an AI Hub pest/disease dataset, we demonstrate that nearly all images share sensor values, with 91.9% of CDD test images having exact sensor duplicates in the training set. Remarkably, an image-free classifier given only timestamps matches or exceeds sensor-driven predictions across all seven evaluated crops, and matches the published macro-F1 of a state-of-the-art CDD fusion model. These results indicate that performance gains on standard random splits cannot be disentangled from session leakage. We propose that multimodal crop studies must evaluate on session-held-out splits and report performance against sensor-free date-time baselines to ensure genuine generalization.
|
| 199 |
MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity
2610.06378
|
cs.CV
|
Hang Wang, Chao Shen, Lei Zhang, Zhi-Qi Cheng |
The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving...The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, temporal regularity in fine-grained visual representations has received limited attention. We find that caption-derived textual representations provide complementary discriminative cues to global visual representations. Our analysis further reveals that AI-generated videos exhibit stronger temporal persistence and lower temporal variability, a pattern we term temporal over-regularity (TOR). Based on these findings, we propose MTOR with a multimodal branch and a TOR component. The multimodal branch integrates global visual and caption-derived textual representations, while the TOR component models temporal over-regularity at three levels: coarse inter-frame continuity, fine-grained token correspondence, and frame-to-video stability. Extensive evaluations on five benchmarks covering 46 generator variants demonstrate state-of-the-art overall performance against 16 representative baselines, while robustness experiments confirm strong resilience to twelve real-world video perturbations. Code and models will be released at https://github.com/hwang-cs-ime/MTOR.
|
| 200 |
Multi-Task Partially Supervised Learning for Super-Resolution and Semantic Segmentation on Earth Observation data
2610.06389
|
cs.CV
|
Ho\`ang-\^An L\^e, Minh-Tan Pham, Solange Lemai-Chenevier, Daniel Greslou |
Super-resolution and semantic segmentation are known to benefit one another, especially in the Earth observation context. However, learning both tasks in a joint model often requires both task annotations, which is impractical and expensive. In this paper, we ...Super-resolution and semantic segmentation are known to benefit one another, especially in the Earth observation context. However, learning both tasks in a joint model often requires both task annotations, which is impractical and expensive. In this paper, we study the multi-task partially supervised learning paradigm for both tasks, where each example is assumed to have only a single-task annotation. To that end, we examine two multi-task architectural variations, the sequential and shared variants, and then propose a hybrid variant and a re-projection loss to benefit from the shared representation and enforce image quality of super-resolution when training with semantic segmentation. Experiments show favorable results compared to the SOTA sequential variant. Source code will be published at https://github.com/lhoangan/munera.
|
| 201 |
Harnessing Multimodal Large Language Models for Training-Free Human-Object Interaction Detection
2610.06394
|
cs.CV
|
Zhaolin Cai, Huiyu Duan, Liu Yang, Yanjun Qin, Bo Ai |
Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising...Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad visual-semantic knowledge and versatile perceptual and reasoning capabilities. However, existing approaches largely invoke these capabilities through loosely coordinated inference stages. This fragmented execution restricts the role of interaction hypotheses in guiding visual exploration, leaving key participants overlooked and local ambiguities unresolved. Furthermore, propagating early semantic assumptions through subsequent visual grounding and relation prediction induces self-reinforcing semantic circularity. To resolve these challenges, we propose HarnessHOI, a training-free framework that transforms passive MLLM inference into an active interaction-centric harness. Specifically, we introduce an interaction-guided perception mechanism that projects emerging interaction hypotheses back into the visual space to discover missing participants and refine ambiguous evidence through targeted observation. Furthermore, a relation-agnostic geometric adjudication module reconciles multi-source evidence to establish a unified spatial basis for grounded interaction reasoning across multiple actions and semantic roles. Extensive experiments on HICO-DET and V-COCO demonstrate that HarnessHOI achieves state-of-the-art performance among training-free methods, confirming the effectiveness of the proposed harness for complex interaction understanding. Code will be released upon publication.
|
| 202 |
SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs
2610.06413
|
cs.CVcs.CLcs.AI
|
Rafael Teixeira Sousa, Vin\'icius Paulo Lopes de Oliveira, Elisa Ayumi Masasi de Oliveira, Luiza Martins de Freitas Cintra, Fernanda Bufon F\"arber |
Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a data...Thinking-enabled vision-language models (VLMs) report ever-higher accuracy on spatial benchmarks, yet final-answer scores cannot reveal whether a correct prediction reflects faithful spatial reasoning or a linguistic shortcut. We introduce SpatialChain, a dataset of 28,350 training and 899 test examples pairing spatially-oriented GQA questions with scene-graph-grounded reasoning chains, retained only when the generated answer matches the symbolic ground truth, and a two-axis evaluation combining objective chain-overlap metrics with a scene-graph-aware LLM judge that scores faithfulness and completeness independently of the final answer. Applied to nine thinking-enabled VLMs, the protocol surfaces three findings invisible to standard accuracy: (i) four of nine models achieve $\geq$79% VQA accuracy while exhibiting shortcut rates above 39%, i.e., correct answers whose reasoning the judge marks as unfaithful; (ii) chain quality significantly predicts answer correctness for seven of nine models, but the two exceptions (Claude Sonnet 4.6, InternVL3.5-8B) reveal qualitatively distinct failure modes, terse output vs. verbose-decorative reasoning, that benchmark accuracy alone conflates; (iii) SFT on SpatialChain improves Qwen3-VL-8B by +6.2 pp in-domain and reduces its shortcut rate to 22%, while a stylistic specialization effect on external benchmarks motivates replay-augmented training as mitigation. The faithfulness judge is validated against 198 human-annotated items, where judge-human agreement matches human-human agreement, and against a second judge from a different provider, which preserves the model ranking ($\rho$ = 0.88). Data, generation scripts, and evaluation code are released at https://github.com/spatialchain/SpatialChainBenchmark.
|
| 203 |
Toward Reliable Infant Pose Estimation: A Training-Dynamics Approach to Noisy Annotation Detection
2610.06423
|
cs.CVcs.AI
|
Emanuele Cardinale, Sara Moccia, Alessandro Cacciatore, Lucia Migliorelli |
Spontaneous movement analysis in preterm infants relies increasingly on markerless pose estimation (PE) to derive clinically relevant motion biomarkers directly from video recordings. Training accurate infant PE models requires large sets of manually annotated...Spontaneous movement analysis in preterm infants relies increasingly on markerless pose estimation (PE) to derive clinically relevant motion biomarkers directly from video recordings. Training accurate infant PE models requires large sets of manually annotated keypoints, and human annotation is inherently prone to error. Noisy keypoints (i.e., keypoints mislocalized with respect to their true anatomical position) are especially problematic in this clinical setting, since they can propagate as artificial artifacts into the reconstructed joint trajectories. Building on the small-loss hypothesis and training-dynamics-based sample selection established in the noisy-label learning literature, we propose a novel framework for detecting noisy keypoint annotations. A hybrid convolutional-attention model is trained to predict the anatomical category of each keypoint from its spatial coordinates and local visual features; the resulting cross-entropy training dynamics are then used to derive per-keypoint descriptors, which are partitioned into clean and noisy subsets via unsupervised clustering. We validate the approach on NeoPose, a newly collected dataset of 65 hospitalized preterm infants, under two realistic noise scenarios (random positional perturbation and left-right swapping) across multiple noise levels. Results show that the proposed approach achieves an F1-score of up to 91.9% in noisy-keypoint detection. The framework further generalizes to the heterogeneous COCO benchmark, where filtering CE-detected noisy keypoints from the training set also yields measurable improvements (up to 7.4 AP points) in downstream pose estimation accuracy at moderate-to-high noise levels.
|
| 204 |
MaRO-GS: Mask-Robust Object-Centric Gaussian Splatting from Inconsistent Multi-view Masks
2610.06472
|
cs.CV
|
Eunji Kim, Gahyeon Kim, Gianella Cravioto, Dong-hun Lee, Chaewon Moon |
We address the challenge of accurate 3D object reconstruction from multi-view images in Gaussian Splatting. Existing object-level 3DGS methods reconstruct the entire scene rather than directly optimizing the target object, even when only the target object is n...We address the challenge of accurate 3D object reconstruction from multi-view images in Gaussian Splatting. Existing object-level 3DGS methods reconstruct the entire scene rather than directly optimizing the target object, even when only the target object is needed, which incurs substantial computational overhead. They also rely on 2D segmentation masks to associate Gaussians with objects, but these masks are often inconsistent across views. Such inconsistencies corrupt Gaussian optimization and produce incorrectly supervised Gaussians that degrade object reconstruction fidelity. To overcome these limitations, we propose MaRO-GS, a 3DGS framework that directly optimizes target-object Gaussians from object-masked multi-view images and remains robust to inconsistent supervision. For reliable supervision, mask-reliability view filtering excludes unreliable views. Object-supported Gaussian density control suppresses Gaussians irrelevant to the target object and prevents background densification, while Silhouette-aligned Object Loss maintains object-focused optimization. Extensive experiments across diverse datasets demonstrate that MaRO-GS improves PSNR, segmentation accuracy, and computational efficiency, with the largest PSNR gain of 2.05 dB on the small-object LERF-Mask dataset.
|
| 205 |
Topology-Informed Prompt-Conditioned Universal Segmentation of Uterine Structures from Ultrasound and MRI
2610.06494
|
cs.CVcs.AI
|
Yongheng Sun, Yuexi Gu, Jingwen Sun, Maureen Kohi, Mingxia Liu |
Multi-structure segmentation of the uterus is important for computer-assisted screening, diagnosis, and treatment planning of uterine diseases, where ultrasound and MRI provide complementary clinical information. However, developing a unified model across thes...Multi-structure segmentation of the uterus is important for computer-assisted screening, diagnosis, and treatment planning of uterine diseases, where ultrasound and MRI provide complementary clinical information. However, developing a unified model across these modalities is challenging due to their substantially different image appearances, anatomical contexts, spatial resolutions, and label spaces. Moreover, existing datasets often define different segmentation targets, making joint learning challenging and potentially leading to negative transfer across heterogeneous tasks. To this end, we propose a Topology-informed Prompt-conditioned Universal Segmentation (TPUS) framework for segmenting multiple uterine structures across ultrasound and MRI. TPUS introduces a graph-based multi-dataset backbone comprising modality-specific stems and a modality-shared graph-based encoder-decoder to support modality-sensitive input adaptation, structural feature reasoning, and joint representation learning across heterogeneous uterine segmentation tasks. In addition, TPUS uses task-aware class prompts to condition the segmentation process for different datasets and label spaces, a dynamic convolutional adaptation module to generate task-specific output responses, and a topology-informed loss to encourage anatomically consistent predictions. Experiments on a uterine ultrasound dataset and a T2-weighted uterine myoma MRI dataset demonstrate that TPUS achieves Dice scores of 0.898 and 0.693 on the two held-out test sets, respectively, outperforming several generic and universal segmentation baselines. Source code can be accessed at https://github.com/YonghengSun1997/TPUS.
|
| 206 |
NeuroCBIR: A Fast and Accurate Image Retrieval System for Whole-Brain and Region-Specific MRI
2610.06502
|
cs.CVcs.LG
|
Felix Nieto-del-Amor, Jingru Fu, J. -Sebastian Muehlboeck, Eric Westman, Daniel Ferreira |
Content-based image retrieval (CBIR) in neuroimaging enables the identification of structurally similar brain scans, supporting diagnosis, prognosis, and treatment planning; however, existing methods are often limited to small datasets, single brain regions, o...Content-based image retrieval (CBIR) in neuroimaging enables the identification of structurally similar brain scans, supporting diagnosis, prognosis, and treatment planning; however, existing methods are often limited to small datasets, single brain regions, or coarse class labels, thereby restricting their clinical utility and generalizability. Here, we present NeuroCBIR, a framework for fast and flexible retrieval of both whole-brain and region-specific 3D T1w MRI scans. A total of 103 cortical and subcortical regions are extracted to enable both whole-brain and region-level queries. NeuroCBIR leverages latent representations learned by a variational autoencoder (VAE) combined with contrastive learning, producing scan-specific embeddings that capture anatomical patterns. These embeddings were evaluated for subject re-identification, zero-shot age prediction, and zero-shot multi-class pathology stratification. Re-identification performance was high across both whole-brain and brain-region levels (mean average precision across the top-5 retrieved images (mAP@5) >= 98.4%), with robust generalization across datasets and acquisition conditions. While NeuroCBIR is not trained for age prediction or pathology stratification, zero-shot evaluations for these two tasks demonstrate that the embeddings encode meaningful information for downstream tasks. Embedding extraction on a 4-core CPU required approximately 18.7 s per scan, whereas similarity search was effectively instantaneous (less than 0.01 s). NeuroCBIR is publicly available for brain MRI with more than 26,000 precomputed T1w MRI embeddings. It supports reproducible research, region-specific flexibility, and clinically meaningful personalized diagnostic support. The software is available at https://github.com/minnelab/NeuroCBIR.
|
| 207 |
Harmful Content Generation in Text-to-Image Models: Capabilities and Moderation Limitations
2610.06503
|
cs.CVcs.AI
|
Paschalis Giakoumoglou, Manos Schinas, Symeon Papadopoulos |
Text-to-image generative models can produce highly realistic imagery but also raise concerns about harmful misuse. While safety mechanisms exist, systematic evaluations of their effectiveness against realistic attacks remain limited. We present a systematic ev...Text-to-image generative models can produce highly realistic imagery but also raise concerns about harmful misuse. While safety mechanisms exist, systematic evaluations of their effectiveness against realistic attacks remain limited. We present a systematic evaluation of harmful content generation across five open text-to-image models using an automated pipeline that transforms legitimate news captions into unsafe prompts targeting sexually explicit content, violence/gore, harmful stereotypes, self-harm, and hate speech. We evaluate both standard models with built-in safety mechanisms and community fine-tuned variants that bypass content restrictions. A human evaluation of 1,500 generated images shows high harmful-content generation rates: 89.2% for gore-related prompts, 47.6% for sexually explicit content, 43.6% for harmful stereotypes, 46.0% for hate speech, and 34.5% for self-harm, predominantly through graphic violence. Models show substantial capability for generating violent and stereotypical content, while community fine-tuned variants are particularly vulnerable to sexually explicit prompts. Generation quality is largely preserved under harmful prompting, producing imagery of sufficient fidelity to pose risks for disinformation and abuse; FLUX.1-dev produces clearly realistic harmful images in 30.9% of cases. We further evaluate automated moderation systems and find substantial detection gaps that allow unsafe images to evade filtering. Finally, we assess synthetic image detectors and show that models trained only on benign datasets perform worse on explicit content, while more diverse training data improves detection, highlighting semantic distribution gaps in current approaches. These findings expose limitations in current generation safeguards, moderation systems, and synthetic image detection, highlighting the need for stronger defenses against misuse at scale.
|
| 208 |
A General Pipeline for Dense Illuminant Estimation via Physically Based Synthetic Data
2610.06508
|
cs.CV
|
Luca Cogo, Gianmarco Corti, Simone Bianco, Raimondo Schettini |
Illuminant estimation is a fundamental problem in computational photography, as it enables the correction of color shifts induced by varying lighting conditions. While learning-based methods have demonstrated strong performance, their progress is hindered by t...Illuminant estimation is a fundamental problem in computational photography, as it enables the correction of color shifts induced by varying lighting conditions. While learning-based methods have demonstrated strong performance, their progress is hindered by the limited availability of large-scale datasets with accurate illuminant ground-truth. In this work, we propose a general and reusable pipeline to derive dense illuminant chromaticity maps from physically based 3D-rendered scenes. By repurposing an existing 3D scene collection, our approach enables the systematic generation of pixel-wise illuminant annotations under controlled lighting conditions, effectively lowering the barrier to data acquisition for learning-based illuminant estimation. Using this pipeline, we generate a large-scale synthetic set of 74,321 images, which we employ for pre-training both single- and multi-illuminant estimation models. Extensive experiments with state-of-the-art architectures show that synthetic pre-training consistently improves performance, with gains of up to 28% for single-illuminant estimation and up to 57% for multi-illuminant estimation, particularly in data-scarce regimes. These findings demonstrate that synthetic data generation pipelines offer an effective and scalable solution for the pre-training of illuminant estimation methods.
|
| 209 |
BrainTRACE: Tracing Longitudinal, Multimodal, and Volumetric Evidence in Brain MRI Clinical Reasoning
2610.06571
|
cs.CVcs.AI
|
Qizhen Lan, Mengchen Fan, Hang Zhang, Jingwei Duan, Moule Lin |
Brain MRI interpretation is a longitudinal clinical reasoning problem: radiologists compare serial studies, integrate information across MRI sequences, localize findings within volumetric anatomy, and translate this evidence into report-grounded assessments. E...Brain MRI interpretation is a longitudinal clinical reasoning problem: radiologists compare serial studies, integrate information across MRI sequences, localize findings within volumetric anatomy, and translate this evidence into report-grounded assessments. Existing medical VQA and 3D imaging benchmarks capture important parts of this workflow, but often evaluate brain MRI through isolated images, static volumes, or ungrounded report-style answers, thereby obscuring failures in the evidence chain that support clinical validity. We introduce BrainTRACE, a report-grounded benchmark for evaluating whether vision-language models can trace the evidence structure required for longitudinal brain MRI interpretation. BrainTRACE contains 7,273 scored VQA instances derived from 1,778 longitudinal patients, 7,299 MRI studies, and approximately 29k co-registered 3D MRI sequence volumes. The benchmark is organized by five levels of clinical reasoning, from acquisition recognition to case-level synthesis, and by evidence demands covering longitudinal comparison, report-grounded references, multi-sequence integration, and volumetric spatial evidence. BrainTRACE supports rendered inputs compatible with standard VLM interfaces, a 3D-evidence condition, and a decomposed case-reasoning track that audits six steps in a longitudinal evidence chain. Evaluation of 20 VLM configurations shows that current systems can identify isolated visual cues but rarely compose them into grounded longitudinal interpretations. We release the benchmark specification, evaluation lists, scoring implementation, scoring rubrics, and audit-record format to support reproducible progress in brain MRI VLM evaluation.
|
| 210 |
FrontVeg V2: A Training-Free Software Framework for Foreground-Aware Zero-Shot Plant Trait Segmentation in High-Resolution Images of Trellised Crops
2610.06575
|
cs.CV
|
Abdoul Djalil Ousseini Hamza (IRHS-IMHORPHEN, DPT SPE), Herearii Metuarea (IRHS-IMHORPHEN, DPT SPE), Corentin Lothod{\'e} (IRHS-IMHORPHEN |
FrontVeg V2 is an open-source, training-free software framework for foregroundaware zero-shot segmentation of plant traits in high-resolution images of trellised crops. The pipeline combines monocular depth estimation, automatic foreground extraction using Val...FrontVeg V2 is an open-source, training-free software framework for foregroundaware zero-shot segmentation of plant traits in high-resolution images of trellised crops. The pipeline combines monocular depth estimation, automatic foreground extraction using Valley-Aware Depth Thresholding, tiled zero-shot segmentation, Graph-Based Mask Assembly, and geometry-aware fusion. This design enables plant organs and disease symptoms to be segmented while reducing detections arising from neighboring vegetation rows. The current implementation integrates Depth Anything V2 (DAV2) and SAM3 and can be used through both command-line batch processing and a Napari graphical interface. FrontVeg V2 provides a reusable framework for multi-crop, multi-trait digital phenotyping without task-specific model retraining.
|
| 211 |
Keepsake: Selective Spatial Memory for Long-Horizon Video Generation
2610.06588
|
cs.CV
|
Abdul Mohaimen Al Radi, Kunyang Li, Yuzhang Shang, Mubarak Shah, Yu Tian |
Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage ...Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage and retrieval costs. Selective-construction approaches reduce redundancy, but make one-time retention decisions that are never revisited, even as an observation's value changes with the evolving memory bank. Both strategies leave a shared question unresolved: as the generated history evolves, which stored observations should still remain in memory? Our key insight is that the value of a stored observation is not fixed, but relational: it depends on the alternatives currently available in the memory bank. A view supported by many geometrically and visually similar substitutes can be relinquished with little loss of coverage, whereas an observation with few viable alternatives should remain regardless of age. We introduce Keepsake, an online, training-free controller for fixed-capacity spatial memory. At each update, Keepsake constructs a pose-appearance graph over retained and newly generated observations, combining camera-pose proximity with visual similarity. A retention priority jointly captures the number of strong substitutes and the similarity of the closest alternative, allowing Keepsake to continually reassess memory value, preserve observations with little alternative support, and evict highly replaceable ones under a fixed budget. The controller modifies only the persistent-memory update; the host generator, denoising schedule, and retrieval rule remain unchanged. Across MemCam and WorldMem, Keepsake improves FVD and LPIPS under a fixed memory budget. On 180-second MemCam trajectories, it retains only 32 of 5,397 frames while reducing FVD by 35.1%.
|
| 212 |
VGGT-Bridge: Beyond Sequential Pose Graphs via Coarse-Stride Skip Edges
2610.06594
|
cs.CV
|
Sungjae Choi, Hanna Bae, Sunghyun Baek, Junmo Kim |
Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences ...Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.
|
| 213 |
Analysis of SWIR Imaging Detection Performance Under Adverse Environmental Conditions for Autonomous Driving Systems
2610.06596
|
cs.CV
|
Rohan Mehra, Alexandre Riffard, Yannis Loumouamou, Mathieu Labussi\`ere |
Short-wave infrared (SWIR) imaging has emerged as a promising modality for autonomous driving, yet its practical benefits over RGB remain poorly characterized across diverse conditions. This paper presents a systematic comparative study of paired RGB and SWIR ...Short-wave infrared (SWIR) imaging has emerged as a promising modality for autonomous driving, yet its practical benefits over RGB remain poorly characterized across diverse conditions. This paper presents a systematic comparative study of paired RGB and SWIR object detection on the RASMD dataset, covering four weather conditions and two real-time detection architectures, with various fine-tunings evaluated against a unified ground truth. Overall, RGB demonstrates comparable or superior performance in most scenarios, while RF-DETR exhibits greater robustness across varying conditions. Beyond aggregate metrics, we propose a sensor-dominance mining framework that combines multi-model agreement with targeted manual inspection to identify scenarios where one sensing modality provides more reliable detections using largely unannotated paired data. This analysis reveals that SWIR offers clear advantages in four safety-critical situations, including windshield glare, water droplets on the windshield, low-contrast object visibility, and long-range vehicle detection. The findings suggest that SWIR should be viewed as a complementary modality that enhances perception in rare but challenging conditions. The datasets will be available upon request, and all code and trained model weights are publicly released at https://github.com/comsee-research/swir-adverse-env-analysis.
|
| 214 |
Multitask Conditional Generative Adversarial Network Enables Automatic Whole Knee Cartilage and Menisci Segmentation and Reliable T1\r{ho} and T2 Quantification Without High-Resolution Morphological Images
2610.06602
|
cs.CV
|
Ahmed Tahseen Minhaz, Richard Lartey, Zhiyuan Zhang, Jeehun Kim, Kunio Nakamura |
Early osteoarthritis detection through quantitative MRI (qMRI) requires accurate cartilage and meniscus segmentation, traditionally necessitating time-consuming, costly 3D high-resolution Double Echo Steady-State (DESS) MRI scans. This study developed a multi-...Early osteoarthritis detection through quantitative MRI (qMRI) requires accurate cartilage and meniscus segmentation, traditionally necessitating time-consuming, costly 3D high-resolution Double Echo Steady-State (DESS) MRI scans. This study developed a multi-task conditional generative adversarial network (MT-cGAN) to simultaneously synthesize DESS-like images and segment tissues directly from qMRI echo images. This retrospective study evaluated 508 knee MRI volumes from 361 subjects (mean age: $40.4 \pm 12.2$ years; 179 female) across three cohorts. Ground truth segmentation masks were generated from DESS images using a pretrained model with manual correction, and $T_{1\rho}$ and $T_2$ maps were computed from magnetization-prepared angle-modulated partitioned $k$-space spoiled gradient echo snapshots (MAPSS) echo images. MT-cGAN was trained to jointly synthesize DESS-like images and segment cartilage and meniscus directly from echo images. Model performance was evaluated using Dice score for segmentation accuracy and coefficient of variation (CV) for $T_{1\rho}$ and $T_2$ quantification. MT-cGAN achieved the highest segmentation performance, mean Dice score 0.84 (range: 0.80--0.86) across all cartilage and meniscus compartments and significantly outperformed the state-of-the-art conditional GAN model with transfer learning (mean Dice, 0.82; $p < 0.001$, Wilcoxon signed-rank test). For relaxometry quantification, MT-cGAN demonstrated the highest consistency with the reference DESS protocol, yielding the lowest CV ($T_{1\rho}$: 1.84%, $T_2$: 1.81%). The proposed MT-cGAN accurately segmented cartilage and menisci while providing reliable $T_{1\rho}$ and $T_2$ quantification directly from echo images. By eliminating the need for separate morphological DESS scans, this workflow reduces required scan times to facilitate the clinical translation of qMRI.
|
| 215 |
Lens3D: Target-Conditioned Visual Foveation for Fine-Grained 3D Understanding
2610.06611
|
cs.CV
|
Junming Huang, Zini Chen, Shuaiying Hou, Chi Wang, Qiang Dai |
Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visu...Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.
|
| 216 |
Video Encoders Built on Image Representations
2610.06616
|
cs.CV
|
Jusheng Zhang, Wenhao Wang, Longqi Cai, Liangzhe Yuan, Yuxiao Wang |
The design of a video encoder determines when frames begin to interact and which frame-specific visual evidence remains accessible to the language model. Native video pathways couple neighboring frames during visual encoding, whereas image pathways preserve in...The design of a video encoder determines when frames begin to interact and which frame-specific visual evidence remains accessible to the language model. Native video pathways couple neighboring frames during visual encoding, whereas image pathways preserve independently computed frame representations but incur a much larger visual-token cost when all image tokens are forwarded. We ask a basic question: whether a compact video encoder can instead be built on image representations. To answer this question, we separate three operations that are often coupled: per-frame representation, cross-frame token allocation, and temporal interaction. A frozen image encoder first produces frame-specific candidates. A question-aware selector then allocates a fixed token budget across frames using relevance, diversity, and cross-frame correspondence, after which a lightweight learned refiner reads neighboring-frame context and writes residual updates only to the retained anchors. This preserves source positions and keeps the visual output at the fixed budget. Across 13 benchmarks and three vision-language backbones, the resulting pathway matches full-image aggregate performance while using only about 28%-35% of its visual tokens. Specifically, on Qwen3-VL-8B, it achieves a 13-benchmark macro-average of 62.75 with 1,535 visual tokens, compared with 62.58 for the full Image pathway at 4,424 tokens and 59.49 for native Conv3D at 2,212 tokens. On Qwen3-VL-32B, it reaches a 13-benchmark macro-average of 66.28, compared with 66.09 for Image, while providing a 2.16x end-to-end speedup. These results show that compact video encoding does not require early temporal mixing: frame-specific evidence can be preserved first, allocated jointly, and temporally contextualized after selection.
|
| 217 |
RealtimeWAM: One-Step Asynchronous World Action Models
2610.06617
|
cs.CVcs.LG
|
Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong, Shiqiao Gu |
World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action exp...World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, $<1\%$ drop) across these benchmarks while delivering significant end-to-end speedup (\eg, $\sim25\times$ on H100). Our code and checkpoints are available via this \href{https://github.com/ModelTC/LightX2V/tree/main/examples/realtimewam}{link}.
|
| 218 |
Talk Like You: Imitating How You Speak in Real-Time Talking Head Generation
2610.06658
|
cs.CV
|
Baiqin Wang, Zhixing Ding, Jijie Li, Jiankuo Zhao, Zhen Lei |
In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronizat...In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: https://github.com/BQ-Wang0511/TalkLikeYou
|
| 219 |
Cross-dataset harmonization for robust endoscopic image analysis
2610.06663
|
cs.CV
|
Romil Imtiaz, Dimitris K. Iakovidis |
A significant problem in endoscopic image analysis is that the machine learning (ML) models used for this purpose usually underperform when applied on images acquired from endoscopes that are different from those used to acquire the images of their training se...A significant problem in endoscopic image analysis is that the machine learning (ML) models used for this purpose usually underperform when applied on images acquired from endoscopes that are different from those used to acquire the images of their training set. The main difference of the images originating from different endoscopes is their color distributions, which depend both on the image sensors and the light sources used. Although previous studies have highlighted this challenge, to the best of our knowledge it has not been previously explicitly tackled. This study focuses on this problem and proposes very simple but impactful method. It implements a reference-based image harmonization that reduces global appearance differences between endoscopic datasets. Specifically, it extracts global color statistics from a chosen reference dataset in the CIE-Lab color space and applies a statistical channel-wise transformation to map each target image toward the appearance of the images of the reference dataset. The method is evaluated in the context of polyp detection in both flexible colonoscopy and capsule endoscopy datasets using a dataset-level cross validation protocol. The results indicate that the proposed harmonization consistently improves cross-dataset performance up to 30.7%, outperforming relevant baseline and state-of-the-art methods. The results indicate that a substantial part of the generalization gap is driven by low-level appearance variation that can be mitigated without retraining.
|
| 220 |
VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding
2610.06672
|
cs.CVcs.AI
|
Yucheng Liu, Yufei Yin, Mingxiao Feng, Jiajun Deng, Wengang Zhou |
Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms: query-driven exploration, which is sensiti...Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms: query-driven exploration, which is sensitive to localization errors, and query-independent memory construction, which may omit question-specific details. We introduce VideoTapestry, a training-free multi-agent framework that adapts a preconstructed hierarchical video memory through coarse-to-fine, query-driven refinement. The preconstructed memory organizes video content into three levels, capturing global narrative context, event-level temporal structure, and fine-grained relational evidence, respectively. To support coarse-to-fine localization and observation, we assign a specialized agent to each level, keeping retrieval and refinement within a scale-specific context. Guided by the query, these agents revisit relevant video regions and enrich layer-wise memories with targeted multimodal observations. Their refinements are assembled according to the original hierarchy into a composite query-adaptive memory, preserving global context in a compact form while retaining fine-grained evidence along query-relevant branches for final reasoning. Compared with direct GPT-5.5 inference, VideoTapestry achieves absolute accuracy gains of 17.2%, 14.9%, 9.8%, and 7.0% on LVBench, LongVideoBench (Long), Video-MME (Long), and EgoSchema, respectively, achieving the state-of-the-art results among all competitors.
|
| 221 |
ChronoWorld: Camera-Controlled Consistent 4D World Generation via Spatiotemporal Cues and Geometric Reflections
2610.06687
|
cs.CV
|
Xiaoyu Zhou, Dingwei Xian, Zhenyu Wang, Yajiao Xiong, Yongtao Wang |
While existing camera-controllable video generation models can produce visually compelling sequences, preserving intrinsic 4D spatiotemporal coherence remains challenging. To address this limitation, we propose ChronoWorld, an "Observation--State--Reflection" ...While existing camera-controllable video generation models can produce visually compelling sequences, preserving intrinsic 4D spatiotemporal coherence remains challenging. To address this limitation, we propose ChronoWorld, an "Observation--State--Reflection" framework that leverages spatiotemporal causal cues and reconstruction priors to generate globally consistent, free-view 4D scenes. Given a context video, we introduce a Spatiotemporal Epipolar Causal Attention mechanism that enforces multi-view epipolar constraints and temporal causality throughout the generation process. In addition, we develop a reconstruction-driven geometric reflection pipeline with a 4D retrieval strategy to enable dynamic self-assessment and correction of generated outputs, improving consistency and accuracy. Extensive experiments show that ChronoWorld achieves state-of-the-art performance in spatiotemporally consistent, cinematic-quality 4D scene generation, with strong generalization and high-fidelity geometry across diverse scenarios.
|
| 222 |
GS-Pool: Object-Level Change Detection in 3D Gaussian Splatting
2610.06688
|
cs.CV
|
Boaz Keren-Gil, James Gain, Patrick Marais |
Factories, museums and surveyors photograph the same space months apart and need to know which objects changed. When each visit is reconstructed with 3D Gaussian Splatting (3DGS), a direct comparison of the two reconstructions does not answer this. Training is...Factories, museums and surveyors photograph the same space months apart and need to know which objects changed. When each visit is reconstructed with 3D Gaussian Splatting (3DGS), a direct comparison of the two reconstructions does not answer this. Training is stochastic, so two reconstructions of an unchanged space never coincide, and the second visit is often a quick re-scan with far fewer photographs. We propose GS-Pool, which takes two independently reconstructed Gaussian fields of the same space and returns the changed objects in each, together with their masks. SAM2 masks of each visit's photographs are lifted onto the Gaussians that render them and merged into an object pool, so every decision is taken once per object in 3D. We introduce a photographic carrier, the 3DGS training loss of each input reconstruction against the other visit's photographs, backpropagated to the Gaussians that rendered each pixel. We combine it with GS-Diff's geometry and colour terms and our distilled DINOv3 features. This evidence is compared with that of the objects present in both visits, which sets a change threshold for each scene. On PASLCD, GS-Pool reaches mIoU/F1 scores of 0.751/0.846 against 0.644/0.758 for GS-Diff, the strongest prior method, a gain of 17%/12%. Its mIoU is also 36%, 40% and 57% above that of O-SCD, PlenoCI and MV-3DCD, and it reaches 0.855 mIoU on CL-Splats, 33% above MV-3DCD. Each changed object is returned as a set of Gaussians with the evidence behind its decision, which an inspector can review in 3D.
|
| 223 |
Extending Dynamic World Surface Water Mapping to Sentinel-1 with AlphaEarth Embeddings
2610.06704
|
cs.CV
|
Rohit Mukherjee, Frederick Policelli, Beth Tellman, TC Chakraborty, Jonathan Giezendanner |
Dynamic World (DW) maps land use and land cover globally at 10 m from Sentinel-2 (S2) imagery, but only for cloud-free observations, which limits where and when surface water can be mapped. We use the DW water class as weak supervision for a Sentinel-1 (S1) sy...Dynamic World (DW) maps land use and land cover globally at 10 m from Sentinel-2 (S2) imagery, but only for cloud-free observations, which limits where and when surface water can be mapped. We use the DW water class as weak supervision for a Sentinel-1 (S1) synthetic aperture radar (SAR) model so that DW-like water maps can be produced for every S1 acquisition. Google's AlphaEarth Foundations (AEF) annual embedding supplies spatial context, while S1 backscatter supplies the acquisition-time observation. On 53 globally distributed scenes with independent annotations of 3 m PlanetScope imagery acquired within 48 h of the S1 overpass, the S1-only model already reaches a pooled water intersection over union (IoU) of 0.77, comparable to 0.75 for the operational OPERA DSWx-S1 product, and adding AEF raises it to 0.85. The fused model improves on the S1-only model on 44 of 53 scenes and exceeds OPERA on 48, and on the independent S1S2-Water benchmark it reaches 0.94, compared with 0.87 for OPERA. Optical land-cover products can thus provide scalable training labels for SAR surface water mapping.
|
| 224 |
MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
2610.06801
|
cs.CVcs.AI
|
Jiarui Chen, Zeqiang Lai, Jiangshan Wang, Ziheng Ouyang, Ye Huang |
Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high spa...Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a $1.80\times$ denoising speedup on Minimax-H3-Base and a $2.32\times$ speedup on 3D asset generation, both with negligible quality loss.
|
| 225 |
Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models
2610.06813
|
cs.CV
|
Zhimin Shao, Xijun Liu, Zhaoliang Zhang, Yutao Tang, Abhay Yadav |
Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-s...Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions also can propagate unreliable evidence from occluded or visually similar but geometrically distant views. In this paper, We introduce a Masked Geometric Encoder (MGE), which promotes the learning of robust geometric representations under incomplete cross-view context. During training, MGE strategically drops frame tokens from global attention and distills from a pretrained full-context teacher model. This allows the model to learn an intrinsically richer per-frame representation while providing sufficient intermediate supervision to avoid performance degradation. Through extensive experiments, we show that MGE leads to much stronger performance under occlusion and doppelganger views while retaining high performance on standard benchmarks. Such a richer frame representation also leads to more effective token reduction during inference. To this end, we develop a novel Anchor-Guided Adaptive token merging technique that preserves representative anchor frames while jointly merging redundant tokens from the remaining views. Compared to other efficient inference approaches, we can achieve inference speedup while consistently maintaining higher reconstruction quality, particularly in limited-view settings.
|
| 226 |
TAPDreamer: Transferable Adversarial Patches for World Action Models
2610.06814
|
cs.CVcs.AI
|
Xuanyu Lu, Fengqing Jiang, Kaiyuan Zheng, Yichen Feng, Yaorui Ding |
World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across task...World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.8% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 2.1% and 0.8% on two DreamWAM configurations and to 10.0% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.
|
| 227 |
UniSlider: Perceptually Uniform Sliders for Continuous Image Editing
2610.06831
|
cs.CVcs.AI
|
David Serrano-Lozano, Duygu Ceylan, Yannick Hold-Geoffroy, Iliyan Georgiev, Javier Vazquez-Corral |
Sliders provide an intuitive interface for continuous image editing. In current generative approaches, however, the slider is simply a rescaling of the method's strength parameter, such as an adapter coefficient, a prompt weight, or an interpolation factor. Th...Sliders provide an intuitive interface for continuous image editing. In current generative approaches, however, the slider is simply a rescaling of the method's strength parameter, such as an adapter coefficient, a prompt weight, or an interpolation factor. This strength relates poorly to perceptual change. The image can partially revert as the slider moves, long stretches of the range produce no visible difference, and short intervals transform the image abruptly. Remapping the strength could fix this uneven pace, but only if the trajectory is monotone, which current methods do not enforce. We therefore distinguish the slider from the strength, and require perceptual distance from the input to grow linearly with the slider value. We introduce UniSlider, a lightweight LoRA trained on a few-step editing backbone so that its strength approximates this ideal slider. Few-step sampling lets us impose this objective in pixel space without intermediate ground truth, and the backbone's output is preserved at full strength. However, a low-rank adapter cannot make the strength fully uniform. Our slider is thus an inference-time remapping of the strength, obtained by adaptive sampling. Since training optmizes to make the trajectory monotone, this remapping closes the remaining gap without extra training or parameters. On a new benchmark of 300 continuous edits evaluating uniformity, monotonicity, edit fidelity, and identity preservation, UniSlider outperforms all prior methods and is preferred in a user study.
|
| 228 |
Anatomy-aware Fine-grained Multimodal Fusion for Laryngopharyngeal Cancer T-Staging Prediction Using CT and Radiology Report
2610.06837
|
cs.CV
|
Xingyue Zhao, Yanzhou Su, Fang Zhang, Zhanghexuan Ji, Yirui Wang |
Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor...Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor invasion. Recent computer-aided approaches face two key challenges: 1) Structural relationship modeling: existing methods underrepresent anatomically structured patterns of tumor invasion, as they either process whole CT volumes without tumor-specific anatomical constraints or rely on labor-intensive tumor segmentation. 2) Fine-grained cross-modal alignment: while radiology reports contain organ-specific invasion details, current methods that apply global feature fusion struggle to accurately align individual anatomical structures with their corresponding textual descriptions. To address these issues, we propose an anatomy-aware multimodal framework that integrates organ-level CT context and radiology reports into a unified representation for laryngopharyngeal T-staging. The framework first constructs an Anatomy-Structured Organ Graph (AOG) that captures invasion patterns between primary sites and surrounding organs, then performs Organ-Anchored Cross-Modal Alignment (OCA) so that each organ node aggregates textual evidence from the radiology report, and finally refines this graph representation by injecting organ-specific invasion cues extracted from the report via Report-Enhanced Graph-Refinement (REG), yielding a multimodal organ graph that combines spatial and textual evidence. Extensive experiments demonstrate that the proposed framework achieves superior performance in T-staging of laryngopharyngeal cancer.
|
| 229 |
Learning to Read the Contextual Tokens in Diffusion Transformers
2610.06844
|
cs.CVcs.LGcs.AI
|
Omer Dahary, Etai Sella, Hadar Averbuch-Elor, Daniel Cohen-Or, Or Patashnik |
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well unde...Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading this contextual space through natural-language interrogation. We train a lightweight bottleneck network that maps intermediate contextual tokens into the input space of a frozen Large Language Model (LLM), allowing the LLM to answer questions about the emerging image directly from these hidden representations. Our reader reveals that contextual tokens encode a rich, global representation of the emerging scene: generation-specific semantics, including attributes left underspecified by the prompt, are accessible surprisingly early in denoising, while increasingly fine-grained details become readable over time. Remarkably, this information remains decodable even when the MM-DiT receives an empty prompt, showing that contextual tokens accumulate substantial image-specific information from the evolving visual representation itself. We further find that generations with more readable contextual representations tend to receive higher human-preference scores. Building on these observations, we introduce Contextual Alignment, a training technique that explicitly reinforces the visual-semantic information encoded in the contextual tokens, improving generation quality and distributional coverage. Together, our results establish contextual tokens as both an interpretable view into the internal dynamics of MM-DiTs and an effective target for improving generative models.
|
| 230 |
S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation
2610.06847
|
cs.CV
|
Jeffrey Hu, Daniel Olmeda Reino, Ayush Tewari |
Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diff...Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which performs autoregressive diffusion at high noise before switching to parallel diffusion at low noise. The autoregressive phase provides the serial computation needed to coordinate interdependent events and produce valid state transitions while the parallel phase jointly refines the entire video and reduces sampling time relative to fully serial generation. We implement S2PD with two architectures: a pixel-space diffusion transformer trained from scratch and a pretrained video model adapted through LoRA fine-tuning with causal attention. Across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines and generates videos with greater temporal stability and sampling efficiency than other serial methods.
|
| 231 |
One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline
2610.06852
|
cs.CVcs.AI
|
Shih-Chen Tseng, Chih-Hsuan Chen, Ryan Yang, Hsi-An Chen, Chun-Wei Tuan Mu |
Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any sile...Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any silently broken connection misrepresents the method. We formulate aspect-ratio-adaptive flowchart relayout as a distinct task: given a raster flowchart and a target ratio, produce a structurally faithful, hallucination-free, editable layout. Existing methods fail characteristically: image-to-image models stretch blocks and reject extreme ratios, text-to-image agentic systems hallucinate content, and parse-then-render systems mis-route edges. We propose an agentic pipeline factored into Parse, Style, and Layout stages, each pairing a main agent with a critic that combines deterministic constraint checks with VLM visual feedback so connectivity is explicitly checked and prevented from being silently broken. Outputs are draw.io-editable mxGraph XML. On a curated benchmark of 100 flowcharts at five aspect ratios, evaluated by Gemini 3.1 Pro and validated against human judgments, our method reaches 68.6% Content Fidelity versus 11.2-41.4% for prior work. Project page: https://onefigureeverycanvas.vercel.app/
|
| 232 |
Beyond the Good, the Bad, and the Ugly: Colormap Assessment through Data-Aware Perceptual Metric
2610.03720
|
cs.CV
|
Xi Duan, Yiwei Lin, Shiqing Xin, Aoying Wang, Yucheng Wang |
Continuous colormaps are widely used to visualize scalar fields, and their quality is typically evaluated using measures such as discriminative power and uniformity. Existing measures primarily characterize the intrinsic perceptual properties of the colormap i...Continuous colormaps are widely used to visualize scalar fields, and their quality is typically evaluated using measures such as discriminative power and uniformity. Existing measures primarily characterize the intrinsic perceptual properties of the colormap itself, largely independent of the underlying data distribution. In practice, however, user perception arises not from the colormap alone but from the visualization generated by mapping data values through the colormap. The perceptual differences that users actually experience depend jointly on the colormap and the underlying data. We propose a data-aware formulation that complements existing data-independent colormap assessment approaches. Rather than analyzing the colormap in isolation, we model the color-encoded visualization as a composite mapping from the spatial domain of the data to perceptual color space. This approach yields data-aware counterparts of established measures, including discriminative power, uniformity, and smoothness, as well as additional measures such as perceptual anisotropy and degeneracy. We validate the proposed measures against both the existing data-independent framework and empirical results from perceptual studies, extend the formulation to 2D colormaps, and demonstrate its integration into colormap optimization. An interactive system is provided for exploring colormap assessment under varying data distributions. Results demonstrate that our formulation offers a principled foundation for data-aware colormap assessment and design.
|
| 233 |
Least Squares for Time Series Forecasting
2610.03812
|
cs.CVcs.LG
|
Weiu-qiou Ciang, Yuzhou Hong, Sherry Chen |
A time-series forecast is scored on a future value of the series. A representation loss that regresses the next latent, as in LeNEPA, is a different least-squares problem on the same bottleneck. We write both programs down. The forecast program minimizes the e...A time-series forecast is scored on a future value of the series. A representation loss that regresses the next latent, as in LeNEPA, is a different least-squares problem on the same bottleneck. We write both programs down. The forecast program minimizes the error of a decoded latent on the coordinate that will be reported. For a scalar target and a linear decoder, every latent rank of at least one matches ordinary least squares, and an isotropy constraint is only a rescaling: after the decoder is refit, the forecast does not move. The other program fits the whole next vector at a fixed rank, then freezes the encoder and attaches a head. On a four-dimensional series whose last three coordinates are the same autoregression, that rank-1 fit puts mass $0.9998$ on the repeated coordinate and forecasts the remaining signal at the marginal variance $2.794$. The forecast program puts mass $1$ on the signal and matches the innovation variance $0.992$. Rank $2$ gives the vector fit a second direction, and the two programs agree. Iterating the fitted one-step coefficient $0.803$ raises the open-loop error from $0.992$ at one step to $2.700$ at eight steps.
|
| 234 |
Image-Based Breast Implant Detection for Mammography Dataset Curation and Near-Real-Time Deployment: Comparing Foundation Models and Task-Specific Convolutional Models
2610.03817
|
cs.CV
|
Vasisht Ishwar, Hari Trivedi, Young Seok Jeon, Beatrice Brown-Mulry, Frank Li |
Purpose: To evaluate the performance-feasibility tradeoffs of foundation models (FMs) and task-specific convolutional neural networks (CNNs) trained from scratch for breast implant classification in 2D mammography, with emphasis on suitability for near real-ti...Purpose: To evaluate the performance-feasibility tradeoffs of foundation models (FMs) and task-specific convolutional neural networks (CNNs) trained from scratch for breast implant classification in 2D mammography, with emphasis on suitability for near real-time clinical deployment. Methods: We evaluated four models: two FMs (RAD-DINO and MammoCLIP) and two CNNs trained from scratch for implant prediction (ResNet18 and our lightweight ResNetLite). Using the Emory Breast Imaging Dataset, 5,000 unilateral screening mammograms were used for training/validation and 1,000 manually reviewed unilateral images were held out for testing. For the FMs, global image embeddings from the pretrained encoder were classified using a support vector machine (SVM). The CNNs were trained end-to-end on 2D mammograms, with ResNetLite optimized via grid search over depth and width to balance accuracy and efficiency. Performance was evaluated using AUROC, sensitivity, specificity, accuracy, embedding visualization, and inference-latency. Results: All models demonstrated strong performance on held-out test data (n = 1,000). MammoCLIP achieved the highest AUROC (0.999) with the quickest training time of 493 seconds. RAD-DINO achieved the highest sensitivity (0.980; accuracy 0.989) but had the slowest inference and training times. ResNet18 and MammoCLIP achieved comparable accuracy (0.985). ResNetLite showed no statistically significant difference from ResNet18 (AUROC 0.993; accuracy 0.976) despite using only 1.4% of ResNet18's parameters, and had the fastest inference time. Conclusion: FMs and task-specific CNN models reliably detect breast implants on 2D mammography. Model selection is best guided by deployment context: MammoCLIP for GPU-equipped hospital settings requiring scalable integration, and lightweight CNNs such as ResNetLite for resource-constrained or edge deployments.
|
| 235 |
GOTT: Object-centric Dexterous Manipulation with a Reusable Cross-Embodiment Primitive
2610.03861
|
cs.CVcs.AI
|
Yulin Liu, Lai Wei, Yen-Jen Wang, Akash Sharma, Pieter Abbeel |
Foundation models and large-scale human data provide rich sources of manipulation intent, but translating this intent into multi-fingered robot behavior remains difficult. Dexterous hands still lack a reusable low-level primitive that reliably establishes cont...Foundation models and large-scale human data provide rich sources of manipulation intent, but translating this intent into multi-fingered robot behavior remains difficult. Dexterous hands still lack a reusable low-level primitive that reliably establishes contact across tasks and embodiments. We propose GOTT, a reach-acquire-move framework built around a single cross-embodiment contact-acquisition primitive. Given a robot-agnostic object trajectory and a reach specification, GOTT first brings the hand near a task-relevant contact region. The shared closed-loop primitive then establishes stable contact from this approximate initialization, and a pose-conditioned controller tracks the desired object motion. Reach specifications may come from future-aware planning, external models, or human demonstrations, while the primitive and tracking backend remain unchanged. Simulation and real-world experiments show that GOTT is able to establish robust contact across diverse objects, arm-hand platforms, and seen and unseen hand morphologies. It also consistently improves end-to-end task success over open-loop grasp execution.
|
| 236 |
Selective Backpropagation for Efficient Few-Shot Class-Incremental Learning
2610.04003
|
cs.CVcs.LGcs.AI
|
Eeham Khan, Abdulmoumen Al-Atrash, Ali Ayub |
Few-Shot Class-Incremental Learning (FSCIL) requires models to continuously learn new classes from limited samples while retaining prior knowledge, under strict constraints on compute and memory. Existing approaches lie along a difficult trade-off: simple fine...Few-Shot Class-Incremental Learning (FSCIL) requires models to continuously learn new classes from limited samples while retaining prior knowledge, under strict constraints on compute and memory. Existing approaches lie along a difficult trade-off: simple fine-tuning is computationally efficient but suffers from catastrophic forgetting, replay-based methods mitigate forgetting at the cost of substantial compute and memory, and exemplar-free methods often reduce forgetting by freezing most of the backbone, improving efficiency at the expense of adaptability. We propose Selective Backpropagation (SBP), a deterministic parameter budgeting framework that bridges this gap. SBP restricts gradient updates to a pre-allocated subset of network parameters, freezing past knowledge and preserving unbiased capacity for future learning, enabling rapid adaptation without costly mask optimization. We show that SBP achieves strong performance across standard FSCIL benchmarks while requiring training time close to that of naive fine-tuning and substantially lower training time than prior SOTA methods. Crucially, our experiments expose a limitation of standard FSCIL evaluation: performance on short, distribution-consistent benchmarks does not necessarily predict behavior under distribution shift or over substantially longer learning horizons. We therefore evaluate FSCIL methods in cross-domain settings and over an 80-session ImageNet-1K stream. SBP remains strong across these regimes while maintaining low training cost, providing a favorable stability-plasticity-efficiency trade-off. Our code is available at https://github.com/PaInt-Lab/sbp-main-public.
|
| 237 |
SUAVE: Unified Video-Action Models via Masked Diffusion
2610.04009
|
cs.CVcs.LG
|
Rhythm Syed, Jean Mercat, Sedrick Keh, Kushal Arora, Paarth Shah |
Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before...Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion backbones can imagine but treat language as frozen conditioning on a continuous latent space. Unified models bring these modalities into one architecture, but they either decode autoregressively, one token at a time, or keep video continuous with an auxiliary action head. In this work, we present SUAVE, a Single vocabulary Unified Action-Video modEl in which a masked diffusion transformer generates video and actions conditioned on language, with all three modalities represented as discrete tokens in a shared sequence. Choosing which tokens to mask at inference turns the same network into a world model, a robot policy, or a video-action model. For action-free co-training, the action positions of unlabeled video are filled with mask tokens and excluded from the loss. Simulation and real-world experiments demonstrate two findings. First, a single SUAVE model predicts long-horizon video and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improves policy performance and zero-shot robustness to distribution shift. Together, these results show that masked diffusion is a practical and versatile foundation for unified video-action modeling.
|
| 238 |
Grounded in Time: A Multi-Source Dataset and Benchmark for Temporal Grounding in Robotic Manipulation
2610.04255
|
cs.CV
|
Yi Wang, Yang Yang, Guangqi Xu, Sumin Lin, Ning Kang |
Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) mod...Robotic manipulation often requires inferring task-relevant states from past interactions when the current observation alone is insufficient to determine the appropriate action. Despite progress in benchmarking memory-augmented vision-language-action (VLA) models, application-oriented tasks requiring history-dependent semantic inference remain underrepresented. We introduce GiT (Grounded in Time), a dataset and benchmark for grounding manipulation decisions in past events across biolaboratory, household, and industrial scenarios. It includes real-robot and Universal Manipulation Interface (UMI) style demonstrations covering 18 bimanual tasks, together with simulation data and a ManiSkill-based evaluation suite covering nine tasks. Fine-grained subtask annotations and annotated counterfactual task pairs, in which similar current observations require different actions depending on prior events, support policy learning and targeted evaluation of history use. Evaluations of representative end-to-end VLA models in simulation and on selected real-world tasks reveal substantial room for improvement in history-dependent manipulation. The dataset and benchmark are available at the project page.
|
| 239 |
FloVMos: Optical Flow-based Medical Video Mosaicking
2610.04258
|
cs.CV
|
Jinyang Liu, Sandesh Ghimire, Chaman Singh, Jennifer Dy, Milind Rajadhyaksha |
Biomedical imaging modalities often require a trade-off among resolution, field of view (FOV), and acquisition speed. Video mosaicking offers a strategy to overcome this limitation by computationally stitching sequential high-resolution frames into a wide-FOV ...Biomedical imaging modalities often require a trade-off among resolution, field of view (FOV), and acquisition speed. Video mosaicking offers a strategy to overcome this limitation by computationally stitching sequential high-resolution frames into a wide-FOV composite. However, existing methods struggle with non-rigid deformations, and modality-specific artifacts arising in clinical and research imaging. Here, we present FloVMos, a generalizable, optical-flow-based deep learning framework for real-time video mosaicking across diverse biomedical imaging modalities. FloVMos achieves robust, pixel-level registration by fine-tuning an optical flow model on synthetic training data with ground-truth deformation fields. We introduce a pipeline for generating this training data, simulating realistic tissue motion and imaging distortions from existing mosaics or raw videos. Our automated synthetic data generation and optical flow model training based on this data allow users to adapt FloVMos to different imaging modalities. To demonstrate this, we applied FloVMos to seven diverse imaging modalities: reflection confocal microscopy, open-top light-sheet microscopy, fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy, and endoscopy. FloVMos outperforms conventional baselines in accuracy, robustness, and speed for all the tested modalities. This adaptable and training-efficient framework enables large-area visualization with real-time performance and may support broader use of video-based biomedical imaging in research and clinical workflows.
|
| 240 |
Asynchronous Tracking, Optical Communication and 3D Motion Capture using Event-based Sensors
2610.04342
|
cs.CV
|
Ziwei Wang, Angus Apps, Holly Battisson, Iain Guilliard, Timothy Molloy |
Inter-robot communication is a key enabling technology in cooperative robotics applications. While wireless communication is ubiquitous, simple, and effective, it has inherent limitations: the receiving robot cannot easily localise the source of an incoming si...Inter-robot communication is a key enabling technology in cooperative robotics applications. While wireless communication is ubiquitous, simple, and effective, it has inherent limitations: the receiving robot cannot easily localise the source of an incoming signal, clock synchronisation is challenging, and signals are broadcast to all nearby devices rather than targeted recipients. Optical communication provides a complementary channel that is inherently directional and spatially localised, directly addressing these limitations. Event cameras are bio-inspired dynamic vision sensors that respond to changes in image intensity with high temporal resolution, high dynamic range, and low latency, making them well-suited as receivers for high-rate optical communication in cooperative robotic systems. In this paper, we propose the Asynchronous Tracking and Optical Communication (ATOC) system, which integrates LED smart-beacon modulation, event-based detection, optical tracking, and event-data demodulation in a single pipeline. By simultaneously tracking and demodulating multiple LED smart beacons, ATOC transforms a conventional visual marker into a robust communication channel suitable for a wide range of high-impact robotics applications. We validate ATOC in a suite of laboratory studies and two 'applications': a smart-city demonstration and a 3D motion-capture system.
|
| 241 |
UnAct: Gradient-Free Unlearning via Targeted Activation Intervention
2610.04426
|
cs.CVcs.LG
|
Saeed Abdul Muizz, Aayat Rafiq, Iqra Altaf Gillani, Janibul Bashir |
Machine unlearning seeks to remove the influence of designated training data from a trained model without retraining from scratch. Retrain-free methods such as Selective Synaptic Dampening (SSD) and its label-free variant LFSSD avoid full retraining but still ...Machine unlearning seeks to remove the influence of designated training data from a trained model without retraining from scratch. Retrain-free methods such as Selective Synaptic Dampening (SSD) and its label-free variant LFSSD avoid full retraining but still require backpropagation and parameter importance computed over the entire dataset. We ask: what happens when a deletion request arrives with only a few images of the class to be forgotten? To answer this question, we introduce UnAct, a gradient-free class-unlearning method that needs only forward passes over the forget images. UnAct scores late-layer units by their responses, attenuates the most responsive connections, and repeats this for up to 20 rounds using no gradients, no labels, and no retained data. On ResNet-18 trained with CIFAR-10, CIFAR-20, and CIFAR-100, UnAct is competitive with SSD and LFSSD when forgetting entire classes and, unlike them, never collapses the network when forget data is scarce. On ResNet-18, across all tested sizes, UnAct's retain accuracy stays within 2.5 points of retraining, while SSD and LFSSD, at their full-class operating points, lose up to 86 points on some classes. With five forget images on CIFAR-10, UnAct's distance to retraining is 0.21 points, against 67 for LFSSD and 90 for SSD, and re-selecting SSD's threshold at each size with an oracle does not close the gap. In preliminary transfer to ViT-B/16, UnAct's distance to retraining is 11.5 against 33.7 for SSD, and a request is 19x faster than SSD when SSD computes its importance at request time. The code is available at https://github.com/abdulmuizz0903/UnAct
|
| 242 |
Understanding Clustering in Slot Attention via Particle Dynamics
2610.04493
|
cs.CVcs.LGcs.AI
|
Vasudev Joy, Rajat Rasal, Avinash Kori, Anthea Monod, Ben Glocker |
Studying attention through the lens of interacting particle dynamics has shown how token clustering can emerge from the underlying dynamics. We extend this perspective to slot attention, a method for object-centric image segmentation and representation learnin...Studying attention through the lens of interacting particle dynamics has shown how token clustering can emerge from the underlying dynamics. We extend this perspective to slot attention, a method for object-centric image segmentation and representation learning in which learned components obscure how much of the clustering behaviour is intrinsic to the attention dynamics. We therefore introduce simplified slot attention (SSA), a parameter-free variant whose dynamics are connected to soft $k$-means clustering and which provides a straightforward mechanistic explanation for the emergence of object-centric representations. On the Pascal VOC dataset, SSA achieves performance comparable to that of slot attention, demonstrating that competitive object-centric segmentation can be achieved without learned neural-network components.
|
| 243 |
Cross-Modal Solar Image Synthesis: Adapting the Surya Foundation Model from He I 10830 {\AA} to EUV Translation and Coronal Hole Segmentation
2610.04553
|
cs.CVcs.AI
|
Marco Marena, Andr\'es Mu\~noz Jaramillo, Qin Li, Haodi Jiang, Jinghao Cao |
The long observational record of He I 10830 {\AA} offers a means to investigate solar morphology before modern extreme-ultraviolet (EUV) imaging. We adapt the Surya solar foundation model to predict Solar Dynamics Observatory/Atmospheric Imaging Assembly (SDO/...The long observational record of He I 10830 {\AA} offers a means to investigate solar morphology before modern extreme-ultraviolet (EUV) imaging. We adapt the Surya solar foundation model to predict Solar Dynamics Observatory/Atmospheric Imaging Assembly (SDO/AIA) 94, 193, and 304 {\AA} images and a coronal hole (CH) probability map from full-disk helium observations. A convolutional input adapter, low-rank backbone updates, and dedicated output decoders learn from temporally paired, geometrically registered observations, with Spatial Possibilistic Clustering Algorithm (SPOCA) catalog polygons providing CH supervision. On observations held out from downstream fine-tuning, the selected dedicated models achieve disk-restricted correlation coefficients (CCs) of 0.8196, 0.8885, and 0.8398 for the three AIA channels, respectively; the CH model achieves an intersection over union (IoU) of 0.4360. The predictions recover broad solar structure, although local agreement varies substantially by channel. An optional residual refiner addresses spatial detail, and its application on pre-SDO dates improves the correlation of synthetic AIA 304 with Solar and Heliospheric Observatory/Extreme-ultraviolet Imaging Telescope (SOHO/EIT) 304 references from 0.6511 to 0.7085. Together, these results support the feasibility of helium-conditioned EUV morphological proxies and motivate their use in historical reconstruction, within the scope of the downstream test and cross-instrument evaluation.
|
| 244 |
ForeAct3D: Policy-Grounded Future World Modeling for VLA Policies
2610.04607
|
cs.CV
|
Zhe Tao, Feiran Wang, Gaowen Liu, Ramana Rao Kompella$, Yan Yan |
Robots need to anticipate how their actions will change the world, since manipulation success hinges on the resulting contacts and object motions. However, existing Vision-Language-Action (VLA) policies that predict future observations from shared features lea...Robots need to anticipate how their actions will change the world, since manipulation success hinges on the resulting contacts and object motions. However, existing Vision-Language-Action (VLA) policies that predict future observations from shared features leave the forecast decoupled from the actions the policy will actually execute, and impose no physical constraints on how the scene may evolve. We introduce ForeAct3D, a framework for policy-grounded future world modeling within VLA policies. Learnable geometric queries decode depth, semantic segmentation, and camera pose from the policy representation into current and future semantic 3D scene states, and the future queries are conditioned on the policy-generated action chunk to ground the forecast in the planned interaction. A physical-consistency closure relates the two states through background staticity and instance-level rigidity, and anchors the wrist-camera pose to end-effector kinematics. These objectives shape the shared representation used for action generation during training, and no future prediction is required at inference. Without robot pretraining, ForeAct3D achieves 98.3\% average success on LIBERO and an average task length of 3.73 on CALVIN, outperforming its base policy on every suite. Ablations show that semantic 3D supervision, physical consistency, and action conditioning each improve manipulation performance, and that action conditioning substantially improves future object localization. Real-world experiments on spatial placement, object insertion, and sequential manipulation further raise average success from 6.7\% to 37.8\% over the base policy. The project page and code are available at https://github.com/anthonytao80-crypto/ForeAct3D.
|
| 245 |
PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training
2610.04616
|
cs.CV
|
Mingyu Liu, Chonghao Sima, Tianjian Feng, Hanqing Wang, Cong Chen |
A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar nou...A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training even after the verb changes, and a gripper that closed on nothing may lift anyway. We call these dependencies modality shortcuts: regularities in successful demonstrations make visual, lexical, or motor cues sufficient to predict expert actions without the task evidence needed for the underlying decision. More demonstrations of the same kind can raise task success while leaving these shortcuts intact. We propose Perturbot which makes task-relevant evidence easier to use and shortcuts insufficient on their own: it applies task-preserving wrist-view perturbations, enriches instructions with decision-relevant captions, and adds random and failed trajectory segments relabeled with the behavior they contain. It complements scaling by changing what is scaled, and leaves inference unchanged. Moreover, we propose GroundingFscore, an offline score that diagnoses how severely a policy relies on modality shortcuts. Task success rate shows whether a policy improves, while GroundingFscore reveals whether the policy scales healthily, relying on task evidence rather than shortcuts. Together, Perturbot and GroundingFscore provide a training-and-evaluation framework for disentangling VLA decisions from shortcut priors while preserving responsiveness to task-relevant evidence.
|
| 246 |
Learning Discriminative Geometry for Drifting Models
2610.04703
|
cs.CVcs.LG
|
Doudou Zhang, Wenwen Hou, Yilin Chen, Qi Chen |
Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drif...Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately $82-95\%$ over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.
|
| 247 |
KALEIDO: Input-Space Adaptation of a Vision Model for Time-Series Forecasting Through Gated Fold Geometries
2610.04786
|
cs.CVcs.LG
|
Xiangyu Shi, Qinghua Liu, Sam Heshmati, Zubin Abraham |
Time-series foundation models buy zero-shot forecasting with large temporal corpora; a vision model needs none, since a natural image implicitly embeds the patterns a forecaster must model, and an ImageNet-pretrained masked autoencoder forecasts a series by in...Time-series foundation models buy zero-shot forecasting with large temporal corpora; a vision model needs none, since a natural image implicitly embeds the patterns a forecaster must model, and an ImageNet-pretrained masked autoencoder forecasts a series by inpainting a rendering of it. A rendered series is not a natural image, however, and closing that gap takes temporal-aware adaptation. We show that the rendering geometry - how the series is folded and drawn - is a controllable, mixable axis for it. Kaleido detects the dominant periods, renders a rule-generated set of fold geometries, combines the inpaintings with a convex per-position gate fit on validation only, and fuses the result with the zero-shot output at one fixed share, with no per-dataset hyperparameter beyond the baseline's published settings. Training only LayerNorm (0.05%), Kaleido lowers MSE by 13% against the published zero-shot baseline on LTSF and, frozen, by 6.6%; on GIFT-Eval it improves the baseline by 7.4% in MASE and 19.3% in CRPS.
|
| 248 |
Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness
2610.04792
|
cs.CVcs.AI
|
Songyuan Sui, Zhen Tan, Mohan Zhang, Rana Muhammad Shahroz Khan, Xia Hu |
Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal ...Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It rotates the principal input subspace toward a unimodal reference along a geodesic path. We prove that this correction minimizes the distance to the reference within a fixed subspace-distance budget. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.
|
| 249 |
MedImageOSWorld: Benchmarking GUI Agents for Medical Image Consoles
2610.04800
|
cs.CVcs.AI
|
Ziyang Long, Xinqi Li, Lujing Xing, Hsin-Jung Yang |
Graphical consoles offer a practical interface for medical acquisition assistance, allowing agents to work through the controls and visual feedback used by human operators. Reliable assistance requires linking on-screen anatomy to acquisition decisions that de...Graphical consoles offer a practical interface for medical acquisition assistance, allowing agents to work through the controls and visual feedback used by human operators. Reliable assistance requires linking on-screen anatomy to acquisition decisions that determine what image evidence becomes available next. We introduce MedImageOSWorld, a benchmark for evaluating this capability in simulated CT, MR, and ultrasound consoles. Using screenshots and mouse-and-keyboard actions, agents configure protocols, plan acquisitions, inspect the resulting images, and make corrective adjustments across seven capability levels, from console operation to feedback-driven control. Evaluation combines task-specific workflow checks with hidden anatomical ground truth to assess procedural completion and acquisition outcomes separately. A common evaluation protocol specifies episode conditions and interaction budgets, while recorded trajectories support analysis of how agents observe, act, and respond to acquisition feedback. Across eleven open-weight agents, success rates range from 3.0 to 25.0 on a 0-100 scale while workflow-progress rates reach 21.5-74.5: agents complete much of the console workflow but rarely acquire the intended anatomy. Success collapses between perception and planning, from 79-90% at the lowest three levels for the best agent to at most 12% for millimetre-level planning and 0% for closed-loop control among open-weight agents. Two proprietary agents reach success rates of about 34 and exceed the best open-weight agent mainly in acquisition quality (61 versus 42). MedImageOSWorld provides a controlled setting for studying whether general-purpose GUI agents can translate visual observations into effective medical acquisition decisions.
|
| 250 |
ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations
2610.04805
|
cs.CV
|
I-Chun Arthur Liu, Jason Chen, Gaurav S. Sukhatme, Daniel Seita |
Three-dimensional perception is critical for robotic manipulation, particularly for high-precision tasks, as recovering metric depth and precise 3D object positions from monocular RGB observations is inherently ill-posed. However, many Vision-Language-Action (...Three-dimensional perception is critical for robotic manipulation, particularly for high-precision tasks, as recovering metric depth and precise 3D object positions from monocular RGB observations is inherently ill-posed. However, many Vision-Language-Action (VLA) models rely solely on RGB observations for perception. Leveraging recent advances in foundation models for stereo matching, we introduce ExStereo, a stereo module that augments pre-trained 2D VLAs with 3D perception. ExStereo reconstructs scene geometry from stereo image pairs and renders multi-view observations as an explicit stereo representation for stereo feature extraction. The action tokens from the action expert selectively attend to the resulting stereo tokens through our proposed action-stereo cross-attention mechanism, enabling the policy to generate robot actions conditioned on 3D scene information. To learn robust 3D representations, we introduce a mid-training stage before task-specific post-training, using a self-supervised learning objective on large-scale stereo data. We validate our approach by fine-tuning two publicly available VLAs, $\pi_{0.5}$ and SmolVLA, and evaluate them in simulation and on a real-world bimanual PiPER platform. Across both settings, VLAs fine-tuned with ExStereo consistently outperform baselines, demonstrating the effectiveness of stereo perception for robotic manipulation. Our project website is at: https://exstereo-vla.github.io/ExStereo/.
|
| 251 |
VideoResearchAgent: Grounded Task Synthesis and Sim-to-Real RL for Open-Web Video Research
2610.04911
|
cs.CVcs.AI
|
Yuhang Zhou, Fei Li, Yuxi Wu, Bin Zhu, Jingjing Chen |
Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover...Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover relevant videos, navigate their temporal content, and ground answers in visual evidence. Training such agents at scale is challenging as live video interaction is slow and unreliable, whereas fixed local simulation can induce retrieval-specific shortcuts that fail to transfer to the open web. We introduce VideoResearchAgent, a scalable training framework to address these challenges. First, we introduce controllable task synthesis pipeline to synthesize multi-hop research tasks from timestamped visual evidence while filtering text-only shortcuts. Second, we build a field-aligned local video simulator that preserves deployment-facing search and watch interactions while accelerating video search by a factor of 34.5-64.6. Third, we introduce Retrieval-Domain-Randomized GRPO (RDR-GRPO), which diversifies candidate rankings, distractors, metadata, and result structure during training to reduce overfitting to simulated retrieval. On Video-BrowseComp, the VideoResearchAgent trained using Qwen3.5-4B achieves 40.48% accuracy, comparable to Gemini-3-Flash-Preview, while reducing cumulative API-token consumption by 74.9% relative to the untrained model. Together, these results establish an accurate and efficient training recipe for open-web video research.
|
| 252 |
FreeLoc: Online Floorplan Localization via Diffusion-Aided Pose Refinement
2610.05011
|
cs.CV
|
Haocheng Peng, Boyang Zhou, Jiarui Hu, Xiyue Guo, Ziyang Zhang |
Floorplans provide compact and widely available geometric maps for indoor localization, but existing high-performing floorplan-based methods still convert them into dense scene-specific offline databases, tying accuracy, storage, and runtime to the sampling re...Floorplans provide compact and widely available geometric maps for indoor localization, but existing high-performing floorplan-based methods still convert them into dense scene-specific offline databases, tying accuracy, storage, and runtime to the sampling resolution of the discretized pose space. We present FreeLoc, an online RGB-based floorplan localization framework that treats the floorplan as a directly queryable geometric map. FreeLoc introduces an efficient online geometric querying and diffusion-aided refinement scheme, which retrieves plausible pose anchors through on-the-fly floorplan ray querying and refines them into accurate continuous pose estimates. For sequential localization, FreeLoc develops an online likelihood construction strategy that bridges single-frame localization and probabilistic temporal fusion by constructing likelihoods from coarse-sampled candidates and refined pose hypotheses, enabling histogram-filter-based temporal fusion without offline databases. Experiments demonstrate real-time online inference and state-of-the-art performance in both single-frame and sequential localization, while real-world results validate practical deployability in indoor robotic localization scenarios.
|
| 253 |
TIRMamba: A Thermal-Prior-Modulated State-Space Network for Sub-Million-Parameter Infrared Image Super-Resolution
2610.05182
|
cs.CVcs.LG
|
Chun-An Lin, Tsung-Jung Liu, Yen-Chieh Ouyang |
Infrared image super-resolution is currently led by Mamba-based networks with 26 to 37 million parameters, which are difficult to deploy on the airborne and handheld platforms where thermal imaging is most needed. This paper presents TIRMamba, a network with 8...Infrared image super-resolution is currently led by Mamba-based networks with 26 to 37 million parameters, which are difficult to deploy on the airborne and handheld platforms where thermal imaging is most needed. This paper presents TIRMamba, a network with 896K to 910K parameters for single-channel thermal imagery. A Thermal Prior Highway computes gradient, local-contrast and spectral cues once at the input and, through one adapter per residual group, modulates a weight-tied bidirectional state-space trunk and gates its dual-scale detail branch; a tri-path reconstruction adds the learned residual to a bicubic radiometric baseline. Because the standard benchmark provides only 265 infrared training images and evaluates fusion products on full images, we train with a replay strategy: grayscale DIV2K pre-training followed by fine-tuning on 64-pixel patches drawn with equal probability from the infrared and natural corpora. At scale factor 4, TIRMamba matches the strongest protocol-trained methods on both official test sets with 29 to 40 times fewer parameters and 2.8 to 9.4 times lower latency; at scale factor 2 it gives the highest SSIM on both. A variant with prior-conditioned selectivity, TIRMamba-Rad, corrects a 3 dB raw-thermal failure of an intermediate size-invariant design and gives the best results at scale factor 4 on raw-thermal, unmanned-aerial-vehicle and independent-sensor test sets. Code and trained models will be released at https://github.com/julian135707/TIRMamba upon acceptance.
|
| 254 |
FACET: Factorized Asymmetric Conditioning for Efficient Transport in High-Fidelity Fluorescence Microscopy Synthesis
2610.05353
|
cs.CVcs.LG
|
Sazan Mahbub, Caleb N. Ellington, Eric P. Xing |
Fluorescence microscopy reveals where proteins localize, but only a limited number of proteins can be imaged in the same cell; generating these images from amino-acid sequence and the cell's morphological context enables in silico localization of unimaged prot...Fluorescence microscopy reveals where proteins localize, but only a limited number of proteins can be imaged in the same cell; generating these images from amino-acid sequence and the cell's morphological context enables in silico localization of unimaged proteins. The two conditions, however, play asymmetric roles: morphological context is spatially aligned with the target, whereas sequence is non-spatial and must specify protein-dependent localization within it, with recurring coarse patterns shared across proteins and finer protein-specific variation. Existing generators condition on both jointly, without separating what each explains. We introduce FACET (Factorized Asymmetric Conditioning for Efficient Transport), a probabilistic generative framework that encodes this structure as an explicit inductive bias: sequence semantics are learned from what context leaves unexplained, coarse localization regularities are shared across proteins through a semantic memory, and protein-specific variation is a bounded residual around them. A variance-preserving state projection further lets FACET perform continuous stochastic transport through a pretrained diffusion predictor with minimal parameter overhead. On held-out proteins, FACET improves spatial overlap by 34.3% on the Human Protein Atlas and 14.0% on OpenCell over a backbone-matched baseline, and reduces FID by 27.2% and 46.5%, respectively, with 75% fewer network evaluations. It also substantially improves protein-association structure recovery and yields better-calibrated predictions, while detailed ablations show complementary contributions from its design choices. These results identify factorized asymmetric conditioning, rather than generator capacity alone, as a key lever for high-fidelity, efficient, and biologically meaningful cellular image synthesis.
|
| 255 |
CASE: Cost-Aware Stopping for Efficient Long-Video Agents
2610.05400
|
cs.CVcs.AI
|
Yiming Du, Chenghao Liu, Zhiyuan Liu, Fangxing Zheng, Zhao Wang |
Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned s...Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to answer? We propose CASE, a plug-in termination framework that frames this decision as policy-conditioned sequential stopping. At each causal checkpoint, CASE combines an auxiliary multiple-choice assessment of accumulated evidence with the host agent's execution state. From complete native trajectories, we construct a cost-aware target that compares answering now with stopping later along the same search path, accounting jointly for answer correctness and the full cost of continued reasoning. A lightweight Ridge regressor learns this decision gap and produces STOP/CONTINUE decisions. We evaluate three vision-language models with VideoSeek and AVP. On Video-MME, end-to-end accuracy changes by +0.67 percentage points on average while CASE reduces model-token use by 53.63%. The same frozen policies then transfer zero-shot to LongVideoBench and MLVU, with end-to-end accuracy changes of +3.38 and +4.58 points while saving 58.78% and 51.28% of model tokens, respectively. Across all agent-model-benchmark combinations, CASE attains the highest accuracy-efficiency Pareto-frontier coverage among the compared stopping methods (83.3%) at the selected operating points. Online execution preserves this favorable accuracy-efficiency trade-off and additionally reduces measured runtime by 54.1% on average. CASE provides a plug-in termination framework for long-video reasoning agents, enabling them to decide when further evidence acquisition is no longer worthwhile.
|
| 256 |
EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation
2610.05418
|
cs.CV
|
Yuheng Na, Zhide Zhong, Junjie He, Junfeng Li, Haodong Yan |
Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visua...Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7\% on RMBench, 82.0\% on RoboMME and 83.8\% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.
|
| 257 |
The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models
2610.05453
|
cs.CVcs.LG
|
Tobias Braun, Jonas Henry Grebe, Emil Sivic, Patrick Mohr Gordillo, Hossein Shakibania |
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the pr...Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information revealed in one part of a conversation may remain accessible when the model later generates content in another modality. This risk is particularly concerning in settings where users rely on locally deployed models for privacy, assuming that sensitive interactions remain confined to their device. We introduce Privacy-Leaking Watermarks (PLWs): invisible, trigger-dependent watermarks that a malicious model provider can condition on prior chat history. With this adversarial intervention, the usual separation breaks: a sensitive keyword or semantic cue mentioned earlier in the conversation can cause a later, unrelated image to carry a hidden yet detectable watermark. PLWs pose a novel threat to users of unified multimodal models: A poisoned model can retain utility while covertly turning image generation into a channel for privacy leakage, even when deployed locally. Across 13 sensitive-attribute triggers and two model families, PLWs reach up to 100.0% TPR at 1% FPR. For example, across all tested conversational separations, OmniGen2 detects every prior disclosure of depression while falsely flagging only 1% of images generated without such a disclosure.
|
| 258 |
Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs
2610.05463
|
cs.CVcs.AI
|
Renchi Zhang, Joost C. F. de Winter, Dimitra Dodou, Harleigh C. Seyffert, Yke Bauke Eisma |
Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using ide...Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates ($\rho = 0.82$), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.
|
| 259 |
Universal Test-Time Training
2610.05484
|
cs.CVcs.CLcs.LG
|
Zefan Cai, Qinzhe Hu, Ziqiao Ma, Hao Tan, Junjie Hu |
Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories....Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, and introduce Universal Test-Time Training (uTTT), in which all layers read and write one shared memory while retaining layer-specific backbone parameters. The shared memory thus recurs over two dimensions, time and depth, with chunks and layers as their units: a write by a deep layer in one chunk can be read by a shallow layer in the next. We instantiate this idea as uTTT-MoE and uTTT-Dense. uTTT-MoE routes each token head to a few experts in a pool shared by all layers; uTTT-Dense applies the whole shared memory at every layer without routing. In language modeling, uTTT-MoE reaches 15.5 and 27.9 RULER accuracy at 124M and 760M, 2.6 and 2.1 points above its layer-private counterpart at equal state and active compute, the highest among tested bounded-state models, with per-token loss matching or beating full attention. In novel view synthesis, sharing at fixed per-layer compute gains 0.92 dB in view-23 object PSNR in routed models and 0.76 dB in dense models.
|
| 260 |
Robust Surgical Robotic Instrument Tracking via Sequential Multi-Cue Fusion and Sim-to-Real Self-Training
2610.05491
|
cs.CV
|
Hanyang Hu, Zekai Liang, Florian Richter, Michael C. Yip |
Efficient and robust tracking of surgical robotic instruments is important for robot-assisted minimally invasive surgery, yet remains challenging due to the complexity of surgical scenes and the unconventional geometry of surgical instruments. Keypoint-based a...Efficient and robust tracking of surgical robotic instruments is important for robot-assisted minimally invasive surgery, yet remains challenging due to the complexity of surgical scenes and the unconventional geometry of surgical instruments. Keypoint-based approaches are efficient, but their performance depends on reliable feature detection. Improving these detectors with real-world supervision is difficult because accurate real-world annotations are costly to obtain at scale. To address this limitation, we introduce a tracker-guided self-training framework that adapts a model pretrained on synthetic images to unlabeled real-world videos. Given measured robot joint states, an uncertainty-aware EKF recursively corrects the instrument pose and the observable joint angles by comparing projected model features with detected keypoints, shaft boundaries, and mask-derived cues. An RTS smoother subsequently refines the resulting trajectory, which is projected into pseudo-labels for fine-tuning the feature detector without laborious pose annotations. Experiments on real-world videos demonstrate consistent improvements from self-training across all evaluated keypoint metrics, and the resulting model outperforms prior approaches in both accuracy and runtime. The code and data will be released upon publication.
|
| 261 |
Robust 2D Traversability Mapping for Construction AMRs via Failure-Mode-Aware Fusion of LiDAR Geometry and Monocular Semantics
2610.05505
|
cs.CV
|
Manoj Karnekar, Om Mandhane, Gautham Ramkumar |
Autonomous Mobile Robots (AMRs) on active construction sites face severe navigational challenges: geometry-based traversability mapping (e.g., LiDAR) misses visually hazardous but geometrically flat surfaces like wet mud and ponding concrete, while abrupt geom...Autonomous Mobile Robots (AMRs) on active construction sites face severe navigational challenges: geometry-based traversability mapping (e.g., LiDAR) misses visually hazardous but geometrically flat surfaces like wet mud and ponding concrete, while abrupt geometry on drivable speed-breakers and inclines produces phantom obstacles. We propose a real-time, failure-mode-aware multimodal traversability pipeline on an NVIDIA Jetson AGX Orin, where LiDAR is the primary geometric safety estimate and monocular semantics act as a selective, class- and confidence-gated corrective signal. The representation retains distinct traversable classes, namely flat road, terrain, and rocky terrain, while flagging construction hazards. We also release a multimodal construction-site dataset from a custom AMR: four closed-loop ROS 2 sequences from two active sites (RGB, depth, LiDAR, IMU, GPS-RTK, odometry) plus 506 annotated frames across 28 semantic classes. By projecting LiDAR onto dense semantic masks, resolving sparsity via morphological in-painting, and applying failure-mode-aware fusion with Patchwork++, the system corrects complementary geometric failure modes for a local AMR costmap.
|
| 262 |
Rethinking Streaming-Perception Evaluation on Heterogeneous Edge Platforms
2610.05578
|
cs.CV
|
Misun Yu, Jinyoung Moon, Jemin Lee |
Multi-camera streaming perception is increasingly deployed on heterogeneous edge platforms shared with co-resident workloads, yet accelerator placement is often evaluated using isolated single-stream experiments and mean streaming average precision (sAP). Usin...Multi-camera streaming perception is increasingly deployed on heterogeneous edge platforms shared with co-resident workloads, yet accelerator placement is often evaluated using isolated single-stream experiments and mean streaming average precision (sAP). Using two end-to-end pipelines on a single GPU--NPU platform, we show that isolated evaluation can mis-rank deployment-time placement. Although the GPU pipeline is preferred in isolation, GPU-localized contention introduces deadline misses that make detections stale and can reverse the preferred placement before full GPU saturation. The NPU pipeline is less accurate than the GPU pipeline on small and medium objects in isolation, but nearly matches it on large objects. The largest absolute sAP losses in our latency and contention experiments occur for large objects. In our four-stream experiments, the preferred placement depends on which path becomes stale, and increasing GPU-side contention shifts the best placement from All-GPU to All-NPU. Under a GPU-saturating vision--language co-tenant, All-NPU achieves $5.2\times$ the worst-stream sAP of All-GPU. Because mean sAP can hide severe single-stream degradation, evaluation should report contention sweeps, deadline-miss rates on both paths, and worst-stream sAP alongside mean sAP.
|
| 263 |
DREAM: Dynamic Resolution Assignment For Multimodal Multi-agent Debate
2610.05615
|
cs.CVcs.CLcs.AI
|
Khanh-Binh Nguyen, Van Dai Do, Tien Anh Nguyen, Svetha Venkatesh, Hung Le |
Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents t...Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variation in the visual scale needed across samples and agents. In addition, these frameworks frequently suffer from groupthink, a phenomenon where agents prematurely abandon correct deductions to conform with confident but hallucinated peer responses. To address these bottlenecks, we introduce DREAM (Dynamic Resolution Assignment For Multimodal Multi-Agent Debate), which operates via two core components: (1) Dynamic Resolution Assignment, a zero-shot probe round where agents test multiple resolutions, quantify uncertainty using Average Normalized Log-Likelihood (ANLL), and use an adaptive threshold to assign each agent to its empirically optimal resolution; (2) Uncertainty-Guided Rollback Aggregation counters groupthink by tracking each agent's uncertainty over rounds and restoring early low-uncertainty answers overridden by group pressure. On six multimodal datasets, DREAM improves the accuracy-token trade-off over multi-agent debate baselines by 1.5-3.2% accuracy without dataset-specific tuning.
|
| 264 |
Atomic Visual Entailment: Enhancing Zero-Shot Vision-Language Reasoning through Atomic Fact Decomposition and Learned Selection
2610.05630
|
cs.CVcs.CL
|
Nallathambi Vethiappan, Derya Soydaner, Gijs Wijnholds |
Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypo...Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis. Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind. A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit. We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave. We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all. Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model. AVE also localises the visual evidence behind its prediction without region-level supervision. These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.
|
| 265 |
Visual Grounding Safety in Vision-Language Models
2610.05637
|
cs.CVcs.LGcs.AI
|
Erfan Shayegani, Kundan Krishna, Yue Dong, Nael Abu-Ghazaleh, Leon Gatys |
Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We...Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and robots can act on, yet safety alignment of this output channel has not been systematically analyzed. We study visual grounding safety by repurposing three safety benchmarks spanning direct harm (VLSU), social bias (BBQ-V), and situational safety (Asimov-2.0) into 15,401 matched pairs of harmful requests that differ only in the requested output: a free-text answer (VQA) or a grounding (point or bounding box). Across five VLMs, models that refuse a harmful request posed as a question often comply when the same request asks for a grounding: averaged over models, grounding refusal trails VQA refusal by 31-59 percentage points, depending on the domain, and safety system prompts do not close this gap. We propose a fine-tuning approach that combines grounding-form refusals with capability grounding data and self-distilled benign data to counter over-refusal. For Qwen3-VL-8B and VisionReasoner-7B, it improves grounding refusal by 77-95 percentage points on VLSU and BBQ-V and by 64-85 points on the held-out Asimov-2.0 domain, while also improving VQA refusal, preserving grounding capability, and keeping over-refusal limited. Representation analysis shows that fine-tuning moves harmful requests toward each model's refusal direction, most strongly for grounding, while leaving benign requests near the harmless reference.
|
| 266 |
Bayesian Data Augmentation for DNN Retraining with Binomial Outcomes in Vision-Based UAV Landing
2610.05674
|
cs.CV
|
Ashik E Rasul, Hyung-Jin Yoon |
In GPS-denied or cluttered urban environments, vision-based landing is essential for reliable UAV missions. Real-world landing sites are often unstructured and highly variable, requiring strong generalization by the perception system. Deep Neural Networks (DNN...In GPS-denied or cluttered urban environments, vision-based landing is essential for reliable UAV missions. Real-world landing sites are often unstructured and highly variable, requiring strong generalization by the perception system. Deep Neural Networks (DNNs) trained with synthetic data augmentation offer a scalable solution for learning landing-site features across diverse vehicle and environmental states. However, computationally expensive DNN retraining, along with challenging performance validation via test flights, limits exhaustive model fine-tuning and necessitates an optimized retraining pipeline. In this work, we deploy a Bayesian data augmentation framework integrated with a photorealistic simulator featuring high-fidelity vehicle dynamics to iteratively retrain the helipad detector DNN, maximizing landing performance as the objective function. We validate our framework with experiments in a photorealistic simulator under different environmental conditions and vehicle states, demonstrating improved landing performance and tighter confidence intervals on predicted landing outcomes.
|
| 267 |
From Pixels, Without Pre-training: Joint Generative and Self-Supervised Representation Learning in One Model
2610.05711
|
cs.CVcs.LG
|
Vicente Balmaseda, Ching-Long Lin, Tianbao Yang |
Strong image generation models are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders. While effective, generation then depends on supervision or pretraining: labels must be annotated, and encoders o...Strong image generation models are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders. While effective, generation then depends on supervision or pretraining: labels must be annotated, and encoders or autoencoders pretrained for the target domain. We study joint generative and self-supervised representation learning in a single model, enabling self-conditioned generation without labels or pretrained models. This is challenging because the objectives are mismatched: contrastive learning consumes clean augmented views and favors coarse, invariant semantics, while flow matching consumes noisy images and must preserve the fine detail and spatial layout that contrastive learning discards. We propose SCION (Self-conditioned Generation on Self-supervised representation), whose core is a single pixel-space encoder conditioned on the flow timestep and an embedding. For representation learning, this conditioning embedding is a learned global vector shared across images, with the encoder's [CLS] token yielding the semantic representation trained by the contrastive loss. For generative training, the conditioning embedding is the image's own [CLS] representation, while patch tokens pass through a decoder to predict the image. To sample without a reference image at inference, we jointly learn a prior over the embedding. Gradient-norm balancing and stop-gradient mechanisms enable joint optimization in one run. SCION is self-supervised and self-contained, with no labels or pretrained models. On ImageNet 256x256, with the JiT-B recipe and no representation guidance, SCION reaches 8.92 FID, surpassing class-unconditional iREPA, which aligns to pretrained DINOv2 (46.44), and RCG, which conditions on it (14.27). With JiT-L, SCION achieves 5.89 FID without guidance and 3.47 with representation guidance, outperforming RCG with the ADM recipe (6.24).
|
| 268 |
Dual-Rate Force-Image Control with Model-Based Orientation Limits for Robotic Ultrasound
2610.05839
|
cs.CV
|
Tyler Foster, Qiang Zhang, A B M Tahidul Haque, Anh Thu Nguyen |
Robotic ultrasound couples a high-rate contact-force loop with slower, delayed image feedback, so image-guided ultrasound probe rotation can perturb contact force before the resulting image response is observed. We derive a closed-form orientation-rate limit t...Robotic ultrasound couples a high-rate contact-force loop with slower, delayed image feedback, so image-guided ultrasound probe rotation can perturb contact force before the resulting image response is observed. We derive a closed-form orientation-rate limit that bounds the modeled rotation-induced estimated-force excursion over a finite horizon while accounting for disturbance rejection by the fast force loop. The limit depends on local contact stiffness, force-loop gains, a conservative rotation-to-force gain bound, the excursion budget, and the prediction horizon. We implement this model in a dual-rate controller with timestamp-based delay reconstruction and joint-torque-based force estimation, and evaluate it on a curved gelatin phantom using paired controller comparisons and component ablations. Relative to unconstrained image guidance, the proposed rate-limited controller reduced first-second root-mean-square (RMS) estimated-force error by 0.40 N while increasing cue-convergence time by 0.94 s. A fixed rate cap near the analytically predicted ceiling produced no resolvable difference in force error and converged 0.32 s faster, indicating that the principal practical value of the model is the rate-design rule rather than online prediction. Delay reconstruction had no resolvable effect at the tested latency. A single-subject popliteal scan demonstrated feasibility, although the image cue was noise-limited on heterogeneous tissue.
|
| 269 |
On Hyperparameter Tuning on the Test Set
2610.05902
|
cs.CVcs.LGcs.AI
|
Matteo Fregonara, Tom Viering, Jan van Gemert |
"Don't tune hyperparameters on the test set" is often stated in machine learning textbooks. Violating it is considered a cardinal sin that produces misleadingly optimistic results, corrupts benchmark integrity, and thus can even be interpreted as scientific fr..."Don't tune hyperparameters on the test set" is often stated in machine learning textbooks. Violating it is considered a cardinal sin that produces misleadingly optimistic results, corrupts benchmark integrity, and thus can even be interpreted as scientific fraud. Yet evidence suggests that test set hyperparameter tuning does occur in practice, making it all the more important to understand its actual consequences. So how bad is it, really? In this work we question this dogma and put it to an empirical test. We systematically study the magnitude of the performance inflation caused by tuning the hyperparameters on the test set for MNIST-1D, CIFAR-10, and three tasks from the GLUE benchmark. Our experiments show that while the effect is real and significant, it is frequently small relative to other sources of noise. In many cases, we find that tuning on the test set recovers exactly the same model as when tuning on the validation set. Most importantly, we find that the rankings of models remain essentially preserved after tuning on the test set and therefore that consistent test-set tuning may not invalidate benchmarks or model selection. Our results call for a more nuanced view of tuning hyperparameters on the test set, stimulating researchers to openly report test tuning.
|
| 270 |
Beyond Transport Cost: Routing Differences between Flow Matching and Optimal Transport
2610.05921
|
cs.CVcs.LG
|
Eungyeol Han, Jong-Seok Lee |
In generative models, Optimal Transport (OT) is used to improve Flow Matching (FM) by reducing noise-data coupling cost. However, different noise-to-output assignments can yield nearly equal costs, raising a key question. Is cost alone sufficient to guide coup...In generative models, Optimal Transport (OT) is used to improve Flow Matching (FM) by reducing noise-data coupling cost. However, different noise-to-output assignments can yield nearly equal costs, raising a key question. Is cost alone sufficient to guide coupling design? We address this question by separating transport cost from routing, i.e., the destination reached by each noise sample. We show numerically how FM and OT can differ in routing while remaining close in cost. We examine its consequences in learned neural FM. Using the exact FM routing as an oracle, we further construct a routing-aware training coupling and find that it yields a directionally consistent improvement in generation over a cost-matched, cost-only counterpart. Our findings highlight what cost minimization can overlook and motivate using both cost and routing to evaluate the design of OT-based FM couplings. Code will be released upon acceptance.
|
| 271 |
Structural Foundations of Nonlinear Systems with Unknown Inputs: The UID-Induced Normal Form and Minimal-Sensing Structure-from-Motion
2610.05939
|
cs.CV
|
Agostino Martinelli |
This paper establishes the first general structural solution to the problem of state estimation for nonlinear systems driven by unknown inputs. Building upon nonlinear unknown-input observability theory, we show that every such system admits a structurally equ...This paper establishes the first general structural solution to the problem of state estimation for nonlinear systems driven by unknown inputs. Building upon nonlinear unknown-input observability theory, we show that every such system admits a structurally equivalent representation, referred to as the UID-induced normal form. The proposed representation decomposes the information carried by the unknown inputs into two complementary components: unknown-input directions that are structurally decoupled from the observable dynamics and observable quantities that completely represent the unknown-input information affecting the observable dynamics. As a consequence, the UID-induced normal form provides a unified structural solution to unknown-input decoupling and unknown-input reconstruction, without requiring any model or stochastic assumption on the unknown inputs. The practical significance of the proposed framework is demonstrated through a previously unexplored minimal Structure-from-Motion configuration. The proposed representation enables recursive state estimation from only three point features and a single-axis gyroscope, allowing the recovery of the three-dimensional structure and camera motion up to an unknown global scale factor. Experiments on real-world data validate the proposed framework and demonstrate the feasibility of this minimal sensing configuration.
|
| 272 |
MEND: RL For Flow Models via Proximal Velocity Matching
2610.05954
|
cs.CVcs.LG
|
Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik |
Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We intr...Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.
|
| 273 |
Ultrasound Operator Guidance Using World Modeling and Retrieval Based Action Planning
2610.06008
|
cs.CVcs.LG
|
Noortje I. P. Schueler, Hans van Gorp, Ruud J. G. van Sloun |
Ultrasound is widely used, but acquisition quality is heavily dependent on the operator's knowledge and expertise. With demand for examinations outpacing the supply of trained sonographers, operator-guidance systems aim to close this gap by instructing a less ...Ultrasound is widely used, but acquisition quality is heavily dependent on the operator's knowledge and expertise. With demand for examinations outpacing the supply of trained sonographers, operator-guidance systems aim to close this gap by instructing a less trained user how to move the probe toward a target view. In this paper, we propose a retrieval-induced latent transition model for ultrasound acquisition dynamics, formulating ultrasound operator guidance as multi-step planning and retrieval in a world model. Using a V-JEPA 2.1 backbone, observations are first encoded into a latent space where anatomically related views lie close together. We then retrieve similar views from a reference database containing encoded latent states and corresponding probe positions and orientations. Rather than learning a parametric transition function, we directly use physically executed transitions from the database to establish our nonparametric, retrieval-induced transition model that supports receding-horizon planning. At deployment, guidance is generated from the live ultrasound image feed alone, without any probe tracking hardware. Applied to carotid ultrasound, the proposed planner reaches the target view in 86% of retrospective closed-loop episodes, versus 52% and 43% for representative baselines, outperforming both on every target view, including the challenging longitudinal internal and external carotid artery views. A prospective feasibility study on unseen volunteers, run in real time on a CPU using distillation, reaches 83% target-view reachability. Because planning is driven by proximity to any encodable goal latent, the same world model can navigate back to any previously acquired, patient-specific frame, supporting reproducible longitudinal imaging for e.g. perioperative or follow-up monitoring.
|
| 274 |
How well do routinely collected demographic and clinical variables aid point-of-care lung ultrasound TB classification
2610.06034
|
cs.CVcs.LG
|
Joshua M. Jansen van V\"uren, Christiaan M. Geldenhuys, Devendra S. Parihar, V\'eronique Suttels, Trevor Brokowski |
We consider the fusion of lung ultrasound images with routinely-collected clinical and demographic data for the purpose of automated tuberculosis (TB) screening using deep-learning. Such deep-learning based screening tools for TB could meaningfully support the...We consider the fusion of lung ultrasound images with routinely-collected clinical and demographic data for the purpose of automated tuberculosis (TB) screening using deep-learning. Such deep-learning based screening tools for TB could meaningfully support the health care system in Africa, where the burden of disease is severe and resources are constrained. Beginning with an established ResNet baseline for classification of lung ultrasound images, which achieves an area under the receiver operating characteristic (AUROC) curve of 0.91 [0.86,0.96] (95% CI), we consider the incorporation of the clinical and demographic data using three fusion approaches. We find that a simple average-based fusion of the output scores of separately-trained image and clinical data classifiers consistently matches or outperforms a more complex approach where the data is fused earlier and a combined classifier is trained. Fusing the image and the clinical classifiers in this way leads to a classifier with an overall AUROC of 0.95 [0.91,0.99] (specificity of 0.76 at sensitivity 0.93) which is an improvement of 4% absolute over the image-only baseline. We also find that greedy feature selection can be used to reduce the number of clinical and demographic inputs without sacrificing classification performance. Finally, when we differentiate between clinical and demographic data that are self-reported, that require some basic measurement or calculation, and that require a point-of-care (POC) test, we find the inclusion of the POC tests included in this study to be of minimal benefit to classification performance. We conclude that the incorporation of routinely-collected clinical and demographic data is a promising way to improve the performance of lung ultrasound based automatic classification.
|
| 275 |
On Impact of Loss Function on the Performance of Neural Networks in Melanoma Diagnosis
2610.06139
|
cs.CVcs.LGcs.AI
|
Morgan May, Pierpaolo Dondio, Simon Caton |
Melanoma is the deadliest type of skin cancer, whose early diagnosis is crucial for patients' survival. Image classification using deep learning models has shown promising results for melanoma diagnosis. However, the performance of these models on the melanoma...Melanoma is the deadliest type of skin cancer, whose early diagnosis is crucial for patients' survival. Image classification using deep learning models has shown promising results for melanoma diagnosis. However, the performance of these models on the melanoma datasets such as SIIM-ISIC melanoma classification dataset is a challenge due to the class imbalance. One of the methods to deal with this challenge is using loss function modifications. In this work, we have investigated the effect of different loss functions on the performance of deep neural networks. We trained these networks using focal loss, logit-adjusted softmax cross-entropy (CE) loss, and weighted softmax CE loss, and we report different metrics for evaluating performance and uncertainty calibration. Our results suggest that focal loss delivers a good combination of performance in terms of AUC and uncertainty calibration in terms of expected calibration error (ECE) simultaneously.
|
| 276 |
Impact of Data Augmentation on Confidence Calibration in Melanoma Classification
2610.06146
|
cs.CVcs.LGcs.AI
|
Morgan May, Simon Caton, Pierpaolo Dondio |
Accurately quantifying the predictive uncertainty or improving model calibration plays an important role in medical image classification, in particular in melanoma diagnosis, where accurate uncertainty quantification can have significant implications for patie...Accurately quantifying the predictive uncertainty or improving model calibration plays an important role in medical image classification, in particular in melanoma diagnosis, where accurate uncertainty quantification can have significant implications for patient care. One of the methods for calibration improvement is data augmentation. In addition, data augmentation as a method for synthetically increasing the size of the dataset has been proven to improve the performance of models trained on imbalanced datasets. However, the impact of data augmentation, as a transformation of a part of the original data, on calibration of models trained on imbalanced datasets, in particular in melanoma classification is under-explored. We train neural networks on SIIM-ISIC 2020 melanoma classification dataset under two conditions: with and without data augmentation, and compare the differences in AUC and expected calibration error (ECE) in both scenarios. Our results shows improvements in uncertainty calibration using different augmentation methods.
|
| 277 |
Loss-Invariant Projections as Passive Probes of Learned Representations
2610.06195
|
cs.CVcs.LG
|
Akshay Chandrasekhar, Pavlo Melnyk |
Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study this structure using $\textit{passive probes}$ that apply fixed, untrained, property-independent projections to representati...Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study this structure using $\textit{passive probes}$ that apply fixed, untrained, property-independent projections to representations as they evolve during training. We motivate this approach through the task of prediction on $S^2$ where equivalent vector and Hermitian parameterizations reveal an additional loss-invariant trace coordinate. This motivates a general construction in which fixed random projections serve as observers of learned features. Because the observer is loss-invariant and independent of the property being studied, changes in accessibility reflect changes in the representation relative to the fixed observer rather than adaptation of the observer itself. We show that ensembles of passive probes can directly reflect task-relevant information such as target alignment. Under our constructions, the accessibility of eventual difficulty evolves differently across tasks. It increases during training in the regression tasks of surface-normal estimation and image inpainting but remains near its initial level in image classification. Comparisons with learned linear probes further show that recoverability and passive accessibility can evolve differently during training. Together, these results show how passive probes can separately characterize changes in representation geometry and the accessibility of eventual task difficulty.
|
| 278 |
LeAVJEPA: A Minimalist Architecture for Audio-Visual Self-Supervised Learning
2610.06226
|
cs.CVcs.LG
|
Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin |
Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free ob...Prior audio-visual self-supervised learning methods rely on mechanisms such as EMA target encoders, prediction heads, reconstruction decoders, and contrastive losses. We introduce LeAVJEPA, the first audio-visual encoder trained under LeJEPA's collapse-free objective. A single early-fusion Vision Transformer processes audio, video, and joint audio-video inputs. Modality dropout treats a missing modality as another view of the same event, making cross-modal alignment implicit in the objective. The model aligns global embeddings with modality-specific local embeddings, and SIGReg prevents representational collapse. A controlled ablation identifies modality dropout as the key mechanism for audio-visual alignment. Despite the architectural simplicity, LeAVJEPA reaches 36.0 mAP on AudioSet-20K and 91.3% accuracy on ESC-50 under frozen evaluation. After fine-tuning, it reaches 61.1% accuracy on VGGSound, and its embeddings support zero-shot audio-visual retrieval.
|
| 279 |
Dual Variational Autoencoders for Efficient Sim-to-Real Transfer in Low-Cost Robotic Navigation
2610.06327
|
cs.CVcs.LG
|
\'Alvaro D\'iez (Department of Computer Science, Artificial Intelligence, University of Alicante), Fidel Aznar (Department of Computer Science, Artificial Intelligence |
Vision-based autonomous navigation for low-cost robots remains a fundamental challenge, primarily due to the significant gap between simulated training environments and real-world operational conditions. Direct policy transfer from simulation is often ineffect...Vision-based autonomous navigation for low-cost robots remains a fundamental challenge, primarily due to the significant gap between simulated training environments and real-world operational conditions. Direct policy transfer from simulation is often ineffective, while training exclusively on real data is impractical. We propose a hybrid transfer learning framework that effectively bridges the sim-to-real gap by combining domain randomization with feature-level domain adaptation. Our method employs a dual convolutional variational autoencoder architecture with a shared decoder, trained on an extensive set of 45225 simulated images and a minimal set of only 4556 real-world samples. This architecture learns a compact, common latent representation space that aligns the distributions of both domains. The adaptation process is further enhanced by two complementary data augmentation techniques designed to expand the limited real-world data. Experimental evaluation demonstrates that our method achieves an average success rate of almost 91% on image classification tasks for real-world indoor navigation, significantly outperforming both simulation-only and real-world-only training. We validate these findings through a direct, real-world deployment, where the proposed policy successfully guides a low-cost robot in a reactive exploration task. Furthermore, we validate the model's efficiency through a rigorous computational estimation, confirming its suitability for resource-constrained embedded platforms such as the Raspberry Pi 4 and NVIDIA Jetson Nano. This work presents a practical solution for developing effective and efficient navigation policies for low-cost robotic systems.
|
| 280 |
Improving Proactive AI Assistance with Hierarchical Procedural Understanding
2610.06505
|
cs.CVcs.LG
|
Jin-Seop Lee, TaeYeon Won, SeongJun Jung, JungHoon Kim, Boyang Albert Li |
Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust th...Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporting these capabilities requires training and evaluation data that reflect procedural structure and capture how guidance should adapt to task progress and user needs. However, existing datasets either focus on detection-based proactive understanding or provide procedural guidance at a fixed granularity. Fixed-granularity guidance provides limited information about fine-grained progress and broader procedural context, making it difficult to determine completion and adapt guidance granularity. To address these limitations, we introduce the ProactiveCoach suite, comprising ProactiveCoach-Instruct for training, ProactiveCoachBench for evaluation, and fine-tuned VLMs with an adaptive guidance system. ProactiveCoach-Instruct provides hierarchically structured guidance at the phase, step, and action levels for learning task progress and procedural context. ProactiveCoachBench evaluates whether models provide appropriate guidance at the right time across different guidance levels and adapt when the requested level changes. We fine-tune pretrained VLMs on ProactiveCoach-Instruct and demonstrate its effectiveness across backbones. Compared with fixed-granularity supervision, hierarchical supervision improves overall performance across backbones by up to 9.6%p. We further build an adaptive guidance system by combining our fine-tuned model with a lightweight guidance router. Without additional fine-tuning, our system outperforms the in-context adaptation baseline by 57.1%p across four guidance-level transitions. Our project page is available at https://jinsuby.github.io/ProactiveCoach/.
|
| 281 |
SimForcing: Distilling Simulation Motion Priors into Real-Domain Robot World Models
2610.06598
|
cs.CVcs.AI
|
Xiaodong Wang, Tianle Li, Chuanxin Song, Junliang Xie, Zhanmi Zhong |
Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearanc...Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation both as a source of transferable motion knowledge and as a controllable reference for prediction. First, we transfer motion knowledge from a simulation teacher through latent-motion distillation, aligning temporal changes in latent space to internalize motion priors while mitigating the influence of appearance differences. Second, we introduce multi-block simulation conditioning with condition dropout to exploit predicted simulation trajectories without relying excessively on their accuracy. Our simulation-conditioning classifier-free guidance scheme unifies these two ideas by balancing predictions based on internalized motion knowledge with those additionally guided by simulation latents. The jointly trained student generates both simulation conditions and real-domain videos, requiring no additional world model at inference. On Bridge, SimForcing achieves the best PSNR, SSIM, LPIPS, and FVD among the compared methods without external embodied pretraining. Evaluation on InternData-A1 further supports its applicability across robot datasets. Moreover, using our trained world model to initialize a vision-language-action model improves LIBERO success, suggesting its utility for downstream policy learning. \url{https://github.com/Wang-Xiaodong1899/SimForcing}
|
| 282 |
AffordCraft: Scalable Construction of Task-Ready Simulation Assets from Single Images
2610.06643
|
cs.CVcs.AI
|
Haoyun Yang, Xueyang Zhou, Ziyi Xie, Yongchao Chen |
Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing methods recover this structure anew ...Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing methods recover this structure anew for every image: generative models predict parts and joints that mostly fail to settle or move in simulation, and general-purpose agents need a long session of model calls for each photograph. AffordCraft builds such an asset from a single RGB image and a task instruction by retrieval instead of generation: it locates the object and the part to operate, selects a matching entry from a library of articulated assets, and fits it to the image while keeping its parts and joints intact. Without any box or mask marking the object, AffordCraft produces a physically valid asset for 1,703 of 2,000 photographs from 31 categories. Five generative methods pass on at most 45% of the same photographs and, at the median, need 10 to 78 times our GPU time per valid asset. On 50 cluttered images, 162 of 237 annotated objects pass the same physical test after automatic detection. Growing the library from 141 to 11,372 entries needs no change to the method and raises category coverage from 46% to 100% and the share of selections with the requested label from 18% to 51%. We also build manipulation tasks from the constructed assets, both with single objects and in composed scenes; policies trained on scripted demonstrations complete both kinds of tasks from initial states unseen in training.
|
| 283 |
Detecting Nighttime Anomalies from NASA Black Marble Using a Generalized Spatio-Temporally Robust Framework of Machine Leaning Ensembles
2610.06674
|
cs.CVcs.LG
|
Srija Chakraborty |
Nighttime lights from NASA's Black Marble product suite capture thermal and light emission signals from anomalous events including fires, volcanic eruptions, and gas flaring. Existing detection approaches rely primarily on thermal bands, limiting sensitivity t...Nighttime lights from NASA's Black Marble product suite capture thermal and light emission signals from anomalous events including fires, volcanic eruptions, and gas flaring. Existing detection approaches rely primarily on thermal bands, limiting sensitivity to weaker signals. We propose a novel machine learning framework that jointly models Black Marble M-band and Day/Night Band (DNB) signals to derive a generalized, spatio-temporally robust ensemble of anomaly detectors. The framework iteratively builds detectors that scale across regions, seasons, anomaly classes, and extends over land and ocean. Detection sets at varying confidence levels are derived based on relevant bands and detector agreement. The approach improves true detection rate while reducing spurious detections and results demonstrate strong generalizability with applications in natural hazard monitoring and energy extraction.
|
| 284 |
PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data
2610.06825
|
cs.CVcs.CL
|
Yaohui Zhang, Binxu Li, Haoyi Duan, Jiacheng Miao, Yixin Wang |
Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover...Scientific figures often encode quantitative results that are not readily available in machine-readable form, making accurate plot digitization important for verifying and reusing published findings. Yet it remains unclear how accurately current models recover plotted values from real scientific figures, as existing benchmarks rely largely on synthetic charts or cover only a limited range of chart types. We introduce PlotGround, an automated pipeline for building plot digitization benchmarks from real scientific figures and their author-released source data. PlotGround maps figures to source tables, identifies reconstructable panels, and generates quantitative questions with source-grounded reference values. We use PlotGround to construct PlotGround-1k, a human-verified benchmark of 1,119 questions from 1,066 bioRxiv preprints. Across sixteen multimodal models, the best reaches 87.5% accuracy at a $\pm 5\%$ relative-error tolerance. Tightening the tolerance to $\pm 2\%$ lowers every model's accuracy by 11-24 percentage points, revealing a gap between approximate visual reading and precise quantitative recovery. PlotGround's paired figure-source structure lets us compare how accurately the same values are recovered from figures and from source tables. Providing source tables instead of figures raises a coding agent's accuracy from 90.0% to 97.4% while cutting cost by 72%.
|
| 285 |
InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation
2610.06850
|
cs.CV
|
Yucheng Zhang, Sirui Xu, Jinhong Li, Liuyu Bian, Anatulya Nandi |
Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion ...Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.
|
| 286 |
Event-based Continuous Color Video Decompression from Single Frames
2312.00113
|
cs.CV
|
Ziyun Wang, Friedhelm Hamann, Kenneth Chaney, Wen Jiang, Guillermo Gallego |
We present ContinuityCam, a novel approach to generate a continuous video from a single static RGB image and an event camera stream. Conventional cameras struggle with high-speed motion capture due to bandwidth and dynamic range limitations. Event cameras are ...We present ContinuityCam, a novel approach to generate a continuous video from a single static RGB image and an event camera stream. Conventional cameras struggle with high-speed motion capture due to bandwidth and dynamic range limitations. Event cameras are ideal sensors to solve this problem because they encode compressed change information at high temporal resolution. In this work, we tackle the problem of event-based continuous color video decompression, pairing single static color frames and event data to reconstruct temporally continuous videos. Our approach combines continuous long-range motion modeling with a neural synthesis model, enabling frame prediction at arbitrary times within the events. Our method only requires an initial image, thus increasing the robustness to sudden motions, light changes, minimizing the prediction latency, and decreasing bandwidth usage. We also introduce a novel single-lens beamsplitter setup that acquires aligned images and events, and a novel and challenging Event Extreme Decompression Dataset (E2D2) that tests the method in various lighting and motion profiles. We thoroughly evaluate our method by benchmarking color frame reconstruction, outperforming the baseline methods by 3.61 dB in PSNR and by 33% decrease in LPIPS, as well as showing superior results on two downstream tasks.
|
| 287 |
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models
2312.01886
|
cs.CV
|
Xunguang Wang, Pingchuan Ma, Zhenlan Ji, Zongjie Li, Shuai Wang |
Large vision-language models (LVLMs) have demonstrated their incredible capability in visual question answering. However, this rich visual interaction also makes LVLMs vulnerable to adversarial examples. In this paper, we formulate a novel and practical target...Large vision-language models (LVLMs) have demonstrated their incredible capability in visual question answering. However, this rich visual interaction also makes LVLMs vulnerable to adversarial examples. In this paper, we formulate a novel and practical targeted attack scenario that the adversary knows only the vision encoder of the victim LVLM, without the knowledge of its prompts and its underlying large language model. This practical setting poses challenges to the cross-prompt and cross-model transferability of targeted adversarial attack, which aims to confuse the LVLM to output a response that is semantically similar to the attacker's chosen target text. To this end, we propose an instruction-tuned targeted attack (dubbed InstructTA) to deliver the targeted adversarial attack on LVLMs with high transferability. Initially, we utilize a public text-to-image generative model to reverse the target response into a target image, and employ GPT-4 to infer a reasonable instruction $\boldsymbol{p}^\prime$ from the target response. We then form a local surrogate model (sharing the same vision encoder with the victim LVLM) to extract instruction-aware features of an adversarial image example and the target image, and minimize the distance between these two features to optimize the adversarial example. To further improve the transferability with instruction tuning, we augment the instruction $\boldsymbol{p}^\prime$ with instructions paraphrased from GPT-4. Extensive experiments on 6 victim LVLMs demonstrate the superiority of our proposed method in targeted attack performance and transferability. In particular, InstructTA achieves an attack success rate of 51.9% on BLIP-2, outperforming the strongest baseline by 10.5%, and consistently yields the highest attack success rates across all evaluated models. The code is available at https://github.com/xunguangwang/InstructTA.
|
| 288 |
Towards Subject-Oriented Video Captioning via User-Specified Targets
2312.13330
|
cs.CV
|
Yunchuan Ma, Chang Teng, Guorong Li, Yuankai Qi, Laiyu Qing |
Describing video content according to users' needs is a long-held goal. Although existing video captioning methods have made significant progress, the generated captions may not focus on the entity that users are particularly interested in. To address this pro...Describing video content according to users' needs is a long-held goal. Although existing video captioning methods have made significant progress, the generated captions may not focus on the entity that users are particularly interested in. To address this problem, we propose a new video captioning task, Subject-Oriented Video Captioning (SOVC), which aims to allow users to specify the describing target via a bounding box. To support this task, we construct two subject-oriented video captioning datasets based on two widely used video captioning datasets: MSVD and MSRVTT, by annotating subjects in each video for each caption. These datasets pave the way for describing users' interested targets. To tackle this task, we introduce a method tailored to this task, named SOVCNet. It consists of two key components: a subject-oriented sampling module that samples frames related to the subject to minimize irrelevant information; and a subject-oriented encoding module that utilizes the subject areas as hard prompts and integrates learnable soft prompts, enhancing the model's focus on the subject's activities and facilitating adaptation to the downstream generation task. Extensive experimental results demonstrate the effectiveness of our method on this new task.
|
| 289 |
XS-VID: A Large-Scale Benchmark for Small Object Detection and Tracking in Videos
2407.18137
|
cs.CV
|
Jiahao Guo, Ziyang Xu, Lianjun Wu, Fei Gao, Wenyu Liu |
Small object detection and tracking in videos remain critical yet underexplored challenges in computer vision, particularly for applications such as public safety, aerial surveillance, and autonomous driving. Existing benchmarks offer limited support due to li...Small object detection and tracking in videos remain critical yet underexplored challenges in computer vision, particularly for applications such as public safety, aerial surveillance, and autonomous driving. Existing benchmarks offer limited support due to limited numbers of small objects, constrained category diversity, and narrow scene coverage. To address these limitations, we introduce XS-VID, a large-scale video benchmark comprising 223K frames and 1.4M annotated bounding boxes across 374 video sequences spanning diverse scene types. XS-VID provides extensive coverage of small-object scales, particularly for extremely small ($0\sim12^2$ pixels) and small ($12^2\sim20^2$ pixels) objects, which collectively constitute over 55% of all annotations. For systematic evaluation, we establish three dedicated tracks: Detection, multiple object tracking (MOT), and single object tracking (SOT), and extensively test the existing state-of-the-art methods on each. The experimental results indicate that existing methods face significant challenges with XS-VID, mainly stemming from insufficient modeling of spatiotemporal features at small scales. To tackle these challenges, we propose a lightweight, high-precision detection framework dubbed YOLOFT. It enhances small-object feature representation and spatiotemporal integration while preserving high detection speed, thereby achieving improved accuracy and robustness on both the XS-VID and VisDrone benchmarks. Our dataset and code are publicly available at https://gjhhust.github.io/XS-VID/, providing a solid foundation for future research on small-object detection and tracking in videos.
|
| 290 |
Motif Channel Opened in a White-Box: Stereo Matching via Motif Correlation Graph
2411.12426
|
cs.CV
|
Ziyang Chen, Yongjun Zhang, Wenting Li, Bingshu Wang, Yong Zhao |
Real-world applications of stereo matching, such as autonomous driving, place stringent demands on both safety and accuracy. However, learning-based stereo matching methods inherently suffer from the loss of geometric structures in certain feature channels, cr...Real-world applications of stereo matching, such as autonomous driving, place stringent demands on both safety and accuracy. However, learning-based stereo matching methods inherently suffer from the loss of geometric structures in certain feature channels, creating a bottleneck in achieving precise detail matching. Additionally, these methods lack interpretability due to the black-box nature of deep learning. In this paper, we propose MoCha-V2, a novel learning-based paradigm for stereo matching. MoCha-V2 introduces the Motif Correlation Graph (MCG) to capture recurring textures, which are referred to as ``motifs" within feature channels. These motifs reconstruct geometric structures and are learned in a more interpretable way. Subsequently, we integrate features from multiple frequency domains through wavelet inverse transformation. The resulting motif features are utilized to restore geometric structures in the stereo matching process. Experimental results demonstrate the effectiveness of MoCha-V2. MoCha-V2 achieved 1st place on the Middlebury benchmark at the time of its release. Code is available at https://github.com/ZYangChen/MoCha-Stereo.
|
| 291 |
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
2412.10510
|
cs.CVcs.CL
|
Tobias Braun, Mark Rothermel, Marcus Rohrbach, Anna Rohrbach |
The proliferation of disinformation demands reliable and scalable fact-checking solutions. We present Dynamic Evidence-based FAct-checking with Multimodal Experts (DEFAME), a modular, zero-shot MLLM pipeline for open-domain, text-image claim verification. DEFA...The proliferation of disinformation demands reliable and scalable fact-checking solutions. We present Dynamic Evidence-based FAct-checking with Multimodal Experts (DEFAME), a modular, zero-shot MLLM pipeline for open-domain, text-image claim verification. DEFAME operates in a six-stage process, dynamically selecting the tools and search depth to extract and evaluate textual and visual evidence. Unlike prior approaches that are text-only, lack explainability, or rely solely on parametric knowledge, DEFAME performs end-to-end verification, accounting for images in claims and evidence while generating structured, multimodal reports. Evaluation on the popular benchmarks VERITE, AVerITeC, and MOCHEG shows that DEFAME surpasses all previous methods, establishing itself as the new state-of-the-art fact-checking system for uni- and multimodal fact-checking. Moreover, we introduce a new multimodal benchmark, ClaimReview2024+, featuring claims after the knowledge cutoff of GPT-4o, avoiding data leakage. Here, DEFAME drastically outperforms the GPT-4o baselines, showing temporal generalizability and the potential for real-time fact-checking.
|
| 292 |
Open-World Panoptic Segmentation
2412.12740
|
cs.CV
|
Matteo Sodano, Federico Magistri, Jens Behley, Cyrill Stachniss |
Robots need to be able to understand their surroundings in order to operate safely and robustly, and to interact with the surrounding environment. Robots deployed in unconstrained real-world scenarios must additionally be able to deal with novel situations and...Robots need to be able to understand their surroundings in order to operate safely and robustly, and to interact with the surrounding environment. Robots deployed in unconstrained real-world scenarios must additionally be able to deal with novel situations and objects that have never been seen before. In this article, we tackle the problem of open-world panoptic segmentation, i.e., the task of discovering new semantic categories and new object instances at test time, while enforcing consistency among the categories that we incrementally discover. We present Con2MAV, a general method for open-world panoptic segmentation. Experiments across a wide range of datasets, from road scenes to underwater environments, highlight its compelling capabilities in open-world segmentation and its competitive performance on known classes. We will open-source the implementation of our approach upon acceptance. In addition, we propose PANIC (Panoptic ANomalies In Context), a benchmark for evaluating open-world segmentation tasks in autonomous driving scenarios. This dataset, recorded with a multi-modal sensor suite mounted on a car, and then manually annotated, provides high-quality, pixel-wise annotations of anomalous objects at both semantic and instance level. PANIC contains 800 images, more than 50 unknown classes, i.e., classes that do not appear in the training set, and over 4,000 object instances, providing a comprehensive benchmark for evaluating open-world segmentation methods in autonomous driving scenarios. We provide competitions for multiple open-world segmentation tasks on a hidden test set. Our dataset and competitions are available at https://www.ipb.uni-bonn.de/data/panic.
|
| 293 |
MIND: Microstructure INverse Design with Generative Hybrid Neural Representation
2502.02607
|
cs.CVcs.LG
|
Tianyang Xue, Longdu Liu, Lin Lu, Paul Henderson, Pengbin Tang |
The inverse design of microstructures plays a pivotal role in optimizing metamaterials with specific, targeted physical properties. While traditional forward design methods are constrained by their inability to explore the vast combinatorial design space, inve...The inverse design of microstructures plays a pivotal role in optimizing metamaterials with specific, targeted physical properties. While traditional forward design methods are constrained by their inability to explore the vast combinatorial design space, inverse design offers a compelling alternative by directly generating structures that fulfill predefined performance criteria. However, achieving precise control over both geometry and material properties remains a significant challenge due to their intricate interdependence. Existing approaches, which typically rely on voxel or parametric representations, often limit design flexibility and structural diversity. In this work, we present a novel generative model that integrates latent diffusion with Holoplane, an advanced hybrid neural representation that simultaneously encodes both geometric and physical properties. This combination ensures superior alignment between geometry and properties. Our approach generalizes across multiple microstructure classes, enabling the generation of diverse, tileable microstructures with significantly improved property accuracy and enhanced control over geometric validity, surpassing the performance of existing methods. We introduce a multi-class dataset encompassing a variety of geometric morphologies, including truss, shell, tube, and plate structures, to train and validate our model. Experimental results demonstrate the model's ability to generate microstructures that meet target properties, maintain geometric validity, and integrate seamlessly into complex assemblies. Additionally, we explore the potential of our framework through the generation of new microstructures, cross-class interpolation, and the infilling of heterogeneous microstructures. Code and data for this paper are at https://github.com/TimHsue/MIND.
|
| 294 |
PerCoV2: Ultra-Low Bit-Rate Perceptual Image Compression via Query-Based 1D Multimodal Image Tokens
2503.09368
|
cs.CV
|
Nikolai K\"orber, Eduard Kromer, Andreas Siebert, Sascha Hauke, Daniel Mueller-Gritschneder |
Despite recent progress in learned image compression, current image codecs still struggle to maintain realistic reconstructions at low bit-rates, often producing structured artifacts such as grid patterns or repetitive textures, even when trained with perceptu...Despite recent progress in learned image compression, current image codecs still struggle to maintain realistic reconstructions at low bit-rates, often producing structured artifacts such as grid patterns or repetitive textures, even when trained with perceptual or adversarial losses. We introduce PerCoV2, an ultra-low bit-rate perceptual image compression system that unifies semantic tokenization, flow-based generation, and learned entropy modeling within a single framework. Building on the fully open flow-based SANA architecture, PerCoV2 introduces a novel resolution-adaptive 1D query-based tokenizer that produces compact semantic image tokens with a dual role in flow matching: providing a data-dependent reconstruction prior for initialization and a conditioning signal for flow-based refinement. By explicitly decoupling semantic representation from perceptual generation, our dual representation simplifies the flow-based learning objective, leading to more stable optimization and improved perceptual compression performance. PerCoV2 further introduces a dedicated 1D masked entropy model to improve rate efficiency and optional decoder-side multimodal enhancement via a vision-language model (Molmo) without increasing the transmitted bit budget. On MSCOCO-30k, PerCoV2 achieves state-of-the-art statistical fidelity, measured by FID and KID, across ultra-low and extreme bit-rates (0.0015-0.025 bpp). When trained solely on the general-purpose SA-1B dataset, PerCoV2 further demonstrates strong zero-shot generalization to widely adopted high-resolution benchmarks, including DIV2K and CLIC 2020, achieving competitive statistical fidelity with the current leading method, AEIC-ME. Finally, we introduce PerCoV2-distilled, a practical single-step variant derived from multi-step flow matching that accelerates decoding by 5.37x over PerCoV1, while preserving perceptual compression performance.
|
| 295 |
MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation
2503.09950
|
cs.CVcs.LGcs.AI
|
Yuxiang Fu, Qi Yan, Lele Wang, Ke Li, Renjie Liao |
In this paper, we address the problem of human trajectory forecasting, which aims to predict the inherently multi-modal future movements of humans based on their past trajectories and other contextual cues. We propose a novel motion prediction conditional flow...In this paper, we address the problem of human trajectory forecasting, which aims to predict the inherently multi-modal future movements of humans based on their past trajectories and other contextual cues. We propose a novel motion prediction conditional flow matching model, termed MoFlow, to predict K-shot future trajectories for all agents in a given scene. We design a novel flow matching loss function that not only ensures at least one of the $K$ sets of future trajectories is accurate but also encourages all $K$ sets of future trajectories to be diverse and plausible. Furthermore, by leveraging the implicit maximum likelihood estimation (IMLE), we propose a novel distillation method for flow models that only requires samples from the teacher model. Extensive experiments on the real-world datasets, including SportVU NBA games, ETH-UCY, and SDD, demonstrate that both our teacher flow model and the IMLE-distilled student model achieve state-of-the-art performance. These models can generate diverse trajectories that are physically and socially plausible. Moreover, our one-step student model is $\textbf{100}$ times faster than the teacher flow model during sampling. The code, model, and data are available at our project page: https://moflow-imle.github.io
|
| 296 |
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
2505.18757
|
cs.CV
|
Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao, Shijian Lu |
Visual token pruning, which aims to compress and prune redundant visual tokens, plays a critical role in efficient inference with large vision-language models (LVLMs). However, existing methods fail to disentangle intra-modal visual redundancy from cross-modal...Visual token pruning, which aims to compress and prune redundant visual tokens, plays a critical role in efficient inference with large vision-language models (LVLMs). However, existing methods fail to disentangle intra-modal visual redundancy from cross-modal redundancy between vision and language. We show that visual token diversity and task-specific token relevance are two crucial yet orthogonal factors that complement each other in conveying useful information and should therefore be treated separately for more effective visual token pruning. Building upon this insight, we design TODRE, a two-stage and training-free framework that incorporates Token Diversity and task RElevance for effective token compression and efficient LVLM inference. Instead of pruning redundant tokens, we introduce a greedy max-sum diversification algorithm that selects and retains a subset of diverse and representative visual tokens after the vision encoder. On top of that, ToDRE leverages an ``information migration'' mechanism to eliminate task-irrelevant visual tokens within certain decoder layers of the large language model (LLM), further improving token pruning and LVLM inference. Extensive experiments show that ToDRE prunes 90% of visual tokens after the vision encoder as well as all visual tokens in certain LLM decoder layers, leading to a 2.6x speed-up in total inference time while maintaining 95.0% model performance plus excellent model compatibility. The code is available at: \href{https://github.com/Yrdal3910/ToDRE}{this https URL}.
|
| 297 |
Deep Learning Reforms Image Matching: A Survey and Outlook
2506.04619
|
cs.CV
|
Shihua Zhang, Zizhuo Li, Kaining Zhang, Yifan Lu, Yuxin Deng |
Image matching, which establishes correspondences between two images to recover 3D structure and camera geometry, is a cornerstone of computer vision and underpins a wide range of applications, including visual localization, 3D reconstruction, and simultaneous...Image matching, which establishes correspondences between two images to recover 3D structure and camera geometry, is a cornerstone of computer vision and underpins a wide range of applications, including visual localization, 3D reconstruction, and simultaneous localization and mapping (SLAM). Traditional pipelines, composed of a detector-descriptor, a feature matcher, an outlier filter, and a geometric estimator, falter in challenging scenarios. Recent advances in deep learning have substantially improved both their robustness and accuracy. This survey reviews how deep learning has progressively transformed the classical image matching pipeline. Our taxonomy is aligned with the traditional pipeline and covers two aspects: i) replacing individual steps with learnable alternatives, including learnable detector-descriptors, outlier filters, and geometric estimators; and ii) merging multiple steps into end-to-end learnable modules, including middle-end sparse matchers, end-to-end semi-dense/dense matchers, and pose regressors. We first examine the design principles, advantages, and limitations of both aspects, and then benchmark representative methods on relative pose estimation, homography estimation, matching accuracy assessment, visual localization, and 3D reconstruction. Finally, we discuss open challenges and directions for future research. By systematically categorizing and evaluating learning-based methods, this survey offers a clear overview of how image matching is evolving and where further progress is needed. The project repository is available at https://github.com/ZizhuoLi/awesome-image-matching-survey.
|
| 298 |
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
2507.07297
|
cs.CV
|
Chengfei Wu, Ronald Seoh, Bingxuan Li, Liqiang Zhang, Fengrong Han |
Recent advances in large vision-language models have led to impressive performance in visual question answering and multimodal reasoning. However, it remains unclear whether these models genuinely perform grounded visual reasoning or rely on superficial patter...Recent advances in large vision-language models have led to impressive performance in visual question answering and multimodal reasoning. However, it remains unclear whether these models genuinely perform grounded visual reasoning or rely on superficial patterns and dataset biases. In this work, we introduce MagiC, a comprehensive benchmark designed to evaluate grounded multimodal cognition, assessing not only answer accuracy but also the quality of step-by-step reasoning and its alignment with relevant visual evidence. Our benchmark includes approximately 5,500 weakly supervised QA examples generated from strong model outputs and 900 human-curated examples with fine-grained annotations, including answers, rationales, and bounding box groundings. We evaluate 15 vision-language models ranging from 7B to 70B parameters across four dimensions: final answer correctness, reasoning validity, grounding fidelity, and self-correction ability. MagiC further includes diagnostic settings to probe model robustness under adversarial visual cues and assess their capacity for introspective error correction. We introduce new metrics such as MagiScore and StepSense, and provide comprehensive analyses that reveal key limitations and opportunities in current approaches to grounded visual reasoning.
|
| 299 |
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage
2507.09200
|
cs.CV
|
Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren, Alper Yilmaz, Minh-Triet Tran |
The rapid proliferation of video in applications such as autonomous driving, surveillance, and sports analytics necessitates robust methods for dynamic scene understanding. Despite advances in static scene graph generation and early attempts at video scene gra...The rapid proliferation of video in applications such as autonomous driving, surveillance, and sports analytics necessitates robust methods for dynamic scene understanding. Despite advances in static scene graph generation and early attempts at video scene graph generation, previous methods often suffer from fragmented representations, failing to capture fine-grained spatial details and long-range temporal dependencies simultaneously. To address these limitations, we introduce the Temporal Hierarchical Cyclic Scene Graph (THYME) approach, which synergistically integrates hierarchical feature aggregation with cyclic temporal refinement to address these limitations. In particular, THYME effectively models multi-scale spatial context and enforces temporal consistency across frames, yielding more accurate and coherent scene graphs. In addition, we present AeroEye-v1.0, a novel aerial video dataset enriched with five types of interactivity that overcome the constraints of existing datasets and provide a comprehensive benchmark for dynamic scene graph generation. Empirically, extensive experiments on ASPIRe and AeroEye-v1.0 demonstrate that the proposed THYME approach outperforms state-of-the-art methods, offering improved scene understanding in ground-view and aerial scenarios.
|
| 300 |
A Multimodal Sequence-to-Sequence Model for Cross-Subject Prediction of Brain Responses to Naturalistic Stimuli
2507.18104
|
cs.CV
|
Qianyi He, Monica D. Rosenberg, Yuan Chang Leong |
Brain encoding models predict time-resolved neural activity from computational representations of ongoing experience, providing a principled framework for testing how information is represented and transformed across cortical systems. Naturalistic audiovisual ...Brain encoding models predict time-resolved neural activity from computational representations of ongoing experience, providing a principled framework for testing how information is represented and transformed across cortical systems. Naturalistic audiovisual narratives are a particularly rich but challenging testbed for these models, requiring integration of multimodal inputs over long temporal horizons and generalization across individuals with substantial response variability. We introduce a multimodal sequence-to-sequence Transformer with a hybrid cross-subject parameterization that predicts cortex-wide parcel-wise fMRI time series autoregressively from visual, audio, language, and vision--language representations. We evaluate the approach on data from the Courtois NeuroMod project, where four deeply-sampled participants viewed six seasons of Friends and four feature-length films during fMRI. Sequence-to-sequence temporal modeling yields consistent improvements over single-frame prediction across cortical networks, with gains extending to novel stimuli. A hybrid architecture that pairs a shared stimulus encoder with lightweight subject-specific decoder components outperforms both fully shared and fully individual models, indicating complementary advantages of learning shared stimulus representations across subjects and fitting individual neural readouts. Finally, we show that in data-scarce settings, hybrid models can be personalized to new individuals with limited fMRI data, demonstrating that multi-subject pretraining serves as a strong inductive prior for building individual-specific encoding models. Together, these results indicate that combining multimodal sequence modeling with a hybrid cross-subject architecture offers a scalable framework for personalized brain encoding under naturalistic conditions.
|
| 301 |
SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting
2508.01802
|
cs.CV
|
Atom Scott, Ikuma Uchida, Kento Kuroda, Yufi Kim, Keisuke Fujii |
Soccer analytics draws on two kinds of information: spatio-temporal data describing where players and the ball are, and event data describing what they do. Public datasets offer them apart, or together only on broadcast footage that leaves players outside the ...Soccer analytics draws on two kinds of information: spatio-temporal data describing where players and the ball are, and event data describing what they do. Public datasets offer them apart, or together only on broadcast footage that leaves players outside the frame unobserved. SoccerTrack v2 combines continuous full-pitch video, long player trajectories and actor-linked events in one resource: ten university-level matches, 932 minutes of fixed-camera 4K panoramic video, annotated per frame with metric pitch coordinates, jersey numbers and persistent identities, roles and team sides for all players, and with ball action events in twelve classes, linked to the acting players through the same identifiers used in the trajectories. We fix a match-level split and report baselines for two tasks. For game state reconstruction, we run a full pipeline over all twenty halves and find that GS-HOTA scores degrade as sequence length increases. For ball action spotting, we train a model on the player trajectories, with and without the ball track. The data, the split and the evaluation tooling are released so that both tasks can be developed and compared at match length on the same footage.
|
| 302 |
AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference via Dynamical Text Guidance
2508.06084
|
cs.CV
|
Weichen Zhang, Zhui Zhu, Ningbo Li, Shilong Tao, Hongzi Zhu |
Vision-language models (VLMs) have achieved impressive performance on multimodal inference tasks, but the cost remains a significant challenge due to the large number of vision tokens processed during the prefill stage. Existing token pruning methods often rel...Vision-language models (VLMs) have achieved impressive performance on multimodal inference tasks, but the cost remains a significant challenge due to the large number of vision tokens processed during the prefill stage. Existing token pruning methods often rely on utilizing the static attention patterns directly, failing to exploit the dynamic internal signals within VLMs. To address the issue, we propose AdaptInfer, a novel plug-and-play framework for vision token pruning. First, we introduce a dynamic text-guided pruning mechanism that construct soft priors over text-token importance on each pruning layer, allowing more informed scoring of vision tokens at each stage. Second, we observe a highly consistent distribution of cross-modal attention shifts by architecture, which inspires us to introduce a efficient data-driven schedule which determines the pruning locations automatically. Experimental results have verified the effectiveness and generalization of the proposed method. Under the same token budget, AdaptInfer surpasses existing approaches in accuracy. For instance, AdaptInfer maintains averagely 99.4% accuracy on Qwen2-VL while 70% of the prefilling vision token overhead is reduced. The source code is available on: https://github.com/weiczh02/AdaptInfer-base.
|
| 303 |
Continual Action Quality Assessment via Adaptive Manifold-Aligned Graph Regularization
2510.06842
|
cs.CV
|
Kanglei Zhou, Qingyi Pan, Xingxing Zhang, Hubert P. H. Shum, Frederick W. B. Li |
Action Quality Assessment (AQA) quantifies human actions in videos, supporting applications in sports scoring, rehabilitation, and skill evaluation. A major challenge lies in the non-stationary nature of quality distributions in real-world scenarios, which lim...Action Quality Assessment (AQA) quantifies human actions in videos, supporting applications in sports scoring, rehabilitation, and skill evaluation. A major challenge lies in the non-stationary nature of quality distributions in real-world scenarios, which limits the generalization of conventional methods. We introduce Continual AQA (CAQA), which equips AQA with Continual Learning (CL) capabilities to handle evolving distributions while mitigating catastrophic forgetting. Although parameter-efficient fine-tuning of pretrained models has shown promise in continual learning, our empirical study shows that the evaluated adapter-based PEFT setting provides less effective downstream adaptation than FPFT for fine-grained AQA. Our empirical and theoretical analyses reveal two insights: (i) sufficiently expressive backbone adaptation is important for bridging the upstream--downstream representation gap; yet (ii) uncontrolled FPFT may induce overfitting and feature manifold shift, thereby aggravating forgetting. To address this, we propose Adaptive Manifold-Aligned Graph Regularization (MAGR++), which couples backbone fine-tuning that stabilizes shallow layers while adapting deeper ones with a two-step feature rectification pipeline: a manifold projector to translate deviated historical features into the current representation space, and a graph regularizer to align local and global distributions. We construct four CAQA benchmarks from three datasets with tailored evaluation protocols and strong baselines, enabling systematic cross-dataset comparison. Extensive experiments show that MAGR++ achieves state-of-the-art performance, with average correlation gains of 3.6% offline and 12.2% online over the strongest baseline, confirming its robustness and effectiveness.
|
| 304 |
MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models
2510.13276
|
cs.CVcs.CL
|
Keyan Zhou, Zecheng Tang, Lingfeng Ming, Qiguang Chen, Wangjie You |
The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for ...The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness in multimodal settings remain limited to short contexts. To bridge this gap, we introduce MMLongCite, the first benchmark evaluating the faithfulness of LCVLMs via multimodal citation generation. MMLongCite features 2,280 examples across 8 tasks and diverse modalities (image, video, interleaved), with context lengths scaled from 16K to 128K tokens. To test spatial localization capabilities of LCVLMs, we also introduce MMLongCite-HR, evaluating fine-grained visual grounding amidst dense pixel spaces. Through extensive benchmarking of cutting-edge LCVLMs, we provide a systematic analysis of current multimodal citation capabilities. Our results reveal a significant discrepancy between answer correctness and citation faithfulness. We also conduct attention pattern investigations and in-depth error analyses to reveal the underlying phenomena of failures in LCVLMs. MMLongCite establishes a rigorous foundation for diagnosing and advancing the faithfulness of LCVLMs. We hope our findings provide meaningful insights to drive further improvements in the long-context capabilities of LCVLMs.
|
| 305 |
RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Multi-Trajectory Semantic Retrieval
2510.16444
|
cs.CVcs.MM
|
Kunyu Peng, Di Wen, Jia Fu, Jiamin Wu, Kailun Yang |
Who is being described, where are they, and what are they doing? Referring Atomic Video Action Recognition (RAVAR) answers these questions jointly by grounding a natural-language reference to a person and recognizing that person's fine-grained atomic actions i...Who is being described, where are they, and what are they doing? Referring Atomic Video Action Recognition (RAVAR) answers these questions jointly by grounding a natural-language reference to a person and recognizing that person's fine-grained atomic actions in complex multi-person videos. Progress in RAVAR is constrained by limited benchmark scale and weak alignment between fine-grained linguistic and scene cues and temporally coherent visual cues. We address both challenges with a new dataset and model. We introduce RefAVA++, a large-scale dataset comprising 2,950,920$ frames, 75,111 annotated person instances, and 80 atomic action categories. Its references describe appearance and spatial attributes while deliberately omitting action labels, requiring models to infer actions directly from visual cues. We further propose RefAtomNet++, which models complementary semantics at the holistic-sentence, partial-keyword, and scene-attribute levels. Fine-grained semantic cues retrieve aligned visual tokens across time to construct trajectories, which are aggregated through Mamba-based state-space modeling and fused using multi-hierarchical semantic-aligned cross-attention. This design enables accurate joint person localization and multi-label atomic action recognition. Equipped with the InternVideo2.5 backbone, RefAtomNet++ achieves 50.39%/51.14% mIoU, 59.21%/59.55% mAP, and 73.89%/76.59% AUROC on RefAVA, and 49.49%/50.33% mIoU, 62.33%/61.90% mAP, and 76.06%/76.10% AUROC on RefAVA++ validation/test sets, respectively. The dataset and code are available at https://github.com/KPeng9510/refAVA2.
|
| 306 |
EndoWave: 4D Gaussian Splatting with Rational Wavelet for Endoscopic Reconstruction
2510.23087
|
cs.CV
|
Taoyu Wu, Yiyi Miao, Jiaxin Guo, Ziyan Chen, Sihang Zhao |
In robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric inconsistencies, non-rigid tissue...In robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric inconsistencies, non-rigid tissue motion, and view-dependent highlights. Most 3DGS-based methods that rely solely on appearance constraints for optimizing 3DGS are often insufficient in this context, as these dynamic visual artifacts can mislead the optimization process and lead to inaccurate reconstructions. To address these limitations, we present EndoWave, a unified spatio-temporal Gaussian Splatting framework by incorporating an optical flow-based geometric constraint and a multi-resolution rational wavelet supervision. First, we adopt a unified spatio-temporal Gaussian representation that directly optimizes primitives in a 4D domain. Second, we propose a geometric constraint derived from optical flow to enhance temporal coherence and effectively constrain the 3D structure of the scene. Third, we propose a multi-resolution rational orthogonal wavelet as a constraint, which can effectively separate the details of the endoscope and enhance the rendering performance. Extensive evaluations on two real surgical datasets, EndoNeRF and StereoMIS, demonstrate that our method EndoWave achieves state-of-the-art reconstruction quality and visual accuracy compared to the baseline method.
|
| 307 |
Mapping and Classification of Trees Outside Forests using Deep Learning
2510.25239
|
cs.CV
|
Moritz Lucas, Hamid Ebrahimy, Viacheslav Barkov, Ralf Pecenka, Kai-Uwe K\"uhnberger |
Trees Outside Forests (TOF) play an important role in agricultural landscapes by supporting biodiversity, sequestering carbon, and regulating microclimates. Yet, most studies have treated TOF as a single class or relied on rigid rule-based thresholds, limiting...Trees Outside Forests (TOF) play an important role in agricultural landscapes by supporting biodiversity, sequestering carbon, and regulating microclimates. Yet, most studies have treated TOF as a single class or relied on rigid rule-based thresholds, limiting ecological interpretation and adaptability across regions. To address this, we evaluate deep learning for TOF classification using a newly generated dataset and high-resolution aerial imagery from four agricultural landscapes in Germany. Specifically, we compare convolutional neural networks (CNNs), vision transformers, and hybrid CNN-transformer models across six semantic segmentation architectures (ABCNet, LSKNet, FT-UNetFormer, DC-Swin, BANet, and U-Net) to map four categories of woody vegetation: Forest, Patch, Linear, and Tree, derived from previous studies and governmental products. Overall, the models achieved good classification accuracy across the four landscapes, with the FT-UNetFormer performing best (mean Intersection-over-Union 0.74; mean F1 score 0.84), underscoring the importance of spatial context understanding in TOF mapping and classification. Our results show good results for Forest and Linear class and reveal challenges particularly in classifying complex structures with high edge density, notably the Patch and Tree class. Our generalization experiments highlight the need for regionally diverse training data to ensure reliable large-scale mapping. The dataset and code are openly available at https://github.com/Moerizzy/TOFMapper
|
| 308 |
Fine-Grained Caching for Diffusion Transformers with Few Calibration Conditions
2512.05134
|
cs.CVcs.LG
|
Zihao Wu, Bohan Zeng, Yuanxing Zhang |
Diffusion transformers require repeated denoiser evaluations, making image and video generation computationally expensive. We propose a training-free framework that uses a few calibration conditions to construct a fixed, fine-grained module-reuse schedule with...Diffusion transformers require repeated denoiser evaluations, making image and video generation computationally expensive. We propose a training-free framework that uses a few calibration conditions to construct a fixed, fine-grained module-reuse schedule without schedule search. The design is motivated by an empirical effect along deterministic sampling trajectories: with initial noise fixed, condition-dependent deviations in several module outputs evolve similarly across adjacent steps, so temporal differencing attenuates much of their variation. We term this effect Conditional Common-Mode Rejection (CCMR). Motivated by this observation, we rank timestep--layer--module locations using normalized ratios of consecutive feature displacements and construct a Cache Book through two-stage calibration. The second stage updates a scorer-only reference with each provisional bit after recording its indexed score, while retaining full-compute trajectories. At inference, the fixed schedule requires no per-input policy estimation and can be nested within fixed step-wise cache policies. Module caching alone achieves $1.66$--$1.72\times$ speedup on the evaluated image models and $1.54$--$1.68\times$ on the video models. Combining it with MagCache or SeaCache provides further acceleration at a model-dependent cost in paired fidelity.
|
| 309 |
Transform Trained Transformer for Accelerating Native 4K Video Generation
2512.13492
|
cs.CV
|
Jiangning Zhang, Junwei Zhu, Teng Hu, Yabiao Wang, Donghao Luo |
Native 4K (2176$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality....Native 4K (2176$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed T3 ($\textbf{T}$ransform $\textbf{T}$rained $\textbf{T}$ransformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, $\textbf{T3-Video}$ introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an "attention pattern" transformation of a pretrained model using only modest compute and data. Results on $\textbf{4K-VBench}$ show that $\textbf{T3-Video}$ substantially outperforms existing approaches: while delivering performance improvements (+4.29$\uparrow$ VQA and +0.08$\uparrow$ VTC), it accelerates native 4K video generation by more than 10$\times$. Project page at https://zhangzjn.github.io/projects/T3-Video
|
| 310 |
OccStress: Stress-Testing the 4D Occupancy Forecasting Chain
2512.15621
|
cs.CV
|
Yu Zheng, Jie Hu, Jiaqi Xiong, Ruiping Liu, Junwei Zheng |
Occupancy world models use historical occupancy states to forecast future 3D scenes, but their robustness under corrupted temporal inputs remains poorly understood. Existing evaluations primarily emphasize clean forecasting accuracy and provide limited evidenc...Occupancy world models use historical occupancy states to forecast future 3D scenes, but their robustness under corrupted temporal inputs remains poorly understood. Existing evaluations primarily emphasize clean forecasting accuracy and provide limited evidence about how errors enter, persist, and propagate through the occupancy perception-forecasting chain. This paper introduces OccStress, a robustness stress-testing benchmark for the occupancy forecasting chain. OccStress contains 21 corruption families with 61 severity configurations and 10,827 strict temporal anchors across 3 datasets. OccStress covers both 3D occupancy perception and 4D occupancy forecasting through two complementary tracks. This design separates model-mediated pipeline errors under standardized sensor stressors from the intrinsic sensitivity of 4D forecasting models to corrupted occupancy states. OccStress further defines temporal injection protocols to test whether errors in the current state, recent history, or earlier history affect future forecasts differently. OccStress provides aggregate metrics for evaluating robustness along the occupancy forecasting chain. Experiments with five 4D forecasters reveal that current occupancy models are substantially affected by both upstream prediction errors and direct state corruptions, and that clean performance alone is an insufficient description of source- and position-specific robustness. Code, data, and evaluation tools are available at https://insailab.org/OccStress.
|
| 311 |
ReFRM3D: A Radiomics-enhanced Fused Residual Multiparametric 3D Network with Multi-Scale Feature Fusion for Glioma Characterization
2512.22570
|
cs.CV
|
Md. Abdur Rahman, Md Noman Hossain, Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan, Yan Zhang |
Gliomas are among the most aggressive cancers, with complex diagnostic processes. Existing glioma segmentation methods often struggle with high variability in imaging data and inadequate optimization. Furthermore, radiomic analysis is typically applied only af...Gliomas are among the most aggressive cancers, with complex diagnostic processes. Existing glioma segmentation methods often struggle with high variability in imaging data and inadequate optimization. Furthermore, radiomic analysis is typically applied only after segmentation is finished, limiting its ability to inform the segmentation process itself. To address these challenges, we propose a novel radiomics-enhanced fused residual multiparametric 3D network (ReFRM3D) for brain tumor characterization. The framework is based on a 3D U-Net architecture and features multi-scale feature fusion, hybrid upsampling, and an extended residual skip mechanism. Additionally, we introduce a radiomic conditioning mechanism that extracts texture and intensity descriptors from an intermediate coarse segmentation and re-injects them into the decoder to refine the final output. Experimental results on BraTS2019, BraTS2020, and BraTS2021 show strong performance, with mean DSC values of 93.45%, 93.61%, and 92.06%, respectively, across whole tumor, enhancing tumor, and tumor core regions. Compared with recent models, ReFRM3D improves average DSC by 5.79%, 7.25%, and 0.96% on these datasets. Our model also generalized well on the BraTS-Africa dataset with an average DSC of 87.8%, despite differences in scanner field strength and patient demographics.
|
| 312 |
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning
2601.02918
|
cs.CV
|
Guoqiang Liang, Jianyi Wang, Zhonghua Wu, Shangchen Zhou, Chen Change Loy |
Image Quality Assessment (IQA) is a long-standing problem in computer vision. Previous methods typically focus on predicting numerical scores without explanation or providing low-level descriptions lacking precise scores. Recent reasoning-based vision language...Image Quality Assessment (IQA) is a long-standing problem in computer vision. Previous methods typically focus on predicting numerical scores without explanation or providing low-level descriptions lacking precise scores. Recent reasoning-based vision language models (VLMs) have shown strong potential for IQA by jointly generating quality descriptions and scores. However, existing VLM-based IQA methods often suffer from unreliable reasoning due to their limited capability of integrating visual and textual cues. In this work, we introduce Zoom-IQA, a VLM-based IQA model to explicitly emulate key cognitive behaviors: uncertainty awareness, region reasoning, and iterative refinement. Specifically, we present a two-stage training pipeline: 1) supervised fine-tuning (SFT) on our Grounded-Rationale-IQA (GR-IQA) dataset to teach the model to ground its assessments in key regions, and 2) reinforcement learning (RL) for dynamic policy exploration, stabilized by our KL-Coverage regularizer to prevent reasoning and scoring diversity collapse, with a Progressive Re-sampling Strategy for mitigating annotation bias. Extensive experiments show that Zoom-IQA achieves improved robustness, explainability, and generalization. The application to downstream tasks, such as image restoration, further demonstrates the effectiveness of Zoom-IQA.
|
| 313 |
MTV: Revisiting Multi-Task Visual Representation Learning
2601.13886
|
cs.CV
|
Shangzhe Di, Zhonghua Zhai, Weidi Xie |
Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised methods (e.g., MAE, DINO) capture intricate local structures yet struggle with h...Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised methods (e.g., MAE, DINO) capture intricate local structures yet struggle with high-level semantic context. We argue that these paradigms are fundamentally complementary and can be integrated into a principled multi-task framework, further enhanced by dense spatial supervision. We introduce MTV, a multi-task visual pretraining framework that jointly optimizes a shared backbone across vision-language contrastive, self-supervised, and dense spatial objectives. To mitigate the need for manual annotations, we leverage high-capacity "expert" models--such as Depth Anything V2 and OWLv2--to synthesize dense, structured pseudo-labels at scale. Beyond the framework, we provide a systematic investigation into the mechanics of multi-task visual learning, analyzing: (i) the individual and marginal gain of each objective, (ii) task synergies versus interference, and (iii) scaling behavior across varying data and model scales. Our results demonstrate that MTV achieves "best-of-both-worlds" performance, significantly enhancing fine-grained spatial reasoning without compromising global semantic understanding. Our findings suggest that multi-task learning, fueled by high-quality pseudo-supervision, is a scalable path toward more general visual encoders.
|
| 314 |
SyncLight: Single-Edit Multi-View Relighting
2601.16981
|
cs.CV
|
David Serrano-Lozano, Anand Bhattad, Luis Herranz, Jean-Fran\c{c}ois Lalonde, Javier Vazquez-Corral |
We present SyncLight, a method to enable consistent, parametric control over light sources across multiple uncalibrated views of a static scene conditioned on a single view. While single-view relighting has advanced significantly, existing generative approache...We present SyncLight, a method to enable consistent, parametric control over light sources across multiple uncalibrated views of a static scene conditioned on a single view. While single-view relighting has advanced significantly, existing generative approaches struggle to maintain the rigorous lighting consistency essential for multi-camera broadcasts, stereoscopic cinema, and virtual production. SyncLight addresses this by enabling precise control over light intensity and color across a multi-view capture of a scene, conditioned on a single reference edit. Our method leverages a multi-view diffusion transformer trained using a latent bridge matching formulation, achieving high-fidelity relighting of the entire image set in a single inference step. To facilitate training, we introduce a large-scale hybrid dataset comprising diverse synthetic environments -- curated from existing sources and newly designed scenes -- alongside high-fidelity, real-world multi-view captures under calibrated illumination. Though trained only on image pairs, SyncLight generalizes zero-shot to an arbitrary number of viewpoints, effectively propagating lighting changes across all views, without requiring camera pose information. SyncLight enables practical relighting workflows for multi-view capture systems.
|
| 315 |
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
2601.17657
|
cs.CV
|
Taewan Cho, Taeryang Kim, Andrew Jaeyong Choi |
Robotic and autonomous systems need dense spatial cues, yet adding a dedicated depth estimator can duplicate visual processing already performed by a multimodal model. CLIP-based depth methods offer an alternative, but commonly rely on text-derived conditionin...Robotic and autonomous systems need dense spatial cues, yet adding a dedicated depth estimator can duplicate visual processing already performed by a multimodal model. CLIP-based depth methods offer an alternative, but commonly rely on text-derived conditioning or backbone adaptation. We present SPACE-CLIP, a decoder-only framework for supervised monocular depth estimation with a frozen CLIP vision backbone and no text encoder at inference. A FiLM-conditioned semantic pathway combines global image context with multilevel patch features, while a structural pathway supplies separately processed spatial features to a hierarchical fusion decoder. Indoor and outdoor evaluations demonstrate depth reconstruction with this architecture, and controlled component comparisons support the contribution of the structural pathway. Layer-selection experiments and frequency interventions further characterize the structural pathway's contribution to depth reconstruction. A shared-backbone microbenchmark further illustrates the reduction in duplicated computation. SPACE-CLIP provides a modular approach to adding dense depth prediction to compatible visual perception stacks. Code is available at https://github.com/taewan2002/SPACE-CLIP.
|
| 316 |
MedAD-R1: Consistency-Reinforced Policy Optimization for Interpretable Medical Anomaly Detection
2602.01081
|
cs.CV
|
Haitao Zhang, Yingying Wang, Jiaxiang Wang, Haote Xu, Hongyang Zhang |
Medical Anomaly Detection (MedAD) offers a promising direction for medical image analysis with Large Multimodal Models (LMMs). However, progress is limited by fragmented datasets and the tendency of Supervised Fine-Tuning (SFT) to learn superficial image-text ...Medical Anomaly Detection (MedAD) offers a promising direction for medical image analysis with Large Multimodal Models (LMMs). However, progress is limited by fragmented datasets and the tendency of Supervised Fine-Tuning (SFT) to learn superficial image-text correlations rather than verifiable diagnostic reasoning. Consequently, current models often generate fluent explanations that are either insufficiently grounded in the image or inconsistent with their final answers, limiting their reliability in high-stakes medical applications. To address these issues, we introduce MedAD-38K, a large-scale, multimodal, and multicenter benchmark containing structured Visual Question Answering pairs and quality-controlled diagnostic Chain-of-Thought annotations across five core MedAD tasks. Based on this benchmark, we propose a two-stage framework. Cognitive Injection first uses SFT to inject domain-specific medical knowledge and establish a structured think-then-answer format. Consistency Group Relative Policy Optimization (Con-GRPO) then employs an Evidence-Aware Consistency Reward to reinforce reasoning that remains grounded in the image and logically supports the final answer. The resulting MedAD-R1 achieves state-of-the-art performance on MedAD-38K and consistently outperforms all evaluated baselines across five source-disjoint external datasets spanning diverse imaging modalities. Beyond accuracy, it achieves higher reasoning-answer consistency and visual-grounding scores across multiple backbones. With only 0.8B parameters, MedAD-R1 offers practical potential for resource-constrained deployment. These results demonstrate the generality of evidence-aware consistency optimization for interpretable MedAD. Project resources are available at https://github.com/zhtstar/MedAD-R1.
|
| 317 |
MambaVF: State Space Model for Efficient Video Fusion
2602.06017
|
cs.CV
|
Zixiang Zhao, Yukun Cui, Lilun Deng, Haowen Bai, Haotong Qin |
Video fusion aims to integrate complementary information from multiple source videos while preserving temporal consistency. Effective modeling of temporal dynamics is essential to this goal, yet existing methods incur substantial computational overhead from op...Video fusion aims to integrate complementary information from multiple source videos while preserving temporal consistency. Effective modeling of temporal dynamics is essential to this goal, yet existing methods incur substantial computational overhead from optical flow estimation and feature warping. In this paper, we present MambaVF, an efficient video fusion framework that uses state space model (SSM) to achieve temporal modeling without explicit motion estimation. First, by formulating video fusion as a sequential state update process, MambaVF captures long-range temporal dependencies with linear complexity, significantly reducing computation and memory costs. Second, the lightweight SSM-based fusion module eliminates conventional flow-guided alignment. Instead, it introduces a mutual state fusion module and a spatio-temporal bidirectional scanning mechanism to enable information aggregation across video streams. Experiments on multiple benchmarks confirm that MambaVF reaches state-of-the-art performance in different video fusion applications (multi-exposure, multi-focus, infrared-visible, medical), while reducing parameters by >90% and FLOPs by >80%, resulting in >50% shorter runtime. Project page: https://mambavf.github.io
|
| 318 |
Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schr\"odinger Bridge Matching
2602.15396
|
cs.CV
|
Jeongwoo Shin, Jinhwan Sul, Joonseok Lee, Jaewoong Choi, Jaemoo Choi |
Diffusion models often yield highly curved trajectories and noisy score targets due to an uninformative, memoryless forward process that induces independent data-noise coupling. We propose Adjoint Schr\"odinger Bridge Matching (ASBM), a generative modeling fra...Diffusion models often yield highly curved trajectories and noisy score targets due to an uninformative, memoryless forward process that induces independent data-noise coupling. We propose Adjoint Schr\"odinger Bridge Matching (ASBM), a generative modeling framework that recovers optimal trajectories in high dimensions via two stages. First, we view the Schr\"odinger Bridge (SB) forward dynamic as a coupling construction problem and learn it through a data-to-energy sampling perspective that transports data to an energy-defined prior. Then, we learn the backward generative dynamic with a simple matching loss supervised by the induced optimal coupling. By operating in a non-memoryless regime, ASBM produces significantly straighter and more efficient sampling paths. Compared to prior works, ASBM scales to high-dimensional data with notably improved stability and efficiency. Extensive experiments on image generation show that ASBM improves fidelity with fewer sampling steps. We further showcase the effectiveness of our optimal trajectory via distillation to a one-step generator.
|
| 319 |
A Hypertoroidal Covering for Perfect Color Equivariance
2603.04256
|
cs.CV
|
Yulong Yang, Zhikun Xu, Yaojun Li, Christine Allen-Blanchette |
When the color distribution of input images changes at inference, the performance of conventional neural network architectures drops considerably. A few researchers have begun to incorporate prior knowledge of color geometry in neural network design. These col...When the color distribution of input images changes at inference, the performance of conventional neural network architectures drops considerably. A few researchers have begun to incorporate prior knowledge of color geometry in neural network design. These color equivariant architectures have modeled hue variation with 2D rotations, and saturation and luminance transformations as 1D translations. While this approach improves neural network robustness to color variations in a number of contexts, we find that approximating saturation and luminance (interval valued quantities) as 1D translations introduces appreciable artifacts. In this paper, we introduce a color equivariant architecture that is truly equivariant. Instead of approximating the interval with the real line, we lift values on the interval to values on the circle (a double-cover) and build equivariant representations there. Our approach resolves the approximation artifacts of previous methods, improves interpretability and generalizability, and achieves better predictive performance than conventional and equivariant baselines on tasks such as fine-grained classification and medical imaging tasks. Going beyond the context of color, we show that our proposed lifting can also extend to geometric transformations such as scale.
|
| 320 |
RAC: Rectified Flow Auto Coder
2603.05925
|
cs.CVcs.AI
|
Sen Fang, Yalin Feng, Yanxin Zhang, Yihao Quan, Juyi Lin |
In this paper, we propose a Rectified Flow Auto Coder (RAC) inspired by Rectified Flow to replace the traditional VAE: 1. It achieves multi-step decoding by applying the decoder to flow timesteps. Its decoding path is straight and correctable, enabling step-by...In this paper, we propose a Rectified Flow Auto Coder (RAC) inspired by Rectified Flow to replace the traditional VAE: 1. It achieves multi-step decoding by applying the decoder to flow timesteps. Its decoding path is straight and correctable, enabling step-by-step refinement. 2. The model inherently supports bidirectional inference, where the decoder serves as the encoder through time reversal (hence Coder rather than encoder or decoder), reducing parameter count by nearly 41%. 3. This generative decoding method improves generation quality since the model can correct latent variables along the path, partially addressing the reconstruction--generation gap. Experiments show that RAC achieves a Pareto improvement over SOTA VAEs, where even a 10$\times$ parameter-reduced decoder exceeds full-scale VAE performance in both reconstruction and generation quality, validating the effectiveness of our approach.
|
| 321 |
DynaTokens: Controlling Token Dynamics for Continual Video-Language Understanding
2603.06662
|
cs.CVcs.LG
|
Toan Nguyen, Yang Liu, Celso De Melo, Flora D. Salim |
Continual VideoQA with multimodal LLMs remains challenging because sequential adaptation induces task interference, while storing task-specific prompts becomes impractical as task sequences grow. We introduce DynaTokens, a transformer-based token generator tha...Continual VideoQA with multimodal LLMs remains challenging because sequential adaptation induces task interference, while storing task-specific prompts becomes impractical as task sequences grow. We introduce DynaTokens, a transformer-based token generator that dynamically produces fine-tuning tokens on demand, enabling task-adaptive prompt updates through shared generation weights. To mitigate forgetting, we introduce meta-learning-inspired regularisers that look ahead to avoid task-specific sharp update directions while anchoring the evolving generator to prior-task behaviours. We theoretically connect this objective to sharpness-aware optimisation, showing how it favours flatter cross-task minima and improves retention. DynaTokens combines gradient-free routing based on robust pretrained token and visual embeddings with lightweight auxiliary multimodal supervision, reducing router drift during continual adaptation. Across standard continual VideoQA benchmarks, DynaTokens achieves higher average accuracy and substantially lower forgetting than strong baselines. It also improves zero-shot generalisation and remains effective in longer domain-incremental sequences with extended task shifts. Finally, we introduce a challenging ImageQA->VideoQA protocol and show that DynaTokens enables robust cross-modal continual transfer.
|
| 322 |
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
2603.07751
|
cs.CVcs.CL
|
Shaoxiong Zhan, Yanlin Lai, Zheng Liu, Hai Lin, Shen Li |
Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to co...Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to construct coherent 3D mental representations from 2D observations. We uncover this gap via diagnostic analyses showing the bottleneck is a missing view-consistent spatial interface rather than insufficient visual features or weak reasoning. To bridge this, we introduce \textbf{3ViewSense}, a framework that grounds spatial reasoning in Orthographic Views. Drawing on engineering cognition, we propose a ``Simulate-and-Reason'' mechanism that decomposes complex scenes into canonical orthographic projections to resolve geometric ambiguities. By aligning egocentric perceptions with these allocentric references, our method facilitates explicit mental rotation and reconstruction. Empirical results on spatial reasoning benchmarks demonstrate that our method significantly outperforms existing baselines, with consistent gains on occlusion-heavy counting and view-consistent spatial reasoning. The framework also improves the stability and consistency of spatial descriptions, offering a scalable path toward stronger spatial intelligence in multimodal systems.~\footnote{https://github.com/Jasaxion/3ViewSense}
|
| 323 |
$M^2$-Occ: Resilient 3D Semantic Occupancy Prediction for Autonomous Driving with Incomplete Camera Inputs
2603.09737
|
cs.CV
|
Kaixin Lin, Kunyu Peng, Di Wen, Yufan Chen, Ruiping Liu |
Semantic occupancy prediction enables dense 3D geometric and semantic understanding for autonomous driving. However, existing camera-based approaches implicitly assume complete surround-view observations, an assumption that rarely holds in real-world deploymen...Semantic occupancy prediction enables dense 3D geometric and semantic understanding for autonomous driving. However, existing camera-based approaches implicitly assume complete surround-view observations, an assumption that rarely holds in real-world deployment due to occlusion, hardware malfunction, or communication failures. We study semantic occupancy prediction under incomplete multi-camera inputs and introduce $M^2$-Occ, a framework designed to preserve geometric structure and semantic coherence when views are missing. $M^2$-Occ addresses two complementary challenges. First, a Multi-view Masked Reconstruction (MMR) module leverages the spatial overlap among neighboring cameras to recover missing-view representations directly in the feature space. Second, a Feature Memory Module (FMM) introduces a learnable memory bank that stores class-level semantic prototypes. By retrieving and integrating these global priors, the FMM refines ambiguous voxel features, ensuring semantic consistency even when observational evidence is incomplete. We introduce a systematic missing-view evaluation protocol on the nuScenes-based SurroundOcc benchmark, encompassing both deterministic single-view failures and stochastic multi-view dropout scenarios. Under the safety-critical missing back-view setting, $M^2$-Occ improves the IoU by 4.36%. As the number of missing cameras increases, the robustness gap further widens; for instance, under the setting with five missing views, our method boosts the IoU by 6.67%. These gains are achieved without compromising full-view performance. The source code will be publicly released at https://github.com/qixi7up/M2-Occ.
|
| 324 |
COMIC: Agentic Sketch Comedy Generation
2603.11048
|
cs.CVcs.CLcs.AI
|
Susung Hong, Brian Curless, Ira Kemelmacher-Shlizerman, Steve Seitz |
We propose a fully automated AI system that produces short comedic videos similar to sketch shows such as Saturday Night Live. Starting from character references, the system employs a population of agents loosely modeled on roles in real production studios and...We propose a fully automated AI system that produces short comedic videos similar to sketch shows such as Saturday Night Live. Starting from character references, the system employs a population of agents loosely modeled on roles in real production studios and structured to optimize the quality and diversity of ideas and outputs through iterative competition, evaluation, and refinement. A key contribution is the introduction of LLM-based critics aligned with real viewer preferences through the analysis of a corpus of comedy videos on YouTube, enabling automatic evaluation of humor. We further propose multi-island relativistic evolution for script improvement, along with script-conditioned rendering critics that perform shot- and video-level tournaments. In both human and automated evaluations, COMIC receives higher scores than agentic and video generation baselines.
|
| 325 |
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
2603.12064
|
cs.CV
|
Shuo Sun, Unal Artan, Malcolm Mielle, Achim J. Lilienthal, Martin Magnusson |
We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a setting that arises naturally when multiple observers capture a shared event. Prior approaches either handle only singl...We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a setting that arises naturally when multiple observers capture a shared event. Prior approaches either handle only single-camera input or require rigidly mounted, pre-calibrated camera rigs, limiting their practical applicability. We propose a two-stage optimization framework that decouples the task into robust camera tracking and dense depth refinement. In the first stage, we extend single-camera visual SLAM to the multi-camera setting by constructing a spatiotemporal connection graph that exploits both intra-camera temporal continuity and inter-camera spatial overlap, enabling consistent scale and robust tracking. To ensure robustness under limited overlap, we introduce a wide-baseline initialization strategy using feed-forward reconstruction models. In the second stage, we refine depth and camera poses by optimizing dense inter- and intra-camera consistency using wide-baseline optical flow. Additionally, we introduce MultiCamRobolab, a new real-world dataset with ground-truth poses from a motion capture system. Finally, we demonstrate that our method significantly outperforms state-of-the-art feed-forward models on both synthetic and real-world benchmarks, while requiring less memory.
|
| 326 |
UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Affordance Segmentation
2603.23478
|
cs.CV
|
Jiaying Lin, Dan Xu |
Affordance segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. Existing training-free methods typically rely on fragmented pipelines, which introduce visual blind...Affordance segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. Existing training-free methods typically rely on fragmented pipelines, which introduce visual blindness during task parsing and limit accuracy through single-scale spatial and temporal processing. We present UniFunc3D, a unified and training-free framework that treats the multimodal large language model as an active observer. By utilizing a unified MLLM backbone, UniFunc3D performs joint semantic-temporal-spatial reasoning to ground task decomposition in direct visual evidence. Our approach introduces active spatial-temporal grounding with a coarse-to-fine strategy. This allows the model to select correct video frames adaptively and focus on high-detail interactive parts while preserving the global context necessary for disambiguation. On SceneFun3D, our UniFunc3D achieves state-of-the-art performance, surpassing prior training-free methods by a large margin with a relative 59.9\% mIoU improvement, and even outperforming training-based methods without any task-specific training. Code is available on our project page: \url{https://jiaying.link/unifunc3d}.
|
| 327 |
Live Interactive Training for Video Segmentation
2603.26929
|
cs.CV
|
Xinyu Yang, Haozheng Yu, Yihong Sun, Bharath Hariharan, Jennifer J. Sun |
Interactive video segmentation often requires many user interventions for robust performance in challenging scenarios (e.g., occlusions, object separations, camouflage, etc.). Yet, even state-of-the-art models like SAM2 use corrections only for immediate fixes...Interactive video segmentation often requires many user interventions for robust performance in challenging scenarios (e.g., occlusions, object separations, camouflage, etc.). Yet, even state-of-the-art models like SAM2 use corrections only for immediate fixes without learning from this feedback, leading to inefficient, repetitive user effort. To address this, we introduce Live Interactive Training (LIT), a novel framework for prompt-based visual systems where models also learn online from human corrections at inference time. Our primary instantiation, LIT-LoRA, implements this by continually updating a lightweight LoRA module on-the-fly. When a user provides a correction, this module is rapidly trained on that feedback, allowing the vision system to improve performance on subsequent frames of the same video. Leveraging the core principles of LIT, our LIT-LoRA implementation achieves an average 18-34% reduction in total corrections on challenging video segmentation benchmarks, with a negligible training overhead of ~0.5s per correction. We further demonstrate its generality by successfully adapting it to other segmentation models and extending it to CLIP-based fine-grained image classification. Our work highlights the promise of live adaptation to transform interactive tools and significantly reduce redundant human effort in complex visual tasks. Project: https://youngxinyu1802.github.io/projects/LIT/.
|
| 328 |
MultiLoc: Look Around As You Localize for Fast And Robust Visual Re-localization
2603.27170
|
cs.CV
|
Nobel Dang, Bing Li |
Relative camera pose is a basic geometric cue for visual re-localization and scene understanding. When using images alone, visual evidence may not sufficiently support reliable estimation of 3D-consistent motion. The challenge is especially acute in unfamiliar...Relative camera pose is a basic geometric cue for visual re-localization and scene understanding. When using images alone, visual evidence may not sufficiently support reliable estimation of 3D-consistent motion. The challenge is especially acute in unfamiliar and ever-changing environments for real-time applications, where a system must infer spatial structure under tight time constraints from what it sees rather than rely on scene-specific reconstruction or training. This raises a fundamental question: how can live 3D camera motion be estimated accurately for re-localization in unseen environments while drawing on enough scene context to resolve pose ambiguity? We introduce MultiLoc, a multi-view-guided relative pose regressor trained at scale to achieve spatial and geometric consistent representations for robust visual re-localization in unseen scenarios with high inference speed. Specifically, MultiLoc creates a minimal 3D-sub-scene representation of the environment and efficiently transforms it into 3D spatially and geometrically consistent features in a single forward pass, enabling precise pose estimates with high inference speed gains. Across diverse indoor, outdoor, and in-the-wild visual re-localization benchmarks---Indoor6, Cambridge Landmarks and WaySpots---MultiLoc consistently outperforms state-of-the-art relative pose regression methods. We also highlight that MultiLoc, a pose regressor, also performs competitively with inference-heavy and structure-based approaches while retaining sub-second inference and generalizing to unseen scenes. Given a small posed support set, it also surpasses relative pose regression, feature-matching, and non-regression methods on several relative camera pose benchmarks. Code will be released.
|
| 329 |
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
2604.00829
|
cs.CVcs.CL
|
Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, Yova Kementchedjhieva |
Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further...Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further fine-tuning, and existing remedies add adapters or alignment modules that increase architectural complexity and inference cost. We propose LinguDistill, an adapter-free knowledge distillation method that uses the original frozen LM as the teacher during multimodal post-training. To let a text-only teacher supervise vision-conditioned outputs, we introduce layer-wise KV-cache sharing, which exposes the teacher to the student's multimodal representations without changing either architecture. We then apply distillation selectively, on language-heavy data only, so the teacher restores linguistic ability while the student keeps its visual grounding on document and OCR tasks. LinguDistill recovers the language and knowledge performance lost during multimodal fine-tuning, matching the original VLM on average over text-only benchmarks (ARC, HellaSwag) and exceeding it on ScienceQA, while keeping vision-heavy performance close to standard fine-tuning. Since the teacher is dropped after training, the final model adds no parameters and no inference cost. More broadly, our results show that a model's own pre-adaptation backbone is a practical teacher for undoing forgetting, suggesting a simple recipe for keeping language ability intact as models are extended to new modalities.
|
| 330 |
Confidence under Visual Token Pruning: Removed Evidence and Risk-Controlled Token Budgets for MLLMs
2604.12035
|
cs.CV
|
Kaizhen Tan, Yang Feng, Heqing Du, Hanzhe Hong, Siru Tao |
Visual token pruning speeds up multimodal large language models (MLLMs) by keeping a small subset of the visual tokens, and pruning methods are compared by the accuracy they retain. We study what pruning does to the confidence of these models, across common se...Visual token pruning speeds up multimodal large language models (MLLMs) by keeping a small subset of the visual tokens, and pruning methods are compared by the accuracy they retain. We study what pruning does to the confidence of these models, across common selectors, several MLLMs, and different output formats. Pruning errors concentrate on questions whose evidence the selector removed, and confidence does not register the removal. When the queried object loses all of its tokens, accuracy on these questions drops from 59% to 17%, while confidence stays at the unpruned level. Returning a few object tokens to the kept set recovers most of the lost accuracy. Selectors that keep the most attended tokens remove such evidence most often and produce confident errors, which temperature scaling cannot re-rank. Selectors that avoid keeping redundant tokens stay close to the calibration of the unpruned model. We then use the confidence of the pruned model to set a per-question token budget. The model answers with few tokens first and again with all tokens when its confidence is low. Conformal risk control sets the threshold to bound the expected deviation from the unpruned model. With coverage-based selection, this cascade needs about a third of the prefill tokens of the unpruned model, while with FastV it needs more than the unpruned model. The savings come mainly from how well the confidence ranks the answers that differ from the unpruned ones.
|
| 331 |
Perturbation-Regularized Open-Vocabulary Remote Sensing Segmentation with Unified Multi-Domain Evaluation
2604.15652
|
cs.CV
|
Bingyu Li, Tao Huo, Haocheng Dong, Da Zhang, Zhiyuan Zhao |
Open-vocabulary remote sensing image segmentation (OVRSIS) aims to segment text-specified categories beyond a fixed label space. Its key challenge is to maintain reliable pixel--text correspondence under substantial appearance variation in remote-sensing image...Open-vocabulary remote sensing image segmentation (OVRSIS) aims to segment text-specified categories beyond a fixed label space. Its key challenge is to maintain reliable pixel--text correspondence under substantial appearance variation in remote-sensing imagery. Existing methods often rely on specialized vision--language representations or complex adaptation mechanisms, yet still perform deterministic matching between fixed text prototypes and individual visual features. We instead formulate OVRSIS as a \emph{distributional pixel--text alignment} problem and propose \textbf{Pi-Seg}, a lightweight perturbation-based framework. Pi-Seg introduces learnable variations into textual prototypes and dense visual features before alignment. Segmentation supervision encourages perturbations that enlarge the target-to-distractor margin while suppressing harmful feature shifts. Pi-Seg consistently improves performance under the existing \textit{single-source} OVRSISBench protocols. To evaluate OVRSIS beyond source-specific training settings, we further introduce a unified \textit{multi-source} protocol and construct \textbf{GlobalRSOV95K}, containing approximately 95K densely annotated images and 35 semantic categories. Models are trained on this shared dataset and evaluated on an image-disjoint suite of 10 downstream datasets. The complete benchmark covers more than 170K images and 122 categories. Experiments under both single-source and multi-source protocols show that Pi-Seg improves cross-dataset transfer and generalizes effectively to practical geospatial segmentation tasks, while GlobalRSOV95K provides a common foundation for transferable OVRSIS research. \footnote{\url{https://github.com/LiBingyu01/Pi-Seg}}
|
| 332 |
Point-MF: Stabilizing One-Step Mean Flows for Single-Image Point Cloud Reconstruction
2604.24586
|
cs.CV
|
Yuta Baba, Keiji Yanai |
Single-image point cloud reconstruction requires recovering complete object-level geometry, including occluded regions, from a single RGB image. Diffusion- and flow-based reconstructors can model this ambiguity, but their iterative sampling requires many netwo...Single-image point cloud reconstruction requires recovering complete object-level geometry, including occluded regions, from a single RGB image. Diffusion- and flow-based reconstructors can model this ambiguity, but their iterative sampling requires many network function evaluations, while naive one-step point-space updates often produce outliers and density imbalance. We propose Point-MF, a one-step conditional Mean-Flow framework for geometry-only point cloud reconstruction. Point-MF predicts an interval-averaged velocity field directly in point-cloud space and reconstructs the output with a single network function evaluation, without training a point-cloud autoencoder. To stabilize large Mean-Flow jumps, we introduce Denoised Space Anchor (DSA), a set-distance auxiliary loss that anchors the denoised point set implied by the predicted velocity to the ground-truth geometry. On all 13 ShapeNet-R2N2 categories and Pix3D, Point-MF achieves strong reconstruction quality under the aligned point-set protocol, improving average CD and EMD over the evaluated point-cloud reconstruction baselines. It runs in 40.72 ms per sample on an RTX A4000, within the same latency order as the feedforward RGB2point baseline and substantially faster than iterative diffusion baselines such as PC$^2$ and BDM. These results show that denoised-space geometric anchoring enables stable one-step Mean Flow directly in point-cloud space.
|
| 333 |
InsHuman: Towards Natural and Identity-Preserving Human Insertion
2605.07402
|
cs.CV
|
Jie Li, Shulian Zhang, Wenbo Li, Jian Chen, Yong Guo |
Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, they often produce failure cases, including inappropriate human pose in new background, inconsistent number of ...Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, they often produce failure cases, including inappropriate human pose in new background, inconsistent number of people, and modified facial identity. Moreover, publicly available human datasets often lack full-body portraits and realistic physical interaction between humans and their background. To address these challenges, we propose InsHuman for natural and identity-preserving human insertion. Specifically, we propose Human-Background Adaptive Fusion (HBAF), which detects foreground humans to obtain a binary mask and applies region-aware weighting to align the human regions between predicted and ground-truth latents, ensuring the person's pose, count, and overall appearance are coherently adapted to the target background.We further propose Face-to-Face ID-Preserving (FFIP), which detects and matches faces between the generated image and the source image in terms of face recognition features to enforce identity consistency for each face.In addition, we propose Bidirectional Data Pairing (BDP) strategy to construct BDP-InsHuman, a high-quality dataset with realistic human-background interactions. Experiments demonstrate that InsHuman achieves significant improvements in generating plausible images while keeping human identity unchanged.
|
| 334 |
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
2605.13062
|
cs.CV
|
Xuehai Bai, Yang Shi, Yi-Fan Zhang, Xuanyu Zhu, Yuran Wang |
Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, d...Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fail to faithfully reflect human judgment, especially for strong frontier models, due to limited task difficulty and coarse-grained evaluation protocols. In parallel, reward models have become increasingly important for RL-based image editing optimization, yet existing reward model benchmarks still rely on unrealistic evaluation settings that deviate from practical RL scenarios. These limitations hinder reliable assessment of both image editing models and reward models. To address these challenges, we introduce Edit-Compass and EditReward-Compass, a unified evaluation suite for image editing and reward modeling. Edit-Compass contains 2,388 carefully annotated instances spanning six progressively challenging task categories, covering capabilities such as world knowledge reasoning, visual reasoning, and multi-image editing. Beyond broad task coverage, Edit-Compass adopts a fine-grained multidimensional evaluation framework based on structured reasoning and carefully designed scoring rubrics. In parallel, EditReward-Compass contains 2,251 preference pairs that simulate realistic reward modeling scenarios during RL optimization.
|
| 335 |
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
2605.15141
|
cs.CV
|
Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou, Bokai Yan |
Real-time interactive video generation requires low-latency, streaming, and controllable rollout. Existing autoregressive (AR) diffusion distillation methods have achieved strong results in the chunk-wise 4-step regime by distilling bidirectional base models i...Real-time interactive video generation requires low-latency, streaming, and controllable rollout. Existing autoregressive (AR) diffusion distillation methods have achieved strong results in the chunk-wise 4-step regime by distilling bidirectional base models into few-step AR students, but they remain limited by coarse response granularity and non-negligible sampling latency. In this paper, we study a more aggressive setting: frame-wise autoregression with only 1--2 sampling steps. In this regime, we identify the initialization of a few-step AR student as the key bottleneck: existing strategies are either target-misaligned, incapable of few-step generation, or too costly to scale. We propose \textbf{Causal Forcing++}, a principled and scalable pipeline that uses \emph{causal consistency distillation} (causal CD) for few-step AR initialization. The core idea is that causal CD learns the same AR-conditional flow map as causal ODE distillation, but obtains supervision from a single online teacher ODE step between adjacent timesteps, avoiding the need to precompute and store full PF-ODE trajectories. This makes the initialization both more efficient and easier to optimize. The resulting pipeline, \ours, surpasses the SOTA 4-step chunk-wise Causal Forcing under the \textit{\textbf{frame-wise 2-step setting}} by 0.1 in VBench Total, 0.3 in VBench Quality, and 0.335 in VisionReward, while reducing first-frame latency by 50\% and Stage 2 training cost by $\sim$$4\times$. We further extend the pipeline to action-conditioned world model generation in the spirit of Genie3. Project Page: https://github.com/thu-ml/Causal-Forcing and https://github.com/shengshu-ai/minWM .
|
| 336 |
ReactiveGWM: Flexible Control and NPC Reactivity in Game World Models
2605.15256
|
cs.CV
|
Zeqing Wang, Danze Chen, Zhaohu Xing, Zizhao Tong, Yinhan Zhang |
Existing game world models typically adopt role-specific interactions, where player and NPC roles are bound to fixed characters. This limits their flexibility in multi-character games, where different characters may receive external control while NPCs must rea...Existing game world models typically adopt role-specific interactions, where player and NPC roles are bound to fixed characters. This limits their flexibility in multi-character games, where different characters may receive external control while NPCs must react to interactions triggered by players. This setting raises two key challenges: how to flexibly assign control roles to individual characters, and how to support direct player control and reactive NPC behavior within a unified model. These challenges are particularly pronounced in shared-view 2D games, where multiple, potentially visually identical characters share the same viewpoint, making camera cues insufficient to distinguish their roles. To address these challenges, we introduce ReactiveGWM, a reactive game world model that flexibly assigns control modes at initialization and jointly simulates externally controlled players and reactive NPCs. Specifically, ReactiveGWM introduces Spatial Role Binding, which grounds learned character handles to their corresponding regions in the initial frame using instance masks. Building on these handles, Unified Agency Conditioning unifies heterogeneous control signals across characters by encoding player actions and conditional NPC rules into character-specific control-token groups. Each group is then bound to its corresponding character handle, enabling the model to apply each control signal to its designated character. Meanwhile, causal self-attention restricts temporal context to the current and preceding latent frames when generating player actions and NPC responses. Experiments on two multi-character 2D games demonstrate that ReactiveGWM supports flexible character control across different player/NPC role assignments while jointly generating accurate player-controlled behaviors and reactive NPC responses, enabling more configurable and richer multi-character interactions.
|
| 337 |
Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
2605.18599
|
cs.CV
|
Yihang Wu, Yihang Sun, Shaofeng Zhang, Zuxuan Wu, Junchi Yan |
Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Pl\"ucker rays) into a shared feature space. Since Pl\"ucker rays n...Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Pl\"ucker rays) into a shared feature space. Since Pl\"ucker rays naturally carry lattice-like spatial structure, these designs can make the spatial bias interfere with appearance representation and degrade rendering fidelity. To this end, we propose to decouple the representation of feedforward NVS transformers into separate semantic and spatial tokens. The decoupled design keeps semantic and spatial information explicit in their branches while preserving cross-branch interaction through shared attention routing. Built on this design, we introduce optional categorized supervision and bidirectional modulation: the former provides branch-specific training signals, while the latter improves interaction between the two branches. Notably, the base decoupled design introduces virtually zero additional inference latency due to its architectural design. The proposed designs achieve consistent improvements, demonstrating effectiveness across decoder-only and encoder-decoder feedforward NVS models.
|
| 338 |
The TIME Machine: On The Power of Motion for Efficient Perception
2605.23045
|
cs.CVcs.LGcs.AI
|
Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara |
Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of self-supervised models trained on next-frame prediction. While these factors have pushed the bo...Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of self-supervised models trained on next-frame prediction. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations. First, scaling video models can reach prohibitive costs, with recent models needing hundreds of years of video to be trained. Second, learning to predict the next frame (or embedding of the frame) naturally focuses on spatial information and as a result, video models still struggle with temporal understanding. In this paper we propose a novel approach that uses motion as a modality to alleviate both of these core issues. Given the motion in a video in the form of point tracks, we mask some of the tracks and use a masked autoencoder to reconstruct the missing tracks. This allows us to learn a representation in a self-supervised manner which we call TIME (Temporally-Informed Motion Embedding), and is focused on capturing temporal information. Also, as motion is inherently appearance-invariant, TIME needs far fewer examples to generalize well. As a result, without bells and whistles, on temporal tasks TIME performs on par with state-of-the-art models, using up to 4 orders of magnitude less training data. For general tasks, TIME can be used in combination with existing representations, and we observe that it leads to a significant improvement for V-JEPA 2, RVM and VideoMAE on standard benchmarks such as SSV2, EgoExo4D and Diving48. These results point to a new promising video paradigm for both more temporally-aware as well as more scalable models.
|
| 339 |
Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models
2605.26332
|
cs.CVcs.AI
|
Arian Komaei Koma, Seyed Amir Kasaei, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban |
Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been introduced to make the model generate such unlearned concepts. These attacks, nevertheless, do not assume a r...Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been introduced to make the model generate such unlearned concepts. These attacks, nevertheless, do not assume a realistic threat model, i.e. they either assume access to the model weights, or result in gibberish adversarial prompts that could be easily detected even through naive rule-based safeguarding. We aim to address this gap in this paper. We introduce BEAP, a black-box, embedding-aware adversarial prompting attack that leverages a large language model (LLM) to iteratively generate effective adversarial prompts and exploit such hidden vulnerabilities. BEAP performs an embedding-aware search in text space, combining multiple reward signals: unlearned concept presence, text-image alignment, and image quality, to refine generated prompts. Unlike previous attack methods, BEAP keeps its prompts undetectable to safety filters while producing high-quality images. Across five unlearning methods, BEAP achieves a macro-averaged ASR of 97.8% under the held-out OpenNSFW2 evaluation, exceeding the white-box UDA baseline by 41.6 percentage points (56.2% to 97.8%). Counting unsuccessful searches at the full 100-query budget, BEAP uses 32.1 image-generation queries per evaluated prompt on average under this criterion.
|
| 340 |
vSV-ViT: Variable-size SuperVertex Vision Transformer for Cortical Surface Learning in Alzheimer's Disease
2605.26514
|
cs.CVcs.LGcs.AI
|
Geonwoo Baek, Ikbeom Jang |
Learning from 3D meshes is challenging because the data reside on non-Euclidean surfaces embedded in 3D space. This is particularly evident in domains such as brain cortical surface analysis, where existing models typically rely on ROI-agnostic, face-based, or...Learning from 3D meshes is challenging because the data reside on non-Euclidean surfaces embedded in 3D space. This is particularly evident in domains such as brain cortical surface analysis, where existing models typically rely on ROI-agnostic, face-based, or fixed-size patches. Such patches can duplicate boundary vertices, conflate anatomically distinct regions, or incorporate non-cortical vertices such as those of the medial wall. We develop variable-size supervertex (vSV) partitioning, which groups cortical surface vertices into vSV patches similar to superpixels in images. Building on these vSVs, we design Variable-size SuperVertex Vision Transformer (vSV-ViT), a Vision Transformer that processes vSVs through padding, mask-aware patch embedding, and vertex position embedding. We demonstrate the benefits of the proposed framework through a quantitative analysis of vertex assignment in surface partitioning and downstream tasks using Alzheimer's disease brain image datasets. For downstream tasks, we use meshes of cortical thickness and cortical curvature as input to various encoder models and evaluate classification performance. The superior predictive accuracy of vSV-ViT over recent surface-based models supports its effectiveness for cortical surface learning, with AD-related and sex prediction serving as downstream validation. Extending the framework to other cortical surface applications and broader 3D mesh domains remains future work. Code is available at https://github.com/labhai/vSV-ViT.
|
| 341 |
OnceSelect: Reusable Data Selection for Efficient Multimodal Instruction Tuning
2605.26761
|
cs.CV
|
Mingkang Dong, Muxin Pu, Hongyi Cai, JieLi, Jiancheng Pan |
Multimodal instruction tuning is widely used to adapt multimodal large language models (MLLMs), yet the large-scale image-text datasets it relies on are often highly redundant. Existing data selection methods are commonly tied to specific datasets, target mode...Multimodal instruction tuning is widely used to adapt multimodal large language models (MLLMs), yet the large-scale image-text datasets it relies on are often highly redundant. Existing data selection methods are commonly tied to specific datasets, target models, or training states, requiring retraining or recomputation when transferred to new settings. We ask whether a selection signal can instead be learned once and reused across datasets and models. We propose OnceSelect, a reusable data selector trained once and directly applied to unseen datasets. OnceSelect encodes image-instruction pairs in a frozen joint multimodal space, clusters them into pseudo-labels capturing coarse semantic structure, and trains a lightweight selector using a fixed validation macro-accuracy criterion. Low-confidence samples under the learned partition are treated as informative candidates. As the selector is independent of the downstream MLLM, it can be transferred across datasets without retraining or model-dependent scoring, while selected subsets can be reused across model architectures. Using only 15% of LLaVA-625K, OnceSelect retains 99.2% of full-data aggregate performance across nine evaluation metrics. Without retraining, it achieves 103.4% and 102.5% relative performance on unseen Vision-Flan-186K and LRV-Sub-180K. The same 93.7K LLaVA subset further achieves 99.9-106.7% relative performance across Qwen3-VL and InternVL3 models without model-specific scoring or reselection. These results demonstrate efficient and reusable multimodal data selection across datasets and model architectures. Code is available at https://github.com/DMK041218/OnceSelect
|
| 342 |
Brain-IT-VQA: From Brain Signals to Answers
2605.29588
|
cs.CVcs.AI
|
Roman Beliy, Matias Cosarinsky, Oliver Heinimann, Navve Wasserman, Michal Irani |
Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years in visual question answering (VQA...Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years in visual question answering (VQA) from fMRI, performance remains limited. Moreover, although recent models can make increasingly accurate predictions, they have rarely been used as tools for understanding the structure of visual representations in the brain. We present Brain-IT-VQA, a framework for visual question answering from fMRI. Unlike previous methods, which extract a fixed representation from the fMRI signal, our extraction is conditioned on the question itself, so what is decoded from brain activity depends on what is being asked. Our model substantially outperforms previous fMRI-based captioning and VQA approaches. We further introduce NSD-VQA, a new dataset and benchmark for visual question answering from fMRI. Unlike existing image-fMRI VQA datasets, which typically provide only a few broad and weakly controlled questions per image, NSD-VQA provides on average 20 question-answer pairs per image across 20 controlled question categories that disentangle multiple levels of visual understanding. This enables more reliable and interpretable evaluation despite limited fMRI test data. Together, Brain-IT-VQA and NSD-VQA provide both a strong predictive framework and a tool for studying brain representations. Using this benchmark, we quantify which forms of visual and semantic information can be reliably decoded from fMRI responses to natural images. We further analyze the contributions of different brain regions across question types.
|
| 343 |
YoCausal: How Far is Video Generation from World Model? A Causality Perspective
2605.30346
|
cs.CV
|
You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu |
As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data, limiting real-world generalization d...As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data, limiting real-world generalization due to the sim-to-real gap. We present YoCausal, a two-level benchmark inspired by the Violation of Expectation (VoE) paradigm from cognitive science. By temporally reversing real-world videos at zero cost as natural counterfactual samples, YoCausal establishes an arbitrarily extensible evaluation protocol. Level 1 introduces the Reverse Surprise Index (RSI), quantifying arrow-of-time perception via denoising loss. Level 2 introduces the Causality Cognition Index (CCI), which leverages a VLM to stratify datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias. Evaluation of 13 state-of-the-art VDMs reveals that perceiving the arrow of time does not imply understanding causality, and a significant gap persists relative to human-level causal cognition. Project page: https://www.youzhexie.me/papers/YoCausal/
|
| 344 |
Text-to-Image Models Need Less from Text Encoders Than You Think
2606.03715
|
cs.CV
|
Nurit Spingarn, Noa Cohen, Tamar Rott Shaham, Tomer Michaeli |
Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that condition the image generation process. Beyond individual token meanings, text embeddings encode contextual informa...Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that condition the image generation process. Beyond individual token meanings, text embeddings encode contextual information across the full prompt, such as compositionality and attribute binding. However, whether image models actually exploit this richer information remains underexplored. Here, we address the question: Which aspects of text representation are essential for image generation? We show that text-to-image diffusion transformer-based models commonly rely only on two relatively straightforward aspects of text representations: (i) the merging of adjacent tokens into a word representation, for words spanning multiple tokens, and (ii) word order, which is imprinted by the positional embedding of the text-encoder. To show this, we construct a new text embedding that encodes only individual word meanings and order but lacks any contextual information about the full prompt. We find that this bag of position-tagged words representation is sufficient to successfully guide image generation, achieving visual quality and text fidelity that are on par with full text embedding-guided generation. This demonstrates that, contrary to common belief, text-to-image models often do not use the rich information encoded in the text embedding beyond individual word meanings and word order. Instead, the decoding of complex linguistic structures is performed by the image model itself. Project webpage: https://nsping13.github.io/contextless-TTI/
|
| 345 |
GOPAgen: Codec-Aware Agentic Long-Video Understanding with Structured Memory
2606.06532
|
cs.CV
|
Haozhe Chi, Yang Jin, Yadong Mu |
Agentic long-video question answering often relies on sparse RGB sampling and caption retrieval. This reliance can lead to missed brief motion events and repeated processing of irrelevant intervals. We introduce GOPAgen, a framework that integrates codec-nativ...Agentic long-video question answering often relies on sparse RGB sampling and caption retrieval. This reliance can lead to missed brief motion events and repeated processing of irrelevant intervals. We introduce GOPAgen, a framework that integrates codec-native Groups of Pictures (GOPs) and motion vectors into an agentic retrieval pipeline. GOPAgen constructs global textual memory and uses two-stage query-conditioned temporal selection. A global-sufficiency check precedes local evidence extraction. Caption and motion agents represent selected intervals as typed memory pages containing complementary appearance and motion evidence, temporal metadata, and provenance. A GOP-Tree organizes these pages for conditional retrieval when the global evidence is insufficient. GOPAgen achieves 65.3\% accuracy on MotionBench Test and 78.7\% on EgoSchema, exceeding the reported agentic baselines listed for these benchmarks. It also achieves 73.2\% on LongVideoBench validation and 77.3\% on MLVU, while remaining below the strongest listed baseline on LVBench and Video-MME Long. These results support codec-aware structured memory as an effective interface for selective long-video reasoning.
|
| 346 |
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
2606.06696
|
cs.CVcs.AI
|
Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun |
Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual pe...Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual perception. Models need to correctly interpret subtle features in images, and they must do so across diverse biomedical modalities, scales, and contexts. Nevertheless, current benchmarks remain limited. To address these gaps, we introduce the Massive Multimodal Biomedical Understanding (MMBU) benchmark. It is the largest biomedical vision and language benchmark to date, covering 35 submodalities with rich structured metadata. It includes both open and closed versions of ungrounded classification, grounded classification, and object detection, enabling systematic evaluation of model performance across biological scales, clinical settings, and imaging modalities. Evaluating 16 open-weight and 5 frontier VLMs in the main comparison, we find that while medical adaptation provides measurable gains for some models, the high accuracy often reported on established benchmarks can mask deficiencies in visual perception and domain generalization. We further define an open-ended MMBU hard set, split into a released public subset and a held-out private subset, to stress-test frontier models.
|
| 347 |
STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation
2606.07036
|
cs.CVcs.LGcs.AI
|
Won June Cho, Daeky Jeong, Hyeongyeol Lim, Hongjun Yoon |
Synthetic histopathology image generation addresses patient-privacy concerns and the growing data demands of foundation models. Existing state-of-the-art histopathology generative models use pretrained Vision Foundation Models (VFMs) as conditioning signals. W...Synthetic histopathology image generation addresses patient-privacy concerns and the growing data demands of foundation models. Existing state-of-the-art histopathology generative models use pretrained Vision Foundation Models (VFMs) as conditioning signals. We show this yields conditioning-dominated diversity: on TCGA-BRCA, 62-75% of their output diversity is attributable to the conditioning signal rather than the learned latent space, while de novo synthesis still requires a VFM at inference. We instead use histopathology VFMs as the latent space itself: their patch tokens are $\ell_2$-normalized on the unit hypersphere $\mathcal{S}^{d-1}$ with strong angular dominance and intrinsic curvature, motivating a Riemannian formulation. We present STREAM, the first framework to apply Riemannian flow matching in the histopathology domain, in two stages: 1) a bridge-type stochastic perturbation that establishes per-token rectifiability on $\mathcal{S}^{d-1}$ for training a Diffusion Transformer, and 2) a novel decoder training design whose noise covariance is anisotropic in the left-singular basis of the per-token tangent-projected velocity-field Jacobian, spending a large robustness budget on its low-response directions and a small one on its high-response directions. Across TCGA-BRCA and TCGA-COADREAD, STREAM achieves state-of-the-art gFID and ranks first on nearly all histopathology-specific metrics as well. Code and a public gallery of generated images are available at https://chokevin8.github.io/STREAM-Patho/.
|
| 348 |
CultureScore: Evaluating Cultural Faithfulness in Video Generation Models
2606.07311
|
cs.CVcs.AI
|
Anku Rani, Wei Dai, Shravan Nayak, Pattie Maes, Mahdi M. Kalayeh |
As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure visual quality but offer no mechanism for a...As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure visual quality but offer no mechanism for assessing cultural faithfulness. Consequently, a model that replaces a Namaste with a handshake receives the same score as one that generates the gesture correctly. We propose CultureScore, a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions: Identity (who is represented), Context (culturally localized background), and Behavior (normative gestures and interactions). We operationalize this framework through an evaluation suite spanning 10 countries, yielding 6,174 generated videos across three state-of-the-art models. Our evaluation reveals that no current model achieves culturally faithful video generation: the best-performing model reaches only 56.8% overall CultureScore, with Behavior the most challenging dimension; no model exceeds 52.1% on behavior. Furthermore, the highest-scoring model (LTX-2) on visual quality was ranked last by native annotators, while CultureScore's Behavior dimension shows the strongest positive correlation with human cultural judgment among the automatic metrics we evaluate, underscoring that cultural faithfulness is an essential criterion for equitable video generation. Data and code are publicly available.\footnote{\url{https://huggingface.co/datasets/ankurani/CultureScore}}
|
| 349 |
EvoState: Closed-Loop Visual State Management for Long-Form Video Generation
2606.16184
|
cs.CVcs.MM
|
Xinlei Yin, Xiulian Peng, Xiao Li, Zhiwei Xiong, Yan Lu |
Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. Existing methods encode past shots as compressed context, implicit features, or retrieved frames, making it hard to tell whether new v...Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. Existing methods encode past shots as compressed context, implicit features, or retrieved frames, making it hard to tell whether new visual evidence is another view of an existing appearance, an evolved state of the same identity, or a newly introduced entity. We propose EvoState, an agentic framework that formulates multi-shot generation as entity-centric evolving visual state management, following a state-observation-update-propagation cycle. An entity-centric Visual State Memory separates persistent identity anchors from their evolved states and observations, so each shot retrieves the required state, accumulates complementary views, and reactivates earlier appearances without conflating them. A Visual State Analyzer reconciles the text-visual-memory triplet by observing generated visual evidence, updating generation conditions according to realized outcomes and memory states, and propagating inferred state transitions into memory and future prompts. As a result, generation is conditioned on an evolving state history grounded in what has been realized, rather than solely on planned textual descriptions. Experiments on our curated StoryBench benchmark demonstrate substantial improvements in identity preservation, state-transition correctness, and narrative coherence over state-of-the-art methods.
|
| 350 |
LoopVLA: cross-subtask event memory for looped task execution
2606.17463
|
cs.CV
|
Shoujing Zhu, Zhenyang Liu, Fungmiu Wang, Jiafeng Wang, Bo Yue |
Repetitive tasks are common in real-world manipulation, yet current vision-language-action (VLA) policies remain unreliable for repeated actions. We present LoopVLA, a cross-subtask event-memory interface for looped task execution. Its central idea is to compr...Repetitive tasks are common in real-world manipulation, yet current vision-language-action (VLA) policies remain unreliable for repeated actions. We present LoopVLA, a cross-subtask event-memory interface for looped task execution. Its central idea is to compress a just-completed subtask into memory for the next subtask's action expert, making completed execution available for subsequent repetition and stopping decisions. Learnable event latents query the closed segment's tokens, while action features query history and joint history--event context; the two action-side readouts are residually fused and injected directly into the action head. On a memory benchmark, LoopVLA achieves $80.22\%$ success on the repetition task group with online subgoal prediction, bringing evaluation closer to realistic deployment while improving overall success over strong memory baselines.
|
| 351 |
Quantile Adaptive Temperature Scaling for Confidence Calibration
2606.21749
|
cs.CV
|
Omprakash Chakraborty, Leo Fillioux, Ismail Ben Ayed, Jose Dolz |
Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect. Temperature Scaling remains the most widely used posthoc calibration method due to its simplicity and effectiveness, yet...Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect. Temperature Scaling remains the most widely used posthoc calibration method due to its simplicity and effectiveness, yet its global, uniform rescaling of logits fails to correct the highly heterogeneous structure of miscalibration observed across the confidence spectrum. In particular, the largest correctness confidence discrepancies arise in different quantile regions depending on the setting, low confidence predictions, where uncertainty matters most, tend to exhibit the largest correctness confidence discrepancies, which standard TS leaves largely unaddressed. We introduce Quantile Adaptive Temperature Scaling (QaTS), a simple and efficient post hoc calibration method that adapts the temperature as a function of a predictions empirical confidence quantile. By mapping confidences into the quantile space, QaTS normalizes the calibration problem, makes the structure of miscalibration explicit and enables a monotone temperature function that adapts across quantiles while leaving well calibrated high confidence predictions largely unchanged. preserving high confidence behavior. This quantile aware formulation aligns naturally with a reparameterized Expected Calibration Error (ECE) objective and yields a sample wise temperature that is robust across a variety of challenging scenarios, such as class imbalance and distributional shifts. Across a broad range of datasets, architectures, evaluation scenarios and diverse tasks, QaTS consistently, and substantially, outperforms state of the art post hoc calibration methods, delivering more reliable and trustworthy confidence estimates without modifying model predictions.
|
| 352 |
Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation
2606.22197
|
cs.CV
|
Rui Wang, Quentin Lohmeyer, Siyu Tang, Mirko Meboldt |
Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation-based approaches preserve temporal correspondence but suffer from motion over-factorization, oversmoothing high-frequency dynamics. In contras...Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation-based approaches preserve temporal correspondence but suffer from motion over-factorization, oversmoothing high-frequency dynamics. In contrast, 4D-primitive methods capture fine visual details yet incur temporal overparameterization, breaking object identity and leading to severe storage overhead. To resolve this, we introduce Multi4D, a framework for high-fidelity dynamic Gaussian Splatting based on multi-level competitive allocation. Instead of a monolithic representation, we distribute modeling capacity across three structured levels: static structure, persistent dynamic geometry, and transient appearance primitives. Through shared rasterization and residual-driven optimization, these levels dynamically compete to explain photometric error, enabling adaptive specialization without pre-assigned decomposition. This allocation preserves long-term motion consistency while capturing fine dynamic detail, achieving state-of-the-art rendering quality and real-time performance with significantly fewer dynamic primitives. Furthermore, because our representation explicitly tracks compact persistent Gaussians over time, semantic features can be embedded afterward, enabling Multi4D to achieve state-of-the-art 4D segmentation accuracy with an order-of-magnitude speedup. Project page: https://batfacewayne.github.io/Multi4D.io/
|
| 353 |
PHOEBI: An Open-World Benchmark for Multi-Label Bacterial Identification in Phase-Contrast Microscopy
2606.22890
|
cs.CV
|
Aaditya Baranwal, Md Jahid Hasan, Shruti Vyas |
Optical microscopy (OM) enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environmental, and industrial microbiology. Real samples, however, are routinely polymicrobial and may contain...Optical microscopy (OM) enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environmental, and industrial microbiology. Real samples, however, are routinely polymicrobial and may contain organisms never seen during training, and no computer-vision benchmark evaluates multi-label species identification from phase-contrast microscopy (PCM) of such mixtures. We introduce Phase-contrast Optical bEnchmark for Bacterial Identification ($\textbf{PHOEBI}$), a wet-lab-prepared dataset of $120{,}000$ PCM images covering $40$ combinations of six rod-shaped species, together with a leave-combinations-out (LCO) protocol that holds out entire species combinations, mirroring a model trained on catalogued mixtures that must recognise new ones. Under LCO, gradient-trained per-image classifiers, from fine-tuned backbones to attention-based multiple-instance learning, collapse on unseen combinations despite high in-distribution accuracy, and the failure lies in how per-image predictions are aggregated rather than in the visual representation. We propose three lightweight $\textbf{anchor-based}$ decoders that read each species' presence against fixed geometric prototypes over a shared frozen tile-feature pool, and they remain stable under the same shift. Without additional training, the same features also support open-set rejection of unseen species and the discovery of a new class from unlabeled test images, with negligible disruption to the known classes.
|
| 354 |
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
2606.24797
|
cs.CVcs.AI
|
Linpeng Huang, Weixing Chen, Zexin Chen, Yang Liu, Liang Lin |
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on Video Question Answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the faithfulness of predicted ...Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on Video Question Answering (VideoQA). Nevertheless, existing benchmarks are predominantly evaluated through answer correctness, while the faithfulness of predicted evidence supporting those answers remains insufficiently evaluated. This disconnect between answer generation and evidence verification motivates the construction of the Evidence-Grounded Video Question Answering Benchmark (EG-VQA), a large-scale open-ended benchmark in which each QA pair is annotated with temporally localized textual evidence, requiring models to jointly produce answers and verifiable evidence. EG-VQA comprises 2,067 videos and 11,838 QA pairs with fine-grained evidence annotations. To evaluate predicted evidence, Evidence-Grounded F1 (EG-F1) is introduced as a metric that jointly measures temporal alignment and semantic consistency between predicted and ground-truth evidence. Experiments reveal a substantial discrepancy between answer correctness and evidence faithfulness: even strong proprietary models can achieve high answer accuracy while exhibiting gaps in temporal-semantic grounding. To investigate evidence-aware learning under EG-VQA, we develop EG-Reasoner, an evidence-aware model trained with explicit evidence supervision. EG-Reasoner achieves strong evidence grounding performance among open-source models while maintaining competitive answer generation results compared with proprietary systems. These findings highlight the importance of explicitly modeling evidence grounding and suggest that evidence-aware supervision provides an effective direction toward more reliable and interpretable VideoQA systems.
|
| 355 |
Neural Voxel Dynamics: Learning Volumetric Feature Advection for 3D Physics in V-JEPA Latent Space
2606.26410
|
cs.CV
|
Zican Wang, Niloy Mitra |
We present Neural Voxel Dynamics, a self-supervised framework for learning 3D latent dynamics from monocular video. While generative video models now produce visually compelling motion, their predominantly 2D representations provide limited geometric structure...We present Neural Voxel Dynamics, a self-supervised framework for learning 3D latent dynamics from monocular video. While generative video models now produce visually compelling motion, their predominantly 2D representations provide limited geometric structure for modeling and controlling physical interactions. Instead, we learn dynamics in a lifted volumetric latent space. Specifically, we unproject semantic Video Joint-Embedding Predictive Architecture (V-JEPA) features into a voxel grid using monocular depth priors, producing a geometrically grounded representation that retains rich video features. We then introduce Volumetric Feature Advection, an action-conditioned transition model that predicts the evolution of these latent features directly in 3D. Unlike hybrid approaches that assume access to privileged physics-engine states or explicit simulation during training and/or inference, we learn only using video-derived supervision and, if available, actions, allowing a single latent dynamics model to represent heterogeneous material-dependent phenomena, unifying rigid-body and fluid dynamics. We evaluate across multiple datasets (synthetic and real) and unseen scenarios (e.g., new boundary conditions, OOD generalization), measuring both predictive fidelity and 3D geometric consistency. We report physical plausibility with Physics-IQ using videos decoded from predicted latent states and videos generated from predicted optical flow. Our results show that geometrically lifting pretrained video latent representations provides a scalable route toward dynamic world models that learn structured 3D physical dynamics without access to privileged simulator state.
|
| 356 |
Obliviate: Erasing Concepts from Autoregressive Image Generation Models
2606.28643
|
cs.CV
|
Hossein Shakibania, Jonas Henry Grebe, Tobias Braun, Ege Aktemur, Saleh Aslani |
The widespread adoption of generative AI models has intensified concerns about misuse, including the creation of unsafe or disturbing imagery. To mitigate such issues, several concept erasure approaches have been proposed to remove harmful content from multimo...The widespread adoption of generative AI models has intensified concerns about misuse, including the creation of unsafe or disturbing imagery. To mitigate such issues, several concept erasure approaches have been proposed to remove harmful content from multimodal generative models. Yet concept erasure for autoregressive image generation remains largely unexplored, despite the growing relevance of these models in recent trends toward unified multimodal architectures. In this work, we fill this gap by introducing Obliviate, a guidance-based concept erasure method for autoregressive image generation. Our method builds on three key design choices: KL-based supervision over visual token distributions, trajectory-level updates over full autoregressive rollouts, and aligned visual prefixes for stable target construction. We evaluate Obliviate on three state-of-the-art autoregressive text-to-image models, Liquid, Emu3-Gen, and Janus-Pro, covering the erasure of explicit content, graphic violence, and branded imagery. Obliviate consistently outperforms current alternatives, reducing nudity on the defensive RAB benchmark from 91.58 to 3.15 while preserving overall model utility. Our code is available at https://github.com/multimodal-ai-lab/Obliviate.
|
| 357 |
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
2607.00218
|
cs.CVcs.AI
|
Siddhant Panpatil, Arth Singh, Mijin Koo, Chaeyun Kim, Haon Park |
Vision-language models (VLMs) are increasingly proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming...Vision-language models (VLMs) are increasingly proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction obscured by binary safety benchmarks. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, with two evaluation tracks. The situational track (800 scenarios) spans routine, safe-but-suspicious, obvious-hazard, and contextual-hazard scenes. The visual-channel track (400 scenarios) tests whether misleading in-scene text corrupts physical-safety judgments, using matched truthful controls. Both tracks use contrastive ladders: near-identical scenarios differing in a single visible deciding cue, forcing predictions to hinge on that cue. Across ten open- and closed-source VLMs, we find that guards often recognize videos containing hazards yet miss the specific hazardous moments, especially for contextual hazards. Misleading in-scene signs further degrade all tested guards: vulnerable models miss up to a third of hazards, while seemingly robust models often over-intervene on safe content. Matched controls show that apparent robustness can reflect indiscriminate alarming rather than true physical reasoning. A 20-clip real-video sanity check further shows that model rankings transfer beyond synthetic rendering, with Spearman \r{ho} = 0.87 for hazard miss rate and 0.94 for visual-channel mismatch recall.
|
| 358 |
Caption Bottleneck Models
2607.00578
|
cs.CV
|
Seref Baris Cagliyan, Umut Ozdemir, Merve Tapli, Emre Akbas |
Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an optimal concept set for a specific dataset remains an open challenge. Existing approaches rely on expensive ...Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an optimal concept set for a specific dataset remains an open challenge. Existing approaches rely on expensive expert annotations or LLM-generated lists based solely on class names. Even "open-vocabulary" variants typically depend on static concept sets, which restrict discovery and introduce label bias. Furthermore, traditional CBMs often suffer from information leakage, where unmodeled visual features bypass the bottleneck and compromise the integrity of the explanations. To overcome these limitations, we propose Caption Bottleneck Models (CaBM), a framework that circumvents the need for predefined concept sets by replacing rigid concept layers with free-form natural language. By representing images via LMM-generated captions and training a classifier strictly on this text, CaBM ensures a leakage-free architecture by construction. Additionally, by analyzing the text classifier post-training, CaBM autonomously discovers high-quality, dataset-specific concepts. Our results across fine- and coarse-grained benchmarks demonstrate that CaBM achieves competitive accuracy while preserving interpretability without the constraints of external dictionaries or manual labeling.
|
| 359 |
AnyGroundBench: A Multi-Domain Adaptation Benchmark for Video Grounding in VLMs
2607.02269
|
cs.CVcs.AI
|
Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji |
Vision-Language Models (VLMs) have shown strong performance in Spatio-Temporal Video Grounding (STVG), yet they are still evaluated mostly in a zero-shot manner on general-purpose benchmarks of everyday scenes. This creates a critical disconnect from real-worl...Vision-Language Models (VLMs) have shown strong performance in Spatio-Temporal Video Grounding (STVG), yet they are still evaluated mostly in a zero-shot manner on general-purpose benchmarks of everyday scenes. This creates a critical disconnect from real-world applications in specialized domains, where models inevitably encounter rare visual or textual concepts. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains with limited data is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured, expert-annotated videos with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated limited training subsets, enabling systematic evaluation of domain adaptability under limited training data. We benchmark 23 state-of-the-art VLMs in the zero-shot setting and further evaluate five adaptation strategies, spanning training-free and fine-tuning-based approaches, on representative models, assessing their zero-shot generalization and adaptation capacity. Our results show that current VLMs remain far from practical performance in the zero-shot setting, that training-free adaptation produces highly variable effects depending on the model and domain, and that fine-tuning-based adaptation, though more effective, still falls short of real-world requirements, with gains varying markedly across domains. These findings expose fundamental limitations in current VLMs' spatio-temporal reasoning, pointing to concrete directions for future research.
|
| 360 |
Mask-supervised Object-centric Representation Learning with LeJEPA
2607.02404
|
cs.CVcs.LG
|
Jakob Geusen, Ender Konukoglu |
Self-supervised image encoders deliver strong features for downstream tasks but need many images for training. A natural remedy to counter this is to make each image count for more. A scene contains many objects, and given masks from human annotators or an off...Self-supervised image encoders deliver strong features for downstream tasks but need many images for training. A natural remedy to counter this is to make each image count for more. A scene contains many objects, and given masks from human annotators or an off-the-shelf segmentation model, pre-training can focus on aligning per-object rather than image-wide representations, extracting more signal from every image. Existing mask-supervised methods do this through reconstruction or contrastive losses that leverage negative objects. We instead use two separate projection spaces for the alignment. In a \emph{semantic space}, per-object representations from different views are aligned. To avoid collapse, instead of using negative objects, which requires category definitions, we extend the negative-free LeJEPA objective and show that its distributional anti-collapse regularizer ports naturally from whole images to the variable-sized set of objects in a scene. In an \emph{instance space}, a contrastive loss separates per-object representations from their context and co-occurring instances, including those of the same category. To separate object representations from their context, we copy objects and paste them into other contexts, where each pasted copy serves as an additional view of the original object. Trained on COCO with ground-truth masks, our method outperforms image-level and mask-guided baselines on tracking (DAVIS), classification (ImageNet-1k) and re-identification (NAVI), matches the best of them on semantic segmentation (ADE20k) and keeps its lead over image-level LeJEPA and a supervision-matched alternative on COCO fractions down to 256 images.
|
| 361 |
TACoS: Weakly Supervised Learning of Two-Dimensional Materials from Scribble Annotations to Precise Segmentation
2607.07169
|
cs.CV
|
Jiabei Chen, Liping Zhang, Jiang-Bin Wu, Zhongming Wei, Enhao Ning |
The precise pixel-level localization of 2D material flakes is crucial for high-throughput screening. However, traditional fully supervised methods rely on dense annotations, which are costly and time-consuming, severely limiting the practical deployment of seg...The precise pixel-level localization of 2D material flakes is crucial for high-throughput screening. However, traditional fully supervised methods rely on dense annotations, which are costly and time-consuming, severely limiting the practical deployment of segmentation models. This paper proposes TACoS, a specialized scribble segmentation framework tailored for 2D materials. First, we design a unified framework that integrates semi-supervised consistency learning with structured tree energy constraints. This framework comprises two core components: an unlabeled weak-strong distribution alignment module and a tree energy regularization module. The former employs cosine consistency constraints to enhance prediction alignment across views. Meanwhile, the latter utilizes minimum spanning trees to establish pixel affinity relationships and generate structure-aware soft pseudo labels for online semantic guidance. Next, we introduce asymmetric regional contrast learning. This approach fuses high-confidence predictions from the weak augmentation branch with scribbles to form augmented labels, and construct category prototypes in the representation space. Simultaneously, we prioritize contrastive constraints on challenging pixels in boundary-unlabeled regions. This strategy enhances intra-class cohesion and inter-class separation at the representation level, effectively reducing category confusion in low-contrast edges and complex backgrounds. Experiments conducted on the constructed graphene and MoS2 datasets demonstrate that our method TACoS achieves over 96% of fully supervised performance using less than 0.6% annotated data. Furthermore, it exhibits superior structural coherence and boundary stability in scenarios with weakly contrasting edges and complex backgrounds, providing an efficient and scalable solution for automated high-throughput screening of 2D material flakes.
|
| 362 |
REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation
2607.09082
|
cs.CV
|
Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana, Sandesh Hegde, Debarshi Banerjee |
Training-free in-context segmentation enables new object categories to be introduced at inference time from a single annotated reference image, eliminating the retraining and memory overhead of class-incremental learning. Recent approaches achieve this by comb...Training-free in-context segmentation enables new object categories to be introduced at inference time from a single annotated reference image, eliminating the retraining and memory overhead of class-incremental learning. Recent approaches achieve this by combining vision foundation models for semantic correspondence with promptable segmentation networks like SAM. However, their performance is fundamentally limited by the quality of the cross-image similarity map; shared contextual backgrounds between the reference and query systematically elevate similarity in non-target regions, degrading prompt localization. We present REBASE, a training-free framework that explicitly suppresses these spurious contextual correspondences. Our method identifies the low-rank background feature subspace from the reference image and project the reference and query features onto its orthogonal complement in closed form, yielding cleaner semantic matching. We then generate positive point prompts using similarity-weighted farthest-point sampling, paired with a refined dense similarity prior. Without any training or parameter updates, our approach establishes a new state of the art among training-free methods on PACO-Part, FSS-1000, and cross-domain datasets such as ISIC2018, demonstrating that explicit background subspace removal is a highly effective principle for one-shot localization.
|
| 363 |
Physics-aware Masked Diffusion-based Flood Simulation for Urban Fisheye Disaster Detection
2607.15527
|
cs.CV
|
Sodtavilan Odonchimed, Tsogt Enkhbayar, Oyunzul Munkhtamga, Munkhjargal Gochoo |
Physical simulations that predict the behavior of urban disasters, such as climate-related flooding, play a crucial role in disaster prevention and the development of anomaly detection models. However, the severe shortage of flood data in real-world environmen...Physical simulations that predict the behavior of urban disasters, such as climate-related flooding, play a crucial role in disaster prevention and the development of anomaly detection models. However, the severe shortage of flood data in real-world environments, combined with the inherent distortions of fisheye lens images, which are used for urban surveillance, has made high-precision simulations challenging. To address this, we propose a new physical simulation system PhysFlood that leverages Diffusion Models to synthesize realistic floods from just a single image captured by a fisheye lens. Our system not only enables simulation from a single image, but also features the ability to freely control and generate diverse flood scenarios by manipulating physically meaningful variables, such as water levels. In our evaluation experiments, we conducted a qualitative human study and demonstrated that the simulation images generated by PhysFlood exhibit both acceptable realism and robustness.
|
| 364 |
The JEPA Predictor: A Transferable Operator for Occluded Feature Completion
2607.16274
|
cs.CV
|
William Nguyen, Christopher Nguyen |
Joint-Embedding Predictive Architectures (JEPAs) train a predictor jointly with their encoder, but downstream deployment discards the predictor and reads features from the encoder alone. The predictor is, by construction, a learned operator from visible-contex...Joint-Embedding Predictive Architectures (JEPAs) train a predictor jointly with their encoder, but downstream deployment discards the predictor and reads features from the encoder alone. The predictor is, by construction, a learned operator from visible-context features to features at masked positions, the structure a partial-view classifier needs. We show that this operator is portable across encoder families. We first establish that, at heavy mask, retaining the frozen predictor on a JEPA encoder substantially closes the accuracy gap against the strongest non-JEPA discriminative baselines. We then bolt the frozen predictors of I-JEPA and V-JEPA 2 onto four non-JEPA hosts (CLIP, DINOv3, DINOv2, MAE) through a single linear projection between feature spaces, fit in closed form on 500 ImageNet-1k images. Across both ImageNet-9 and Stanford Dogs and across three mask fractions, the lift over each host's masked-encoder baseline grows monotonically with the mask fraction K in every host-donor pair. CLIP paired with the I-JEPA predictor recovers most of the accuracy that masking removed on ImageNet-9 at heavy occlusion, and lifts fine-grained Stanford Dogs from 15.9% to 52.1% (+36 pp). The mechanism is identifiable: the projection pays a fixed cost on visible patches and the predictor provides a growing benefit on masked patches; the benefit dominates the heavy-occlusion regime. At low K on fine-grained classification the projection cost exceeds the benefit, defining the boundary where the linear bridge breaks down. The frozen JEPA predictor functions as a portable operator for occluded feature completion across encoder families, requiring no retraining of either model while fitting matched linear probes per mask fraction.
|
| 365 |
MuViSeg: Multi-View Segment Correspondences from Dense Geometry Priors
2607.17938
|
cs.CV
|
Denis Fatykhoph, Timur Akhtyamov, Konstantin Pakulev, German Devchich, Gonzalo Ferrer |
Object-level mapping and topological navigation reason about whole objects, yet the correspondences they rely on are computed over keypoints or pixels and only aggregated into objects afterwards. A recent line of work removes this detour by matching directly a...Object-level mapping and topological navigation reason about whole objects, yet the correspondences they rely on are computed over keypoints or pixels and only aggregated into objects afterwards. A recent line of work removes this detour by matching directly at the level of instance segments: a class-agnostic segmenter partitions each image, and per-segment descriptors are pooled from large 3D foundation models over the masks. This shifts the open question to how frozen foundation features should be processed once the unit of matching is a segment. We introduce MuViSeg, three learned matching heads: a LightGlue-style attention head on frozen MASt3R descriptors; a DPT-style multi-scale fusion module that exposes layered spatial detail from VGGT before pooling; and, as our main contribution, a joint multi-view head that attends over segments from several views at once, recovering transitive correspondences that pairwise matchers cannot express. In zero-shot evaluation on Replica and Virtual KITTI 2 across 0--180 degree viewpoint changes, our heads consistently improve over a parameter-free Sinkhorn matcher on the same backbone. Across 102 HM3D navigation episodes, direct segment matchers and sparse keypoint aggregation are statistically indistinguishable, while dense matches voted into masks lose 13.7--18.0 SPL.
|
| 366 |
Medical foundation models converge less under label supervision
2607.20274
|
cs.CVcs.CLcs.LGcs.AI
|
Soroosh Tayebi Arasteh, Sebastian Ziegelmayer, Mahshad Lotfinia, Lisa Adams, Sven Nebelung |
Diagnostic classifiers and imaging biomarkers are fitted on the embeddings of medical foundation models. These models are replaced as new versions appear. This practice assumes that different models represent images alike. We tested this assumption with more t...Diagnostic classifiers and imaging biomarkers are fitted on the embeddings of medical foundation models. These models are replaced as new versions appear. This practice assumes that different models represent images alike. We tested this assumption with more than 750,000 images from 14 datasets in five imaging modalities, 18 public models, and 101 models trained on chest radiographs and histopathology that differ in pretraining objective, label type, size, random seed, initialization, or training patients. Agreement was measured as the mutual k-nearest-neighbor overlap, the share of an image's nearest neighbors common to two models. Two public models shared on average 0.091 of the 10 nearest neighbors of a chest radiograph. A randomly initialized network shared 0.036 with them. In the four other modalities, they shared at most 0.300. Convergence depended on the pretraining objective. Models trained with label supervision converged least in all five controlled settings. On chest radiographs, two label-supervised models that differed only in their random seed shared 0.062 to 0.100 of their nearest neighbors. Two models trained with self-distillation, masked image modeling, or contrastive pretraining shared 0.411 to 0.802. On chest radiographs, agreement depended more on the source dataset and the radiographic view than on the findings. It was not significantly associated with accuracy. A linear mapping between the embeddings of two models, fitted on 4,096 unlabeled radiographs, transferred classifiers for 14 chest findings with 0.987 of their original area under the receiver operating characteristic curve. Convergence can therefore be chosen when a model is trained. A chest radiograph classifier can be transferred to a new model without new labels.
|
| 367 |
DART: A Degradation-Aware Recurrent Transformer for Archival Film Restoration
2607.21219
|
cs.CVcs.LG
|
Miko{\l}aj Jastrz\k{e}bski, Wojciech Koz{\l}owski, Kamil Adamczewski |
Archival film restoration is a challenging problem because historical footage contains compound degradations such as scratches, dust, blur, noise, flicker, and photometric aging, while clean reference videos are unavailable. Existing video restoration methods ...Archival film restoration is a challenging problem because historical footage contains compound degradations such as scratches, dust, blur, noise, flicker, and photometric aging, while clean reference videos are unavailable. Existing video restoration methods largely treat these degradations implicitly, reconstructing frames without explicit knowledge of where damage occurs or how severe it is. We propose DART, a degradation-aware recurrent transformer for archival film restoration. DART predicts and propagates a soft defect mask through time, using it to guide temporal fusion and condition the restoration network on both damage location and severity. This makes the restoration process explicitly aware of film artifacts rather than relying only on reconstruction losses. Experiments on real archival benchmarks show that DART improves no-reference perceptual quality over prior restoration architectures while remaining compact and efficient, producing cleaner and more temporally consistent restorations of structured film damage.
|
| 368 |
ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation
2608.01530
|
cs.CV
|
Mohamed Farag, Genc Hoxha, Yahia Maleki, Chris McCool, Ribana Roscher |
Reliable decision support in digital agriculture requires not only accurate predictions but also well-calibrated uncertainty estimates, particularly for dense prediction tasks such as semantic segmentation. Ensembles provide strong uncertainty quantification b...Reliable decision support in digital agriculture requires not only accurate predictions but also well-calibrated uncertainty estimates, particularly for dense prediction tasks such as semantic segmentation. Ensembles provide strong uncertainty quantification but are computationally and memory demanding, while single-model approximations often sacrifice uncertainty quality. We propose ST-LoRA, a parameter-efficient ensemble that builds diverse members from a single training trajectory by combining Low-Rank Adaptation (LoRA) with snapshot ensembling. All members share a frozen pretrained backbone and differ only in lightweight low-rank adapters, which sharply reduces trainable parameters, checkpoint storage, and I/O overhead. We evaluate SegFormer, Mask2Former, and EoMT on GrowliFlower-L (cauliflower, open field) and BUP20 (sweet pepper, glasshouse), covering in-distribution performance, calibration under covariate shift, and near- and far-out-of-distribution (OoD) detection, with BUTom21 (tomato) as near-OoD data. Extensive ablations show that feed-forward layers, not attention projections, are the critical LoRA target for dense prediction, and that the scaling ratio $\alpha/r$ governs an accuracy--calibration trade-off. Against full-rank snapshot ensembles, ST-LoRA is competitive in segmentation quality, with architecture-dependent training time and energy savings. Against MC Dropout, DDU, and six post-hoc calibrators, it achieves the strongest far-OoD image-level detection and near-OoD pixel-level localization with low cross-seed variance, although full-rank ensembles remain better calibrated. These results show that LoRA-based ensembling offers a compelling efficiency--performance trade-off for agricultural vision systems.
|
| 369 |
Site Is Decodable Before Pretraining: Negative Controls for Probing Frozen Brain-MRI Foundation Models
2608.10295
|
cs.CVcs.AI
|
Saman Rahbar |
Brain foundation models are often tested with a probe. The model is frozen, a simple classifier is trained on its output, and the classifier's accuracy is taken as a sign of what pretraining learned. We show that this reading can be wrong unless two controls a...Brain foundation models are often tested with a probe. The model is frozen, a simple classifier is trained on its output, and the classifier's accuracy is taken as a sign of what pretraining learned. We show that this reading can be wrong unless two controls are reported with it. We probed three frozen 3-D brain-MRI models at five depths, on two independent multi-site cohorts (ABIDE-I and ABIDE-II). The site where a scan was acquired could be predicted at about 0.9, on a scale where 0 is chance and 1 is perfect. No clinical or demographic target came close. An untrained copy of the same model predicted site just as well on both cohorts. Untrained models of three architectures, each built with three random seeds, did the same on ABIDE-I. The raw image, shrunk to 12x12x12 voxels and given to no model at all, predicted site at 0.95. Pretraining did not create the site signal. It was already in the scans. We also looked for any clinical target on which pretraining beats an untrained model, and found no reliable case. Adjusting for age, sex and diagnosis left site prediction almost unchanged, and so did a 4.6-fold increase in the number of voxels. Resting-state fMRI from the same people showed the same pattern. Site was the most predictable attribute of brain connectivity (0.81), and it survived five untrained layers while age did not. We recommend reporting two controls with every probe of a brain foundation model: running the probe again on the untrained model, and on the raw input. We release the code for both.
|
| 370 |
MG-VQA: Manipulation Grounded Visual Question Answering with VLMs
2608.17129
|
cs.CV
|
Vineet Bhat, Mikaela Angelina Uy, Siyi Chen, Alex Zook, Xuning Yang |
Vision-language models (VLMs) have shown promising spatial reasoning capabilities from static visual inputs, where the evidence needed to answer a question is available in the provided views. However, in cluttered environments, answer-relevant evidence may be ...Vision-language models (VLMs) have shown promising spatial reasoning capabilities from static visual inputs, where the evidence needed to answer a question is available in the provided views. However, in cluttered environments, answer-relevant evidence may be occluded rather than absent: an object may lie beneath a pile, be covered by another object, or have identifying information facing away from the camera. Answering such questions requires physical interaction to reveal the hidden evidence. We formalize this setting as Manipulation-Grounded Visual Question Answering (MG-VQA), where an agent answers a question about an initially cluttered scene by using manipulation as an intermediate evidence-gathering operation. We introduce MG-VQA-Bench, comprising 600 human-verified questions across four spatial reasoning tasks, and evaluate it in a cluttered tabletop interactive environment. Eight VLMs are evaluated under three levels of environment access: Direct (image only), Perception (pointing, segmentation, and scene graphs), and Manipulation (perception, grasping, and pushing). Across all models, Direct and Perception achieve average success rates of 36.2% and 37.6%, respectively, only slightly above an image-blind chance baseline of 32.7%. In contrast, Manipulation raises average success to 56.8% (40.8-83.0% across models). Stronger tool-calling VLMs, GPT~6~Astra (83.0%) and Gemini~3.8~Flash (69.3%), search persistently and re-ground after unsuccessful interactions, while weaker models often answer without gathering sufficient evidence or stop prematurely after failed actions. Our results highlight the need for VLM agents that use manipulation for persistent, physically grounded evidence gathering and recovery. Code and benchmark is available here: https://github.com/vineet2104/MG-VQA
|
| 371 |
TRACE: Training-time Report-guided and Clinically Ordered Concept Editing
2608.20809
|
cs.CVcs.AI
|
Wentao Yue, Tianyou Lai, Jiayu Luo, Qingyu Mao, Ziying Wang |
Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising alternative, the...Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising alternative, they often assume complete annotations or require multimodal inputs at inference, which significantly limits their real-world applicability. To tackle these issues, we propose Training-time Report-guided and Clinically Ordered Concept Editing (TRACE), a training-time report-guided framework that leverages structured radiology reports as privileged concept supervision while enabling image-only diagnosis at test time. TRACE refines image-derived concepts through a teacher-guided editing mechanism within a malignancy-aware ordered concept space. To address incomplete annotations, we introduce Strategic Concept Missing Training (SCMT) and train an image-only self-editor via edit distillation for autonomous concept refinement. Besides, we introduce BUSC, a concept-enriched benchmark linking images, labels, and structured attributes. Experiments across multiple datasets demonstrate that TRACE achieves superior performance and improved cross-domain robustness compared to existing methods.
|
| 372 |
Bee Detection and Tracking at Hive Entrance using YOLO11 and ByteTrack
2608.23213
|
cs.CV
|
Thi Thu Thao Nguyen, Johannes Reschke |
This work presents an automatic bee entrance monitoring system based on YOLO11 transfer learning and the ByteTrack tracking algorithm. The study investigates the influence of data augmentation, backbone freezing, and tracker parameter optimization on the detec...This work presents an automatic bee entrance monitoring system based on YOLO11 transfer learning and the ByteTrack tracking algorithm. The study investigates the influence of data augmentation, backbone freezing, and tracker parameter optimization on the detection and counting of small, fast-moving bees. The detector with progressive backbone unfreezing strategy achieved about 97.0% precision and 98.7% mAP50, while providing more stable convergence than full fine-tuning. Experiments also showed that light augmentation outperformed heavy augmentation. For tracking, ByteTrack parameters were optimized to improve trajectory continuity under low-confidence detections. On an independent 25 FPS side-view video, the optimized YOLO11-ByteTrack system correctly counted 43 of 47 incoming bees (91.5%) and 7 of 30 outgoing bees (23.3%). Error analysis showed that most counting errors were caused by missed detections due to rapid bee motion and motion blur, while tracking failures became less frequent after parameter optimization. Overall, the results indicate that moderate augmentation, progressive backbone unfreezing, and ByteTrack tuning improve the reliability of automatic bee entrance monitoring under realistic recording conditions.
|
| 373 |
DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models
2608.23850
|
cs.CV
|
Jeong-gi Kwak, Sho Kagami, Yuki Ono, Kwang Moo Yi |
Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view mod...Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models---their internal knowledge of 3D geometry---into a single-view estimator, we can obtain enhanced 3D consistent foundational features. Our key idea is to construct a multi-view teacher by fusing pretrained 2D foundation features with multi-view geometric features, and refining the fused representation with a discriminative ranking objective. Through our discriminative distillation framework, we enforce the learned features to be both 3D consistent and locally distinctive, while keeping them aligned with the feature space of the original foundation model to preserve the semantic structure of the pretrained representation. Consistency and local discriminability are critical for 3D computer vision problems such as forming semantic and geometric correspondences across images. To demonstrate the effectiveness of our method, we perform comprehensive experiments spanning multiple angles: direct feature analysis, dense prediction transfer, and explicit 3D lifting and rendering. Across these evaluations, our method consistently produces stronger 3D-aware foundation features that improve multi-view consistency and local discriminability while preserving the semantic transferability of the original representation.
|
| 374 |
Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs
2608.26684
|
cs.CV
|
Ji Soo Lee, Jinyoung Park, Seohyun Lee, Jongha Kim, Joonmyung Choi |
Recent large language models achieve strong performance on complex reasoning tasks, where reinforcement learning with Group Relative Policy Optimization (GRPO) has emerged as a leading paradigm for optimizing models on self-generated trajectories. However, the...Recent large language models achieve strong performance on complex reasoning tasks, where reinforcement learning with Group Relative Policy Optimization (GRPO) has emerged as a leading paradigm for optimizing models on self-generated trajectories. However, the on-policy nature of GRPO bounds the model to the reasoning skills it can already produce, restricting to learn more advanced capabilities. Prior works inject privileged reasoning traces from a stronger teacher policy to guide training, yet these traces are inherently out of distribution with respect to the student policy. We observe that this mismatch between on-policy and off-policy causes gradient clipping on semantically critical reasoning tokens, ultimately rewarding correct answers while leaving the reasoning that justifies them unlearned. Hence, we propose \textbf{Echo-GRPO}, a framework that lets the model reason in the words it speaks. Rather than imitating low-probability privileged traces from the teacher model, Echo-GRPO rewrites them into the student policy's own \textit{idiolect}, that is, its own characteristic vocabulary and expression patterns, while preserving their semantics via Dual-Reference Decoding. We instantiate this framework as \textbf{VideoEcho-R1} for video reasoning distillation, achieving consistent improvements across three multimodal LLM backbones and five benchmarks. Finally, we show that our idiolectal paraphrasing is a plug-in module that consistently improves both RL and supervised fine-tuning frameworks for reasoning distillation, demonstrating that policy-aligned supervision extends beyond GRPO.
|
| 375 |
Feature-Spectral Fragility in Segmentation: Dataset Dependence, Architecture-Specific Localization, and Spectral Correlates
2608.29167
|
cs.CV
|
Subhash Kashyap |
Robustness of segmentation models is commonly assessed through input-domain perturbations, while dependence on frequency content within learned feature representations remains less understood. We probe this dependence using targeted post-training low-pass inte...Robustness of segmentation models is commonly assessed through input-domain perturbations, while dependence on frequency content within learned feature representations remains less understood. We probe this dependence using targeted post-training low-pass interventions on internal representations of three segmentation architectures, ResNet50-UNet (CNN), VM-UNet (SSM), and Swin-UNETR (Transformer), across CVC-ClinicDB and ISIC2018, with headline evaluations performed on untouched held-out test sets. At cutoff $\rho=0.25$, feature-domain low-pass filtering causes severe degradation on CVC: Dice drops by 100%, 73.2%, and 30.9% for CNN, SSM, and Transformer, respectively, compared with 9.4%, 10.3%, and 0.6% on ISIC. The cross-dataset difference is statistically significant for every architecture. Single-stage interventions further show that the most sensitive stage depends on architecture and dataset: the CNN peaks at the 2nd to 3rd encoder block, whereas the SSM peaks at the 1st to 2nd encoder stage. Native feature-domain spectral measurements show an inverse association between high-frequency energy and fragility on CVC; the relationship is only partial on ISIC and is therefore treated as a candidate correlate rather than a proven mechanism. Finally, Fourier augmentation improves robustness to input-space low-pass filtering but leaves feature-domain degradation essentially unchanged. These results show that feature-spectral robustness is strongly dataset-dependent, architecture-specific, and distinct from input-domain spectral robustness.
|
| 376 |
Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment
2608.30964
|
cs.CV
|
Kaizhen Tan, Yuantao Deng |
Pretrained vision embeddings are widely used to model how people appraise urban scenes, and they are validated almost entirely by how well they predict human ratings. Accurate prediction shows that an embedding contains the information needed to recover the ra...Pretrained vision embeddings are widely used to model how people appraise urban scenes, and they are validated almost entirely by how well they predict human ratings. Accurate prediction shows that an embedding contains the information needed to recover the ratings. It does not show that the embedding arranges scenes as the human visual system does, which is assumed when distances or dimensions in the embedding are read as perceptual. Using openly released EEG recorded while 63 adults viewed and rated street scenes of Berlin, we compared the neural representational geometry of the scenes with the geometry of pretrained vision models that differ in training objective and size, of simple image descriptors and of the ratings themselves, each relative to a noise ceiling given by the agreement between participants. No feature space reached more than about half of the noise ceiling, which corresponds to about a fifth of the reliable variance in the neural geometry. A descriptor of oriented edge energy reached the level of most pretrained models while capturing a different part of the neural geometry, and deeper layers corresponded to later neural responses. The same embeddings predicted held-out ratings well, up to r = 0.87, yet prediction accuracy and neural correspondence were not reliably related across models, and the best predictor was among the least aligned. Predicting how a street is appraised is therefore weak evidence that a model represents the street as the brain does.
|
| 377 |
You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change
2609.00649
|
cs.CV
|
Kaizhen Tan |
Vision-language models (VLMs) are increasingly used to measure urban change from repeated street-level imagery, and the results are often mapped for individual sample points. Such maps assume that a perception score changes only if the street does. We test thi...Vision-language models (VLMs) are increasingly used to measure urban change from repeated street-level imagery, and the results are often mapped for individual sample points. Such maps assume that a perception score changes only if the street does. We test this assumption with repeated Google Street View captures of unchanged streets in five US cities and with controlled experiments that hold the photograph, scene or camera fixed. Re-photographing an unchanged street shifts its perception score by two-thirds of the average difference between two streets in the same city. Much of this shift arises in the scoring and averages out over question orders; the rest comes from the photographs and is not predicted by image statistics of weather and light. A single image measures a place moderately well, but the change at one location barely separates redeveloped from unchanged streets. All models tested score degraded images as worse-looking streets, and in crowdsourced imagery camera differences cause false reports of physical change unless both images are rendered through a common virtual camera. After aggregation, redeveloped streets are scored as wealthier and better maintained, with more enclosure and less greenery. Streets photographed in the same capture campaign share part of their error, which averaging within a district does not remove. For urban research and planning, VLM perception scores can support comparisons between groups of streets, such as redeveloped and unchanged streets photographed in the same campaigns, but not the identification of individual streets that improved or declined.
|
| 378 |
Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning
2609.00658
|
cs.CV
|
Kaizhen Tan, Yang Feng, Heqing Du, Siru Tao, Xin Xu |
Metric questions about video, such as the speed of a moving object, require a vision-language model to convert visual measurements into physical units using a real-world reference supplied in the prompt. We find that current models use this reference only part...Metric questions about video, such as the speed of a moving object, require a vision-language model to convert visual measurements into physical units using a real-world reference supplied in the prompt. We find that current models use this reference only partially. When every world-space quantity in the prompt is multiplied by a common factor, the prompt still describes the same video and the correct answer changes by exactly that factor, but the predictions of eight models change by less, and their accuracy stays concentrated near the scale that the depicted objects usually have. Asked the same physics in a scale-free form, the two models we test recover the closed-form scaling laws on most items, which indicates that the deficit lies in metric grounding and not in knowledge of the physical mechanism. Because the scaling relation is exact, it can serve as supervision without metric annotations. Equivariance Self-Distillation (EquiSD) projects a model's own prediction onto the functions that satisfy this relation and fine-tunes the model on the resulting targets, with one query per training question and no ground truth. Trained on synthetic video only, EquiSD brings a 3B model close to the exact relation on held-out simulated videos, also at scales not seen in training, and improves its accuracy across scales. Without adaptation, it also improves accuracy across scales on the QuantiPhy benchmark, where its gain reaches 93% of that obtained by supervision with exact simulator answers.
|
| 379 |
FujinSplat: Seeing Through Smoke with RAW-Domain Gaussian Splatting
2609.06017
|
cs.CV
|
Gengjia Chang, Ziteng Cui, Shuhong Liu |
The appearance of a smoky scene is shaped by two processes that a camera records together: the participating medium alters scene radiance in a view-dependent way, and the image signal processor (ISP) then remaps the result through a nonlinear tone and color tr...The appearance of a smoky scene is shaped by two processes that a camera records together: the participating medium alters scene radiance in a view-dependent way, and the image signal processor (ISP) then remaps the result through a nonlinear tone and color transformation. Recovering a clean 3D scene requires separating both. Per-view sRGB dehazing acts only after the ISP has entangled them; standard 3D reconstruction ignores the medium and absorbs it into scene geometry and radiance. FujinSplat addresses the problem in the RAW domain, where the two processes remain separable. A per-scene Base ISP is fitted from the scene's hazy RAW captures to its own camera renderings and then frozen, providing a fixed photometric anchor that performs no dehazing. Analyzing expert corrections reveals a compact, low-dimensional correction space identifiable from RAW alone. FujinSplat therefore fits per-view action answers at the training poses and trains a single scene-agnostic controller to regress them from RAW; the corrected views supervise one static 3D Gaussian representation, jointly with a bounded per-view residual that reconciles cross-view photometric inconsistencies. On the RealX3D real-world smoke benchmark FujinSplat clearly outperforms the strongest comparable baseline, ahead of both physics-based reconstruction and restoration-then-3DGS pipelines. Code is available at https://github.com/I2WM/FujinSplat.
|
| 380 |
FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute
2609.08848
|
cs.CV
|
Hongchi Xia, Tianhang Cheng, Wei-Chiu Ma, Shenlong Wang |
We present FIRE3D, a unified framework that transforms a single RGB image or casual video into interactable 3D scene assets for games and interactive applications in under a minute for up to twelve textured instances including preprocessing. At the core of FIR...We present FIRE3D, a unified framework that transforms a single RGB image or casual video into interactable 3D scene assets for games and interactive applications in under a minute for up to twelve textured instances including preprocessing. At the core of FIRE3D is a feed-forward inference pipeline that predicts a compositional scene representation from posed RGB-D observations estimated from the RGB capture, including the 6-DoF pose, bounding box, mesh, and texture for detected objects. By modeling the scene as a collection of discrete entities, FIRE3D produces amodal object assets that can be independently edited and used in interactive applications. Our framework requires no test-time optimization, runs substantially faster than the compared reconstruction systems, and provides object-level completeness beyond existing feed-forward 3D approaches. We demonstrate leading detection accuracy, strong geometry reconstruction, and competitive rendering quality on the evaluated benchmarks, with substantial runtime gains.
|
| 381 |
HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation
2609.12151
|
cs.CV
|
Imad Ali Shah, Imran Mehmood, Enda Ward, Martin Glavin, Edward Jones |
The HSI-Road dataset provides paired RGB and 25-channel NIR (600-960nm) images with binary masks but no surface-level labels. This paper introduces a manually labeled six-class taxonomy: Background, Asphalt, Concrete, Dirt, Water, and Grass, and an RGB-to-NIR ...The HSI-Road dataset provides paired RGB and 25-channel NIR (600-960nm) images with binary masks but no surface-level labels. This paper introduces a manually labeled six-class taxonomy: Background, Asphalt, Concrete, Dirt, Water, and Grass, and an RGB-to-NIR registration pipeline with corresponding annotations. Six semantic-segmentation models are evaluated under four input configurations: original-resolution RGB, registered low-resolution RGB (RGB$_{\text{reg}}$), NIR, and channel-stacked RGB$_{\text{reg}}$-NIR (RGBN$_{\text{stk}}$). The comparison quantifies the effect of spatial-resolution reduction on RGB, along with evaluation of NIR and RGBN$_{\text{stk}}$, with results reported using per-class and mean IoU and F1 scores. The original-resolution RGB achieves the highest overall performance but contains 12$\times$ more pixels and incurs a 15.5-20.6$\%$ latency penalty compared to the reduced-resolution inputs. At the common 192$\times$384 resolution, RGBN$_{\text{stk}}$ outperforms NIR for all six models and RGB$_{\text{reg}}$ for five of six, with the most consistent gains for the Water class. These results highlight the importance of spatial resolution while showing that NIR provides complementary information to RGB.
|
| 382 |
CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models
2609.18462
|
cs.CVcs.AI
|
Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma |
FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and ...FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.
|
| 383 |
Fragment-Aware Vision Transformers for Fresco-Fragment Style Classification
2609.21012
|
cs.CVcs.LG
|
Sara Miketek, Biagio Barchielli, Nadeem Iqbal Kajla, Sinem Aslan |
Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconographic structure. In archaeological settings, however, artworks often survive only as fragmented remains, forcin...Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconographic structure. In archaeological settings, however, artworks often survive only as fragmented remains, forcing recognition from incomplete, irregular, and context-limited visual evidence. We study fresco-fragment style classification using a progressive transformer-based framework. Starting from a ViT-B/16 baseline, we introduce foreground-guided masking to suppress background-only tokens, inpainting-based geometric regularisation to align irregular fragment supports with the ViT patch grid, and a supervised contrastive objective that operates on predictive distributions through a Kullback-Leibler similarity and consistently improves every branch. We combine the branches with a deliberately simple learnable logit ensemble. Experiments on CLEOPATRA and POMPAAF show that fragment-aware modelling improves over the standard ViT baseline, with the ensemble increasing accuracy from 0.604 to 0.656 and macro-F1 from 0.596 to 0.648 on CLEOPATRA, and outperforming the best single branch in four of six fragmentation settings on POMPAAF. We additionally evaluate a more complex graph-fusion variant and find that it matches the simple ensemble on POMPAAF while offering only a small, dataset-specific gain on CLEOPATRA, which does not justify its added complexity. Beyond these empirical gains, our contribution is twofold: a distribution-level contrastive objective that consistently sharpens single-branch recognition, and an interpretability analysis that verifies the models exploit genuine painted evidence, while quantifying that the inpainting-based branch draws part of its attribution from the synthesised surround.
|
| 384 |
VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration
2609.21521
|
cs.CVcs.AI
|
Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito |
While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching,...While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.
|
| 385 |
VideoGen-Agent: Reinforcing Video Generation Agents
2609.24997
|
cs.CV
|
Binxu Li, Haoyi Duan, Yuhui Zhang, Yaohui Zhang, Zihao Lin |
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. ...Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools. Project page: https://andyca111.github.io/VideoGen_Agent/
|
| 386 |
FFM-CP: Cross-Backbone Fusion of Vision-Language Foundation Models for Few-Shot Computational Pathology
2609.27710
|
cs.CVcs.LG
|
Anh-Tien Nguyen, Trung DQ. Dang, Nghiem Tuong Diep, Bui Ngoc Han Nguyen, Tan-Ha Mai |
Pathology vision-language foundation models vary in performance across diseases and tasks, with no single model consistently performing best. The high cost of expert pathology annotation can also limit the labeled data available for task-specific adaptation. C...Pathology vision-language foundation models vary in performance across diseases and tasks, with no single model consistently performing best. The high cost of expert pathology annotation can also limit the labeled data available for task-specific adaptation. Combining complementary pretrained representations is a potential approach to these limitations, yet learning an effective fusion from few labeled examples remains challenging. We introduce Few-shot Fusion Foundation Models of Computational Pathology (FFM-CP), which is a framework that combines multiple pathology vision-language models in the few-shot learning setting. The framework first aligns heterogeneous representations using a closed-form Orthogonal Procrustes transformation estimated from corresponding support images. This alignment preserves within-model feature geometry without training an additional alignment network. Within the aligned space, a unified graph enables information exchange across backbones by jointly refining support-image features and visual and textual class prototypes. These refined representations support complementary text-prototype and case-retrieval branches that capture semantic class knowledge and within-class visual variation, respectively. Each branch learns to combine predictions from all ordered backbone pairs, allowing queries encoded by one model to draw on evidence represented by another. We evaluate three backbone combinations on six histopathology datasets at 4, 8, and 16 shots per class. FFM-CP achieves higher mean macro-F1 than the strongest individually adapted member of each fused set in 50 of 54 comparisons. These findings suggest that combining complementary pretrained representations can improve histopathological classification when annotations are limited.
|
| 387 |
FLIP: Final Layer Inference-Time Probing for Vision-Language Models
2609.30993
|
cs.CVcs.AI
|
Drandreb Earl O. Juanico, Rowel O. Atienza |
We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under intern...We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise flooring to the final normalized hidden state before logit computation, leaving parameters, prompts, and decoding unchanged. On a controlled detection/counting probe, sweeping intervention strength reveals three regions: negligible change, a bounded interior regime in which detection recall at IoU 0.50 ($R_{50}$) improves while tolerant counting error ($\mathcal{E}_{\mathrm{count}}$) falls, and over-suppression. We formalize a four-criterion probe-and-sweep protocol for disciplining the interpretation of intervention effects: regime structure, grounding-proxy alignment, feature-coherence dependence, and failure to reproduce the same positive regime on a performance-based negative control. The post-normalization state passed to the output head is the logit-facing instantiation of this test; under a non-targeted flooring sweep it satisfies the full protocol. Raw decoder-layer interventions, including the last-block output before final normalization, and the singleton-pair left/right control fail to reproduce the Final-site signature, while same-site operators and multiple VLMs replicate it. FLIP is therefore a validation step for intervention-based mechanistic interpretability, not a steering method. Repo at https://github.com/earl-juanico/flip-qwen3vl/
|
| 388 |
Towards Transparent Diagnostics: Investigating Architectural Trade-offs and Explainability in Malaria Detection
2609.31682
|
cs.CV
|
Suman Kunwar, Avishek Dangol |
More than 80 countries have reported malaria cases with 610 thousand deaths and are projected to increase. Identifying malaria early and accurately helps save lives and effective way to diagnose malaria is through microscopic methods that are labor intensive a...More than 80 countries have reported malaria cases with 610 thousand deaths and are projected to increase. Identifying malaria early and accurately helps save lives and effective way to diagnose malaria is through microscopic methods that are labor intensive and require experts with special equipment. Deep learning (DL) has shown promising results in medical diagnosis. Here, we explored various DL models: ResNet18, MobileNetV2, EfficientNet-B2, VGG19 and proposed model ResNet18+TTA (ResNet18 backbone with modified classification head and test time augmentation) for detecting malaria presence using blood smears taken from the NIH Malaria dataset. Our experiment shows MobileNetV2 achieved 96.85 % accuracy with smallest model size (8.49 MB) and fastest inference (1.35 ms). The ResNet18+TTA model achieved 97.96 % accuracy, 0.996 AUC with longest inference time (13.32 ms). Larger architecture outputs a larger model size with moderate accuracy. Upon further pruning, ResNet18+TTA model gained a slight improvement in accuracy and reduced inference time. GRAD-CAM, SHAP and LIME provide explainable AI (XAI) insights into model predictions, using explanation agreement and divergence to evaluate predictive reliability.
|
| 389 |
Levy-Driven Correspondence Estimation for Registration
2609.32612
|
cs.CV
|
Qianliang Wu, Jiaqi Yang, Wankou Yang, Le Hui, Jin Xie |
Finding reliable point correspondences is difficult when point clouds have low overlap or undergo non-rigid deformation. Iterative refinement can correct uncertain matches, but costly network evaluations limit the number of updates. We present LevyMatch, a L\'...Finding reliable point correspondences is difficult when point clouds have low overlap or undergo non-rigid deformation. Iterative refinement can correct uncertain matches, but costly network evaluations limit the number of updates. We present LevyMatch, a L\'evy-driven method that uses random jumps to refine a soft matching matrix. At each step, a network uses the current matching state and geometric information to predict a target matching matrix. A Brownian reference bridge gives an explicit formula for the update toward this target. A Gamma random clock sets the time step for each update. The updated matches provide new geometric feedback for the next target prediction. We further propose a fixed front-loaded Gamma policy that assigns more expected clock time to early updates and less to later ones, without retraining or extra network evaluations. Reordering the same sampled Gamma increments shows that placing larger increments early gives higher accuracy than placing them late. On 4DMatch and 4DLoMatch, our method improves both non-rigid feature matching recall (NFMR) and inlier ratio (IR) over the compared methods. The front-loaded policy achieves 93.09% NFMR and 92.11% IR on 4DMatch, and 82.79% NFMR and 79.07% IR on 4DLoMatch.
|
| 390 |
ReVision3D: Attribution-Guided Recursive Self-Improvement for 3D Medical Perception
2609.32984
|
cs.CV
|
Ho Hin Lee, Yuyin Zhou, Yannan Yu, Shi Gu, Yifan Wu |
Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe,...Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe, or downstream inference, while self-generated feedback and logged trajectories provide little guidance on which component should change. We introduce ReVision3D, an RSI system that leverages 3D volumes with spatially grounded annotations to determine where visual evidence is lost and recursively improve the corresponding visual capability. A frozen language-model designer proposes revisions to acquisition, perception, training, or inference, while the verifier and system-level objective remain fixed. Our key insight is that an annotated volume forms an exact replay world for view rendering and spatial verification: unvisited views can be rendered on demand, and localized predictions can be checked directly against reference masks. This grounded feedback directs targeted revision, while only changes that improve beyond measured seed noise are retained. Each accepted change triggers renewed attribution, allowing the dominant bottleneck to shift across rounds. On abdominal CT, attribution identifies perception as the dominant remaining limitation. Revising that level enables ReVision3D to achieve 79% liver recall and 83% kidney recall at under 0.4 false positives per patient, outperforming the evaluated frozen multimodal foundation models, with the largest gains on small lesions. Our framework is public available with interactive demo in this project page: \url{https://leeh43.github.io/ReVision3D/}
|
| 391 |
Octree-based Video Representation
2609.33100
|
cs.CV
|
Rungui Zhou, Chuanzhi Zhou, Yuk-Kit Hou, Peng-Shuai Wang |
Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight s...Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.
|
| 392 |
VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking
2609.34390
|
cs.CV
|
Zhizhen Li, Zan Wang, Huidong Peng, Bohan Tan, Shimin Shan |
Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly...Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly supporting broad animal coverage, large-scale video data, and extensive within-video multi-instance association. To address this gap, we introduce VastMAT, which has four key characteristics: (1) Large scale. It comprises 2,947 videos with 1,002,562 annotated frames, totaling 27.85 hours. (2) Broad category coverage. These videos cover 337 animal categories with diverse morphologies and motion patterns. (3) Extensive instance annotations. It provides 3,663,248 bounding boxes and 22,883 identity trajectories---to our knowledge, the largest numbers of both among dedicated MAT benchmarks. (4) High-quality annotations. To ensure reliability, annotations undergo iterative expert review and correction, and quality is assessed through an independent reannotation audit. To systematically assess tracking performance and cross-category generalization, we establish Seen-category and category-disjoint Unseen-category protocols, and evaluate eight representative MOT methods under both protocols. Under these protocols, the highest baseline HOTA scores are 66.37\% and 52.90\%, respectively, highlighting the challenge of tracking unseen animals. To address the low-overlap association challenge revealed by our analysis, we propose Center-Distance-Augmented Association (CDA), a lightweight module that adaptively combines IoU with center similarity normalized by the boxes' own scales. Without additional training, CDA improves TrackTrack's HOTA by 1.58 and 1.31 percentage points under the two protocols, respectively. To facilitate further MAT research, we will publicly release our benchmark and code.
|
| 393 |
Revisit to Segment: Working Memory Distillation for Reasoning Segmentation
2609.34863
|
cs.CV
|
Cilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu, Xiaolong Jiang |
Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when r...Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when revisiting the same image and query. Our exploration reveals that MLLMs benefit from using this self-generated working memory as context, leading to enhanced reasoning segmentation. Motivated by this finding, we seek to strengthen the backbone model's reasoning segmentation capabilities by distilling the guidance gained from revisiting prior attempts, enabling it to benefit with or without working memory at inference time. To this end, we propose Reasoning Segmenter with Working Memory (SWiM), a working-memory distillation framework for reasoning segmentation. Specifically, SWiM selects rollouts based on segmentation quality to construct working memory and uses the memory-conditioned model as a teacher. The teacher provides token-level distributional supervision along student-generated trajectories, while the student receives only the original image and query. Joint optimization of on-policy self-distillation and outcome-based reinforcement learning combines working-memory guidance with direct feedback on segmentation quality. Extensive experiments on reasoning segmentation benchmarks demonstrate that SWiM achieves state-of-the-art performance, validating the effectiveness of working-memory distillation.
|
| 394 |
CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes
2609.36024
|
cs.CV
|
Shuzhao Xie, Lelin Wang, Guying Lin, Zhi Wang, Minchen Li |
Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depend...Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object-level generated meshes guide the agent toward detailed, compact geometry; articulated rigid objects are decomposed into movable parts with explicit joints. For deformables, category-wise agent sessions reconstruct curves as centerlines with radii, surfaces as manifold shells with thickness, and volumes as watertight solids for volumetric meshing. Reusable simulator skills initialize compatible physical models and parameters, while agent-guided behavioral tests expose mismatches and trigger targeted revisions of motion, geometry, numerics, or material modeling. On evaluated Replica and ScanNet++ scenes, CoDimRecon achieves competitive compositional reconstruction accuracy while additionally producing deformable assets for rod, shell, and solid simulation. We further demonstrate robot interactions across all three representations, including a controlled paper-folding case in which behavioral testing motivates plastic bending.
|
| 395 |
Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models
2609.36531
|
cs.CV
|
Estela Monserrat Arriaga Santana, Julian Rosas Scull, Eh\'ecatl Sacamch'en N\'u\~nez Rico, Hugo Jair Escalante |
Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgme...Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.
|
| 396 |
SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence
2609.36545
|
cs.CV
|
Gyeonggwan Lee, Eunsoo Im, Seunghwan Hong, Junghun Suh |
Dense feature matching between 360$^\circ$ panoramas underpins omnidirectional pose estimation, 3D reconstruction, and SLAM. Such panoramas are stored in the equirectangular projection (ERP), which unrolls the viewing sphere onto a flat chart and thereby intro...Dense feature matching between 360$^\circ$ panoramas underpins omnidirectional pose estimation, 3D reconstruction, and SLAM. Such panoramas are stored in the equirectangular projection (ERP), which unrolls the viewing sphere onto a flat chart and thereby introduces three distinct distortions -- a longitudinal seam (topology), latitude-dependent stretch (metric), and non-uniform pixel area (area) -- that the coarse stage of perspective-trained dense matchers does not model, so these matchers degrade systematically on ERP. We show that correcting the three distortions at the coarse-stage interfaces where they arise -- pairwise distortions in attention, per-pixel distortion in covisibility gating -- improves PCK@1$^\circ$ from 0.230 to 0.275 on Matterport3D under a fixed coarse scaffold, with the refiner architecture unchanged -- our central result. Concretely, SCCM (Spherically Consistent Coarse Matching) augments a chart-naive cross-attention/dual-softmax coarse matcher with two sphere-derived priors: Spherical Positional Attention (SPA) pairs a yaw-periodic RoPE (topology) with a tangent-plane bias (metric), and Area-Aware Covisibility (AAC) applies a pre-sigmoid log-area correction (area). The chart-naive scaffold serves as a controlled reference, separating the scaffold-replacement effect from the spherical-prior effect. Instantiated in the RoMa V1 framework with the same frozen encoder, refiner architecture, and loss, SCCM also outperforms the ERP-native EDM (0.163) and an ERP-retrained RoMa V1 (0.198) under a unified ERP dense matching protocol, while perspective-trained matchers largely fail on ERP. It further transfers zero-shot to Stanford2D3D and, when trained on outdoor Holo360D, leads there as well.
|
| 397 |
ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking
2609.36677
|
cs.CV
|
Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie |
Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and pr...Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This representation preserves uncertainty about alternative trajectories through successive observations. This belief predicts the next camera, arrival time, and entry region, while appearance and language evidence guide association. By training across successive handoffs, the model learns to retain uncertainty that remains useful for later predictions and identity decisions. ReWorld-Track achieves HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, with improved identity continuity across repeated handoffs. On MTMMC, its structured posterior update gains 0.50 HOTA points over a similarly sized generic updater and 0.94 points over fixed-moment soft association, raising next-camera accuracy from 86.03% to 87.41% and reducing median arrival-time error from 0.78 s to 0.71 s for subsequent target returns.
|
| 398 |
EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding
2609.36803
|
cs.CV
|
Yuwei Miao, Xuesheng Zhang, Wenhao Zou, Jixia Zhang, Jianwei Lv |
Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher rece...Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.
|
| 399 |
S4VY: Segment Anything in Feed-Forward 4D Visual Geometry
2609.36875
|
cs.CV
|
Jingdong Zhang, Xin Li, Jan Kautz, Wenping Wang, Chris Choy |
Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, wh...Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.
|
| 400 |
Why MLLMs Struggle to Count: Overcoming Individuation and Aggregation Bottlenecks with ConvStack
2609.37096
|
cs.CV
|
Liwei Che, Yihao Quan, Sen Fang, Hongyi Wang, Ranjay Krishna |
Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inh...Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we reveal an individuation bottleneck stemming from image patchification: because Vision Transformers process patches independently, they struggle to group fragmented geometric features across boundaries into distinct object representations. Second, we identify a collapse in the subsequent counting aggregation process, where representation separation rapidly diminishes as numerosity increases due to attention compression. Identifying and formalizing these twin bottlenecks constitutes our first major contribution. To overcome them, we propose ConvStack, a lightweight architecture that operates directly in the visual token space to explicitly aggregate and inject local spatial structures via zero-initialized residual connections. By explicitly addressing the individuation bottleneck, ConvStack provides unambiguous geometric evidence for downstream aggregation. Remarkably, by fine-tuning exclusively on counting tasks, the model achieves substantial improvements in dense object counting and broader spatial understanding benchmarks, without compromising on general visual capabilities.
|
| 401 |
AESOP: Asymmetric Human-Camera Generation with Translation-Intensity Control
2609.37229
|
cs.CV
|
Jingzhong Lin, Zhanke Wang, Heng Li, Wenxiang Liu, Zhao Zhang |
Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human mo...Human motion defines an action, while a camera trajectory determines how it is presented. Camera generation for a given human motion and joint human-camera generation are usually treated as separate tasks, although both share an asymmetric dependency: human motion can be generated independently, whereas the camera responds to the realized action. We introduce AESOP, a unified framework with an independent human pathway and a shared human-conditioned camera generator. Its asymmetric architecture serves both tasks while preserving the human output during camera generation. Although human context anchors the shot to the action and camera text describes its movement, translation intensity remains underspecified. We therefore construct trajectory pairs that differ in camera translation magnitude while sharing human motion and camera text, then use these pairs to learn an explicit intensity condition. Experiments on the PulpMotion dataset demonstrate strong camera distributional and framing quality in both tasks and effective control over camera translation intensity.
|
| 402 |
Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models
2609.37537
|
cs.CV
|
Arian Komaei Koma, Seyed Amir Kasaei, Aida Aryafar, Matin Ghiasi, Ali Aghayari |
Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack...Machine unlearning has emerged as a critical post-hoc safety measure to erase sensitive concepts from Text-to-Image (T2I) models without prohibitive retraining. However, we reveal that current state-of-the-art (SOTA) approaches are brittle due to a severe lack of robustness to noise initialization. We call this phenomenon ``probabilistic forgetting'': suppressed concepts re-emerge under specific random initial noise conditions, despite appearing unlearned on other initializations. We trace this failure to the misalignment between standard Gaussian sampling during unlearning and the unlearning objective. Since the target concept manifests only in specific initial noise regions throughout the unlearning phase, uniform random sampling yields sparse, uninformative gradient updates that fail to drive robust erasure. To overcome this issue, we propose an adaptive, concept-conditioned sampling strategy that dynamically concentrates gradient updates on regions where the target concept manifests, down-weighting uninformative areas. We integrate our framework with six distinct SOTA unlearning methods across four diffusion backbones and evaluate it across safety, object, and artistic-style unlearning, as well as under black-box and white-box adversarial attacks. Our method reduces the conditional nudity re-emergence rate across random initializations by 67.2% on average over four baselines and lowers attack success rates across both adversarial evaluations. Across concept domains, Adaptive Noise Sampling strengthens adversarial robustness and non-target retention while preserving competitive generative quality and target-erasure performance.
|
| 403 |
DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
2609.39096
|
cs.CV
|
Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou, Zihan Ding |
Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local atte...Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's value in long-term retention: tokens with larger step-to-final discrepancies tend to carry visual evidence that is less predictable from the retained context. DeCoPrune measures each current-chunk token's denoising difficulty using the discrepancy between its intermediate clean prediction and final denoised value, retaining high-discrepancy tokens in the long-term cache while pruning those with low discrepancy. To evaluate information retention, we introduce CMBench, comprising 58 approximately one-minute generated or real-world context episodes and 116 Reappear or Revisit continuation tasks that require recalling specific previously observed objects or scenes. Experiments with LingBot World v2 show that DeCoPrune preserves near-FullKV long-range recall while pruning over 85% of historical KV tokens and accelerating continuation generation by over $4\times$, substantially outperforming the evaluated compression baselines at comparable budgets. These results indicate that denoising consistency can serve as a model-intrinsic signal for retaining long-range information while reducing autoregressive inference cost. Our project homepage is https://decoprune.github.io. The code is available at https://github.com/DeCoPrune/CMBench, and the benchmark at https://huggingface.co/datasets/Aoraku/CMBench.
|
| 404 |
Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM-Based Urban Sensing
2610.00031
|
cs.CV
|
Kaizhen Tan |
Street-view imagery is increasingly analysed with vision-language models (VLMs) to infer urban attributes, but predictive accuracy alone does not show how much a photograph contributes beyond data already available for the same place. Using three VLMs, we comp...Street-view imagery is increasingly analysed with vision-language models (VLMs) to infer urban attributes, but predictive accuracy alone does not show how much a photograph contributes beyond data already available for the same place. Using three VLMs, we compare image-based predictions with existing urban data for seven attributes drawn from five public resources. Each urban unit is evaluated with imagery, with location or text context, and against non-image predictions from nearby observations or public records. We also replace images and add conflicting records to test which source the predictions follow. Nearby observations or public records matched or exceeded image-only predictions for road damage, curb ramps, population, and house price. Images were more informative for building type, building function, and the floor count of low-rise buildings. For floor count, the advantage of images over nearby OpenStreetMap labels increased by 5.7 percentage points per doubling of distance to the nearest labelled building and declined for tall buildings whose rooflines often fell outside the frame. When images and records disagreed, predictions usually moved towards the supplied record. Released OpenFACADES floor annotations, generated with OpenStreetMap floor values as input, showed the same dependence: their agreement with the reference increased with building height, whereas that of image-only reruns decreased. The value of street-view imagery therefore depends on whether an attribute is visible and how well the place is already covered by existing data. Because machine-derived labels are often reused as references, these comparisons also bear on how urban datasets are documented and evaluated.
|
| 405 |
PACT: End-to-End Learning of Human Pose, Contacts, and Forces from Video
2610.00451
|
cs.CVcs.LGcs.AI
|
Rikhat Akizhanov, Yangsong Zhang, Nikolai Kaliazin, Peter Wolf, Yoshihiko Nakamura |
Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate...Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.
|
| 406 |
RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation
2610.00970
|
cs.CVcs.AI
|
Minsu Kim, Jaesung Choe, Jiwoo Lee, Yu-Chiang Frank Wang, Seon Joo Kim |
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations be...Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
|
| 407 |
CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment
2610.01166
|
cs.CVcs.AI
|
Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme |
Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measureme...Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.
|
| 408 |
EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes
2610.01210
|
cs.CV
|
Hongming Fu, Jingcheng Shi, Wenjia Wang, Binhua Zuo, Bo Zhao |
Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate han...Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.
|
| 409 |
PROWBench: Do Video Models Render What the Program Specifies?
2610.02205
|
cs.CV
|
Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu, Fengbo Lan |
Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing bench...Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.
|
| 410 |
Sphere Encoder 2
2610.02208
|
cs.CV
|
Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia |
Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equato...Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at https://github.com/kaiyuyue/sphere2.
|
| 411 |
SCION: Scene Composition with Instanced Neural Primitives
2610.02322
|
cs.CV
|
William Koch, Amogh Joshi, Cyrus Vachha, Cheng Zheng, Felix Heide |
Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstrac...Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hierarchical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this representation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives. The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts. Project webpage: https://light.princeton.edu/SCION
|
| 412 |
Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation
2610.02914
|
cs.CV
|
Yunseung Ok, Hyunsoo Kim, Minseo Kim, Suhyun Kim |
Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use ...Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30s than causal image-to-video and reference-to-video models, while generating each frame 9.5-28.5 times faster than these long-video baselines.
|
| 413 |
LAS-CLIP: A Lightweight Adapter Steering Approach for CLIP's Visual Encoder
2610.03370
|
cs.CV
|
Anh-Khoa Dinh-Duc, Duc-Tai Dinh, Tam V. Nguyen, Minh-Triet Tran |
CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, ...CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation.
|
| 414 |
Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries
2208.08697
|
cs.CVcs.LG
|
Manaar Alam, Shubhajit Datta, Debdeep Mukhopadhyay, Arijit Mondal, Partha Pratim Chakrabarti |
The security of deep learning (DL) systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks. Despite overwhelming promises, the deep learning systems ...The security of deep learning (DL) systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks. Despite overwhelming promises, the deep learning systems are vulnerable to crafted adversarial examples, which may be imperceptible to the human eye, but can lead the model to misclassify. Protections against adversarial perturbations on ensemble-based techniques have either been shown to be vulnerable to stronger adversaries or shown to lack an end-to-end evaluation. In this paper, we attempt to develop a new ensemble-based solution that constructs defender models with diverse decision boundaries with respect to the original model. The ensemble of classifiers constructed by (1) transformation of the input by a method called Split-and-Shuffle, and (2) restricting the significant features by a method called Contrast-Significant-Features are shown to result in diverse gradients with respect to adversarial attacks, which reduces the chance of transferring adversarial examples from the original to the defender model targeting the same class. We present extensive experimentations using standard image classification datasets, namely MNIST, CIFAR-10 and CIFAR-100 against state-of-the-art adversarial attacks to demonstrate the robustness of the proposed ensemble-based defense. We also evaluate the robustness in the presence of a stronger adversary targeting all the models within the ensemble simultaneously. Results for the overall false positives and false negatives have been furnished to estimate the overall performance of the proposed methodology.
|
| 415 |
A Unified Deep Learning Framework for Motion Correction in Medical Imaging
2409.14204
|
cs.CV
|
Jian Wang, Razieh Faghihpirayesh, Danny Joca, Polina Golland, Ali Gholipour |
Deep learning has shown significant value in medical image registration for motion correction; however, current techniques are either limited by the type and range of motion they can handle or require iterative inference and/or retraining for new imaging data....Deep learning has shown significant value in medical image registration for motion correction; however, current techniques are either limited by the type and range of motion they can handle or require iterative inference and/or retraining for new imaging data. To address these limitations, we introduce UniMo, a Unified Motion Correction framework that uses deep neural networks to correct various types of motion in medical imaging. UniMo uses an alternating optimization scheme with a unified loss function to train an integrated model of 1) an equivariant neural network for global rigid motion correction and 2) an encoder-decoder network for local deformations. It features a geometric deformation augmenter that 1) enhances the robustness of global motion correction by addressing local deformations, whether caused by non-rigid motion or geometric distortions, and 2) generates augmented data to improve training. As a hybrid model that uses both image intensities and shapes, UniMo is robust to appearance variations and generalizes to various imaging modalities without retraining. We trained and tested UniMo for motion tracking in fetal magnetic resonance imaging, which is challenging due to 1) both large rigid and non-rigid motion and 2) large variations in image appearance. We then tested the trained model, without retraining, on three public datasets: MedMNIST, lung CT, and BraTS. UniMo surpassed existing motion correction methods in accuracy and, notably, enabled one-time training on a single modality while maintaining high stability and adaptability across multiple unseen imaging datasets. By offering a unified solution to motion correction, UniMo marks a significant advance in challenging applications with a mixture of bulk motion and local deformations. Code is available at https://github.com/IntelligentImaging/UNIMO
|
| 416 |
Optimizing Breast Cancer Detection in Mammograms: A Comprehensive Study of Transfer Learning, Resolution Reduction, and Multi-View Classification
2503.19945
|
cs.CVcs.AI
|
Daniel G. P. Petrini, Hae Yong Kim |
Mammography, an X-ray-based imaging technique, remains central to the early detection of breast cancer. Recent advances in artificial intelligence have enabled increasingly sophisticated computer-aided diagnostic methods, evolving from patch-based classifiers ...Mammography, an X-ray-based imaging technique, remains central to the early detection of breast cancer. Recent advances in artificial intelligence have enabled increasingly sophisticated computer-aided diagnostic methods, evolving from patch-based classifiers to whole-image approaches and then to multi-view architectures that jointly analyze complementary projections. Despite this progress, several critical questions remain unanswered. In this study, we systematically investigate these issues by addressing five key research questions: (1) the role of patch classifiers in performance, (2) the transferability of natural-image-trained backbones, (3) the advantages of learn-to-resize over conventional downscaling, (4) the contribution of multi-view integration, and (5) the robustness of findings across varying image quality. Beyond benchmarking, our experiments demonstrate clear performance gains over prior work. For the CBIS-DDSM dataset, we improved single-view AUC from 0.8153 to 0.8343, and multiple-view AUC from 0.8483 to 0.8658. Using a new comparative method, we also observed a 0.0217 AUC increase when extending from single to multiple-view analysis. On the complete VinDr-Mammo dataset, the multiple-view approach further improved results, achieving a 0.0492 AUC increase over single view and reaching 0.8511 AUC overall. These results establish new state-of-the-art benchmarks, providing clear evidence of the advantages of multi-view architectures for mammogram interpretation. Beyond performance, our analysis offers principled insights into model design and transfer learning strategies, contributing to the development of more accurate and reliable breast cancer screening tools. The inference code and trained models are publicly available at https://github.com/dpetrini/multiple-view.
|
| 417 |
ZeST: an VLM-based Zero-Shot Traversability Navigation for Unknown Environments
2508.19131
|
cs.CVcs.AI
|
Shreya Gummadi, Mateus V. Gasparino, Gianluca Capezzuto, Marcelo Becker, Girish Chowdhary |
The advancement of robotics and autonomous navigation systems hinges on the ability to accurately predict terrain traversability. Traditional methods for generating datasets to train these prediction models often involve putting robots into potentially hazardo...The advancement of robotics and autonomous navigation systems hinges on the ability to accurately predict terrain traversability. Traditional methods for generating datasets to train these prediction models often involve putting robots into potentially hazardous environments, posing risks to equipment and safety. To solve this problem, we present ZeST, a novel approach that treats repeated VLM outputs as stochastic measurements and fuses them into an uncertainty-aware posterior. Our approach not only performs zero-shot traversability and mitigates the risks associated with real-world data collection but also accelerates the development of advanced navigation systems, offering a cost-effective and scalable solution. To support our findings, we present navigation results, in both controlled indoor and unstructured outdoor environments. As shown in the experiments, ZeST provides safer navigation with 90-100% success rate with up to 4s inference delays when compared to other state-of-the-art methods, constantly reaching the final goal.
|
| 418 |
The Universal Weight Subspace Hypothesis
2512.05117
|
cs.CVcs.LGcs.AI
|
Prakhar Kaushik, Shravan Chaudhari, Ankit Vaidya, Rama Chellappa, Alan Yuille |
We show that deep neural networks trained across diverse tasks exhibit remarkably similar low-dimensional parametric subspaces. We provide the first large-scale empirical evidence that demonstrates that neural networks systematically converge to shared spectra...We show that deep neural networks trained across diverse tasks exhibit remarkably similar low-dimensional parametric subspaces. We provide the first large-scale empirical evidence that demonstrates that neural networks systematically converge to shared spectral subspaces regardless of initialization, task, or domain. Through mode-wise spectral analysis of over 1200 models - including 500 Mistral-7B LoRAs, 500 Vision Transformers, and 50 LLaMA-8B models - we identify universal subspaces capturing majority variance in just a few principal directions. By applying spectral decomposition techniques to the weight matrices of various architectures trained on a wide range of tasks and datasets, we identify sparse, joint subspaces that are consistently exploited, within shared architectures across diverse tasks and datasets. Our findings offer new insights into the intrinsic organization of information within deep networks and raise important questions about the possibility of discovering these universal subspaces without the need for extensive data and computational resources. Furthermore, this inherent structure has significant implications for model reusability, multi-task learning, model merging, and the development of training and inference-efficient algorithms, potentially reducing the carbon footprint of large-scale neural models.
|
| 419 |
MTRACE: Multilingual Retrieval-Augmented Generation for Temporally Diverse Text Corpora
2512.12694
|
cs.CV
|
Souhail Bakkali, Anthony Mudet |
Large multilingual knowledge bases expose temporally diverse information, yet retrieval quality remains sensitive to lexical variation and cross-lingual terminology shifts. We develop and evaluate MTRACE (Multilingual Temporal Retrieval-Augmented Generation wi...Large multilingual knowledge bases expose temporally diverse information, yet retrieval quality remains sensitive to lexical variation and cross-lingual terminology shifts. We develop and evaluate MTRACE (Multilingual Temporal Retrieval-Augmented Generation with evidence grounding), a pipeline designed to test whether query expansion and multi-query fusion mitigate vocabulary mismatch in temporally layered text corpora, on the French and English subsets of MIRACL. Our approach integrates: (i) semantic query expansion (SQE) and multi-query fusion via Reciprocal Rank Fusion (RRF), targeting retrieval stability under query variation; (ii) a generation prompt enforcing strict grounding in retrieved evidence and explicit abstention when evidence is insufficient; and (iii) a modular architecture enabling systematic component evaluation. Ablation studies on Named Entity Recognition (NER) and embedding model selection demonstrate the importance of syntactic coherence in entity extraction and of self-retrieval and efficiency measurements for retriever selection. Our end-to-end evaluation over 50 constructed queries shows faithful answers for well-supported queries, correct abstention on unanswerable questions, and no re-scored similarity gains from multi-query fusion over single-query dense retrieval. By scoping our claims to a clean, text-only baseline, we separate these effects from OCR-noise confounds; direct measurement of diachronic lexical drift is left to future work. Code and configurations are available at \url{https://anonymous.4open.science/r/MIRAGE-8EAA/
|
| 420 |
De-occluding broadband metalens
2601.19403
|
cs.CVcs.AI
|
Seungwoo Yoon, Dohyun Kang, Eunsue Choi, Sohyun Lee, Seoyeon Kim |
Obstructions such as raindrops, fences, or dust degrade captured images, especially when mechanical cleaning is infeasible. Conventional solutions to obstructions rely on a bulky compound optics array or computational inpainting, which compromise compactness o...Obstructions such as raindrops, fences, or dust degrade captured images, especially when mechanical cleaning is infeasible. Conventional solutions to obstructions rely on a bulky compound optics array or computational inpainting, which compromise compactness or fidelity. Metalenses composed of subwavelength meta-atoms promise compact imaging, but simultaneous achievement of broadband and obstruction-free imaging remains a challenge, since a metalens that images distant scenes across a broadband spectrum cannot properly defocus near-depth occlusions. Here, we introduce a learned split-spectrum metalens that enables broadband obstruction-free imaging. Our approach divides the spectrum of each RGB channel into pass and stop bands with multi-band spectral filtering and learns the metalens to focus light from far objects through pass bands, while filtering focused near-depth light through stop bands. This optical signal is further enhanced using a neural network. Our learned split-spectrum metalens achieves broadband and obstruction-free imaging with relative PSNR gains of 32.29% and improves object detection and semantic segmentation accuracies with absolute gains of +13.54% mAP, +48.45% IoU, and +20.35% mIoU over a conventional hyperbolic design. This promises robust obstruction-free sensing and vision for space-constrained systems, such as mobile robots, drones, and endoscopes.
|
| 421 |
TRACER: Texture-Robust Affordance Chain-of-Thought for Deformable-Object Region Grounding
2601.20208
|
cs.CV
|
Wanjun Jia, Kang Li, Fan Yang, Mengfei Duan, Wenrui Chen |
The central challenge in robotic manipulation of deformable objects lies in aligning high-level semantic instructions with physical interaction points under complex appearance and texture variations. Existing vision-based affordance prediction methods often su...The central challenge in robotic manipulation of deformable objects lies in aligning high-level semantic instructions with physical interaction points under complex appearance and texture variations. Existing vision-based affordance prediction methods often suffer from boundary overflow and fragmented functional regions. To address these issues, we propose TRACER, a Texture-Robust Affordance Chain-of-Thought for Deformable-Object Region Grounding framework that maps hierarchical semantic reasoning to appearance-robust and physically consistent interaction regions. Specifically, a Tree-structured Affordance Chain-of-Thought (TA-CoT) decomposes high-level task intentions into hierarchical affordance-semantic instructions. A Spatially-Constrained Boundary Refinement (SCBR) mechanism suppresses prediction spillover and guides responses toward valid object regions. Furthermore, an Interactive Convergent Refinement Flow (ICRF) refines dispersed affordance responses into coherent regions, improving spatial continuity and physical plausibility. Experiments on the Fine-AGDDO15 dataset and a real-world robotic platform demonstrate that TRACER improves affordance grounding precision across diverse textures and patterns. It also enhances long-horizon manipulation success, bridging high-level semantic reasoning and low-level physical execution. The source code and dataset will be made publicly available at https://github.com/Dikay1/TRACER.
|
| 422 |
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
2602.10719
|
cs.CV
|
Sining Ang, Yuguang Yang, Chenxu Dang, Canyu Chen, Cheng Chi |
Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, but how VLM representations differ from vision-only encoders after policy learning, and whether such differences matter for planning, remains unclea...Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, but how VLM representations differ from vision-only encoders after policy learning, and whether such differences matter for planning, remains unclear. Under a unified VLM-hidden + diffusion-policy paradigm, we compare multiple VLM families/scales (InternVL3 and Qwen3VL) with standard vision-only encoders (ResNet, ViT, and EVA-CLIP) while keeping the downstream planner fixed. We study representation, behavior, and system design. CKA/CCA and Shared--Unique SAE show that policy learning enlarges a common decision subspace, but both branches retain non-transferable residual factors. We further replicate this shared-plus-unique representation pattern on nuPlan using the AsyncDrive planning stack. Latent-intervention policies and scenario-level analysis further show that these residuals are behaviorally meaningful: vision-only encoders are stronger in simple geometry-dominant scenes, whereas VLMs are more effective in semantically complex and interaction-heavy long-tail cases. The two branches also exhibit distinct progress--braking and path-choice tendencies, and an oracle best-of-two VLM+ViT selector reaches 93.58 PDMS on NAVSIM. We convert this complementarity into two lightweight systems: HybridDriveVLA, which selects from a compact cross-model candidate set using a learned trajectory scorer and improves PDMS from 90.80 to 92.10, and DualDriveVLA, a fast--slow variant that invokes the VLM in only 15% of scenarios, achieving 91.00 PDMS with about $1.9\times$ lower latency than the VLM baseline. Code will be available at https://github.com/WilliamXuanYu/HybridDriveVLA.
|
| 423 |
SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
2603.07865
|
cs.CVcs.SDeess.AS
|
Ayush Barik, Sofia Stoica, Nikhil Sarda, Arnav Kethana, Abhinav Khanduja |
Text-to-audio (T2A) diffusion models generate high-quality audio but require tens of neural function evaluations (NFEs), resulting in substantial inference latency. Existing acceleration methods primarily optimize each generation trajectory independently, even...Text-to-audio (T2A) diffusion models generate high-quality audio but require tens of neural function evaluations (NFEs), resulting in substantial inference latency. Existing acceleration methods primarily optimize each generation trajectory independently, even when reusable acoustic structure exists in prior examples. We introduce SoundWeaver, the first training-free framework for compositional cross example acoustic reuse. Rather than generating from pure noise, SoundWeaver constructs a prompt-conditioned acoustic prior from segments distributed across cached audios and warm-starts diffusion from an intermediate noise level. We introduce an Acoustic Composition Graph (ACG), which uses neural codec representations and segment-level identification to identify candidate cross-cache transitions and structured path search to jointly optimize prompt relevance and acoustic continuity. We further establish a warm-start error bound showing how prior quality and diffusion SNR jointly govern the admissible skip depth, motivating an adaptive contextual-bandit Skip Gater that selects the warm-start level for each request. Across T2A diffusion backbones, SoundWeaver achieves 1.48-3.74x generation speedup with only a 1K-clip cache while improving generation quality.
|
| 424 |
EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
2604.09568
|
cs.CVcs.CL
|
Tianfu Wang, Leilei Ding, Ziyang Tao, Yi Zhan, Zhiyuan Ma |
High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often l...High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intuitive flexibility. To bridge this gap, we introduce EvoDiagram, an agentic framework that generates object-level editable diagrams via an intermediate canvas schema. EvoDiagram employs a coordinated multi-agent system to decouple semantic intent from rendering logic, resolving conflicts across heterogeneous design layers. Additionally, we propose a design knowledge evolution mechanism that distills execution traces into a hierarchical memory of domain guidelines, enabling agents to retrieve context-aware expertise adaptively. We further release CanvasBench, a benchmark consisting of both data and metrics for canvas-based diagramming. Extensive experiments demonstrate that EvoDiagram exhibits excellent performance and balance against baselines in generating editable, structurally consistent, and aesthetically coherent diagrams. Our code is available at https://github.com/AuraX-AI/EvoDiagram.
|
| 425 |
FLASH: Efficient Visuomotor Policy via Sparse Sampling
2605.15492
|
cs.CV
|
Jiaqi Bai, Jindou Jia, Yuxuan Hu, Gen Li, Xiangyu Chen |
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present Fast Legendre-p...Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy), which replaces discrete action-chunk generation with continuous Legendre polynomial trajectory representation. Specifically, by fitting expert demonstrations under sparse temporal sampling, FLASH enables a single inference to cover a significantly extended action horizon. To further accelerate generation, FLASH initiates the flow matching process from history polynomial coefficients rather than uninformative Gaussian noise, shortening the transport distance and enabling accurate single-step inference. Moreover, analytic polynomial differentiation directly provides desired velocity feed-forward signals to the torque controller without numerical approximation. Extensive experiments on five simulated and two real-world manipulation tasks demonstrate that FLASH achieves state-of-the-art success rates ($\ge 92\%$ across all tasks), a per-episode inference time of $31.40\,ms$ (up to $175\times$ faster than diffusion policies and $18\times$ faster than prior flow matching policies), up to $4\times$ faster training convergence than ACT, and $5\times$ to $7\times$ reduction in controller tracking error compared to discrete-action baselines.
|
| 426 |
QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning
2605.16813
|
cs.CV
|
Yiheng Zhang, Zhe Zhu, Tingrui Shen, Zhuojiang Cai, Tianxiao Li |
The generation of production-ready quad-dominant meshes is a cornerstone of modern 3D content creation. Generating anisotropic quad-dominant meshes from point clouds is challenging, as existing methods are typically limited to producing either pure triangular ...The generation of production-ready quad-dominant meshes is a cornerstone of modern 3D content creation. Generating anisotropic quad-dominant meshes from point clouds is challenging, as existing methods are typically limited to producing either pure triangular meshes or pure quadrilateral meshes with isotropic densities. In this paper, we present QuadLink, a unified framework consisting of three stages for quad-dominant mesh generation by linking points into structured faces. QuadLink formulates polygonal mesh generation as a hybrid centroid-conditioned vertex linking model: it first predicts a unified set of anchors (vertices and face centroids), then learns centroid-conditioned links that associate vertices with face centroids, and finally assembles polygonal faces with a quad-first strategy guided by robust geometric verification strategies. This link-based formulation enables efficient generation of sparse and anisotropic quad-dominant meshes with coherent edge flow and meanwhile supporting hybrid polygonal topology. To construct training data for this model, we further introduce a Tri-to-Quad Operator that converts artistic triangle meshes into quad-dominant training data via global merge selection. Extensive experiments show that QuadLink produces production-ready quad-dominant meshes from point clouds and achieves improved geometric fidelity and topological quality compared to prior baselines. Our method natively supports hybrid polygonal topology, generalizing to arbitrary n-gon meshes without architectural changes.
|
| 427 |
Adaptive Fused Prior Transfer for Controllable Generative Image Compression
2605.16817
|
cs.CV
|
Yifei Pei, Ying Liu, Nam Ling |
At very low bitrates, image compression discards fine textures and local structures, while distortion-oriented reconstruction often produces over-smoothed images. Generative codecs synthesize missing details, but existing codebook-based controllable designs ge...At very low bitrates, image compression discards fine textures and local structures, while distortion-oriented reconstruction often produces over-smoothed images. Generative codecs synthesize missing details, but existing codebook-based controllable designs generally rely on single-codebook reconstruction priors. We propose Adaptive Fused Prior Transfer for Controllable Generative Image Compression (AFP-GIC), which transfers an image-adaptive fused prior from a frozen pretrained AdaCode model. Encoder-side prior features guide latent formation, while the decoder predicts a compatible fused prior from the compressed representation and control variables, without transmitting the prior itself. A motivating analysis shows that better decoder-side prior alignment tightens a reconstruction-error upper bound and that the fused-prior family includes single-codebook choices as special cases. A single pretrained model supports five evaluated bitrate operating points. Under the unified benchmark, AFP-GIC achieves 18.1% lower decoder latency and uses 31.10 million (20.5%) fewer inference parameters than DC-VIC. Experiments on Kodak, CLIC2020, and DIV2K show competitive PSNR and SSIM, with the clearest naturalness gains in NIQE scores and very-low-bitrate visual comparisons. Code: https://github.com/yifeipet/AFP_GIC.
|
| 428 |
Sample-wise Targeted Adversarial Attacks on Test-time Adaptation
2605.23411
|
cs.CVcs.LG
|
Phuc Duc Nguyen, Quang Duc Nguyen |
Test-time adaptation (TTA) mitigates distribution shifts by adapting models to unlabeled test inputs, but also exposes them to adversarial manipulation. Existing class-wise targeted attacks remain suboptimal for stealthy exploitation in this setting: since TTA...Test-time adaptation (TTA) mitigates distribution shifts by adapting models to unlabeled test inputs, but also exposes them to adversarial manipulation. Existing class-wise targeted attacks remain suboptimal for stealthy exploitation in this setting: since TTA operates on batches, forcing a subset of samples toward a target label unintentionally pulls similar benign samples along, resulting in a conspicuously high frequency of the target label that is easy to detect. To capture a more realistic threat, we introduce a sample-wise targeted attack. Unlike prior approaches, the attacker aims to misclassify only inputs carrying an attacker-chosen trigger, while preserving a benign-like prediction distribution to evade detection. To achieve this, we propose a meta-learning-based attack with a novel priority-aware gradient alignment strategy that explicitly prioritizes attack success. The strategy formulates the gradient update as an ellipsoidal trust-region problem, mitigating gradient conflict between the attack and stealth objectives, while providing theoretical guarantees for effective optimization of the attack objective in the presence of gradient misalignment. Extensive experiments on CIFAR-10-C, CIFAR-100-C, and ImageNet-C across TTA protocols demonstrate that our method achieves high targeted success rates while maintaining prediction behavior close to benign adaptation under multiple label-free stealth metrics, making it difficult to detect in unlabeled TTA deployment scenarios. Furthermore, we demonstrate that our attack shows strong robustness against existing defenses.
|
| 429 |
DrawVideo: Grounded and Faithful Multi-Shot Video Generation from Storyboard Keyframe Sketches
2605.23508
|
cs.CVcs.AIcs.MM
|
Chuanzhi Xu, Huiqi Liang, Bang Shi, Huiming Zhang, Guangcheng Lin |
Long video generation requires high-fidelity visual synthesis, coherent narrative organization, shot-level structure, and explicit user control. Existing text-to-video methods typically generate videos from a single long-form prompt, making it difficult for cr...Long video generation requires high-fidelity visual synthesis, coherent narrative organization, shot-level structure, and explicit user control. Existing text-to-video methods typically generate videos from a single long-form prompt, making it difficult for creators to directly control character pose, camera composition, spatial layout, and local motion. We propose DrawVideo, a sketch-guided and storyboard-driven framework that grounds multi-shot video generation in creator-provided spatial structure, appearance, and motion. Each shot is specified by a black-and-white sketch, an appearance prompt, and a motion prompt. DrawVideo first generates a structure-aligned reference keyframe, expands the motion description into derivative keyframes representing ordered action states, and then synthesizes local video clips between adjacent keyframes. We further introduce SketchLongVideo, to the best of our knowledge the first evaluation dataset of ordered multi-shot storyboards with aligned sketch, appearance, and motion conditions. Extensive experiments demonstrate strong structural controllability, appearance consistency, intra-shot visual stability, and motion alignment, providing an effective solution for director-oriented video creation.
|
| 430 |
Flash-WAM: Modality-Aware Distillation for World Action Models
2606.05254
|
cs.CVcs.LG
|
Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen |
World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation ha...World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods break down in the joint video-action setting because video and action streams use different SNR-shifted noise schedules and reach training with substantially different marginal noise distributions, an asymmetry that single-modality distillation methods cannot accommodate. We introduce \textbf{Flash-WAM}, a modality-aware step-distillation framework inspired by consistency distillation that selects the consistency function for each modality to match its noise regime: a linear-gradient-scaling parametrization for the action stream's low-noise regime, paired with a variance-preserving parametrization for the video stream's high-noise regime, grounded in a structural analysis of the consistency-function family that characterizes the achievable gradient scaling under the consistency boundary condition. Instantiated on LingBot-VA, Flash-WAM compresses inference to a single step in each modality. On RoboTwin 2.0, this reduces per-chunk latency from $8.1$ seconds to $348$ ms on NVIDIA L40S, a $23{\times}$ speedup that enables real-time inference. Flash-WAM preserves task success on simulation benchmarks ($85.5\%$ RoboTwin 2.0, $95.7\%$ LIBERO) and substantially recovers real-world performance ($60\%$ average on a Unitree G1 humanoid robot), while naive consistency distillation drops to $24\%$ at the same step budget.
|
| 431 |
Gefen: Optimized Stochastic Optimizer
2606.13894
|
cs.CVcs.CLcs.LGcs.AI
|
Nadav Benedek, Tomer Koren, Ohad Fried |
AdamW is a default optimizer for deep learning, but its moment states add two parameter-sized buffers to training memory, increasing the cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment est...AdamW is a default optimizer for deep learning, but its moment states add two parameter-sized buffers to training memory, increasing the cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first moment using a learned codebook. Gefen reduces AdamW's optimizer memory footprint by up to 8x while maintaining performance, saving 6.5 GiB per billion parameters. Prior work shares second moments across parameters grouped along the Hessian's block-diagonal structure, but relies on hand-specified architectural rules and leaves unexplained why such grouping works. We prove that large mixed Hessian entries constrain the ratio of squared gradients toward one, explaining why shared second moments are accurate when the squared gradients they pool are similar. The Hessian need not be computed: its block structure is inherited by squared gradients, allowing blocks to be found directly. Gefen therefore infers block structure from initial squared gradients, requiring no architecture-specific metadata or user-tuned hyperparameters beyond AdamW defaults. Gefen learns an exact histogram-based dynamic-programming quantization codebook and reuses the blocks for first-moment scaling. Across diverse pretraining experiments, Gefen achieves the lowest peak optimizer memory among compared methods that maintain AdamW-level performance. In single-machine and distributed training, the reduced footprint enables larger microbatches and substantially improves throughput over AdamW, making Gefen a drop-in replacement that can train larger models or use larger global batch sizes. We provide the complete Python implementation, including fused CUDA kernels at https://github.com/ndvbd/Gefen
|
| 432 |
FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision
2606.14049
|
cs.CVcs.SD
|
Shiyao Wang, Xijuan Zeng, Hui Wang, Shiwan Zhao, Feng Deng |
We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized, versatile audio synthesis for diverse tasks. Existing VTA methods either have mu...We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized, versatile audio synthesis for diverse tasks. Existing VTA methods either have multi-modal control but weak temporal alignment or strong alignment but lack reference audio conditioning and semantic precision. FoleyGenEx fills this gap via three core innovations: a conditional injection mechanism for audio-controlled VTA and Foley extension, a multi-modal dynamic masking strategy preserving training synchronization, and an adverb-based data augmentation algorithm leveraging signal processing and large language models to enhance textual supervision with nuanced semantics. Experiments on AudioCaps, VGGSound, and Greatest Hits demonstrate its competitive controllable VTA performance against existing methods. Demo samples are available at https://foleygenex.github.io/FoleyGenEx.
|
| 433 |
Edge-Aligned Beam Placement in Scanning Probe Tomography via Reconstruction-Free Sequential Design of Experiments
2606.21713
|
cs.CV
|
San Dinh, Zichao Wendy Di, Matt Menickelly |
In X-ray scanning probe tomography, reconstruction quality generally improves with larger numbers of projections. However, additional projections increase experiment costs, acquisition time, and the radiation dose imparted to the sample. One mitigation to thes...In X-ray scanning probe tomography, reconstruction quality generally improves with larger numbers of projections. However, additional projections increase experiment costs, acquisition time, and the radiation dose imparted to the sample. One mitigation to these trade-offs is to adopt a sequential design of experiments, in which each subsequent measurement is determined as a function of previously acquired data in order to maximize information gain. In scanning probe tomography, a widely used heuristic to maximize information is to align beams with the edges of the sample. A key challenge, however, is that the true sample is unknown, so identifying edge-aligned beams typically requires reconstructing the sample based on available measurements. This work proposes a novel sequential design method that identifies edge-aligned measurements directly from the sinogram, bypassing any reconstruction, thereby improving computational efficiency and reducing the experimental design's susceptibility to reconstruction errors. Our method dynamically selects the next set of measurement beams by maximizing an acquisition function that balances exploration and exploitation over the domain of all possible measurements, improving reconstruction quality while reducing measurement redundancy.
|
| 434 |
AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression
2606.24286
|
cs.CVcs.CL
|
Yijing Chen, Wenhui Tan, Xiaoyi Yu, Yuyue Wang, Xin Cheng |
Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, w...Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video understanding in Omni-modal Large Language Models. AVOC introduces a learnable token compression module between the modality encoders and the LLM backbone. We reframe multimodal token compression as a top-$K$ retrieval problem: given a fixed context budget, the module must retrieve a compact subset of tokens that best supports answering the user query. We draw inspiration from three classical Information Retrieval criteria for selecting informative units from a large candidate pool: relevance, importance, and diversity. AVOC instantiates each criterion as a tailored mechanism for audio-video understanding, and integrates them into a unified retrieval-style compression pipeline. Experiments show that AVOC achieves state-of-the-art performance on long-form audio-video benchmarks, surpassing the second-best model by 4.9 and 5.5 points in average accuracy on OmniVideoBench and LVOmniBench, respectively. Moreover, AVOC maintains robust performance on Audio-Video Needle-in-a-Haystack task at durations up to one hour. Code and model are at github.com/YJCX330/AVOC.
|
| 435 |
Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds
2606.26964
|
cs.CVcs.AI
|
Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang, Zhi Wang |
As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D ...As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as Narrative-Grounded World Visual Attention, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose Look-Before-Move, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on StoryBlender, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.
|
| 436 |
Mask2Real-WM: Controllable Dexterous World Models via Segmentation Masks as a Sim-to-Real Bridge
2607.04546
|
cs.CVcs.LGcs.AI
|
Riccardo O. Feingold, Davide Liconti, Chenyu Yang, Robert K. Katzschmann |
Learning action-conditioned world models for dexterous manipulation that are genuinely controllable requires capturing complex, high-dimensional hand kinematics from limited real-world data. We present Mask2Real-WM, a two-stage world model that improves contro...Learning action-conditioned world models for dexterous manipulation that are genuinely controllable requires capturing complex, high-dimensional hand kinematics from limited real-world data. We present Mask2Real-WM, a two-stage world model that improves controllability for dexterous hands under a limited real-data budget by decoupling pixel prediction into a dynamics model and a rendering model. The dynamics model predicts future segmentation masks from past masks and a high-dimensional action sequence. The rendering model converts the predicted masks into photorealistic RGB images. This design allows us to train the dynamics model on a large synthetic dataset spanning the full range of hand motions and interactions. We then fine-tune the dynamics model and train the rendering model on only 2.5 h of real demonstrations to obtain a controllable world model. We compare five models with matched training recipes on a 23-DoF robotic system, measuring per-DoF controllability both with random target commands and with sinusoidal per-DoF actuation, judged blind by human raters. The decoupled model achieves the strongest controllability, with further gains from simulation data. Beyond improved controllability, Mask2Real-WM produces sharper frames while maintaining comparable perceptual quality.
|
| 437 |
GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory
2607.05543
|
cs.CV
|
Hu Zhu, Bohan Li, Xianda Guo, Yanlun Peng, Hongsi Liu |
Embodied agents exploring indoor environments require reliable semantic occupancy memory that persists across observations and revisits. Building such memory is challenging because each observation provides incomplete and uncertain geometric and semantic evide...Embodied agents exploring indoor environments require reliable semantic occupancy memory that persists across observations and revisits. Building such memory is challenging because each observation provides incomplete and uncertain geometric and semantic evidence. We introduce GEM-Occ, a Gaussian Evidence Memory framework that consolidates evidence accumulated over time into persistent semantic occupancy memory. Local predictions are converted into occupied semantic Gaussians and free-space ray evidence. Confidence- and visibility-aware causal updates integrate supporting observations, suppress occupancy contradicted by observed free space, and preserve previously observed structures through occlusion. A hierarchical memory organization supports continued mapping and efficient queries across connected indoor spaces. To evaluate this capability, we introduce HIOcc, a unified benchmark for embodied semantic occupancy memory. HIOcc establishes a shared semantic label space and evaluation framework spanning local prediction, room-level online mapping, and building-level mapping, while accommodating perspective and panoramic observations. Experiments on HIOcc demonstrate that GEM-Occ outperforms existing methods, enabling accurate semantic occupancy prediction and consistent online mapping across spatial scales with efficient memory usage and fast occupancy queries.
|
| 438 |
UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
2607.06306
|
cs.CVcs.AI
|
Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen, Yenchi Tseng |
Large language models (LLMs) have demonstrated growing competence in generating web pages from UI screenshots, which convey both visual structure and cues to application behavior. Yet most screenshot-to-code benchmarks emphasize visual fidelity, while interact...Large language models (LLMs) have demonstrated growing competence in generating web pages from UI screenshots, which convey both visual structure and cues to application behavior. Yet most screenshot-to-code benchmarks emphasize visual fidelity, while interactive generation benchmarks often supply behavioral specifications or demonstrated transitions. Whether models can infer and realize interactions from static screenshots alone remains insufficiently evaluated. We introduce UI2App to evaluate interaction inference: inferring and realizing application behavior from static visual cues without added behavioral guidance. UI2App comprises 600 screenshots organized into 95 state-coherent sets for runnable multi-route web applications. Our end-to-end pipeline evaluates each artifact along three dimensions: executability, visual fidelity, and interaction inference. The interaction metric (IIS) assesses functional correctness and state-management complexity, crediting valid implementations rather than requiring a match to a single reference. Experiments on six frontier vision-language models reveal a marked mismatch between visual fidelity and interaction realization: the visual-fidelity leader scores only 8.1 on IIS, ranking fourth, while the IIS leader achieves 4.5 times that score. High-complexity interactions such as cross-route state persistence remain a major bottleneck, with five of the six models scoring at most 3.5 on this dimension. Overall, these results highlight interaction inference as a key challenge in generating functional web applications from static screenshots.
|
| 439 |
Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning
2607.12112
|
cs.CVcs.LGcs.AI
|
Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li |
Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting,...Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired knowledge across visual, linguistic, and cross-modal representations. Addressing this challenge is especially critical for autonomous networked AI operating in safety-sensitive domains, such as content moderation, where reliable retention of prior knowledge underpins system integrity. To overcome this, we propose Federated Continual Multimodal Learning (FedCMM), a framework that embeds continual-learning safeguards into the federated optimization loop at three complementary levels. At the parameter level, modality-aware elastic weight consolidation computes separate Fisher information matrices for the vision encoder, language backbone, and cross-modal projector, providing granular, asymmetry-aware protection against modality-specific forgetting. At the data level, each client trains a lightweight local generative replay module to synthesize raw-data-free embedding-level multimodal replay tuples without any raw data sharing. At the aggregation level, Task-similarity-aware gradient aggregation autonomously filters and reweights client updates by gradient cosine similarity, suppressing conflicting directions and stabilizing the global learning trajectory. Extensive experiments on two benchmarks demonstrate that FedCMM consistently outperforms recent baselines on accuracy and backward transfer, confirming that holistic, modality-aware optimization enables robust evolutive adaptation across heterogeneous networked AI deployments.
|
| 440 |
The Label Complexity of Useful Class-Conditional Prediction Sets under Distribution Shift
2607.18088
|
cs.CVcs.LG
|
Weijia Han, Lisha Qu, Tianxin Zhou, Zhenda Li, Liying Liang |
Prediction sets can make deployed classifiers safer by returning several plausible labels when a single prediction is uncertain. Their value depends on classwise reliability: average coverage can meet its target while rare or difficult classes fail repeatedly....Prediction sets can make deployed classifiers safer by returning several plausible labels when a single prediction is uncertain. Their value depends on classwise reliability: average coverage can meet its target while rare or difficult classes fail repeatedly. This concern is sharper after distribution shift, when calibration labels come from a source environment but reliability is needed on the target. We ask what labeled source data and unlabeled target inputs reveal about class-conditional prediction sets, and when target labels are necessary. Under unrestricted joint shift, two target laws can produce the same observable data while requiring different classwise thresholds; any label-free rule covering both must enlarge its sets on one law. We give a labeled target audit that estimates the missing quantiles with a simultaneous guarantee. Probability-scale error is invariant to increasing score transformations, and threshold recovery follows under local regularity. At fixed confidence, achieving threshold tolerance $\varepsilon$ with fixed, nonadaptive class-stratified labeled pairs has total complexity $\Theta(K\varepsilon^{-2}\log K)$, or $\Theta(\varepsilon^{-2}\log K)$ labels per class under equal allocation. Class imbalance creates a separate acquisition cost; for foreground class probabilities of order $1/K$, the mixed-stream label complexity is also $\Theta(K\varepsilon^{-2}\log K)$ at fixed confidence. Experiments on action-recognition and image shifts show that marginal coverage can conceal severe class failures and that source classwise calibration depends on the shift. The results connect the information available at deployment to the target labels needed for useful class-conditional prediction.
|
| 441 |
GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors
2608.00946
|
cs.CVcs.LGcs.AI
|
Jibao Yuan, Yuhui Zhao, Yinzhen Lv, Chao Xu, Shun Li |
Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis on GraspNet-1Billion shows that detector confidence is often poorly aligned with grasp quality, leaving successful grasp candidates at low ranks. Motiv...Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis on GraspNet-1Billion shows that detector confidence is often poorly aligned with grasp quality, leaving successful grasp candidates at low ranks. Motivated by this observation, we study whether learned re-ranking can improve candidate ordering while keeping detector parameters and grasp candidates unchanged. We propose GraRe, which estimates grasp quality from candidate attributes, shell-stratified local geometry, and object context. Candidate attributes condition the local geometric and object-context representations, and a Transformer fuses all three feature types. The predicted quality is combined with detector confidence to produce the final ranking. Experiments on GraspNet-1Billion with five frozen detectors show consistent improvements, with gains of up to 13.56 points in Average AP. Real-robot experiments further demonstrate robust grasping in cluttered scenes. These results show that improving candidate ranking provides a practical way to enhance frozen 6-DoF grasp detectors. Project code is available at https://minakanmi-yuki.github.io/grare/.
|
| 442 |
Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
2608.14430
|
cs.CVcs.LG
|
Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He |
Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized li...Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout samples. This paper shows that these seemingly different losses arise from a single path-space principle. Starting from the regularized diffusion-RL objective, we use importance sampling between sampling SDEs to obtain an explicit policy-gradient estimator on trajectory space. The estimator contains the stochastic It\^o integral underlying Flow-GRPO-type updates; we derive an equivalent variance-reduced value-gradient form that recovers the forward-matching structure of AWM and DiffusionNFT. This identifies the empirical gap between these method families as a variance-reduction effect rather than a difference in RL principle. The derivation yields a unified design space organized by value-gradient estimation, weight functions, and sampling choices. Within this space, we propose a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones. Experiments on SD3.5-M and Qwen-Image models validate the variance-reduction explanation and show that the resulting recipe improves over prior diffusion-RL baselines.
|
| 443 |
RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion with Matrix-Valued Trust for Multimodal Prediction under Modality Uncertainty
2609.10798
|
cs.CVcs.LG
|
Yingfan Xu, Tieming Liu, Ye Liang |
Multimodal prediction from images and structured metadata requires integrating complementary evidence whose reliability can vary across samples and latent directions. A single confidence weight per modality cannot capture this directional variation. We propose...Multimodal prediction from images and structured metadata requires integrating complementary evidence whose reliability can vary across samples and latent directions. A single confidence weight per modality cannot capture this directional variation. We propose RiVaT-Fuse, a reliability-calibrated variational tensor fusion framework that formulates fusion as sample-wise latent-state estimation. For each image-metadata pair, the fused representation minimizes a quadratic objective combining agreement with modality embeddings, structured cross-modal interactions, and regularization. Positive-definite, low-rank-plus-diagonal trust matrices are conditioned on learned state descriptors and metadata completeness, allowing modality contributions to vary across latent directions. Additive, multiplicative, and relational interactions model cross-modal dependencies within the latent estimation objective. The resulting system admits a unique solution computed through a differentiable linear solve. A first-order analysis with fixed trust operators relates latent sensitivity to system conditioning and perturbations in modality embeddings and interactions. The framework further incorporates a state-binned entropic surrogate for conditional distributionally robust learning and task-coupled quadratic prediction heads. We instantiate RiVaT-Fuse on mBRSET, pairing retinal images with clinical and demographic metadata for diabetic retinopathy grading, diabetic macular edema detection, and referable-status prediction. Comparisons with unimodal and representation-level fusion baselines assess the predictive utility of the complete framework across these related clinical tasks.
|
| 444 |
Rolling-WAM: World Action Models with Rolling Imagination
2609.30247
|
cs.CVcs.AI
|
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang |
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting c...World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
|
| 445 |
Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL
2609.32665
|
cs.CVcs.LGcs.AI
|
Qinwei Ma, Jingzhe Shi, Simin Fan, Ling Li, Alex Lamb |
ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to...ELBO-based reinforcement learning offers a sampler-agnostic approach to fine-tuning flow matching models with reward feedback. Timestep weighting in ELBO-based RL has large impact on performance, and it also provides a unified view (as we show in this work) to understand prediction losses heuristically chosen in prior work, yet it remains under-researched and is often chosen to inherit pretrain configs. We investigate impacts and dynamics of timestep weighting in ELBO-based RL. We show that effective weighting depends on both the reward landscape and stage of learning. (1) Through experiments on controlled CIFAR image generation, complemented by robotics, we investigate how weighting impacts reward-driven updates across noise levels. (2) Through gradient analysis, we reveal distinct patterns of cross-noise coordination across tasks and their evolution during training. These findings motivate the hypothesis that useful weighting depends on the gap between the policy's current behavior and the behavior favored by the reward. (3) Guided by this analysis, we study simple static weighting, budgeted profile selection, and dynamic schedules that improve performance beyond conventional target choices. Our results establish timestep weighting as an important design choice for flow-matching RL and motivate further research into methods that choose and adapt it throughout learning.
|
| 446 |
VehicleArena: A Realistic Urban Environment for Multi-Agent Driving
2609.35916
|
cs.CVcs.CL
|
Jie Yang, Jiajun Chen, Jiazheng Zhou, Mianqiu Huang, Yining Zheng |
Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction ...Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decisions can reshape traffic flow, delays, risks, and subsequent observations for surrounding agents. The benchmark provides 112 evaluation tasks spanning single-agent and multi-agent driving. Across nine evaluated models, the highest arrival rates reach only 65.0% on single-agent tasks and 65.6% on multi-agent tasks, while strong passenger-request or cabin scores do not reliably translate into successful trip completion. Moreover, in matched multi-agent runs, every tested focal policy reduces the arrival rate of surrounding vehicles relative to the simulator's native traffic controller, revealing measurable externalities beyond the focal vehicle itself.
|
| 447 |
Recovering Off-Policy Supervision for Speculative Decoding
2609.38795
|
cs.CVcs.CLcs.LG
|
Jungseob Lee, Chanjun Park, Sugyeong Eo, Hyeonseok Moon |
Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in sev...Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary components. The first component, Anchor-Label Relabelling (ALR), replaces corpus labels with distributions from greedy target rollouts, restoring valid supervision across all predicted slots. The second component, In-Rollout Anchors (IRA), places draft blocks directly inside these rollouts to expose the drafter to target-generated context, reusing precomputed rollout features at no additional target cost. Across fixed vision-language and text corpora, our framework increases greedy accepted length by up to 36.5% over DFlash and consistently outperforms erasing baselines. Notably, a single epoch of our method surpasses the best erase schedules. After three epochs, it matches the acceptance length of training on target-regenerated responses. These results show that our framework provides an effective and compute-efficient approach for training speculative drafters on fixed corpora without modifying the original text. Code is available at https://github.com/js-lee-AI/ALR-IRA.
|
| 448 |
Reliability-Aware Checkpoint Selection for Domain Generalization
2609.39934
|
cs.CVcs.LG
|
Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian |
Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accurac...Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We identify an empirical selection opportunity within fixed training trajectories: reselecting among checkpoints with near-optimal source accuracy can improve mean target probability quality with small observed changes in mean target accuracy. We study accuracy-constrained reliability selection (AC), which retains checkpoints within a tolerance of the best source-validation accuracy and ranks them by source reliability. Our reference rule aggregates within-set normalized negative log-likelihood (NLL) and class-wise calibration error (CwECE) using $D_\infty$. AC uses no target data and requires neither additional training nor weight averaging. We evaluate five domain generalization training algorithms on three benchmarks, using PACS to develop the objectives and a 0.5-percentage-point tolerance. In exploratory aggregation comparisons on 360 OfficeHome and TerraIncognita runs, the reference rule reduces mean target soft-bin squared-gap ECE and CwECE by 0.240% and 0.182%, respectively, and NLL by 0.030 relative to Source-Acc. Mean target accuracy changes by +0.213 percentage points. These results identify opportunities for reliability-aware reselection, while the additional benefit of joint over single-objective ranking remains unresolved.
|
| 449 |
The Effect of Tissue Detection on False Positives of Diffusion-Based Artifact Detection in Histopathology
2609.40083
|
cs.CV
|
Konstantinos Moutselos, Ilias Maglogiannis |
One-class artifact detectors for whole-slide images learn normal tissue from a clean training pool and flag departures from it. The pool is built by a preprocessing pipeline whose tissue-detection step is usually treated as neutral. We tested whether it is. On...One-class artifact detectors for whole-slide images learn normal tissue from a clean training pool and flag departures from it. The pool is built by a preprocessing pipeline whose tissue-detection step is usually treated as neutral. We tested whether it is. On 16 annotated slides from The Cancer Genome Atlas, we rebuilt the clean pool of a diffusion-based detector with different tissue detection methods and compared the resulting models in a four-fold cross-validation. Per-slide saturation-Otsu detection excluded normal tissue, chiefly tissue with large clear spaces such as adipose tissue and alveolar parenchyma, and on slides with thick marker ink kept the ink while excluding ordinary tissue. Replacing it with entropy-based detection reduced the false-positive fraction on held-out clean slides from 0.102 to 0.016, in every fold and with a second training seed, without loss of sensitivity; the gain came from the composition of the pool, not its size. Across three tissue detection methods, false positives followed the fraction of such clear-space tissue in the pool, a statistic that needs no labels or training (0.103, 0.016 and 0.009). The effect did not carry over at the same size to a nearest-neighbour detector on foundation-model features. On an external cohort, the curated pool lowered clean-control false positives by about 20%, far less than on the development slides, and the remaining cross-center loss was not explained by stain differences. For one-class quality control, tissue detection decides what the model learns as normal and should be chosen and reported accordingly.
|
| 450 |
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
2610.02196
|
cs.CV
|
Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui |
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a ...We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.
|
| 451 |
Lightweight and Resource-Efficient Perception for Robotic Guide Dogs
2610.03187
|
cs.CV
|
Jinse Kwon, Yoojin Lim, Choonghan Lee, Yongseung Yu, Yongin Kwon |
Robotic guide dogs should understand their surroundings, objects, and potential risks. Prior research has focused on raw sensor data from cameras and 2D or 3D LiDAR, which precisely measure distance points rather than provide a semantic understanding of the sc...Robotic guide dogs should understand their surroundings, objects, and potential risks. Prior research has focused on raw sensor data from cameras and 2D or 3D LiDAR, which precisely measure distance points rather than provide a semantic understanding of the scene. While these physical measurements are effective for robot-centric collision avoidance and robot safety, they are not suitable for human-centric guidance. The system should recognize the type and relevance of obstacles and explain them, clearly and actionably, in terms of their spatial relation to the user. We present complete on-device perception modules that fuse a 360 camera and a 2D LiDAR for reliable collision avoidance, with moving-object detection and tracking for human-centric guidance. Finally, in walking-impossible situations, a vision--language model delivers pathway explanations as a safety mechanism to reduce user anxiety. In experiments, verification of fused 360 camera--LiDAR depth shows reliable near-range perception but inherent mid-range bias, while the system as a whole sustained real-time performance under 55 W. On the real-world egocentric GuideDogQA benchmark, our system achieved 83.8\% accuracy, compared with 67.1\% for GPT-4o. These results demonstrate that practical human-centric guidance with real-time on-device inference is feasible even on quadrupeds.
|
| cs.LG 726 papers | ||||
| 855 |
Bayes-Sufficient Compression Is Not Enough: How Does Communication Help Multi-Agent Systems?
2610.03769
|
cs.LG
|
Yi Xie, Zhanke Zhou, Yi Fan, Yong Ge, Bo Han |
Multi-agent LLM systems pair a sender with broad context and an executor with a limited local view. We study when a short message improves the executor's next decision, when raw context is preferable, and when a stronger sender helps. Our framework, \emph{rece...Multi-agent LLM systems pair a sender with broad context and an executor with a limited local view. We study when a short message improves the executor's next decision, when raw context is preferable, and when a stronger sender helps. Our framework, \emph{receiver-relative bounded coordination}, expresses message utility as receiver gain minus protocol tax. Compression beats raw context when tax savings exceed losses from omitted information and decoder mismatch. Even \emph{Bayes-sufficient} compression can fail when a bounded executor cannot use its surface form. A three-stage decomposition separates externalization, absorption, and \emph{action closure}, explaining how errors remain after the correct content reaches the receiver. Under a single-crossing condition, sender upgrades help above a receiver-burden threshold. Across six benchmarks, the same Qwen protocol raises ContextBench joint accuracy from $0.633$ to $0.775$ but lowers ToolSandbox from $0.889$ to $0.653$. Fixed-message replay reveals closure failures despite correct artifact recovery. These results guide an inference-time selector that improves the accuracy-cost frontier on the evaluated communication regimes.
|
| 856 |
Memory-State Critic for Asymmetric Actor-Critic with Application to Vision-Based Pursuit-Evasion
2610.03830
|
cs.LG
|
Arthur Louette, Alejandro S\'anchez Roncero, Gaspard Lambrechts, Pascal Leroy, Julien Hansen |
In partially observable Markov decision processes, the optimal policy generally depends on the history of observations and past actions. Asymmetric actor-critic methods have become popular to learn such policies when additional information, such as the true st...In partially observable Markov decision processes, the optimal policy generally depends on the history of observations and past actions. Asymmetric actor-critic methods have become popular to learn such policies when additional information, such as the true state of the environment, is available during training. The critic, which is not needed at execution, is given access to the state. A critic conditioned on the state alone is generally ill-defined and yields biased policy gradients. Conditioning on the state and the history, the history-state critic restores both. In this paper, we show that conditioning the critic on the state and the policy's own memory, i.e., the internal representation of the history through which the policy selects its actions, is already well-defined and gives unbiased policy gradients, removing the need for a second recurrent approximator of the history. We call it the memory state critic. It follows that a critic based on the policy's memory need not backpropagate its loss into that memory, even though the memory is a lossy encoding of the history. We evaluate the memory-state critic in a vision-based pursuit-evasion environment between two quadrotors across two arena types. The pursuer is the learning agent, and the evader is sampled per episode from a fixed pool of heuristic behaviours. The results show that the memory-state critic outperforms the history-state critic and converges faster. In addition to being unbiased compared to the state-only critic, it maintains a slight edge in the wall arena, where the actor's history carries information that the privileged state alone does not.
|
| 857 |
Where Does Jagged Competence Come From?
2610.03831
|
cs.LG
|
Ioannis Tsiokos |
Capable systems often show jagged competence: low average error alongside failures on particular inputs. We ask where it comes from in a task built from two known layers. A lower layer A computes five per-slot sums from records; an upper layer B uses the slot-...Capable systems often show jagged competence: low average error alongside failures on particular inputs. We ask where it comes from in a task built from two known layers. A lower layer A computes five per-slot sums from records; an upper layer B uses the slot-1 sum and a mode carried over from earlier boards to predict the next board's category, so B cannot be computed from the current A alone: B is a strict extension of A. We train small recurrent networks on B and measure, against exact ground truth, what they acquire of A. The central finding is that learning B gives the network a jagged version of A, measured through probe readability and supervised outputs. It is readable where B needs it (slot 1, 97-98% by a linear probe) and becomes less readable where B does not (the other slots, 5-10%); category misreads concentrate near the slot-1 cutoffs; and when A is trained explicitly it comes out only approximately right. The networks show jagged competence in B: natural KL below $3\times10^{-4}$ bits coexists with a maximum law TV of about 0.25 on constructed histories. Matched experiments show that even a good A is not enough: connecting a learned A to B cuts misreads 1.3 to 17-fold, the same exact A gives fewer misreads as one-hot inputs (0-2) than as numerical inputs (10-317), one B update makes an exact A inexact unless the B gradient is blocked, and exact A still leaves some B failures. In this task, uneven acquisition, access and preservation of the lower theory explain part of the jagged competence of a network that learns the theory built on it; some failures remain unexplained. Whether the same holds in larger systems is a test to run.
|
| 858 |
LLM-enhanced spatio-temporal learning for grid-level docked bike sharing demand prediction
2610.03834
|
cs.LG
|
Xuxilu Zhang, Francesc Soriguera |
Short-term bike-sharing demand forecasting is complicated by spatial-temporal non-stationarity and the practical difficulty of incorporating unstructured external text into numerical pipelines. Conventional approaches rely on historical flow sequences and fixe...Short-term bike-sharing demand forecasting is complicated by spatial-temporal non-stationarity and the practical difficulty of incorporating unstructured external text into numerical pipelines. Conventional approaches rely on historical flow sequences and fixed graph structures, thereby constraining their accuracy when anomalous social events perturb normal travel patterns. We propose a forecasting framework in which a Large Language Model (LLM) drives a semantic shockwave mechanism that converts free-form urban text, such as municipal event schedules, local news, and transit bulletins, into quantified spatial-temporal perturbation fields. The LLM extracts three physically interpretable parameters per event (intensity, spatial reach, and temporal lag), from which Gaussian decay fields are constructed and injected into a Zero-Inflated Adaptive Spatio-Temporal Graph Convolutional Network (ZI-ASTGCN). To handle the pronounced sparsity of grid-level measurements, the model couples a dual-branch output head with a multi-task zero-inflated loss that jointly trains a gating probability and a conditional flow intensity. Experiments on the operational Barcelona Bicing dataset show that ZI-ASTGCN outperforms established neural baselines, with particularly strong gains during high-demand periods, validating the utility of physics-grounded semantic signals in spatial-temporal mobility forecasting.
|
| 859 |
KVE-KD: Key Visual Evidence-Guided Knowledge Distillation for Vision-Language Models
2610.03842
|
cs.LG
|
Jianbin Zhang, Xin Sun, Shanwen Wang, Wei Ye, Susanto Rahardja |
Knowledge distillation is crucial for deploying vision-language models on resource-constrained devices. However, existing methods typically impose uniform supervision across visual tokens or rely on static token selection, which confuses task-relevant cues wit...Knowledge distillation is crucial for deploying vision-language models on resource-constrained devices. However, existing methods typically impose uniform supervision across visual tokens or rely on static token selection, which confuses task-relevant cues with background noise and degrades cross-modal reasoning. To address this limitation, we propose Key Visual Evidence-guided Knowledge Distillation (KVE-KD), a framework that dynamically focuses feature distillation on task-relevant visual tokens identified by the teacher model. Specifically, KVE-KD appoints the final pre-generation textual token as a unified semantic anchor and identifies the target cross-modal fusion layer by analyzing changes in the anchor representation through iterative visual-token contribution removal. Within this layer, KVE-KD ranks visual tokens via the anchor-conditioned attention distribution and selects the most informative visual tokens as key visual evidence with normalized entropy. The key visual evidence subsequently guides focused visual feature distillation, making the student align closely with the teacher's task-relevant visual representations while suppressing irrelevant background information. Extensive experiments on six benchmarks demonstrate that KVE-KD outperforms state-of-the-art cross-modal distillation methods, with particularly pronounced gains on tasks requiring complex reasoning and fine-grained visual understanding. Importantly, these improvements are achieved without introducing any inference-time overhead. The source code is available at https://github.com/zhangjianbin07/KVE-KD.
|
| 860 |
BAT-NO: A Boundary-Condition-Aware Transformer Neural Operator for Crashworthiness Prediction of Vehicle Components
2610.03854
|
cs.LG
|
Haoran Li, Yingxue Zhao, Haosu Zhou, Mustapha Ziane, Pierre Culiere |
High-fidelity finite-element simulations provide accurate crashworthiness predictions, but their cost limits iterative design exploration. Deep learning surrogates can reduce this cost, but many component-level models are developed under a single prescribed bo...High-fidelity finite-element simulations provide accurate crashworthiness predictions, but their cost limits iterative design exploration. Deep learning surrogates can reduce this cost, but many component-level models are developed under a single prescribed boundary condition, limiting generalisation to boundary variations. This work proposes a Boundary-Condition-Aware Transformer Neural Operator (BAT-NO) for autoregressive prediction of transient displacement fields and scalar crashworthiness responses under variations in geometry and boundary conditions. A B-pillar simulation framework evaluates generalisation across variations in geometry, impact position and velocity, and support stiffness. BAT-NO combines recurrent mesh processing with latent-grid Fourier operator processing. Boundary-condition information is transferred to the latent grid through a hybrid local--global mechanism. Slice-based attention models interactions among physically related regions, while direct boundary-to-grid projection preserves local spatial structure. Across the validation sets for the shape-only, shape-and-loading, and shape-loading-boundary cases, BAT-NO achieves the lowest mean final-step mean nodal Euclidean displacement error among the evaluated baselines. In the most challenging case, it reduces the mean error by 32.6% relative to the second-best model. Hyperparameter tuning reduces the validation error from 0.451 to 0.269 mm, with a comparable error of 0.267 mm on 300 unseen test simulations sampled within the investigated design space. An attention-based scalar decoder jointly predicts six response trajectories with a mean relative error of 2.46%. Most derived crashworthiness indicators have median errors below 3%. These results show that explicit local and global boundary-condition representations improve crashworthiness prediction over expanded component-level design spaces.
|
| 861 |
Retrieval-Centric Deep Learning in Growing Nonparametric Neural Networks
2610.03858
|
cs.LGcs.AI
|
Maximilian Schlegel, Rajai Nasser, Seijin Kobayashi, Yanick Schimpf, Oliver Sieberling |
We investigate a general-purpose layer for deep learning that, instead of compressing arbitrary-size training data into fixed-size weight matrices, stores a new pair of key-value representations for every data point during training, and retrieves and recombine...We investigate a general-purpose layer for deep learning that, instead of compressing arbitrary-size training data into fixed-size weight matrices, stores a new pair of key-value representations for every data point during training, and retrieves and recombines these representations through an attention mechanism at inference time - resulting in a growing neural net (NN). While Irie et al. (arXiv:2202.05798) have put forward this perspective from the classic duality expressing any linear layer in a deep NN trained by gradient descent as linear attention (LA) over the training data points, replacing LA by more powerful attention functions, as they suggest, turns out to be non-trivial: we show that naively applying learning rules from the LA case to advanced kernels does not lead to principled optimization. Here we fill this gap and develop functional gradient-based learning rules for kernelized attention layers, based on radial basis function (RBF) and softmax-like kernels - establishing the principled "retrieval-centric deep learning" (RCDL) paradigm. Empirically, we demonstrate the promising performance and learning-efficiency of RCDL on image classification and synthetic teacher-student learning tasks. Moreover, we show that replacing LA in the dual form of NNs by advanced LA variants, namely MesaNet/DeltaNet, yields a formal connection to recently proposed optimizers for conventional fixed-size NNs, offering a novel perspective on deep learning optimization.
|
| 862 |
Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking
2610.03880
|
cs.LG
|
Pawan Prakash, Philipp H\"ollmer, Addis Fuhr, Peter Hirschfeld, P. Ganesh |
Inverse materials design is a long-standing goal of computational materials discovery. Generative models for crystalline materials are typically trained to match the distribution of a structure database, while nothing in their training objective points them at...Inverse materials design is a long-standing goal of computational materials discovery. Generative models for crystalline materials are typically trained to match the distribution of a structure database, while nothing in their training objective points them at specific design goals such as targeted properties. We use group-relative policy optimization (GRPO) to align a generative model based on stochastic interpolants and discrete flow matching with general black-box reward functions through reinforcement learning. Atom types are generated by a discrete flow and the policy gradient of our generalization of GRPO directly acts on the likelihoods of the atom-type transitions, which differentiates our work from previous reinforcement-learning approaches for diffusion and flow-based generative models of crystalline materials. We introduce a reward function that raises the yield of metastable, unique and novel structures (mSUN) from 13.4% for the pretrained model to 45.5% for the reinforced model, as evaluated by a community benchmark. Our reward also improves the performance of a reinforcement learning framework for crystalline materials based on latent denoising diffusion models. At the same time, we find that directly reinforcing atom-type transition likelihoods enables reward exploitation that has to be prevented with explicit guards. The same analysis also exposes a gap in the community metric. Single-element structures in distinct packings are counted as metastable, unique and novel materials and inflate mSUN without yielding any new compounds. A stability claim is only as good as its reference hull. We report every result split by the number of reference phases behind it and argue that benchmarks should do the same.
|
| 863 |
AdaEva: Accelerating LLM-Driven Algorithm Design with Adaptive Partial Evaluation
2610.03896
|
cs.LG
|
Tai Nguyen, Fei Liu, Phong Le, Carola Doerr, Nguyen Dang |
Large Language Models (LLMs) are increasingly used for automated algorithm design. However the computational cost of evaluating the generated algorithms can be excessive. We consider the common LLM-driven automated algorithm design (LLM4AD) setting in which a ...Large Language Models (LLMs) are increasingly used for automated algorithm design. However the computational cost of evaluating the generated algorithms can be excessive. We consider the common LLM-driven automated algorithm design (LLM4AD) setting in which a candidate algorithm is evaluated by aggregating its performance over a shared set of training instances. This instance-wise structure raises a natural question: must every candidate be evaluated on the entire instance set before deciding whether it remains competitive? Taking inspiration from algorithm configuration, we introduce AdaEva, a drop-in adaptive partial-evaluation framework that progressively evaluates candidates on larger subsets of the same instance pool and eliminates unpromising candidates as evidence accumulates. Importantly, AdaEva leaves the underlying LLM4AD procedure and per-instance evaluator unchanged and requires no prior knowledge about instance difficulty. We instantiate this idea using successive halving (AdaEva-S) and statistical racing (AdaEva-R), and evaluate both mechanisms across three representative LLM4AD frameworks, multiple LLM backbones, and optimization domains spanning combinatorial and continuous black-box optimization. Under matched evaluation budgets, AdaEva more reliably balances evaluation effort across candidates than fixed partial-evaluation strategies, yielding strong search efficiency and anytime performance together with improved held-out generalization across the evaluated settings.
|
| 864 |
COVER: Learning to Accept More in Selective Sleep Staging
2610.03911
|
cs.LG
|
Yukai Song (Department of Electrical and Computer Engineering, University of Pittsburgh), Yangfan Deng (Department of Electrical and Computer Engineering, University of Maryland, College Park) |
Traditional sleep-staging methods apply the same model to every EEG epoch. Such uniform deployment expends computation on epochs that a smaller model could handle reliably, motivating cascades in which a primary classifier accepts its reliable predictions and ...Traditional sleep-staging methods apply the same model to every EEG epoch. Such uniform deployment expends computation on epochs that a smaller model could handle reliably, motivating cascades in which a primary classifier accepts its reliable predictions and defers the remainder to a more capable model. In this paper, we study the first stage of such a cascade: maximizing the coverage of fixed primary predictions subject to a prescribed accepted-risk target. We propose COVER (COVerage-oriented Error Ranking), which integrates two key innovations: (i) auxiliary-informed primary-error learning, which replaces maximum softmax probability (MSP) with a learned error score while preserving the primary labels, and (ii) fixed-scale scorer refinement, which builds on this score to directly maximize coverage under an empirical accepted-risk constraint rather than error-prediction accuracy over all epochs. We evaluate COVER on Sleep-EDF-20 at a 5% accepted-risk target, with subjects held out from all fitting and selection. Auxiliary-informed error learning raises mean subject coverage from 31.5% for MSP to 48.8% at similar subject-equal risk. At equal acceptance volume, with MSP accepting the same number of epochs as the learned scorer in each subject (20,639 in total), errors fall from 1,467 to 867. Fixed-scale refinement then adds 1.7 percentage points of coverage over its initialization in nested development, and COVER attains the highest mean coverage among eight evaluated scorers, 50.4% at 4.5% subject-equal risk, above the selective-ranking baseline SELE (49.3%) and the probability-fusion comparator DuoF (43.6%). To the best of our knowledge, this is the first work to combine auxiliary-informed primary-error learning with fixed-scale coverage refinement for selective sleep staging, offering a basis for reliability-aware allocation of computation in cascaded sleep staging.
|
| 865 |
TreeWalker: Partial Evaluation for Grouped Tree-Ensemble Inference
2610.03939
|
cs.LG
|
Durmus Karatay, Richard Newman |
Many inference workloads evaluate a trained tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into $G$ time steps, click-through-rate models score every item in a search session, and scenario analyses vary...Many inference workloads evaluate a trained tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into $G$ time steps, click-through-rate models score every item in a search session, and scenario analyses vary a few inputs while holding the rest fixed. Standard inference treats each row independently and repeats the shared work $G$ times. We present TreeWalker, which applies partial evaluation to grouped inference: constant features are static, varying features dynamic. It walks each tree once per group, partitions a row bitmask at varying splits, and skips empty subtrees. Training is unchanged: TreeWalker reads standard LightGBM and XGBoost models. We prove a structural work decomposition: per-tree work splits into the constant-projected subtree size $|T_c|$, $G$ leaf writes, and a predicate-mask provisioning cost $Q$. For the trace evaluator, per-row work approaches a $(d_v+1)/(d+1)$ fraction of a row-independent walk as $G \to \infty$. On Intel, TreeWalker is 2.5-3.2$\times$ faster than a row-independent traversal at the reference configuration ($T=500$, $L=8$) and 6.8-7.8$\times$ faster at $G=128$ on the survival datasets, with larger gains on Arm. On a scenario-analysis benchmark it is faster in all 16 configurations on both architectures. For f64 models, outputs match treelite's GTIL up to summation order; for f32 models, f64 accumulation is closer to a Kahan-compensated reference than native f32 on 99.98% of rows and never farther.
|
| 866 |
DePICT: Decision-Preserving Interface for Constrained Downstream Tasks
2610.03945
|
cs.LG
|
Utkarsh Grover, Ravi Ranjan, Agoritsa Polyzou, Wyatt T. Mackey, J. Morris Chang |
A constrained optimization problem may involve a parameter in its objective and active constraints, yet the final decision may remain insensitive to small changes in that parameter. This raises a fundamental question: which inputs does a decision making system...A constrained optimization problem may involve a parameter in its objective and active constraints, yet the final decision may remain insensitive to small changes in that parameter. This raises a fundamental question: which inputs does a decision making system truly depend on? Building on this question, we introduce DePICT, a procedure for constructing decision preserving interfaces by ranking context directions according to the optimizer's solution sensitivity and aggregating them across an operating regime. We study this problem in a high dimensional setting where primitive context parameterizes a constrained task and the downstream agent observes only a selected subset of context directions. For locally regular constrained programs, we derive a Karush Kuhn Tucker (KKT) based characterization of when a context direction is optimizer relevant. Our analysis shows that appearing in the active optimization problem does not necessarily imply that a variable affects the final decision. Some context directions can alter the KKT conditions while leaving the optimal solution unchanged because their effect is absorbed by the dual variables. DePICT is designed to remove exactly these directions. In a controlled diagnosis, it recovers the decision relevant interface exactly and reduces linear predictor regret to 0.009, compared with 0.475 for the strongest competing baseline.
|
| 867 |
Synthesizing Physics Formulae with Transformers
2610.03947
|
cs.LG
|
Shuwei Wang, Vadim Bulitko, Michael Youngblood, Ramon Lawrence, William Yeoh |
Finding a compact formula that fits a set of input-output pairs and predicts outputs on unseen inputs is a fundamental problem in science. Symbolic regression automates the search for such formulae: search-based methods explore the space of possible formulae d...Finding a compact formula that fits a set of input-output pairs and predicts outputs on unseen inputs is a fundamental problem in science. Symbolic regression automates the search for such formulae: search-based methods explore the space of possible formulae directly, while transformers pre-trained on synthetic data produce formulae of comparable quality substantially faster. Existing transformers, however, are prone to overfitting --- they find formulae that fit the training data well but do not extrapolate to input ranges unseen during training. We address this by shaping the set of formulae used to train a transformer, and show that the resulting formulae extrapolate substantially better. Fine-tuning the transformer on data with noise-corrupted target values further makes the synthesized formulae robust to noise in the observations. On SRBench and LLM-SRBench our transformer synthesizes a formula in about ten seconds and extrapolates better than all evaluated methods at a comparable budget. Search-based methods surpass our accuracy only when given one to three orders of magnitude more time.
|
| 868 |
Evolving LLM-Generated Features for Interpretable Classification
2610.03951
|
cs.LG
|
Jack Butler, Zainab Afolabi, Nikita Kozodoi |
Large language models (LLMs) are increasingly used as classifiers, yet they operate as opaque systems whose decisions are difficult to interpret, which complicates their use in regulated domains such as credit scoring or medical diagnosis. We propose an evolut...Large language models (LLMs) are increasingly used as classifiers, yet they operate as opaque systems whose decisions are difficult to interpret, which complicates their use in regulated domains such as credit scoring or medical diagnosis. We propose an evolutionary framework that iteratively discovers natural language feature definitions (rubrics) for interpretable classification. An LLM generates candidate binary features, evaluates each sample against them, and the resulting vectors can be used to train a transparent classifier such as logistic regression. The feature set evolves over multiple iterations guided by classification errors, per-class activation rates, and feature ablation scores. We evaluate across three benchmarks, including a credit risk dataset representative of regulated domains, comparing single-shot LLM rubrics, evolved rubrics, and direct zero-shot LLM classification. Evolved features improve over single-shot rubrics by +2.9 pp on average and outperform zero-shot LLM classification on two of three tasks, while providing fully auditable decision logic. On the credit risk task, the zero-shot LLM performs at chance (50.7%) with a strong bias toward a single class, whereas evolved features achieve balanced, interpretable predictions. Crucially, this failure is invisible in aggregate accuracy and surfaces only under per-class auditing. Analysis reveals that evolution is most effective when label boundaries cannot be inferred from category names alone or when the LLM lacks reliable domain-specific reasoning.
|
| 869 |
One-Step Curvature Probes Miss the Fitting Operator: Retained Capacity and Terminal Null-Space Correction for Continual Learning
2610.03952
|
cs.LG
|
Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Ibne Farabi Shihab, Md Najmus Swaqeeb |
A one-step curvature probe evaluates an initial direction, whereas continual learners are judged after reaching comparable new-task fit. In an overparameterized linearization, projected gradient descent converges to $\Delta_P=PJ^\top(JPJ^\top)^{-1}r$, and its ...A one-step curvature probe evaluates an initial direction, whereas continual learners are judged after reaching comparable new-task fit. In an overparameterized linearization, projected gradient descent converges to $\Delta_P=PJ^\top(JPJ^\top)^{-1}r$, and its squared-displacement inflation is exactly the reciprocal of the retained fitting capacity $c_P(r)$. More generally, the terminal old-task quadratic ratio factorizes as $1/[c_P(r)G_{\rm end}(P,r)]$, where $G_{\rm end}$ compares curvature along endpoint directions. In the rank-one case, $G_{\rm end}$ equals the one-step probe gain; for multiple outputs, the two gains can differ. The terminal quadratic also separates into a curvature-optimal fitting floor and an algorithm-dependent null-space excess, motivating terminal null-space correction, which preserves linearized new-task outputs on its Jacobian batch. Controlled checks validate the local quadratic and show that projection can greatly reduce matched-norm curvature while barely changing terminal forgetting. Across common-threshold configurations, projection yields the larger signed old-task loss change in 73/99 matched pairs, retained fitting capacity falls with rank, and fixed-rank comparisons separate terminal forgetting even when probe gain is approximately matched. On Permuted MNIST and Split CIFAR-100, terminal correction decreases signed old-task loss in 166/180 method--dataset--seed pairs under the all-seed intention-to-correct analysis; because acceptance and outcome reporting use the same held-out split, this result is descriptive and test-conditioned. Overall, terminal cost depends jointly on retained fitting capacity, endpoint-direction curvature, and the null-space component selected by optimization.
|
| 870 |
Probabilistic Algorithms for Ising Machines from Optimization to Generative AI
2610.03972
|
cs.LG
|
Corentin Delacour, Xiuqi Zhang, Abdelrahman S. Abdelrahman, Saleh Bunaiyan, Kyle Lee |
Ising machines have emerged as promising hardware accelerators for intractable optimization and sampling problems, yet their practical impact increasingly hinges on the co-design of algorithms and hardware, where algorithmic demands shape new architectures and...Ising machines have emerged as promising hardware accelerators for intractable optimization and sampling problems, yet their practical impact increasingly hinges on the co-design of algorithms and hardware, where algorithmic demands shape new architectures and new hardware capabilities inspire entirely new algorithms. In this Review, we survey probabilistic algorithms designed for portability across diverse Ising platforms, advocating a top-down perspective that prioritizes principled methods with provable guarantees. We cover foundational methods such as simulated annealing and parallel tempering, including two-dimensional extensions that natively encode hard constraints, and examine approaches that expand the scale of solvable problems from cluster mean-field methods to variational samplers. We highlight the Probabilistic Approximate Optimization Algorithm (PAOA), a classical analog of QAOA that emerged directly from probabilistic hardware development, and explore how generative AI and Ising machines might reinforce each other: learned models propose global moves to accelerate optimization, while probabilistic techniques improve inference in large language models. Much as quantum computing has seen algorithms co-evolve with hardware, probabilistic and Ising computing stand at a similar inflection point. We outline a co-design framework for accelerating the capabilities and adoption of next-generation Ising machines.
|
| 871 |
Learning Latent Protein Languages for Autoregressive Generation
2610.03978
|
cs.LG
|
Mahdi Pourmirzaei, Farzaneh Esmaili, Amir Ziashahabi, Mohammadreza Pourmirzaei, Dong Xu |
Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates requi...Autoregressive transformers remain comparatively weak for protein sequence and structure generation. We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework. We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, with one token per residue. Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while retaining decoding to backbone coordinates. We separately pretrain autoregressive transformer models on PLL and SLL tokens using next-token prediction, yielding PLLM and SLLM. Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model. In unconditional sequence generation, PLLM reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across sampling temperatures. For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training. For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in our measurements. In backbone generation, SLLM compares favorably with other generative models on diversity and novelty. We also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample. These results position learned latent protein languages as a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.
|
| 872 |
$\tilde{O}(\sqrt{T})$ Regret and Polylogarithmic Constraint Violation for COCO
2610.03983
|
cs.LG
|
Dhruv Sarkar, Abhishek Sinha |
We study constrained online convex optimization with adversarial convex losses and constraints ($\mathsf{COCO}$). At each round \(t\in[T]\), a learner selects \(x_t\) from a \(d\)-dimensional convex decision set \(\mathcal X\), after which an adaptive adversar...We study constrained online convex optimization with adversarial convex losses and constraints ($\mathsf{COCO}$). At each round \(t\in[T]\), a learner selects \(x_t\) from a \(d\)-dimensional convex decision set \(\mathcal X\), after which an adaptive adversary reveals a convex cost function \(f_t\) and constraint function \(g_t\). Consequently, the learner incurs cost \(f_t(x_t)\) and constraint violation \(\max\{0,g_t(x_t)\}\), and aims to simultaneously minimize regret and cumulative constraint violation ($\mathsf{CCV}$) over the entire horizon. Existing algorithms achieve \(O(\sqrt{T})\) regret and \(\widetilde O(\sqrt{T})\) $\mathsf{CCV}$. We show that an online policy can achieve \(O(\sqrt{T\log T})\) regret and \(O(\log^2 T)\) $\mathsf{CCV}$, reducing the $\mathsf{CCV}$ from polynomial to polylogarithmic while retaining near-optimal regret. Our approach combines continuous Hedge with elimination on shrinking feasible sets. The key observation is that whenever the mean of the Hedge distribution violates a constraint, Gr\"unbaum's inequality guarantees that a constant fraction of the Hedge probability mass is eliminated. We use an adaptive learning-rate schedule and a potential function coupling the surviving volume with the learning rate to convert this probability-mass reduction into a bound of \(O(\log^2 T)\) on the $\mathsf{CCV}$.
|
| 873 |
LD-EnFF: Latent-Dynamics Ensemble Flow Filtering for Data Assimilation with Sparse Observations
2610.04034
|
cs.LG
|
Ziyu Tian, Kaichen Shen, Wenbo Hao, Phillip Si, Peng Chen |
Data assimilation combines model forecasts with noisy, incomplete observations to estimate the evolving state of a dynamical system. Existing methods face two compounding challenges: high-dimensional nonlinear dynamics make repeated forward simulation computat...Data assimilation combines model forecasts with noisy, incomplete observations to estimate the evolving state of a dynamical system. Existing methods face two compounding challenges: high-dimensional nonlinear dynamics make repeated forward simulation computationally expensive, while sparse observations provide limited direct information about the full state. To address these challenges, we propose the Latent-Dynamics Ensemble Flow Filter (LD-EnFF), a sequential Bayesian filtering framework that performs both forecast propagation and filtering updates in a compact latent space. LD-EnFF combines a latent dynamics surrogate for ensemble propagation with a variational autoencoder (VAE)-based observation model that evaluates a state-dependent observation likelihood in latent space. At each assimilation step, an ensemble filtering update based on flow matching uses the forecast ensemble and this likelihood to generate posterior samples, jointly updating latent states and uncertain parameters. This design avoids repeated full-state simulation during forecasting and full-field reconstruction during likelihood evaluation. LD-EnFF substantially outperforms a broad range of data assimilation algorithms on benchmarks spanning Kolmogorov flow, tsunami propagation, and atmospheric modeling, all featuring complex dynamics and sparse, noisy observations.
|
| 874 |
Protecting Sensitive Data in Image Synthesis via PAC-Private Adaptation for Diffusion Models
2610.04038
|
cs.LG
|
Boming Miao, Tao Zhang, Netanel Raviv, Murat Kantarcioglu, Bradley A. Malin |
Synthetic data are increasingly used as an alternative to sharing sensitive records. However, synthetic data generation does not guarantee privacy, as diffusion models trained or adapted on sensitive data remain susceptible to reconstruction attacks. Moreover,...Synthetic data are increasingly used as an alternative to sharing sensitive records. However, synthetic data generation does not guarantee privacy, as diffusion models trained or adapted on sensitive data remain susceptible to reconstruction attacks. Moreover, while approaches that use differential privacy (DP), such as DP-SGD, achieve provably private diffusion model training, the repeated gradient clipping and noise injection they require result in significant utility loss. An important limitation of DP-based privacy is that, although it has a provable relationship to reconstruction privacy (RP), that relationship is indirect. RP is defined in terms of limiting how much an adversary's posterior distribution over sensitive data differs from the prior, whereas DP provides guarantees by bounding the sensitivity of outputs to changes in individual records. This indirection is an important source of the utility loss. To address this, we propose a PAC-private diffusion model adaptation to achieve reconstruction privacy. Since PAC-privacy is defined directly with respect to posterior advantage over the prior, it directly implicates RP. To obtain scalable PAC privatization in high dimensions, we first learn a compact data-dependent diffusion model component using LoRA or Textual Inversion, and then calibrate anisotropic Gaussian noise from the covariance of repeated mechanism outputs. Unlike DP-SGD, our method perturbs the learned component only once after optimization, thereby avoiding privacy composition across gradient updates. We evaluate the framework on few-shot concept personalization and full-dataset image synthesis, and show that the proposed approach better preserves subject identity, generation quality, and downstream classification accuracy than DP while achieving the same reconstruction privacy.
|
| 875 |
On architectural choices for interpretability and thermodynamic consistency in Physically Recurrent Neural Networks in the low-data regime
2610.04067
|
cs.LG
|
M. A. Maia, K. A. Meyer, A. M. C. M. van Gils, I. B. C. M. Rocha, F. P. van der Meer |
In this paper, we unravel the effect of different decoder architectures on the interpretability of the latent space of the Physically Recurrent Neural Network. Particular emphasis is given to a new weight normalization constraint, which acts as a regularizatio...In this paper, we unravel the effect of different decoder architectures on the interpretability of the latent space of the Physically Recurrent Neural Network. Particular emphasis is given to a new weight normalization constraint, which acts as a regularization technique and enables robust training in the low-data regime. A brief visual exploration illustrates how these changes impact the latent space and how the fictitious stress can align with the true state of the RVE without explicit training. Reaping the benefits of a meaningful latent space, a case study illustrates how information from the microscopic level can be retrieved and incorporated into a multi-task approach that does not require extra parameters or larger training sets. Another key contribution shows that a specific architectural choice can naturally lead to a thermodynamically consistent formulation. By enforcing an adjoint encoder-decoder structure with positive scalar contributions, this modification ensures energy consistency across scales and non-negative dissipation, leading to even lower training requirements. This alternative completes the study on interpretability, inductive bias, and thermodynamic consistency, and demonstrates that data efficiency can be improved with careful architectural choices rooted in the underlying physics.
|
| 876 |
Articulatory Entrainment and Coordination Complexity in Spontaneous Autistic and Non-autistic Dialogue
2610.04071
|
cs.LG
|
Thanushi Withanage, Carol Espy-Wilson, Elizabeth Redcay, Desi Jones, Noah Sasson |
Articulatory entrainment, the adaptation of vocal tract coordination to facilitate interaction remains underexplored in spontaneous dialogue, particularly among autistic speakers. Many prior studies have utilized task-based, phoneme-level analyses with invasiv...Articulatory entrainment, the adaptation of vocal tract coordination to facilitate interaction remains underexplored in spontaneous dialogue, particularly among autistic speakers. Many prior studies have utilized task-based, phoneme-level analyses with invasive measurement techniques. Here, we introduce a speaker-independent framework to quantify articulatory entrainment in spontaneous dyadic conversations among autistic and non-autistic adults using acoustic-to-articulatory inversion and coordination complexity metrics. We examine temporal changes in articulatory coordination across interaction. Non-autistic dyads exhibit increasing coordination complexity and stronger entrainment over time, autistic dyads show moderate effects, and mixed dyads demonstrate the least alignment. Greater articulatory entrainment correlates with higher self-reported conversational success, indicating increase in coordination complexity as a marker of effective social interaction.
|
| 877 |
DUET: Co-Evolving Solver and Grader Agents
2610.04087
|
cs.LGcs.AI
|
Fengyu Gao, Sourav Pal, Austin Z. Henley, Arjun Radhakrishna, Gustavo Soares |
Agentic workflows are increasingly used across domains such as technology, finance, and enterprise operations. As these agents become more widely deployed, continually improving them becomes increasingly important. This raises an immediate challenge: How shoul...Agentic workflows are increasingly used across domains such as technology, finance, and enterprise operations. As these agents become more widely deployed, continually improving them becomes increasingly important. This raises an immediate challenge: How should the agent evolve? This evolution requires effective evaluation that can assess outcomes and provide useful feedback for optimization. As the agent evolves, its behaviors and failure modes may also change, making a fixed evaluator increasingly inadequate. Another fundamental question: How should we evaluate an evolving agent? These two challenges are inherently coupled; changes in agent behavior can expose limitations of the current evaluator, while a stronger evaluator provides more informative feedback for improving the agent. Motivated by this interaction, we introduce DUET, a framework that jointly optimizes a solver agent and a grader agent to improve both. DUET iteratively selects training tasks, executes them with the solver, evaluates the resulting outcomes with the grader, and uses a tool-using update module to revise the solver and the grader, alternating between the two across rounds. By updating the grader within the optimization loop, DUET turns evaluation from a fixed source of feedback into a first-class optimization objective that adapts alongside the solver. Experiments across four agent benchmarks show that DUET improves both solver and grader performance and consistently outperforms baselines that optimize the solver with a fixed grader.
|
| 878 |
Pareto-Dominant Clarification: Post-Training Coding LLMs via PPO-Lagrangian Budget Constraints
2610.04089
|
cs.LG
|
Abhinav Rajput, Acey Vogelstein |
Coding agents operating under ambiguous instructions or user prompts must decide whether to ask clarifying questions or attempt a solution directly. While clarification from the user may improve the correctness of the agent's solution, each back-and-forth inte...Coding agents operating under ambiguous instructions or user prompts must decide whether to ask clarifying questions or attempt a solution directly. While clarification from the user may improve the correctness of the agent's solution, each back-and-forth interaction incurs user and system costs, forming an explicit accuracy vs. efficiency tradeoff. Existing works study clarification behavior but do not train policies under enforceable clarification budgets; penalty-based approaches typically require separate coefficient tuning swept across all clarification budget levels. We formulate clarification as a Constrained Markov Decision Process (CMDP) and post-train Qwen2.5-Coder-7B-Instruct with PPO-Lagrangian to optimize coding accuracy, subject to an expected question-budget constraint. Evaluated on HumanEvalComm with a GPT-4o-mini oracle simulator, the resulting policies reveal that untuned clarification behavior is Pareto-inefficient: budget-constrained policies can simultaneously achieve higher accuracy and lower clarification rates than the baseline model. Across budget levels, we observe a log-shaped Pareto frontier with diminishing returns to additional clarification. Gains arise not from simply asking more questions overall, but from improved question targeting and better code generation under ambiguity. Without explicit supervision, trained policies learn to allocate clarification budget non-uniformly, asking more frequently on tougher (multi-degradation) tasks. These results suggest that unconstrained interactive LLM systems may systematically use clarification inefficiently.
|
| 879 |
Robust blind unmixing: A geometric approach to overcoming basis variation
2610.04091
|
cs.LGcs.AI
|
Dumitru Mirauta, Vladimir V. Gusev, Michael W. Gaultois, Matthew J. Rosseinsky, Yannis Goulermas |
Signal separation problems are common in science. A prominent example of this occurs during the use of diffraction or spectroscopy to identify the individual components of a mixture by measuring it. In the simplest case, the measured signal is a linear combina...Signal separation problems are common in science. A prominent example of this occurs during the use of diffraction or spectroscopy to identify the individual components of a mixture by measuring it. In the simplest case, the measured signal is a linear combination of basis patterns corresponding to the constituent parts. The unmixing problem is to infer all or some of these basis patterns and abundances of components from measurements of distinct mixtures. One of the core challenges of this task is the variation of the basis from mixture to mixture due to noise and the exact physics of the measurement process. This is usually addressed with tailored model-based and parametric methods that are then limited in use to specific application domains by the nature of the assumptions made. We propose a novel geometric approach to unmixing problems which views the generation of data during measurement through a metric space lens, thereby shifting the focus from parametrised models to a general relationship between basis transformations and the corresponding geometry. We take advantage of the optimal transport distances to capture commonly occurring basis variations, and use minimisation of in-class variance of candidate solutions to drive the optimisation. We pay special attention to the one-dimensional case due to its practical importance and availability of efficient distance and transport map routines. The effectiveness of our approach is demonstrated on a range of unmixing tasks using random Gaussian mixture models, simulated powder X-ray diffraction, and laboratory hyperspectral imaging datasets.
|
| 880 |
Beyond Masked Sparsity: SNACK Enables Truly Sparse Neural Networks on GPU
2610.04093
|
cs.LG
|
Jafar Badour, Maurice van Keulen, Elena Mocanu |
Deep neural networks continue to grow in parameter count, driving up training and inference cost on GPUs. Sparse neural networks and Dynamic Sparse Training (DST) promise to reduce these costs, but most implementations rely on binary masks over dense tensors a...Deep neural networks continue to grow in parameter count, driving up training and inference cost on GPUs. Sparse neural networks and Dynamic Sparse Training (DST) promise to reduce these costs, but most implementations rely on binary masks over dense tensors and recover little of the theoretical compute, memory, or energy savings. We propose SNACK, a truly sparse GPU layer that stores and computes only non-zero connections. SNACK exposes a simple PyTorch API for restructuring connections and backpropagating gradients entirely in the sparse paradigm, and ships SNACK-COO, a custom COO-format SpMM CUDA kernel with a batch-to-Streaming-Multiprocessor mapping tuned for the small-batch, high-sparsity regime typical of large-model training and single-stream inference. At the kernel level, SNACK is up to 7x faster than the masked dense baseline (Dense+Mask) and competitive with cuSPARSE, Sputnik, and FlashSparse at 95% sparsity. At 90% sparsity, a single SNACK layer accelerates training by 8x and 3.7x, and inference by 4x and 2x, over Dense+Mask and fully dense layers, respectively, while using 72% less memory than dense and substantially less energy. End-to-end, SNACK reduces GPT-2 peak training memory by up to 40% and graph-style inference latency by 4.8x over Dense+Mask at 99% sparsity.
|
| 881 |
Progressive Multi-Ancestor Bit-Depth Distillation
2610.04100
|
cs.LG
|
Adil Mubashir Chaudhry, Osama Ahmad, Zubair Khalid, Murtaza Taj |
Model compression strategies are widely employed to reduce memory footprint and network complexity, particularly for devices with constrained computational, memory, and energy resources. Prior works that rely on simultaneous conversion from floating-point high...Model compression strategies are widely employed to reduce memory footprint and network complexity, particularly for devices with constrained computational, memory, and energy resources. Prior works that rely on simultaneous conversion from floating-point high-precision (FP32) to integer low-precision (INT4) representations and distillation into smaller models suffer from unstable training and drastic degradation of prediction performance. To address these limitations, we propose a unified framework, known as \textbf{P}rogressive \textbf{M}ulti-\textbf{A}ncestor \textbf{B}it-depth \textbf{D}istillation (PMABD), that progressively compresses the network while transferring knowledge through a growing pool of higher-precision ancestor teachers. PMABD generates a sequence of intermediate teachers that each learn from all higher-precision ancestors and jointly supervise the final target student. This multi-ancestor, multi-stage design stabilizes ultra-low-bit quantization by lowering quantization noise profiles across training and ensuring stable quantization. Experiments on CIFAR-10/100 with ResNet-20/32/18, and Tiny-ImageNet with MobileNetV2 show that PMABD outperforms state-of-the-art compression frameworks, results in 1.06$\%$ increase in performance of W2A2 (ResNet-18/CIFAR-100) student model. We show that a saturation-based stopping criterion contributes to improve the performance of our final student.
|
| 882 |
Physics is the Best Teacher: Consistency Learning for Time-Invariant Operators of Chaotic Dynamics
2610.04108
|
cs.LG
|
Lufang Chiang, Jiachen Yao, Thomas Y. L. Lin, Anima Anandkumar |
Accelerating the prediction of long-term behavior in chaotic systems is crucial in scientific computing. However, existing methods rely on numerical solvers or autoregressive models that advance one small step at a time, which makes long horizons expensive. We...Accelerating the prediction of long-term behavior in chaotic systems is crucial in scientific computing. However, existing methods rely on numerical solvers or autoregressive models that advance one small step at a time, which makes long horizons expensive. We instead view this problem as learning the system's time-invariant evolution operator, which jumps the state across a large time span in a single evaluation. To this end, we derive the consistency equations a time-invariant operator must satisfy, with differential and compositional objectives in physical time. These equations also connect the learned operator to the physics-prescribed instant dynamics, enabling physics embedding in consistency learning. Across five chaotic systems, we find that physics-distilled consistency makes both short-term trajectories and long-term statistics more accurate. The learned operator survives temporal extrapolation and requires one-tenth as many evaluations as autoregressive rollout, offering an efficient route to long-term simulation of chaotic dynamics.
|
| 883 |
Sharp Convergence and Sample Complexity of Policy Mirror Descent for Average-Reward MDPs
2610.04117
|
cs.LG
|
Enes Arda, Atilla Eryilmaz |
Policy mirror descent (PMD) has a mature finite-time theory in discounted Markov decision processes (MDPs), but less is known in the average-reward setting, a more natural objective for many control applications. We give a finite-time, finite-sample analysis o...Policy mirror descent (PMD) has a mature finite-time theory in discounted Markov decision processes (MDPs), but less is known in the average-reward setting, a more natural objective for many control applications. We give a finite-time, finite-sample analysis of PMD in ergodic average-reward MDPs built around a single master recursion that governs convergence for any critic, without external regularization. Its specializations yield linear rates for exact, inexact-tabular, and linear function approximation (LFA) updates, with a superlinear regime for exact PMD. We complement these convergence results with end-to-end sample complexities of order $t_{\mathrm{mix}}^3/\varepsilon^2$ in both tabular ($|S||A|$-dependent) and LFA ($d$-dependent) settings. Our LFA sample complexity sharpens the prior best $t_{\mathrm{mix}}^5$ mixing dependence to $t_{\mathrm{mix}}^3$, and matching information-theoretic lower bounds establish that the critic's $t_{\mathrm{mix}}^3/\varepsilon^2$ sample complexity is unimprovable in both settings.
|
| 884 |
Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction
2610.04137
|
cs.LG
|
Som Sagar, Shasha Li, Hejie Cui, Ransalu Senanayake, Sercan \"{O}. Ar{\i}k |
Agent harnesses specify the roles, instructions, tools, and communication structure used to solve a task, and the right harness depends on the query. Because the value of each design choice is observable only through execution, tailoring a harness to each quer...Agent harnesses specify the roles, instructions, tools, and communication structure used to solve a task, and the right harness depends on the query. Because the value of each design choice is observable only through execution, tailoring a harness to each query has required either executing alternatives at inference time or costly manual design. We introduce SHIFT, which moves execution out of the per-query search loop. A local LLM architect learns a policy over harness-building actions from search, and a value function that predicts, from measured executions, a utility balancing accuracy against execution cost. For each query, Monte Carlo tree search uses these predictions to construct a harness. Across 9,193 tasks in six benchmarks, from math to document and general-assistant tasks, with a Gemini 3.5 Flash executor, SHIFT attains the highest mean accuracy, about 80%, outperforming 17 baselines that span prompting, prompt optimization, and workflow search, and exceeding the strongest baseline by 7.2 percentage points. A cheaper mode of SHIFT also attains a higher mean accuracy than every baseline while using 32% fewer execution tokens than the strongest baseline. We further show that choosing structure, instructions, and tools jointly beats choosing only instructions or only tools by up to 9.1 percentage points, and that learned value selection identifies more accurate harnesses with lower execution cost from candidate pools.
|
| 885 |
Dual-Scale Relational Graph Transformers for Ecosystem-Aware Fraud Detection
2610.04138
|
cs.LGcs.AI
|
Mohsen Nayebi Kerdabadi, Xinrou Li, Yao Xiao, Zijun Yao, Xin Sun |
Account takeover (ATO) fraud is a growing threat to digital banking, requiring effective detection while minimizing friction for legitimate customers. Production systems predominantly rely on tabular models that score sessions in isolation, discarding the rela...Account takeover (ATO) fraud is a growing threat to digital banking, requiring effective detection while minimizing friction for legitimate customers. Production systems predominantly rely on tabular models that score sessions in isolation, discarding the relational structure of the underlying interaction network. Although graph-based models exploit relationships among sessions and network entities, they primarily reason over local neighborhoods and therefore capture only part of the problem: fraud risk depends jointly on the local relational structure surrounding a session and the evolving global state of the fraud ecosystem. We present HERMES (HEterogeneous Relational Micro--macro graph transformer Encoder for high-risk Sessions), a dual-scale architecture that jointly models these complementary scales of information. Micro-GT captures local heterogeneous graph structure through structured, relation-aware attention over a temporally safe session neighborhood. Complementing this local representation, Macro-GT models ecosystem-level context using non-anticipative climate tokens that summarize fraud dynamics, platform shifts, and infrastructure reuse, together with adaptive class prototypes that track representative fraud and benign session patterns over time. Evaluated on more than 130 million high-risk transaction sessions from a leading U.S. financial institution, HERMES consistently outperforms production and strong graph-based baselines, achieving a 44.44% relative reduction in customer friction and a 24.66% relative improvement in fraud recall over the production system. Ablation and temporal-stability analyses further demonstrate complementary gains from local relational modeling and global ecosystem context across changing fraud regimes.
|
| 886 |
Ideal Paths for Approximating Logistic Gradient Descent Trajectories at Large Initialization
2610.04142
|
cs.LG
|
Junjie Xiao, Huiwen Jia |
Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected...Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected by prolonged training, but this direction alone does not provide information regarding the intermediate behavior. We address this question through a geometric approximation of full-batch logistic gradient descent (GD) trajectories on strictly linearly separable data, with large initialization of scale $R$ motivated by prior training. From any limiting normalized initial position, we use minimum-norm projection rules to construct a unique continuous ideal path consisting of finitely many linear segments. The path has two stages: negative-margin correction followed by minimum-margin growth. We prove that, after an explicit two-stage time reparameterization, the fixed-step GD trajectory divided by $R$ converges uniformly to this path on every fixed parameter interval as $R\to\infty$. Further, our quantitative error bounds account for initialization perturbations and the transition between stages. This approximation provides asymptotic formulas for peak evaluation loss and cumulative training loss. In particular, peak evaluation loss can grow linearly in $R$ even when both endpoint losses tend to zero. The cumulative losses in the correction and margin-growth stages, normalized by $R^2$ and $R$, respectively, converge to explicit limits. Experiments on controlled geometries and fixed image features complement our theoretical results.
|
| 887 |
Consideration Circuits: Depth Separation and Universality Beyond a Single Softmax
2610.04143
|
cs.LG
|
Junjie Xiao, Huiwen Jia |
Most feature-based choice models, classical and deep, score items and apply a single softmax. We introduce consideration circuits (CC), feature-based models of multi-stage choice defined by directed acyclic graphs of multinomial logit (MNL) units. Source units...Most feature-based choice models, classical and deep, score items and apply a single softmax. We introduce consideration circuits (CC), feature-based models of multi-stage choice defined by directed acyclic graphs of multinomial logit (MNL) units. Source units assign probabilities to menu items, and internal units combine predecessor distributions using MNL weights computed from their probability-weighted feature summaries. On a three-item compromise task with fixed non-collinear features, menu-independent random-utility models (RUM), including a single MNL unit, suffer an error bounded away from zero. For CC, in contrast, we establish a sharp depth--norm separation: increasing depth from $2$ to $3$ reduces the optimal maximum taste-vector norm for error $\epsilon$ from $\Theta(\log(1/\epsilon)/\epsilon)$ to $\Theta(\log(1/\epsilon))$. The depth-$2$ lower bound holds for arbitrary width and menu-independent routing biases, while a five-node depth-$3$ circuit with zero routing biases attains the logarithmic rate. More generally, we characterize two geometric conditions that are necessary and sufficient for approximating arbitrary deterministic choice tables on finite menu families. Under these conditions, depth $3$ suffices, while depth $4$ achieves optimal logarithmic norm scaling whenever the family contains a non-singleton menu. In experiments, standalone tree circuits with fewer than $600$ parameters attain the lowest mean test negative log-likelihood (NLL) among the evaluated models on four fixed-pool benchmarks and the Expedia temporal split. As output heads, CC generalize the linear MNL readout and lower mean test NLL for every tested encoder on Expedia and Trivago.
|
| 888 |
You May Be Running the Wrong Inception Crop
2610.04147
|
cs.LG
|
Jason Chuan-Chih Chou |
A decade after its inception, Inception crop has become the standard crop-based data augmentation method for training deep vision models. Not only is its practice of uniformly sampling crop scale and aspect ratio widely adopted, but also its lower and upper bo...A decade after its inception, Inception crop has become the standard crop-based data augmentation method for training deep vision models. Not only is its practice of uniformly sampling crop scale and aspect ratio widely adopted, but also its lower and upper bounds, with the scale lower bound being the sole exception that is sometimes tuned. It is therefore surprising that the standard implementation in the TensorFlow / JAX ecosystem samples crop scale with probability density function $f(A) \propto \frac{1}{\sqrt{A}}$ unlike the PyTorch counterpart, which follows the original description. Motivated by this discovery, we train 522 ViT-S/16 models on the ImageNet-1k dataset with various training budgets and crop scale distributions. We reach $78.78\pm0.09$ top-1 val. accuracy with 90 epochs of training budget and find that 1. Higher training budget requires stronger augmentation; 2. Lower tail of the distribution of the crop scale determines the augmentation strength of Inception crop; 3. Models trained with higher training budget exhibit sparser saliency, regardless of the crop scale distribution or weight decay. Based on 2. we revisit the performance of Beta crop, whose softer cutoff allows it to optimize model performance across training budgets with less compromise. We replicate 1. and 3. with Scion optimizer in addition to AdamW, suggesting that the results may be general.
|
| 889 |
How RL Reshapes LLM Reasoning: Transferability, Coverage, and Scaling Laws
2610.04158
|
cs.LGcs.AI
|
Ziheng Cheng, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi |
Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve performance in other domains, yet gains in Pass@1 can coincide with lower Pass@$N$ than the base...Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve performance in other domains, yet gains in Pass@1 can coincide with lower Pass@$N$ than the base model. This raises a fundamental question: does RL expand an LLM's reasoning boundary, or merely reweight its existing reasoning space? We revisit these phenomena across Qwen and Gemma model families, showing both cross-domain gains and forgetting, while coverage at large sampling budgets increases on some tasks and decreases on others. Detailed analysis of solution traces before and after RL indicates a shift in the reasoning strategies the model employs, motivating a two-stage autoregressive policy model that separates \emph{strategy selection} from problem-specific execution. Within this framework, we prove how RL's implicit bias reshapes strategy preferences, allowing gains on some tasks while suppressing strategies required by others. This mechanism can also broaden or narrow coverage at a given sampling budget even without expanding strategy support. We further provide theoretical justifications for log-sigmoid and log-linear scaling laws in RL compute, and evaluate their predictive power. Together, these results connect changes in strategy selection to cross-domain transfer, reasoning coverage, and compute scaling.
|
| 890 |
FTD-GNO: Memory-Efficient Graph Neural Operators through Functional Tensor Decomposition of the Kernel
2610.04212
|
cs.LGcs.AI
|
Xiaomin Zhang, Boyue Wang, Junbin Gao, Yongli Hu anbd Baocai Yin |
Graph Neural Operators (GNOs) provide flexible surrogate models for learning solution operators of partial differential equations (PDEs). However, standard GNOs typically parameterize the integral kernel with a monolithic neural network and evaluate kernel int...Graph Neural Operators (GNOs) provide flexible surrogate models for learning solution operators of partial differential equations (PDEs). However, standard GNOs typically parameterize the integral kernel with a monolithic neural network and evaluate kernel interactions over graph edges, leading to substantial computational and memory overhead at high resolutions or with large neighborhoods. To address these limitations, we propose Functional Tensor Decomposition Graph Neural Operator (FTD-GNO), a memory-efficient GNO framework that decouples the high-dimensional continuous integral kernel into low-dimensional mode-wise functions. By instantiating the kernel with classical tensor decomposition formats, including CP, Tensor-Train, and Tucker decompositions, FTD-GNO enables algebraic reconstruction of the integral operator without explicitly materializing full edge-wise kernel tensors. This factorized formulation reduces the memory footprint of kernel evaluation and aggregation while retaining the continuous operator-learning structure of GNOs. Theoretical complexity analysis shows that FTD-GNO substantially lowers parameter and activation-memory costs associated with high-dimensional kernel construction. Experiments show lower peak memory than the corresponding unfactorized graph-integral baselines, with shorter recorded training times. Fourier-graph experiments further demonstrate that FTD can improve the efficiency of a graph-integral layer within a hybrid operator and has good scalability.
|
| 891 |
PaLoRA: Paced Low-Rank Adaptation for Continual Learning
2610.04226
|
cs.LG
|
Yuxuan Li, Fanhu Zeng, Hao Tang |
LoRA-based continual learning methods mitigate catastrophic forgetting through various mechanisms, yet nearly all complement these with small learning rates as a heuristic to restrict gradient scaling magnitude. Such fixed heuristics lack theoretical guidance ...LoRA-based continual learning methods mitigate catastrophic forgetting through various mechanisms, yet nearly all complement these with small learning rates as a heuristic to restrict gradient scaling magnitude. Such fixed heuristics lack theoretical guidance on how the strength of this restriction should evolve as tasks accumulate. We reveal that even under directional constraints such as nullspace projection, finite-precision updates inevitably leak into the subspace of accumulated prior knowledge along multiple directions. While small learning rates attenuate such leakage, they cannot prevent the accumulated forgetting from intensifying as the effective rank of historical knowledge grows. We show that the optimal magnitude restriction should adaptively increase with this effective rank to balance stability and plasticity, i.e., preservation of previous knowledge and acquisition of new task information. Under an anisotropic leakage model, we derive a pacing law $s^*=\sqrt{R/c}$ that characterizes the optimal scaling of gradient steps, i.e., the magnitude restriction itself, where $R$ is the effective rank of past updates. Based on this insight, we propose PaLoRA, which compresses historical knowledge via adaptive SVD truncation, projects gradients onto the nullspace of prior tasks, and applies rank-aware adaptive pacing. Experiments demonstrate consistent improvements over prior methods, with particularly strong performance in long-horizon settings, achieving substantial gains of 4% accuracy on challenging 50-task ImageNet-A and ImageNet-R benchmarks.
|
| 892 |
Stochastic Adaptive Fourier Decomposition for Operator Learning
2610.04241
|
cs.LG
|
Pengqing Shi, Liming Zhang, Tao Qian, Stephen Tierney, Jie Yin |
Fourier Neural Operators (FNOs) offer an efficient paradigm for solving partial differential equations (PDEs). However, FNOs rely on a fixed Fourier basis and hard frequency truncation, which inherently limit their ability to model non-periodic, localized, and...Fourier Neural Operators (FNOs) offer an efficient paradigm for solving partial differential equations (PDEs). However, FNOs rely on a fixed Fourier basis and hard frequency truncation, which inherently limit their ability to model non-periodic, localized, and fine-scale solution structures. We propose the Stochastic Adaptive Fourier Decomposition Neural Operator (SAFDNO), a spectral neural operator that replaces predefined Fourier modes with an adaptive Takenaka-Malmquist (TM) orthonormal system derived from the theory of Stochastic Adaptive Fourier Decomposition (SAFD). Instead of performing an expensive greedy pole search in classical SAFD, SAFDNO amortizes stochastic pole selection through a neural pole predictor and constructs the adaptive TM system directly from latent features. The resulting operator performs learned filtering on coefficients of analytic branches in the Hardy space with respect to the input-adaptive TM system, preserving the global receptive field of spectral operators while providing a more flexible representation than using fixed spectral bases. Across nine PDE benchmark problems, SAFDNO achieves the best performance on all six regular-grid problems among strong neural operator baselines, with especially notable gains on Darcy, Burgers, and Navier-Stokes, where fixed Fourier modes are often less effective at modeling localized oscillations, sharp transitions, and multiscale structures. SAFDNO also exhibits stronger zero-shot super-resolution performance and shows less performance degradation when deployed on finer discretizations. These results suggest that the input-adaptive TM system provides a promising alternative to fixed spectral representations for neural operator learning.
|
| 893 |
Optimizer Geometry Sets the Pace: Spectral Learning Dynamics in Matrix Factorization
2610.04249
|
cs.LG
|
Mahalakshmi Sabanayagam, Simon Lucey |
Recent successes of matrix- and curvature-based optimizers have renewed interest in how update geometry shapes learning. These methods normalize or precondition updates, changing how different components progress during training. In deep matrix factorization, ...Recent successes of matrix- and curvature-based optimizers have renewed interest in how update geometry shapes learning. These methods normalize or precondition updates, changing how different components progress during training. In deep matrix factorization, the geometry that slows gradient descent (GD) favors low-rank solutions by delaying the emergence of small singular modes. This raises the question of what remains of that spectral bias when normalization or curvature correction weakens or removes the slowdown. We study how optimizer geometry and depth jointly govern this behavior through a common framework for singular mode dynamics. Under explicit balance and alignment assumptions, we derive birth, saturation, and decay laws for Euclidean GD, coordinate-wise updates of SignGD and an instantaneous Adam approximation, spectral updates of Muon and cumulative Shampoo, and a block curvature model of K-FAC. The resulting picture is not a simple ordering from stronger to weaker low-rank bias. To highlight, SignGD and ideal Muon eliminate the divergent birth barrier and drive unsupported modes to zero in finite time. Cumulative Shampoo initially retains GD's depth-dependent barrier, then accumulated gradients produce a catch-up phase while making previously active modes increasingly persistent. Undamped K-FAC cancels the factorization-induced slowdown while preserving the ordering of the target singular values, whereas positive damping introduces a spectral threshold below which the slow GD phase laws reappear. These results give normalization, accumulated state, damping, and depth a direct interpretation as controls determining when modes emerge, persist, and disappear during training.
|
| 894 |
A Hand-Checkable Proof That Two Hidden ReLU Layers Compute the Maximum of Six Numbers
2610.04256
|
cs.LG
|
Dimitrios Myrisiotis |
Exactly computing the maximum function is a standard test case for studying depth in ReLU networks. Two hidden layers are known to suffice for up to twelve inputs through computer-assisted constructions. For six real inputs, we give an explicit hexagon identit...Exactly computing the maximum function is a standard test case for studying depth in ReLU networks. Two hidden layers are known to suffice for up to twelve inputs through computer-assisted constructions. For six real inputs, we give an explicit hexagon identity whose local structure yields a self-contained analytical proof of this depth bound. The identity was found by computer-assisted search; we prove it through explicit cancellations that can be checked entirely by hand, without executing a verification program. The identity also yields an explicit network with hidden widths $17$ and $41$, zero biases, and rational weights.
|
| 895 |
Cross-Trait Transfer in Subliminal Learning
2610.04260
|
cs.LG
|
Xingyu Zhao, Yiqiao Zhong |
Subliminal learning is a phenomenon where a student language model acquires a teacher model's behavioral traits by training on semantically unrelated outputs. It is a subtle statistical phenomenon as trait transmission relies on weak statistical patterns in th...Subliminal learning is a phenomenon where a student language model acquires a teacher model's behavioral traits by training on semantically unrelated outputs. It is a subtle statistical phenomenon as trait transmission relies on weak statistical patterns in the generated data. To understand trait transmission between teacher-student pairs, we study cross-trait transfer: how data generated under one teacher trait changes the student's preferences of other traits. To this end, we introduce a directed trait-transfer matrix that quantifies these effects using log-probability gains for student answers. We find that the trait-transfer matrix reveals clusters of related traits, with students sometimes developing preferences for traits similar, but not identical, to the teacher's trait. Such cross-trait structure can be partially captured by output distribution metrics and representation-based metrics. Further, we analyze trait development and interaction: learning dynamics shows a progression from broad shared shifts toward more trait-specific transfer, and multi-trait experiments suggest that opposed traits can enhance such differentiation. Together, our findings reveal salient statistical structures over trait transfer and competition, thus providing a broader view of how hidden preferences are transmitted in subliminal learning.
|
| 896 |
Adaptive Bregman Alternating Projections for Feasible Gromov-Wasserstein Learning
2610.04264
|
cs.LG
|
Aoran Zhang, C\'esar A. Uribe |
The Gromov-Wasserstein (GW) problem compares structured distributions without requiring a shared feature space or known correspondences, but its nonconvex objective and coupled marginal constraints make computation challenging. Bregman alternating projected gr...The Gromov-Wasserstein (GW) problem compares structured distributions without requiring a shared feature space or known correspondences, but its nonconvex objective and coupled marginal constraints make computation challenging. Bregman alternating projected gradient (BAPG) uses inexpensive alternating row and column updates, yet its fixed-penalty relaxation leaves a persistent feasibility gap. We propose Adaptive KL-BAPG (A-KL-BAPG), which combines a finite fixed-penalty burn-in with a guarded increasing-penalty phase. At each tail iteration, the method reuses BAPG's alternating updates and backtracks a delayed-power step until a Sinkhorn-inspired projective-diameter safeguard is satisfied. We prove finite termination of the backtracking at each iteration and show that the feasibility gap vanishes asymptotically. We further establish a best-iterate $O(1/\log N)$ bound for the weighted squared corrected residual and, under a support regularity condition, the existence of a stationary accumulation point for the original GW problem. This distinguishes A-KL-BAPG from fixed-penalty BAPG, whose stationarity guarantees are given for the relaxed problem. Experiments show that A-KL-BAPG achieves a favorable balance of accuracy, objective value, feasibility, and stationarity relative to BAPG variants, projection-based methods, and task-specific baselines. For synthetic and real graph alignment problems, it closely matches the accuracy and objective value of fixed-penalty KL-BAPG while reducing the marginal feasibility gap by 62-99% and the projected stationarity residual by 28-98%. Heterogeneous domain adaptation experiments show a similar pattern: A-KL-BAPG maintains comparable target accuracy and objective values while achieving better feasibility and stationarity than fixed-penalty KL-BAPG.
|
| 897 |
Self-Reflection Fine-Tuning: Enhancing Agent Security against Prompt Injection Attacks from Failure Experience
2610.04269
|
cs.LG
|
Zixuan Wang, Hao Li, Fengyu Gao, G. Edward Suh, Yi Zeng |
Large language model (LLM) agents are increasingly deployed in tool-augmented environments, but their reliance on external inputs makes them highly vulnerable to prompt injection attacks that can hijack task objectives. Existing safety alignment methods rely o...Large language model (LLM) agents are increasingly deployed in tool-augmented environments, but their reliance on external inputs makes them highly vulnerable to prompt injection attacks that can hijack task objectives. Existing safety alignment methods rely on static expert trajectories or preference optimization, limiting their ability to generalize to adaptive attack patterns. In this work, we propose Self-Reflection Fine-Tuning (SRFT), a training framework that enables agents to improve robustness by learning from their own failure experiences under adversarial conditions. Instead of passively imitating expert behaviors, SRFT exposes the agent to compromised trajectories constructed via injected attacks, and leverages an expert model to generate structured self-reflection reasoning that contrasts unsafe and optimal actions. This reflective supervision teaches the agent to identify malicious instructions, reason about their consequences, and maintain alignment with the original user intent. We instantiate this framework in SR-Agent, built on Llama-3.1-8B-Instruct and Qwen3-8B, and evaluate it on both static and adaptive prompt injection benchmarks. Experimental results show that SRFT substantially reduces attack success rates while preserving task performance, and demonstrates strong generalization under adaptive attacks. These findings suggest that learning from failure via self-reflection is a promising direction for building robust and secure LLM agents. Our code is released at https://github.com/Eden-Wang1710/srft-repo.
|
| 898 |
Integrated Imputation-Classification for Supervised Learning with Missing Data
2610.04273
|
cs.LG
|
Yue Liu, Ben Liang, Ali Tizghadam, Ilijc Albanese |
We study supervised classification problems with missing feature values. Existing approaches often decouple imputation from classification, producing imputations that may be plausible but uninformative for prediction. Instead, we propose the Integrated Imputat...We study supervised classification problems with missing feature values. Existing approaches often decouple imputation from classification, producing imputations that may be plausible but uninformative for prediction. Instead, we propose the Integrated Imputation and Classification Network (IICN), which jointly trains an imputer and an ${(n{+}1)}$-classdiscriminator adversarially with a single class supervised classification objective, where the discriminator learns to distinguish among the $n$ true classes and an additional ``imputed" class. We prove that at the global optimum, the imputer and discriminator together implement marginalization over missing coordinates and yield a Bayes-optimal classifier. We evaluate IICN on FashionMNIST, CIFAR-10, and tabular datasets with naturally occurring missingness. IICN outperforms classical impute-then-classify pipelines and recent generative baselines, showing strong robustness and accuracy in challenging settings.
|
| 899 |
Do RUL explanations hold up? Faithfulness and stability of attributions on C-MAPSS
2610.04278
|
cs.LG
|
Manh Hien Nguyen, Ngoc Thanh Nguyen, Isabella Mendoza Cortes, Tam Khuat, Thanh Pham |
Deep remaining-useful-life (RUL) models on NASA C-MAPSS are now routine, and so are heatmaps that colour sensors and timesteps. A heatmap that looks mechanical is not the same as an explanation an engineer can act on. We train three standard architectures - a ...Deep remaining-useful-life (RUL) models on NASA C-MAPSS are now routine, and so are heatmaps that colour sensors and timesteps. A heatmap that looks mechanical is not the same as an explanation an engineer can act on. We train three standard architectures - a 1D CNN, an LSTM, and a small Transformer encoder - on the official FD001 and FD003 splits with the piecewise RUL cap of 125 cycles and the official PHM08 asymmetric score. We then attach three attribution maps (Integrated Gradients, occlusion, last-layer attention) and evaluate them with the checks the XAI-for-PdM literature still under-reports: deletion/insertion faithfulness, Spearman stability under sensor-scale noise, agreement across training seeds, and cosine consistency inside RUL bins. Prediction error is a prerequisite, not the claim. The headline is which explanation method moves the RUL output when its top cells are removed, and which map survives a 5% input perturbation. Integrated Gradients and occlusion are similarly faithful on the LSTM; Transformer attention is cheap and temporally smooth but weakly faithful. All three maps are almost unchanged under 5% input noise, yet IG/occlusion agree only moderately across two LSTM seeds - stability to sensor jitter is not the same as stability to retraining. A secondary tabular check on the AI4I 2020 failure dataset shows the same deletion pattern for tree importances. We recommend occlusion or IG for any C-MAPSS-style report that will be read by a maintenance engineer, and we treat raw attention weights as a visualisation only.
|
| 900 |
What to Preserve in Recursive Computation: A Local Predictive Sufficiency Principle
2610.04303
|
cs.LGcs.AI
|
Peilin Wang, Feng Shiyang, Hongfu Gao, Cencheng Zhao, Di Yuan |
Recursive computation repeatedly compresses or reuses intermediate states, creating a simple tension: information that must remain useful across longer recursive paths is also exposed to more opportunities for loss before reaching the final prediction. Existin...Recursive computation repeatedly compresses or reuses intermediate states, creating a simple tension: information that must remain useful across longer recursive paths is also exposed to more opportunities for loss before reaching the final prediction. Existing reconstruction or local-prediction objectives provide tractable supervision, but do not ensure that the retained information remains sufficient for subsequent recursive computation. We identify local predictive sufficiency with recursive predictive closure: controlling local predictive deficiencies at individual interfaces controls the resulting discrepancy at the root. We then turn this principle into a tractable training procedure. Starting from a variational characterization, we derive finite predictive tests and an empirical predictive deficiency that measures predictive value retained across compression. Its predictive sensitivities define margin-relaxed half-space constraints on parameter updates, and we project the host optimizer's proposed update onto their intersection only when predictive preservation would otherwise be violated. Across temporal graphs, language memory, vision-language-action control, and recursive self-improvement, the method matches or improves the corresponding host models under matched compression budgets, with larger gains under heavier recursive or memory demands, while better preserving predictive information across successive transformations. Crucially, the same task-agnostic predictive-preservation principle is instantiated across all four settings through host-compatible interventions while keeping the endpoint task, backbone, and evaluation protocol fixed. These results establish predictive preservation at recursive interfaces as a general training principle for recursive compression.
|
| 901 |
Modeling Deletion Requests in Machine Unlearning
2610.04310
|
cs.LG
|
Christian Cianfarani, Aloni Cohen |
Machine unlearning is seen as a promising approach to enable users to exercise the "right to erasure" in the context of AI models. We ask how users might influence the behavior of models when exercising this right. We define two types of behaviors that users m...Machine unlearning is seen as a promising approach to enable users to exercise the "right to erasure" in the context of AI models. We ask how users might influence the behavior of models when exercising this right. We define two types of behaviors that users might adopt when requesting the deletion of their data: adaptivity and collectivity. Drawing connections between the goals of users in this context and results in stochastic optimization, we demonstrate theoretical gaps between the potential effects of groups of users who do and do not display these behaviors. We then show how techniques from data valuation might be used to design deletion requesters that can significantly alter model behavior in realistic settings. In experiments on computer vision tasks, we demonstrate the differential effects of different models of user behavior and attempt to isolate the impacts of adaptivity and collectivity.
|
| 902 |
LyapuFlow: Controlling Generative Flows with Lyapunov Feedback for Inverse Problems
2610.04326
|
cs.LG
|
Minseon Gwak, Hans Hao-Hsun Hsu, Danielle C. Maddix, N. Benjamin Erichson |
Pretrained flow models are now widely used as generative priors in science and vision, where inference-time guidance enables test-time constraints without retraining. Existing methods use projection, posterior sampling, or iterative optimization of the generat...Pretrained flow models are now widely used as generative priors in science and vision, where inference-time guidance enables test-time constraints without retraining. Existing methods use projection, posterior sampling, or iterative optimization of the generative trajectory. We propose LyapuFlow, an alternative based on Lyapunov feedback control. At each sampling step, LyapuFlow predicts the terminal sample towards which the current flow is evolving, and evaluates the constraint violation on this prediction. Then, we compute the minimum-norm control that satisfies a prescribed Lyapunov decrease condition. The resulting control remains inactive when the uncontrolled dynamics already reduce the constraint violation at the prescribed rate. Otherwise, it provides a corrective update within a feedback trust region that prevents the control from dominating the pretrained dynamics. We demonstrate LyapuFlow in both data and latent spaces, outperforming alternatives spanning different mechanisms for test-time constraint enforcement in scientific machine learning and image inverse problems.
|
| 903 |
S$^3$N: A Spherical Spiral Scanning Network for Weather Forecasting
2610.04338
|
cs.LG
|
Fan Yan, Chen Hui, Weisi Lin, Haiqi Zhu, Feng Jiang |
Machine learning-based weather prediction (MLWP) has achieved strong performance in global weather forecasting. Recent Hierarchical Equal Area isoLatitude Pixelation (HEALPix)-based methods use the HEALPix (HP) grid to avoid area distortion near the poles of c...Machine learning-based weather prediction (MLWP) has achieved strong performance in global weather forecasting. Recent Hierarchical Equal Area isoLatitude Pixelation (HEALPix)-based methods use the HEALPix (HP) grid to avoid area distortion near the poles of conventional latitude-longitude (LL) grids. However, existing HP-based approaches often use pointwise mapping methods and process HP pixels within separate base faces or local windows. Consequently, the mapping may introduce reconstruction errors and cross-face communication depends on handcrafted boundary handling or shifted windows. We propose the Spherical Spiral Scanning Network (S$^3$N) to address both limitations. First, L2Proj provides a bidirectional method for mapping atmospheric fields between the LL and HP grids through an $L^2$ projection of their continuous finite-element representations. Second, the Attention-Guided Quad-Spiral State-Space Scanning (AQSS) block uses cross-latitude attention to guide selective state-space updates along four global pole-to-pole spiral paths. This design enables continuous information propagation across HP base-face boundaries without additional boundary-processing mechanisms. Experiments show that S$^3$N achieves better results at 4-, 7-, and 10-day lead times, and exhibits slower error growth in long-range forecasting.
|
| 904 |
Checkable NTK Positivity and Finite-Width Gradient Descent for Scalar- and Vector-Valued PINNs with Strong-Form, Weak-Form, and Nonlocal Linear Constraints
2610.04357
|
cs.LG
|
Zifan Lyu |
We give checkable positive-definiteness criteria for the limiting neural tangent kernel (NTK) and high-probability finite-width gradient-descent guarantees for scalar- or vector-valued physics-informed neural networks (PINNs) with linear constraints. The const...We give checkable positive-definiteness criteria for the limiting neural tangent kernel (NTK) and high-probability finite-width gradient-descent guarantees for scalar- or vector-valued physics-informed neural networks (PINNs) with linear constraints. The constraints may be strong-form differential rows of any fixed finite order, including coupled systems, with any linear initial or boundary conditions; weak-form residual and boundary functionals; or finite-measure nonlocal observations such as integral, nonlocal-diffusion, and Dirac-data rows, all in any dimension. For each class, positive definiteness of the limiting NTK is equivalent to a rank condition on a finite coefficient, functional, or moment matrix of the fixed design: a certificate computed before training that detects structural zero modes. The model hypotheses are those of an ordinary two-layer PINN, a smooth nonpolynomial activation with bounded symmetric initialization, and are met by standard choices such as $\tanh$ with uniform initialization. Given a certificate, explicit width and step-size conditions ensure, with high probability, that the empirical NTK retains at least half the limiting gap and that full-batch gradient descent decreases the training loss geometrically at every iteration. Controlled experiments check the certificates and the finite-width mechanism.
|
| 905 |
A Bird's-Eye View of Iterative Reward Design
2610.04364
|
cs.LGcs.AI
|
Logan Mondal Bhamidipaty, Lauren Robson, Linda Petrini, Shengrui Lyu, Kamal Ndousse |
Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback. However, these methods a...Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback. However, these methods are often hard to compare because they differ in implementation details, feedback assumptions, and evaluation environments. To address this, we introduce a Benchmark for Iterative Reward Design (BIRD) that expresses existing methods in a unified configuration and evaluation space. This lets us compare algorithms directly, ablate individual design choices, and prototype new components under matched feedback conditions and policy-training budgets. Across MuJoCo, Meta-World, Assistax, and HumanoidBench, we identify a small set of simple design choices that consistently improve performance. Combining these choices yields significantly better performance than the evaluated methods from prior work. Our results highlight the strength of simple baselines and motivate further study of when additional algorithmic complexity improves iterative reward design. Code is available at https://github.com/safety-research/bird.
|
| 906 |
A multi-stage probabilistic framework to estimate gas-fired generator performance during extreme winter weather
2610.04368
|
cs.LG
|
Sajjad Uddin Mahmud, Anamika Dubey |
Extreme winter weather has repeatedly disrupted gas-fired power generation in the United States, yet the plant-level data needed to systematically quantify outage risk remain proprietary. Using publicly available weather and electricity demand data together wi...Extreme winter weather has repeatedly disrupted gas-fired power generation in the United States, yet the plant-level data needed to systematically quantify outage risk remain proprietary. Using publicly available weather and electricity demand data together with anonymized generator contingency records from the North American Electric Reliability Corporation (NERC), we develop a three-stage Bayesian probabilistic framework for estimating winter-driven generator performance. Applied to New York State (2013--2022), the framework sequentially estimates: the hourly probability of a generator contingency event, the expected net available capacity conditioned on an event occurring, and the event duration. Colder conditions and higher electricity demand are associated with higher failure probability, lower retained capacity, and longer event duration. Under the most severe observed stress conditions, estimated mean hourly event probability reaches 24\% , while expected mean net available capacity falls to 13\% of nameplate rating. Full outage events have a median duration of 12.7 hours, while partial derating event duration increases from 2.4 to 7.1 hours with capacity loss severity. The proposed framework establishes a transferable baseline that utilities with access to plant-level records can directly extend to obtain more precise reliability estimates for operational planning and resource adequacy assessment.
|
| 907 |
CEENs: Causality-enforced evolutional networks for solving time-dependent partial differential equations
2610.04405
|
cs.LG
|
Jeahan Jung, Heechang Kim, Hyomin Shin, Minseok Choi |
Despite the growing popularity of physics-informed neural networks (PINNs), their applicability in the long-time integration of partial differential equations (PDEs) remains constrained. We argue that this problem stems from the lack of consideration of tempor...Despite the growing popularity of physics-informed neural networks (PINNs), their applicability in the long-time integration of partial differential equations (PDEs) remains constrained. We argue that this problem stems from the lack of consideration of temporal causality in the original PINN formulation, resulting in a bias towards satisfying governing equations at later times before learning the initial condition and hence leading to erroneous solutions. To this end, we propose a novel method that seamlessly integrates temporal causality into the training process. Drawing inspiration from classical numerical methods where the temporal causality is reflected, we divide the time domain into nonoverlapping subintervals, assign a unique neural network to each subinterval, and construct a loss function founded on the integral form of PDEs within these subintervals. The proposed networks undergo sequential training, beginning with the initial time step. Our method demonstrates significant improvement in accuracy for long-time simulations of various PDE problems where the original PINN method fails while it requires less computational cost and memory compared to the PINN method. A parallelization algorithm is provided to further enhance the computational efficiency, showing a significant speedup for solving time-dependent PDEs.
|
| 908 |
Specific Algorithmic Interpretability of Neural Networks: A Case Study on Textures
2610.04413
|
cs.LGcs.AI
|
Yanglin Zhang, Anneke von Seeger, Gilad Lerman, Ron Levie |
We develop a principled framework for constructing neural networks whose specific parameter realizations admit an explicit algorithmic interpretation. Existing algorithm-inspired architectures can explain the computational structure of a network, yet after sta...We develop a principled framework for constructing neural networks whose specific parameter realizations admit an explicit algorithmic interpretation. Existing algorithm-inspired architectures can explain the computational structure of a network, yet after standard training the learned parameters need not retain a clear relation to the motivating algorithm. We address this gap as follows. First, we model each data point as a sample of a class-dependent stochastic process and assume that statistics of this process can be estimated from a single sample and these statistics are sufficient to distinguish the classes. We then construct a neural network whose initial parameters exactly implement an algorithm for estimating these statistics, making the network fully interpretable. To account for mismatch between the idealized model and real data, we fine-tune this network while controlling its deviation from the algorithmic initialization. The trained network hence roughly retains the interpretation of the initial network. A PAC-Bayesian analysis yields a uniform generalization bound whose complexity term scales with the fine-tuning radius, providing a statistical motivation for our approach. We instantiate the framework for texture classification using the scattering transform to estimate the discriminative statistics.
|
| 909 |
MaDeL: Manifold-Decomposed Feature Losses for Generative Modeling
2610.04419
|
cs.LGcs.AI
|
Beomsu Kim, Jong Chul Ye, Kwanyoung Kim |
Generative models are often trained with isotropic objectives such as mean-squared error. For data concentrated near a low-dimensional manifold, however, such losses conflate displacement along the manifold, which may represent valid variation, with displaceme...Generative models are often trained with isotropic objectives such as mean-squared error. For data concentrated near a low-dimensional manifold, however, such losses conflate displacement along the manifold, which may represent valid variation, with displacement away from it, which produces invalid samples. This mismatch is especially problematic in sparse, highly constrained domains, where ambient-space regression can encourage off-manifold interpolation. We ask whether a generative objective can distinguish manifold-parallel variation from manifold-orthogonal deviation directly from data, without explicitly estimating the manifold. We introduce a manifold-decomposed feature loss (MaDeL) that learns complementary representations from corrupted observations: one is trained to recover the clean sample, while the other is trained to recover the corruption. We show that, under a feature bottleneck, their Jacobians align with the tangent and normal spaces, exactly for linear manifolds and locally for smooth manifolds. Together, these representations define an anisotropic objective that separately measures intrinsic variation and off-manifold deviation. Across synthetic, Earth and climate science, and torsion-angle benchmarks, MaDeL improves support recovery and average angular $W_1$ under single-step sampling; on protein backbones, it reduces steric clashes across one- and few-step sampling budgets.
|
| 910 |
Measuring Effective Data Resolution in Guided Diffusion Posteriors
2610.04422
|
cs.LG
|
Ridham Patel, Defu Cao, Jiacheng Pang, Yan Liu |
Guided diffusion samplers are increasingly used to reconstruct physical fields from sparse observations, but standard diagnostics do not say how much of the reconstruction was actually determined by the data. We introduce effective data resolution for black-bo...Guided diffusion samplers are increasingly used to reconstruct physical fields from sparse observations, but standard diagnostics do not say how much of the reconstruction was actually determined by the data. We introduce effective data resolution for black-box generative posteriors: a comparison between the resolution warranted by the inverse problem, $\mathrm{dof}_{\mathrm{ref}}$, and the resolution realised by the sampler, $\mathrm{dof}_{\mathrm{samp}}$. A perturbation estimator measures $\mathrm{dof}_{\mathrm{samp}}$ and the spatial map $R(x,x)$ from sampler queries alone. We validate the estimator against exact references and use it to study guided diffusion. The resulting measurements show that guidance weight can strongly alter apparent information transfer, that mean, spread and resolution are not jointly corrected by one weight even with an exact prior and score, and that resolution fidelity does not follow reliably from the apparent principledness of a guidance rule.
|
| 911 |
On the Trade-off Between Information Loss and Generalization in Sparse Attention
2610.04424
|
cs.LG
|
Zhongqi Fan, Zheng Tan |
To mitigate the quadratic complexity bottleneck of the Transformer, sparse attention has emerged as a pivotal technology. Despite the extensive empirical success of sparse Transformers, the theoretical understanding of sparse attention remains fragmented. In p...To mitigate the quadratic complexity bottleneck of the Transformer, sparse attention has emerged as a pivotal technology. Despite the extensive empirical success of sparse Transformers, the theoretical understanding of sparse attention remains fragmented. In particular, two fundamental questions remain unclear: (1) How does sparsification affect the information fidelity of attention mechanisms? (2) How does this information loss interact with the generalization behavior of the model? To bridge this gap, this paper proposes a systematic analysis of the Jensen-Shannon (JS) divergence and of the generalization gap of sparse attention mechanisms. Specifically, we first characterize the approximation error via the JS divergence. Through an order-statistics-based concentration analysis of the truncation mass alpha --- where the attention scores are assumed to be independent and identically distributed sub-Gaussian random variables with parameter sigma --- the JS divergence between the full attention distribution and the sparse attention distribution is shown to admit the closed form log 2 + ((1 - alpha)/2) log(1 - alpha) - ((2 - alpha)/2) log(2 - alpha). Subsequently, we derive a generalization bound through Rademacher complexity, quantified by O(gamma * sqrt(M/n) * (sqrt(log(3eL/M)) + sqrt(pi)/2)). Furthermore, building on a mutual-information-based generalization bound together with an entropy and covering-number analysis of the sparse hypothesis class, we obtain the sparsity-dependent generalization bound O(sqrt((M/(2n)) * (log(eL/M) + log(1 + 2/epsilon)))). Our analysis shows that sparsity reduces the complexity of the considered hypothesis class while introducing approximation error that can be quantified by the JS divergence. These findings provide a theoretical characterization of the trade-off between information fidelity and generalization in sparse Transformer architectures.
|
| 912 |
COSMOS: Soft Mechanism Mixtures with Verifiable Routing for Long-Horizon PDE Forecasting
2610.04427
|
cs.LGcs.AI
|
Anupam Rawat, Manikandan Padmanaban, Jagabondhu Hazra |
Neural operators offer an efficient alternative to classical PDE solvers, but most learn a monolithic map per equation family and discretization. Real systems are compositional: transport, diffusion, wave, and reaction processes can act simultaneously. Existin...Neural operators offer an efficient alternative to classical PDE solvers, but most learn a monolithic map per equation family and discretization. Real systems are compositional: transport, diffusion, wave, and reaction processes can act simultaneously. Existing mixture-of-experts operators typically use sparse top-$K$ routing, although concurrent physics is naturally a blend rather than a discrete choice. We propose COSMOS (Cooperative Operator Specialists with Mechanism-level Operator Soft-routing), a soft mechanism-mixture neural operator. Four process-biased specialists remain active at every step and are continuously mixed by a learned gate, with their features fused by a small network. Specialists share a coarse latent grid, while a zero-initialized full-resolution residual restores detail lost through the bottleneck. We also introduce an operator-splitting compositional benchmark with known mixture weights $w^\star$ per trajectory. Against family-tuned FNO under 20-step rollouts, an initial 3-seed evaluation suggested gains on diffusion--reaction, parity on Navier--Stokes, and weaker shallow-water performance. An 11-seed audit showed that the diffusion--reaction gain was unstable, motivating caution in small-seed rollout comparisons and precluding a reliable accuracy-win claim on these families. Ablations show that uniform routing or removing the specialist mixture substantially degrades stable-regime rollouts. On the labeled benchmark, dense soft routing yields $2.2\times$ lower error than hard top-1 routing at identical fusion; using generator weights $w^\star$ at inference further lowers rollout error to $0.034$, diagnosing limitations of the learned gate. However, routing labels do not align with the specialists' intended mechanisms: COSMOS supports compositional accuracy, not mechanism identity.
|
| 913 |
LocusRL: Diagnosing LLM Reward and Policy Interventions in Competitive Games
2610.04441
|
cs.LG
|
Chengyu Luan, Bo Xin, Songyan Guo, Yuxiang Zuo, Ahmed Yazdan |
Large language models can intervene in reinforcement learning through both reward design and action selection, yet aggregate performance offers an incomplete account of what these interventions actually do. Similar returns can conceal different learning mechan...Large language models can intervene in reinforcement learning through both reward design and action selection, yet aggregate performance offers an incomplete account of what these interventions actually do. Similar returns can conceal different learning mechanisms, while plausible rewards can induce undesirable behavior. We introduce LocusRL, a diagnostic framework that connects controlled reward-policy comparisons with audits of reward judgments, signal delivery, optimization objectives, and executed actions. The framework traces performance differences to testable explanations and checks targeted corrections through executable rules and counterfactual replay. Across two evaluation batches covering ten Connect Four training seeds, we uncover seed-dependent reversals in intervention effects and show how tracing actual updates changes their interpretation: historical Qwen training operates through reward-weighted teacher-action likelihood. A separate matched three-seed reward-direction experiment distinguishes sensitivity to a learning signal from its usefulness. With terminal rewards held fixed, a sign-reversed dense oracle yields a 2.8% aggregate win rate, compared with 57.2% for terminal-only training and 46.7% for the positive dense oracle. Thus, a reward can strongly influence learning without improving performance. At the decision level, counterfactual replay verifies a winning correction to a diagnosed action error. Complementary experiments in Leduc and reward-validation studies in Goofspiel extend the analysis to imperfect-information settings, revealing how reference-label definitions and validation-data exposure affect intervention assessment. Together, these findings show why evaluating LLM interventions requires tracing how their outputs become learning signals and actions. LocusRL turns aggregate outcomes into actionable diagnoses and verifiable corrections.
|
| 914 |
One-Step Generation via Riemannian Wasserstein Gradient Flows
2610.04454
|
cs.LG
|
David Li, Chanhyuk Lee, Jaehoon Yoo, Nikita Gushchin, Floor Eijkelboom |
Recently, Drifting Models and Wasserstein Gradient Flows have attracted substantial attention because they move iterative distributional refinement to training and amortize it into a generator, enabling fast inference. However, existing formulations have been ...Recently, Drifting Models and Wasserstein Gradient Flows have attracted substantial attention because they move iterative distributional refinement to training and amortize it into a generator, enabling fast inference. However, existing formulations have been developed largely for continuous Euclidean domains, such as image spaces, where particles admit unconstrained additive updates. On constrained spaces, these updates can leave the valid domain or ignore its geometry, making them unsuitable targets for training. Recent work has adapted updates to these spaces, but has focused on particular fields or offered limited empirical comparison. We derive and compare several geometry-aware fields within a common training framework for one-step generators. We test the method on data with different structures and obtain competitive one-step results in each setting. The best-performing field varies by task, showing why the choice of objective matters in practice.
|
| 915 |
Target-free Latent Safety Alignment
2610.04467
|
cs.LGcs.AI
|
Luoyu Chen, Weiqi Wang, Chenhan Zhang, Zhiyi Tian, Yuxian Huang |
Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then...Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improvements in safety alignment. However, these methods typically construct adversarial samples either by encouraging fixed harmful target completions or by performing targeted activation ablation derived from fixed benign--harmful data pairs. As a result, the generated adversarial samples tend to induce homogeneous harmful behaviors that poorly reflect the diversity of behaviors elicited by real-world jailbreak attacks. This behavior-level narrowness fundamentally limits their robustness. To address this issue, we propose a target-free adversarial training framework that generates adversarial samples in an unsupervised manner. By amplifying and diversifying behavior-level shifts in the model's latent space, our approach produces semantically diverse adversarial samples that induce a wide range of harmful behaviors. This expanded behavioral coverage exposes more diverse failure modes and thereby improves safety alignment. To quantify this effect, we use semantic entropy as an output-level measure of adversarial behavioral diversity. Empirically, our method elicits diverse harmful behaviors in the target model, substantially mitigating behavioral narrowness and improving robustness to jailbreak attacks.
|
| 916 |
BARQ: Balanced Codebook Refinement for Low-Bit LLM Quantization
2610.04490
|
cs.LG
|
Chenhang Cui, Xu Xie, Linrui Xu, Xiaohao Liu, Xingyu Zhu |
As large language models (LLMs) grow in parameter count, model storage and parameter memory traffic have become major bottlenecks to efficient deployment. Codebook-based weight quantization reduces these costs, but imbalanced nearest-codeword assignments durin...As large language models (LLMs) grow in parameter count, model storage and parameter memory traffic have become major bottlenecks to efficient deployment. Codebook-based weight quantization reduces these costs, but imbalanced nearest-codeword assignments during fitting can leave some codewords insufficiently updated, limiting effective codebook utilization. To address this limitation, we propose Balanced Assignment Refinement for Quantization (BARQ), which improves quantization quality through balanced fitting of existing codebooks. Specifically, we first compute joint soft assignments between weight blocks and codewords through entropically regularized optimal transport with uniform marginals and curvature-weighted reconstruction costs, ensuring equal positive fitting mass for every codeword in the exact solution. We then refine the codewords through an assignment-weighted barycentric update, which we prove minimizes the fitting objective for fixed assignments. For finite Sinkhorn iterations, the implemented update retains this optimality provided all codeword masses exceed the denominator floor. Finally, we discard the soft assignments and use the refined codebook for standard hard nearest-codeword encoding, with our analysis establishing sufficient conditions for reducing hard-quantization distortion and evaluation loss. Across multiple LLMs, BARQ achieves lower perplexity and higher mean zero-shot accuracy than the evaluated baselines at comparable bit budgets. The code is available at https://github.com/chenhangcuisg-code/BARQ.
|
| 917 |
LoRA's Second Descent Extends Beyond Parameter Parity
2610.04507
|
cs.LG
|
Yueran Ma |
Double descent has sparked considerable interest, with recent work relating it to the data, the model and the learning configuration. Practical fine-tuning commonly involves training a small adapter on top of frozen pretrained weights, as in low-rank adaptatio...Double descent has sparked considerable interest, with recent work relating it to the data, the model and the learning configuration. Practical fine-tuning commonly involves training a small adapter on top of frozen pretrained weights, as in low-rank adaptation (LoRA). The adapter's rank is the hyperparameter that sets its capacity, yet how this rank relates to double descent has not been well explored. We quantify this relation under label noise on four vision backbones and a 7B language model with a module-matched rank sweep (MMRS), which extends past full rank and compares every rank with dense fine-tuning of the same modules, paired by seed. On DeiT-Tiny, risk is lowest at rank one and rises sharply as the adapter becomes able to fit the noisy labels, forming an interpolation cliff. Past the peak, risk falls again, but every tested post-peak rank that still saves parameters remains above dense risk. Rank-one LoRA outperforms dense fine-tuning on three of the four vision backbones, consistent with strong regularization at small rank. LoRA thus exhibits a second descent, but matches dense risk only after losing its parameter advantage, first on DeiT-Tiny at four times dense's projection weights. Code is available at https://anonymous.4open.science/r/lora-second-descent-C7B5.
|
| 918 |
Length Generalization Needs Proper Regularization
2610.04518
|
cs.LG
|
Pavlo Vasylenko, Matthias Lindemann, Andr\'e F. T. Martins, Marcos Treviso |
Length generalization is the ability of sequential models to perform well on context lengths unseen during training. In this work, we show that the challenge of achieving length generalization is related not only to architectural choices such as positional enc...Length generalization is the ability of sequential models to perform well on context lengths unseen during training. In this work, we show that the challenge of achieving length generalization is related not only to architectural choices such as positional encoding and the attention mechanism but also to the training procedure itself. We study how regularization affects length generalization and find that weight decay can hinder extrapolation. In contrast, dropout improves extrapolation when its placement within the architecture is reconsidered. In that regard, we show that the standard placement of dropout before layer normalization introduces a systematic distributional mismatch, and that applying dropout just before the linear projection resolves this issue. For example, a modified SmolLM3 with sliding window attention, continually pre-trained with dropout, can extrapolate perfectly to 64$\times$ on Needle-in-a-Haystack and far beyond the pre-training context size on RULER and HELMET. Mamba2 also benefits from dropout, suggesting an architecture-agnostic nature of the problem. We further propose Variance-Preserving Affine Dropout (VPAD), a new dropout strategy that substantially reduces the resulting pre-activation variance mismatch, leading to further extrapolation improvements in transformers.
|
| 919 |
Proximal Causal Learning under Unmeasured Confounding
2610.04519
|
cs.LG
|
Ying Tang, Yi Wang |
Estimating treatment effects from observational data typically relies on the No Unmeasured Confounding Assumption (NUCA), which rarely holds in practice. Proximal causal learning (PCL) addresses unmeasured confounding via proxy variables, yet existing methods ...Estimating treatment effects from observational data typically relies on the No Unmeasured Confounding Assumption (NUCA), which rarely holds in practice. Proximal causal learning (PCL) addresses unmeasured confounding via proxy variables, yet existing methods require the proxy variables to be pre-specified. Thus, we propose PCL-U, a framework that learns proxy variables directly from observed covariates. PCL-U uses neural encoders to decompose covariates into treatment-inducing, outcome-inducing, and shared proxies, guided by minimax mutual information objectives, and obtains causal estimates through a practical moment-based risk function. Experiments on benchmarks show that PCL-U matches or outperforms existing baselines. Besides, there are two types of synthetic datasets with varying dimensions and confounding strengths that illustrate that our method maintains stable estimation accuracy.
|
| 920 |
Weight Decay and Neuron Condensation: A Three-Stage Analysis of Two-Layer ReLU Networks
2610.04533
|
cs.LGcs.AI
|
Cheng Xu, Pengxiao Lin, Zhangchen Zhou, Zhi-Qin John Xu |
Weight decay is widely used as a regularization technique in neural network training, yet its role in neuron condensation (parameter direction alignment) remains unclear. Starting from a parameter initialization in the neural tangent kernel regime, we characte...Weight decay is widely used as a regularization technique in neural network training, yet its role in neuron condensation (parameter direction alignment) remains unclear. Starting from a parameter initialization in the neural tangent kernel regime, we characterize training dynamics under weight decay through three stages: rapid fitting, amplitude compression, and neuron condensation. Using a two-layer ReLU network, we analyze a residual correlation field that governs both neuron amplitudes and directions. During rapid fitting, the residual approaches a quasi-static equilibrium maintained by weight decay while the tangent kernel remains nearly unchanged. In the early stage of amplitude compression, kernel decay amplifies the residual correlation field, whose isolated attracting extrema provide candidates for condensation directions. As neuron amplitudes stabilize, we bound the drift of attracting extrema and demonstrate contraction of neuron directions around them, leading to neuron condensation. This staged analysis provides a dynamical understanding of how weight decay promotes a condensed representation, beyond reducing parameter norms.
|
| 921 |
PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory
2610.04537
|
cs.LG
|
Seoyoon Yum, Sehoon Kim |
On-device assistants run GPU-based LLM inference alongside CPU retrieval on unified-memory systems. Under a saturated local-retrieval workload, four concurrent retrieval workers raise 95th-percentile (p95) decode latency by 60-61% on two M4 systems, whereas pr...On-device assistants run GPU-based LLM inference alongside CPU retrieval on unified-memory systems. Under a saturated local-retrieval workload, four concurrent retrieval workers raise 95th-percentile (p95) decode latency by 60-61% on two M4 systems, whereas prefill latency rises by only 5.7-6.9%. We study LLM phase as an admission signal for independent CPU retrieval under controlled LLM workloads. PHASEGATE calibrates separate concurrency limits for prefill and decode, selecting four and one on our base-M4 configuration. Under a backlogged queue, it achieves 2.0 times the aggregate retrieval throughput of the best tested feasible fixed policy, with both p95 LLM latency metrics within 1.25 times their no-retrieval baselines in all seven held-out runs. A phase-blind control, TimeGate, uses the same two limits on a calibration-derived schedule without observing LLM phase. It achieves similar retrieval throughput but violates the output-token latency limit in every run. M2 and M2 Pro Mac minis reproduce the policy ordering, while output-length sweeps show that the advantage narrows as decode occupies more of each request.
|
| 922 |
TMAML: Temporal Model-Agnostic Meta-Learning for Cold-Start Time Series Forecasting
2610.04547
|
cs.LG
|
Wannes Janssens, Matthias Bogaert, Dirk Van den Poel |
Cold-start forecasting, the task of forecasting a time series with little to no historical data, is a common challenge. Addressing it requires approaches that learn quickly from few datapoints and leverage information from related series, typically through sta...Cold-start forecasting, the task of forecasting a time series with little to no historical data, is a common challenge. Addressing it requires approaches that learn quickly from few datapoints and leverage information from related series, typically through static covariates, to generalize well to unseen series. While some global forecasting models can generate cold-start predictions by leveraging information shared across multiple series, they are not optimized for out-of-train-set generalization or adaptation from short histories. In this work, we formulate cold-start forecasting as a few-window learning problem and introduce Temporal Model-Agnostic Meta-Learning (TMAML), which tailors the model-agnostic meta-learning algorithm, originally developed for few-shot adaptation of neural networks, to deep time series forecasting. TMAML constructs meta-tasks as temporally consistent support-query windows and pairs them with a temporal meta-training and meta-testing procedure, yielding forecasting models that are explicitly optimized for cold-start forecasting. We instantiate TMAML on the Temporal Fusion Transformer (TFT) and present an initial empirical analysis of forecast accuracy and calibration across three cold-start scenarios: TMAML consistently outperforms or matches a standard ERM-trained TFT, yields better-calibrated forecasts than naive on two of the three scenarios, but does not consistently outperform naive on probabilistic forecast accuracy.
|
| 923 |
Coupling Noisy Pairwise Knowledge to the DAG Posterior for Causal Discovery
2610.04559
|
cs.LGcs.AI
|
Guoliang Xu, James E Corter |
External causal reports can improve structure learning from limited observations, but their reliability varies across sources and variable pairs. We introduce HB-NoisyKG, a Bayesian framework that combines observational data with repeated causal reports from s...External causal reports can improve structure learning from limited observations, but their reliability varies across sources and variable pairs. We introduce HB-NoisyKG, a Bayesian framework that combines observational data with repeated causal reports from sources such as large language models. Each report is a noisy observation of a direct pair state implied by one DAG. A feature-conditioned Beta prior pools information about pair reliability, and a shared error matrix captures systematic mistakes. Alternating inference uses the graph posterior to refine reliability estimates, which determine how reports influence subsequent graph updates. The report likelihood uses only graph pair-state marginals, so the same observation layer supports discrete and continuous likelihoods in graph-only and joint inference. Against an 80-restart no-KG baseline, HB uses at most 80 total restarts and lowers mean SHD from 22.39 to 16.06 on five discrete benchmarks. On a physical light tunnel with random variable IDs and retained descriptions, HB lowers SHD from 39.00 for no-KG to 27.30. On continuous Sachs, graph-only BGe raises AUROC by 0.121 over no-KG Top-K. In a controlled synthetic study, continued updating also reduces mean reliability estimation error and held-out report log loss compared with one-time estimation.
|
| 924 |
Only Project Once: Projection-Adaptive Loss for Exact Constraint Satisfaction
2610.04572
|
cs.LGcs.AI
|
Tim Aebersold, Soheyl Massoudi, Mark Fuge |
Precise constraint satisfaction is a prerequisite to deploying learned models in many areas, motivating methods that repair raw neural predictions with a repair procedure. Current methods unroll multiple repair steps in training and softly penalize constraint ...Precise constraint satisfaction is a prerequisite to deploying learned models in many areas, motivating methods that repair raw neural predictions with a repair procedure. Current methods unroll multiple repair steps in training and softly penalize constraint violations that remain after the unroll. This is compute- and memory-intensive, lacks robustness when the repair fails to converge, and surrenders most of the constraint satisfaction work to the repair. Our central finding is that, contrary to common practice, a single detached projection step suffices in training. We accomplish this with a Projection-Adaptive Loss (PAL), which uses the constraint residual after this single step to adaptively weigh constraint penalties on the raw prediction. In experiments, PAL is the only method that retains virtually perfect feasibility on extremely nonlinear constraints, and matches or outperforms current methods on synthetic and engineering benchmarks. Because it only requires a single detached projection step, PAL trains 2.5x faster than the canonical repair-based method (DC3) on its own ACOPF benchmark. PAL can also be trained when constraints are expensive to evaluate (e.g., via neural surrogates), a setting where current unrolled methods are memory-intractable.
|
| 925 |
DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation
2610.04596
|
cs.LG
|
Karn Tiwari, Varnith Chordia, Prathosh A P |
On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objec...On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but suffers from sparse rewards and coarse credit assignment. We show that OPD and RLVR exhibit complementary blind spots: teacher signals provide dense local guidance but are weakly aligned with rollout correctness, whereas group-relative rewards capture task success but provide coarse token-level credit and vanish on all-failure groups. We introduce DiffGate, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance. Teacher supervision is applied only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher--student discrepancies from dominating optimization. The verifier therefore determines \emph{which trajectories} receive teacher guidance, while the teacher provides dense token-level update directions within those trajectories. Across Qwen3-0.6B and Qwen3-1.7B students, DiffGate improves code avg@8 over matched GRPO by $+1.7$ and $+1.8$ points and pass@8 by $+1.6$ and $+5.7$ points, respectively. On mathematics, avg@8 remains within $0.5$ points of GRPO while pass@8 improves by $+1.1$ and $+3.9$ points. Overall, DiffGate improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under our evaluation protocol.
|
| 926 |
Asymptotically Optimal Best Arm Identification with Fixed-Budget under Differential Privacy
2610.04600
|
cs.LG
|
Keqin Chen, Jie Bian, Yulian Wu, Vincent Y. F. Tan |
Best arm identification under differential privacy is a pure-exploration problem in which both statistical efficiency and privacy protection must be achieved simultaneously. We study fixed-budget best arm identification for bandits under pure $\epsilon$-differ...Best arm identification under differential privacy is a pure-exploration problem in which both statistical efficiency and privacy protection must be achieved simultaneously. We study fixed-budget best arm identification for bandits under pure $\epsilon$-differential privacy, where the learner must recommend an arm after a prescribed sampling budget while protecting the full transcript. We prove that the optimal exponential decay rate of the error probability is upper bounded by an instance-dependent privacy-aware transportation exponent that differs from the analogous quantity used to characterize the stopping time in fixed-confidence analysis by Jourdan and Azize [2025]. Guided by this exponent, we propose AO-Pri-BAI, an adaptive algorithm that maintains private running estimates through Laplace-tree mechanisms and learns a sampling design through a min--max interaction between hard alternatives and arm allocations. We prove that AO-Pri-BAI satisfies pure $\epsilon$-differential privacy. We also establish that the exponent of the failure probability of AO-Pri-BAI matches the privacy-aware benchmark. Numerical studies show that even in the non-asymptotic setting, AO-Pri-BAI outperforms benchmark algorithms on various instances, complementing the theoretical analyses.
|
| 927 |
Anticipating the Consequences of Curriculum Decisions with Large Language Models
2610.04604
|
cs.LG
|
Octavio Pappalardo, Nathan Herr, Tim Rockt\"aschel |
Automatic curriculum learning can improve the effectiveness of reinforcement learning by selecting the training experiences presented to the agent over time. Predicting the consequences of such decisions can, however, be difficult. We analyze automatic curricu...Automatic curriculum learning can improve the effectiveness of reinforcement learning by selecting the training experiences presented to the agent over time. Predicting the consequences of such decisions can, however, be difficult. We analyze automatic curriculum learning as a sequential decision-making problem, highlighting a gap between the quantities that determine the value of curriculum decisions and the information captured by local learning signals commonly used to guide them. We then investigate whether Large Language Models (LLMs) can exploit richer information about the learning problem to better anticipate the consequences of curriculum decisions. We introduce a method that combines online learning-progress estimates with LLM-informed estimates of (i) the potential downstream benefits of learning on each task and (ii) whether direct training on a task is currently likely to produce progress. We evaluate the approach on a custom benchmark of 256 textual goals in Craftax under different curriculum objectives. We observe the strongest gains when optimizing for individual target tasks. When optimizing across the full task set, the benefits vary across learners with different mechanisms for cross-task transfer, ranging from modest improvements in learning speed to larger gains that persist through the end of training.
|
| 928 |
Backward-Consistent Diffusion Sampling for Sparsely Observed PDE Inverse Problems
2610.04624
|
cs.LG
|
Yida Pan, Muhammad H. Ashiq, Chanyong Jung, Yixuan Jia, Jonah M. Miller |
Recovering Partial Differential Equation (PDE) coefficient fields from extremely sparse observations is a severely ill-posed inverse problem for which generative machine learning methods (e.g., diffusion models) have become a leading way to encode the prior. R...Recovering Partial Differential Equation (PDE) coefficient fields from extremely sparse observations is a severely ill-posed inverse problem for which generative machine learning methods (e.g., diffusion models) have become a leading way to encode the prior. Recent state-of-the-art diffusion solvers lift these priors to function spaces, finding a physics-consistent reconstruction in the output space of the diffusion denoiser. We prove that, in a discontinuous PDE setting, output space methods can result in failure to appropriately minimize the unobserved error with the correct coefficient field. Consequently, we propose Function space Backward-Consistent Sampling (FunBCS), an input space optimization approach for solving PDE problems which aims to find the best input such that the denoiser reconstruction is physics-consistent. We then prove that FunBCS appropriately minimizes the unobserved error, unlike output space optimization methods. Per our theoretical analysis, we also provide insights on how to dynamically allocate the number of input space optimization steps used throughout the sampling process. Our evaluations, across four PDE inverse problems (including the discontinuous Darcy flow), demonstrate that FunBCS reduces the reconstruction error by $27$-$64\%$ while running $1.4$-$2.1\times$ faster when compared to the current state-of-the-art.
|
| 929 |
The Numerical Linear Algebra of Large Language Models
2610.04631
|
cs.LG
|
Abdelkader Baggag, Yousef Saad |
Numerical Linear Algebra (NLA) has consistently played a vital role in advancing science by providing tools to solve fundamental problems encountered in scientific and engineering applications. Over the decades, it has continually evolved to meet the demands d...Numerical Linear Algebra (NLA) has consistently played a vital role in advancing science by providing tools to solve fundamental problems encountered in scientific and engineering applications. Over the decades, it has continually evolved to meet the demands driven by successive waves of scientific discovery. For instance, during the 1950s and 1960s, substantial efforts were devoted to developing methods for solving eigenvalue problems that emerged from the rapidly growing field of aerodynamics. This led to the discovery of the LR and QR algorithms. Later the attention turned to the solution of sparse linear systems that were common in applications like computational aerodynamics. Today we are experiencing yet another wave of major scientific advancement and NLA is once more at the heart of its development. This Machine Learning (ML) wave is proving to be utterly disruptive in science and engineering. Many tools in ML particularly Large Language Models (LLMs) are grounded in matrix and tensor methods. As we are approaching Artificial General Intelligence (AGI), it is clear that matrix methods will be called to play an even more significant role. For the numerical linear practitioner the speed of the current change makes it particularly challenging to adapt. This is a survey article that centers on machine learning techniques, with a particular focus on large language models. It has two main objectives. The first is to clarify the core concepts behind Large Language Models in a manner accessible to specialists in numerical methods. The second is to examine the key Numerical Linear Algebra concepts employed by LLM techniques, while also highlighting several significant recent contributions of NLA to the field.
|
| 930 |
Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression
2610.04638
|
cs.LGcs.AI
|
Nianli Peng, Geoffrey J. Gordon, Kiant\'e Brantley |
Actor-critic methods reuse past experience to improve sample efficiency. However, historical data are typically regarded as off-policy samples for the current policy-improvement update. This work introduces Regularized Dual Averaging Actor Critic (RDA2C), whic...Actor-critic methods reuse past experience to improve sample efficiency. However, historical data are typically regarded as off-policy samples for the current policy-improvement update. This work introduces Regularized Dual Averaging Actor Critic (RDA2C), which assigns a distinct role to replay. In regularized dual averaging, the subsequent policy is determined by accumulated policy-improvement feedback, so historical advantage estimates contribute directly to the actor objective rather than solely to the most recent update. RDA2C stores state-action samples with critic-estimated advantage labels, fits a dual score model $Z_\theta$ to the aggregated dataset, and derives the current policy from the accumulated score model using the entropy mirror map. In this way, replay defines an empirical dual objective from which the policy is computed. To analyze RDA2C, we establish a finite-time value-gap decomposition, separating the regularized dual-averaging term from errors due to stale-replay supervised fitting, critic bias, finite-buffer variance, and replay coverage, and stating the assumptions under which each error is bounded. RDA2C accepts advantage labels from any critic. With GAE labels, RDA2C outperforms PPO on six of eight MuJoCo tasks and eight of twelve Atari games. RDA2C also outperforms AAPDA, the closest dual-averaging baseline, on six of eight MuJoCo tasks. With twin-$Q$ labels, RDA2C matches SAC at matched batch size and update frequency.
|
| 931 |
Revisiting the Generalization of Neural Graph Edit Distance Models
2610.04644
|
cs.LG
|
Zhouyang Liu, Ning Liu, Yixin Chen, Jiezhong He, Dongsheng Li |
Neural approaches to Graph Edit Distance (GED) have achieved strong results under standard within-dataset evaluation, but much less is known about how well these models transfer across graph collections. We conduct a systematic study of this problem using exac...Neural approaches to Graph Edit Distance (GED) have achieved strong results under standard within-dataset evaluation, but much less is known about how well these models transfer across graph collections. We conduct a systematic study of this problem using exact GED supervision across diverse graph datasets and a broad set of representative learning-based methods. Our results reveal a pronounced gap between within-collection performance and cross-collection transfer. Models that perform well on their training collections often lose this advantage when evaluated on structurally different data. Training on multiple source collections substantially improves zero-shot transfer and provides a better starting point when limited supervision is available for a new target collection. Further analysis shows that transfer behavior varies with the source--target direction and the structural characteristics of the collections involved. These findings suggest that conventional within-collection evaluation provides only a partial view of the generalization behavior of neural GED models and motivate broader evaluation across heterogeneous graph collections.
|
| 932 |
Gated Target Propagation for Compositional Generalization in Continual Learning
2610.04649
|
cs.LG
|
Abdel Mfougouon Njupoun, Colin Bredenberg, Blake Aaron Richards, Guillaume Lajoie |
Continual learning is typically framed as acquiring new knowledge without catastrophically forgetting previous tasks. However, a flexible continual learner should also be able to reuse and recombine previously acquired knowledge to rapidly solve novel task com...Continual learning is typically framed as acquiring new knowledge without catastrophically forgetting previous tasks. However, a flexible continual learner should also be able to reuse and recombine previously acquired knowledge to rapidly solve novel task compositions. We introduce Gated Target Propagation (GaTaP), a continual learning algorithm in which task-specific gating variables---learned through a closed-form inner loop update---selectively suppress or enhance network modules. Network parameters are learned in a slower timescale outer loop, using the same local difference target propagation error signal as is used for adapting gating variables. We provide tractable experiments on class-incremental learning scenarios for both multilayer perceptron and convolutional network architectures. We show strong performance retention on previously learned tasks, as well as compositional generalization to unseen tasks, achieved through few-shot gain adaptation at inference. We analyze learned gating patterns and find that related tasks exhibit similar gating patterns, suggesting that inferred gates capture meaningful, reusable task structure. Overall, GaTaP provides a powerful framework for jointly ameliorating catastrophic forgetting and enabling few-shot compositional generalization in neural network models.
|
| 933 |
Pareto-Improving Adversarial Attacks with Primal-Dual Regularization
2610.04652
|
cs.LG
|
Yang Dai, Longfei Zhang, Wei Tao, Li Shen, Jincai Huang |
Transferable adversarial attacks are arguably the most practical black-box threat model. Under the same perturbation budget, stronger transfer attacks attain higher attack success rate (ASR), yet their imperceptibility also tends to degrade. Under such a fixed...Transferable adversarial attacks are arguably the most practical black-box threat model. Under the same perturbation budget, stronger transfer attacks attain higher attack success rate (ASR), yet their imperceptibility also tends to degrade. Under such a fixed-budget protocol, transferability and imperceptibility therefore appear to trade off against each other. We argue that this conflict is an artifact of fixed-budget evaluation, not an intrinsic trade-off. When attacks are compared on the ASR--imperceptibility Pareto frontier obtained by sweeping $\epsilon$, stronger transfer attacks already attain better imperceptibility at matched ASR than weaker ones. To exploit this latent advantage, we introduce the stealthy transfer attack ST, a plug-in primal-dual wrapper that adds an $L_\infty$ saturation regularizer to the standard constrained objective and resolves it through a two-step primal-dual update: a projected primal step on the perturbation coupled with an $L_1$-ball projection on a dual variable that absorbs the regularizer through Fenchel duality, requiring no auxiliary models or handcrafted perceptual priors. Empirically, ST extends the Pareto frontier across different base attacks and additional surrogate architectures. At $\epsilon{=}16/255$, average imperceptibility gains over each base attack are $17\%$ on LPIPS and $14\%$ on NIQE while ASR is preserved or improved. At matched high-ASR levels, the strongest ST variants further Pareto-dominate dedicated stealth-oriented transfer attacks, confirming that the latent imperceptibility advantage of strong transfer attacks can be unlocked by a primal-dual optimization wrapper without sacrificing transferability. Code will be made available at \url{https://github.com/AndssY/ST}.
|
| 934 |
Path Laplacian Encodings for Directed Graphs
2610.04657
|
cs.LG
|
Lydia Mezrag, Semih Cant\"urk, Michael Perlmutter, Bastian Rieck, Guy Wolf |
Directed graphs naturally model many real-world systems in which interactions are asymmetric, such as citation networks, web graphs, and information-flow networks. However, graph learning methods commonly rely on message passing with symmetrized graph represen...Directed graphs naturally model many real-world systems in which interactions are asymmetric, such as citation networks, web graphs, and information-flow networks. However, graph learning methods commonly rely on message passing with symmetrized graph representations or positional encodings that only partially exploit edge directionality. We introduce PathLapPE, a novel spectral positional encoding (PE) derived from the path Laplacian on directed graphs. PathLapPE provides node- and edge-level features that encode directional higher-order structure and can be incorporated into standard graph learning architectures. Empirical results on node- and graph-level benchmark tasks show that PathLapPE yields consistent improvements across several architectures, especially when combined with direction-aware message passing. Compared with magnetic Laplacian positional encodings, a widely studied spectral positional encoding for directed graphs, PathLapPE does not require additional fine-tuning of directionality hyperparameters while offering competitive runtime and performance.
|
| 935 |
Low-Fidelity FDM Spectral Guidance for Neural Eigenvalue Solvers
2610.04695
|
cs.LG
|
Aryan Chaudhary, Manikandan Padmanaban, Jagabondhu Hazra |
Operator eigenvalue problems appear throughout science. Classical methods usually discretize the operator into a matrix and then solve the resulting matrix eigenvalue problem. This works well in low dimensions, but fine grids quickly become expensive in both m...Operator eigenvalue problems appear throughout science. Classical methods usually discretize the operator into a matrix and then solve the resulting matrix eigenvalue problem. This works well in low dimensions, but fine grids quickly become expensive in both memory and computation as the dimension grows. Neural network based solvers avoid storing these large grids, but recent state of the art neural methods can require hundreds of thousands of training steps and may struggle to find the desired eigenvalues. We show that the two approaches can help each other. A coarse finite difference method (FDM) calculation acts as a cheap numerical model of the operator spectrum. We use the approximate eigenvalues as fixed shifts during the training of the neural solver, as they only need to locate the relevant part of the spectrum. We also introduce Stabilized Inverse Power Method Neural Network (SIPMNN), a more stable training procedure for higher-dimensional problems. Across five test problems at $d=10$, the combined approach is more accurate overall than the tested fully neural alternatives while using eight to ten times fewer iterations.
|
| 936 |
Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning
2610.04696
|
cs.LG
|
Zeyang Li, Yunan Wang, Risheek Garrepalli, Mohammad Ghavamzadeh, Navid Azizan |
Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unno...Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct samples from it. Many existing methods rely on importance sampling to construct training signals, which can suffer from high variance, increasing computational cost and destabilizing training. We propose Score-Calibrated Flow (SCF), a simple and efficient algorithm for training generative models to sample from unnormalized densities without importance sampling or backpropagation through the sampling trajectory. We learn the desired flow by enforcing self-consistency, bypassing target posterior mean estimation. By jointly exploiting the prescribed target score and the structure of flow matching, we establish these self-consistency requirements as score-calibrated optimality conditions, first for the terminal density and then for the trainable velocity field. We prove that their unique solutions are, respectively, the target density and the ideal flow model that conditional flow matching (CFM) would recover if target samples were available. We formulate the velocity condition as a fixed-point equation and exploit its conditional-expectation structure to construct a stop-gradient objective for enforcing it. The resulting training procedure retains the scalable sample-interpolate-regress structure of CFM despite the absence of target samples, using endpoints generated by the current flow. For online RL, the critic gradient supplies the target score at the generated actions, yielding a direct approach to actor training. Experiments on RL benchmarks demonstrate that SCF matches or improves upon state-of-the-art generative-policy baselines, while substantially reducing training time.
|
| 937 |
Localized Operator Learning with Adaptive Partition-of-Unity Mixture-of-Expert Networks
2610.04708
|
cs.LG
|
Madison Cooley, Ramansh Sharma, Shandian Zhe, Robert M. Kirby, Varun Shankar |
Operator learning methods such as DeepONets and FNOs often struggle with PDE families featuring sharp interfaces, heterogeneous coefficients, and localized multiscale structures. We introduce a partition-of-unity (POU) mixture-of-experts framework for localize...Operator learning methods such as DeepONets and FNOs often struggle with PDE families featuring sharp interfaces, heterogeneous coefficients, and localized multiscale structures. We introduce a partition-of-unity (POU) mixture-of-experts framework for localized operator learning, in which geometry-aware gating networks produce smooth spatial partitions which blend the contributions of local expert networks. Our main contribution is HiRefPOU, a residual-style hierarchical POU architecture for DeepONets that organizes localized representations through nested parent-child partitions while preserving global continuity. We also show that the same POU principle can be incorporated into Fourier Neural Operators to introduce spatial adaptivity without modifying the underlying spectral layers. On heterogeneous Darcy and reaction-diffusion benchmarks, HiRefPOU achieves substantially lower error than global DeepONet and static POU-MoE baselines, while the broader operator-learning experiments show that the benefits of localization depend on the PDE structure and the chosen neural-operator backbone. The learned partitions are interpretable and align with interfaces and regions of rapid solution variation. These results show that explicit geometric localization can improve both accuracy and interpretability in neural operator learning.
|
| 938 |
FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies
2610.04754
|
cs.LG
|
Sara Taheri, Deep Kumar Ganguly, Jan K\v{r}et\'insk\'y, Majid Zamani |
Even a few action perturbations can substantially degrade the performance of a deployed decision policy. Certifying the resulting return loss is challenging in stochastic environments, where returns vary even without an attack. We introduce FoSeRL, a framework...Even a few action perturbations can substantially degrade the performance of a deployed decision policy. Certifying the resulting return loss is challenging in stochastic environments, where returns vary even without an attack. We introduce FoSeRL, a framework for certifying deployed RL policies against precommitted, temporally sparse action attacks. The deployed policy is unchanged, with no smoothing or retraining. Certification requires a resettable simulator supporting shared randomness and independent one-step successor queries, but no analytical dynamics model. FoSeRL certifies that an attacked episode loses no more than a prescribed amount of return relative to the same episode unattacked, with at least a target probability and at a user-specified confidence level. Both runs share the initial state and randomness, so the measured loss reflects the attack, not the episode; carrying the running return gap as a state coordinate makes it the terminal value, reducing trajectory-level certification to terminal safety. Time-dependent barrier conditions on the augmented state bound the terminal failure probability: satisfied exactly, they certify every admissible precommitted attack; learned from sampled trajectories and verified on held-out data, they certify the same guarantee under a specified attack-episode setting. Across six stochastic continuous-control environments and three RL policy families (TD3, SAC, and PPO), FoSeRL certifies non-trivial cardinality--magnitude robustness frontiers, achieves substantially larger certified budgets than policy smoothing, and reveals marked robustness differences among policies with comparable nominal performance.
|
| 939 |
EasyClassifier: Honest, Reproducible Machine-Learning Classification for Researchers Who Do Not Program
2610.04758
|
cs.LGcs.AI
|
Ahmad B. A. Hassanat, Ghada A. Altarawneh |
Machine-learning classification is now used across medicine, the social sciences, economics, education and engineering, very often by researchers who do not program. Two errors recur in such work: preprocessing learned on all rows before cross-validation (data...Machine-learning classification is now used across medicine, the social sciences, economics, education and engineering, very often by researchers who do not program. Two errors recur in such work: preprocessing learned on all rows before cross-validation (data leakage), and reporting the cross-validation score of the classifier that won a comparison (selection bias). Both make published scores optimistic. We present EasyClassifier, an open-source Python package that guides a user from a CSV or Excel file to a finished analysis through a sequence of plain-language, multiple-choice questions, with no code. It makes the correct procedure the default rather than an option: every preprocessing step is learned inside each training fold; the selected classifier receives a separate final score from nested cross-validation (up to 2,000 rows), or an untouched 20% test set; classifiers run with fixed default parameters and fixed random seeds; and every run writes a report with a ready-to-adapt Methods paragraph, the references to cite, and figures prepared for publication. On ten public datasets from seven fields and a random-label control, split in half so that one half served as an untouched external test, the common practice overstated balanced accuracy by 2.9 percentage points on average (up to 9.9 on small data), whereas the score EasyClassifier reports had a mean signed bias of +0.7 points and a mean absolute error of 2.3 points (one-sided Wilcoxon p = 3.2 x 10^{-4}, 30 runs). On random labels it reported 48.8% balanced accuracy against a chance level of 50%. EasyClassifier is available under the MIT license from PyPI (pip install easyclassifier), GitHub, and Zenodo.
|
| 940 |
Variance-Optimal Control Variates for Learning with Black-box Feedback
2610.04766
|
cs.LG
|
Zihao Zhao, Shuhan Zhang, Kai Wang |
Modern models increasingly learn through black-box oracles such as humans, optimization solvers, and external tools that provide feedback without exposing their internal mechanisms. A common remedy is to learn an (action-)value function as a control variate. I...Modern models increasingly learn through black-box oracles such as humans, optimization solvers, and external tools that provide feedback without exposing their internal mechanisms. A common remedy is to learn an (action-)value function as a control variate. In this paper, we first observe that even an exact action-value function can be arbitrarily far from variance-optimal. We show that this gap arises because the value function minimizes the noise in each action's own gradient term, while an action can still affect the rest of the gradient estimator through shared parameters. A simple unbiased correction, at no extra oracle cost, can still reduce its variance by an arbitrarily large factor. Motivated by this, we then prove that the residual variance can be decomposed exactly by actions with no cross terms. This decomposition yields a closed-form variance-minimizing correction for neural-network parameters, which can be computed by a simple projection. Empirically, our correction consistently reduces the variance left by the value function and improves learning across all tasks. The source code for all experiments is available at https://github.com/Zihao-Kevin/black_box_opt.
|
| 941 |
Repeated-Measure Leakage, Distribution Shift, and Reliability under Partial Observation in Patient World Models
2610.04778
|
cs.LG
|
Arjun Subramanian |
Patient world models are increasingly proposed for longitudinal prediction, intervention-aware reasoning, and clinical-trial simulation. Causal or clinical intervention validity is distinct from predictive generalization and reliability; before making stronger...Patient world models are increasingly proposed for longitudinal prediction, intervention-aware reasoning, and clinical-trial simulation. Causal or clinical intervention validity is distinct from predictive generalization and reliability; before making stronger claims, the underlying predictive state should generalize across patients, survive realistic shifts and missing observations, and expose failure through meaningful reliability signals. We evaluate these prerequisites in a deliberately narrow setting: short-horizon digital-biomarker forecasting from PhysioNet GaitPDB, comprising 165 participants, 306 recordings, and 51,129 context-future pairs. Using persistence, ridge, MLP, GRU, Transformer, and a compact JEPA-style predictor, we build an evaluation ladder that progressively removes raw temporal overlap, same-recording familiarity, and same-patient familiarity before testing unseen-patient generalization. For GRU, NMSE rises from 0.1227 under random-window splitting to 0.1393 after eliminating raw train-test overlap and to 0.1961 under patient holdout. Among 54 participants with repeated recordings, exposure to a different recording from the same patient improves GRU NMSE from 0.2177 to 0.1556, while a recording-excluded identity hypothesis is not supported at the participant level. Under participant-held-out evaluation, MLP and Transformer are statistically indistinguishable. Study shift, a four-times-longer prediction gap, and partial observation further degrade performance; under 50% temporal masking, Transformer NMSE rises to 0.611 while MC-dropout predictive variance falls. We do not claim a longitudinal or intervention-aware simulator. Instead, the results support a prerequisite evaluation stack of patient separation, repeated-measure controls, shift, missingness, and uncertainty validation before stronger patient-world-model claims are trusted.
|
| 942 |
DASH: Fast, Valid Counterfactuals for Deep Networks via Batched Directional Search
2610.04783
|
cs.LG
|
Shraman Pal, Gabriel Medeiros, Clayton Escouper das Chagas, Can Li |
Counterfactual explanations are most useful when they can be generated with low latency, remain close to the factual input, and satisfy input-domain, categorical, and actionability constraints. Achieving these objectives simultaneously is challenging for deep ...Counterfactual explanations are most useful when they can be generated with low latency, remain close to the factual input, and satisfy input-domain, categorical, and actionability constraints. Achieving these objectives simultaneously is challenging for deep neural networks. Heuristic methods are often fast but may return invalid counterfactuals, whereas exact methods can certify global proximity but may not finish within practical time limits. We introduce DASH, a batched heuristic search method for finding close, valid, and actionable counterfactuals for deep neural networks under $\ell_1$, $\ell_2$, and $\ell_\infty$ objectives. DASH uses directional Lipschitz bounds and local affine models to generate anchors, then ranks and expands promising regions with batched network evaluations. We compare DASH against nine prior heuristic methods and time-limited exact mixed-integer baselines on four tabular datasets, with network depths from 2 to 32, and evaluate scalability on PBMC3k. Across 9,000 tabular query-norm cases, DASH returns a valid counterfactual within $5\%$ of the best heuristic-observed valid distance in $94.6\%$ of cases, with a median CPU search runtime of $0.061$ s. PGD-bisect, the baseline with the highest pooled within-$5\%$ coverage, meets this criterion in $41.1\%$ of cases, with a median runtime of $0.298$ s. These results show that the proposed search maintains high valid proximity across norms while keeping its search runtime practical.
|
| 943 |
TIMBRE: Teaching Time Series Forecasters to Read, Remember, and Reconcile
2610.04795
|
cs.LG
|
Xinyu Guan, Zhirong Zhang, Hongyuan Liu, Pengcheng Xu, Yu Sun |
Event-informed forecasting requires translating reports and historical responses into changes to a numerical forecast. We propose TIMBRE (Temporal Integration of Memory-Based Responses and Evidence), which combines source-aware representation, state-conditione...Event-informed forecasting requires translating reports and historical responses into changes to a numerical forecast. We propose TIMBRE (Temporal Integration of Memory-Based Responses and Evidence), which combines source-aware representation, state-conditioned response transfer, and reliability-guided fusion before a frozen forecast head. A separate readout adjusts interval widths while preserving the median. In a single-seed, one-epoch development study of 13 tasks, TIMBRE improves MAE over ordinary fusion on eight tasks but over native Chronos-2 on only two. Disabling response transfer in the trained model reduces BTC and AULF MAE by 47.04% and 6.81%, respectively. These findings identify sensitivity to learned response transfer rather than a general forecasting advantage. Missing development-set scores and the absence of retrained ablations limit attribution to individual evidence mechanisms.
|
| 944 |
Multi-Agent Spectrum Sharing
2610.04802
|
cs.LG
|
Job Elliott, Graduate Student Member, IEEE, Justin G. Metcalf, Golnaz Habibi |
This project explores how multiple cognitive radars can learn to share limited wireless spectrum with other radio users without interfering with one another. Using machine learning (ML), each device independently decides where and how widely to transmit within...This project explores how multiple cognitive radars can learn to share limited wireless spectrum with other radio users without interfering with one another. Using machine learning (ML), each device independently decides where and how widely to transmit within a fixed 100 MHz band. The system analyzes real or simulated signal activity to detect which parts of the spectrum are currently in use and which are open. Based on this information, the devices adapt their transmission choices to avoid crowded frequencies while making efficient use of available space. The goal is to develop a flexible, scalable approach to spectrum sharing that could support future wireless communication systems. Experimental results using both over-the-air software-defined radio (SDR) recordings and simulated environments demonstrate that the proposed meta-learning approach consistently balances competing objectives better than conventional reinforcement learning (RL) methods in multi-agent spectrum-sharing scenarios. Across five multi-agent benchmark environments, our proposed method achieved the highest average reward among the primary baseline algorithms while simultaneously maintaining low collision rates and stable transmission behavior.
|
| 945 |
Separating Decision Time from Decision Quality in the Real-Time Gap of Distilled Deciders: Evidence from a Game and a Conveyor Simulator
2610.04810
|
cs.LGcs.AI
|
Chihoon Shin, Junyeong Lee, Kihyeok Jeong, Wonok Kwon |
Real-time agents are often evaluated with the world paused, or with decision time rounded to whole ticks (tick conversion). Tick conversion predicts almost no loss for a learned decider whose hit accuracy falls by 14.8 percentage points (pp) on the same games ...Real-time agents are often evaluated with the world paused, or with decision time rounded to whole ticks (tick conversion). Tick conversion predicts almost no loss for a learned decider whose hit accuracy falls by 14.8 percentage points (pp) on the same games in asynchronous play. We instead split the decider's gap to a zero-latency teacher into time and quality components. In ViZDoom, we train a small decider by imitating a scripted teacher, run the game on wall-clock time, and add a control in which the teacher waits for a call to the decider's server before deciding. On a Windows host, three preregistered studies put the time component at 6.5-8.7 pp and the quality component at 6.3-8.6 pp. On a Linux host that skips 0.08-0.09% of ticks, the time component disappeared (-0.1 pp; paired 95% bootstrap interval [-0.4, 0.0]) while the quality component remained (5.6 pp [3.7, 7.5]). Retraining on teacher-labelled states from the decider's own play (DAgger) improved it on new games on both hosts (2.6 pp [0.4, 4.9] and 3.1 pp [1.1, 5.2]), a preregistered partial success. Imposing one delay schedule on every arm in live play left a quality component of 5.8 pp [4.0, 7.5] on 100 new games, and retraining cut the decider's disagreement with the teacher on its own states from 22.7% to 14.9%. In a conveyor simulator with a 400 ms deadline, delay past it erased the quality component. A latency-matched control whose delays match the decider's shows which remedy to try.
|
| 946 |
GRAM: Correcting Frozen Time-Series Foundation Models via Graph-Retrieved Amplitude Memory
2610.04827
|
cs.LG
|
Xiaoyun Yu, Xiangfei Qiu, Yonggui Huang, Shixiang Tang, Nanqing Dong |
Time-series foundation models (TSFMs) enable zero-shot forecasting through large-scale cross-domain pretraining, while retrieval augmentation further improves their performance by leveraging historical information. However, existing methods typically correct T...Time-series foundation models (TSFMs) enable zero-shot forecasting through large-scale cross-domain pretraining, while retrieval augmentation further improves their performance by leveraging historical information. However, existing methods typically correct TSFM forecasts using the ground-truth futures of similar historical windows, which contain both predictive components already captured by the foundation model and sample-specific random fluctuation that is difficult to transfer. In contrast, recurring systematic model bias within prediction errors more directly characterizes the failure modes of a frozen TSFM and therefore provides more valuable correction signals. Effectively exploiting such model bias, however, poses two challenges: prediction errors at different numerical levels are difficult to compare due to scale differences, and the recurring bias must be extracted from prediction errors contaminated by random fluctuation. To address these challenges, we propose GRAM, a general retrieval-augmented framework for frozen TSFMs. GRAM first introduces an Amplitude Memory Module (AMM) that scales prediction errors by amplitude and aggregates them into retrievable prototypes. It then employs a Prototype Graph Module (PGM) to model relations among prototypes to aggregate consistent bias information while suppressing random fluctuation. During online forecasting, GRAM retrieves and expands prototypes relevant to the current query and generates per-horizon corrections to refine the original TSFM forecast. Experiments across multiple datasets and foundation models demonstrate consistent forecasting improvements.
|
| 947 |
Which Preferences to Train On? End-to-End Multi-Objective Alignment with an Adversarial Preference Distribution
2610.04845
|
cs.LG
|
Minjae Lee, Kyunghyun Cho, Sangdon Park |
Aligning large language models (LLMs) with human values is important for safe, efficient, and beneficial AI deployment. However, human values are multifaceted: helpfulness, harmlessness and humor trade off against one another, and different users want differen...Aligning large language models (LLMs) with human values is important for safe, efficient, and beneficial AI deployment. However, human values are multifaceted: helpfulness, harmlessness and humor trade off against one another, and different users want different trade-offs. Multi-objective alignment (MOA) addresses this by training a policy that can provide any point of the Pareto front, but existing methods either train one model per preference, interpolate a few separately aligned experts post hoc, or train a single conditioned model without considering which preferences it should be trained on. Since the hard regions of the preference simplex depend on the objectives at hand, existing methods leave them under-trained and do not get the most out of a single model. Therefore, we propose MAESTRO (Multi-objective Alignment via End-to-end STeering and Robust Optimization), which formulates MOA as a minimax problem over preference distributions and trains a single prompt-conditioned policy end-to-end with RL against an adversarial preference distribution: a Dirichlet distribution updated by online mirror descent toward the preferences the current policy serves worst, rather than on a fixed one. On HH-RLHF, BeaverTails and a summarization task, with up to three objectives, MAESTRO attains the best Pareto front on most tasks in a single training run, at the lowest training cost among the compared methods. The largest margins appear in the hard regions that a fixed preference distribution leaves under-trained, confirming that a single prompt-conditioned model is capable of covering the objective trade-offs on its own.
|
| 948 |
PIT-GCL: Protein Interaction using Topological Graph Contrastive Learning
2610.04850
|
cs.LGcs.AI
|
Jae Won Choi, Ryoonki Hong, Alan Liang, Manjula Adiveppa Wader, Bingsong Zeng |
Protein binding prediction is central to target identification, therapeutic binder design, and large scale screening, yet remains challenging because binding depends on sequence, three dimensional geometry, and global structural organization. Recent folding mo...Protein binding prediction is central to target identification, therapeutic binder design, and large scale screening, yet remains challenging because binding depends on sequence, three dimensional geometry, and global structural organization. Recent folding models such as AlphaFold3 and Boltz-2 have substantially improved structure prediction, but their confidence outputs (pLDDT, pTM, ipTM) are not specifically designed for binary binding prediction, and dedicated structure aware predictors often require bound complex structures that are unavailable at screening scale. We introduce PIT-GCL, a dual tower structure aware framework that encodes each protein independently from its amino acid sequence, C{\alpha} point cloud, and a global persistent homology descriptor. Each tower combines residue ESM-2 embeddings with a topological summary computed from the H0 and H1 persistence landscapes of a Vietoris-Rips filtration, and processes the resulting tokens with a structure aware Transformer in which pairwise C{\alpha} distances enter as a learned attention bias. A bidirectional cross attention module then performs latent space soft docking between the two per-protein representations, and the model is trained with a combined binary cross entropy and NT-Xent contrastive objective. On three binary interaction prediction benchmarks, general PPI on PPIRef, TCRpMHC binding on STAG, and whole chain pairs on PPB-Affinity, PIT-GCL outperforms representative sequence based, structure aware, and task specific baselines on general PPI under our evaluation, and is the only method above chance on PPB-Affinity; on TCR-pMHC it leads at a fixed decision threshold but is outranked by a task specific sequence model. Because each protein is encoded independently in the first phase, its representation can be precomputed and reused across candidate pairs, which is convenient for large scale screening.
|
| 949 |
Greedy Local Learning for Language Model Pretraining: Gaps and Objective Design
2610.04867
|
cs.LGcs.AI
|
Jihwan Moon, Sheir A. Zaheer, Jinmyoung Lee, Gunhee Kim, Chan Y. Park |
Greedy block-wise local learning splits a network into gradient-isolated blocks trained by local auxiliary losses, deleting the backward pass between blocks: inter-stage communication becomes forward-only and every block can step its optimizer independently, p...Greedy block-wise local learning splits a network into gradient-isolated blocks trained by local auxiliary losses, deleting the backward pass between blocks: inter-stage communication becomes forward-only and every block can step its optimizer independently, properties directly relevant to decentralized model-parallel training. Local learning is competitive with end-to-end backpropagation on image classification, and on small Transformers it is known to trade a worse best loss for parallel speedup. How this loss gap behaves in autoregressive language model (LM) pretraining at larger scale, and which auxiliary designs reduce it, has not been measured. We present a token-budget-matched empirical study at 125M and 400M parameters with $K \in \{1,2,4\}$ blocks at Chinchilla-optimal budgets, factorizing the auxiliary design into network architecture and training objective. We observe: (i) the gap to end-to-end training more than doubles from $K=2$ to $K=4$, but at $K=4$ shrinks from 125M to 400M; (ii) replacing an MLP auxiliary with a Transformer-based one is a strong network-side intervention, recovering 22-41% of the gap; (iii) a multi-token-prediction (MTP) auxiliary objective helps at the first block boundary, whereas adding it at deeper boundaries hurts, and restricting it to the first block yields the best $K=4$ configuration ($+0.062$ vs. $+0.075$ nats at 400M); and (iv) deployment-style per-block execution reduces activation memory by up to $2.2\times$. We frame these results as an empirically grounded method direction rather than a finalized method: local objectives should apply future-predictive pressure selectively across boundaries while resisting shortcuts that bypass predictive content.
|
| 950 |
Your Temporal Link Predictor Is Blind to Who Is Active: A Missing Factor That Transfers Across Models
2610.04869
|
cs.LG
|
Ji Zhang, Zixin Liu, Yiran Ding, Jiayi Wang, Yilu Du |
An interaction has two parts: someone decides to act, and then chooses whom to act on. Temporal link prediction has concentrated on the second, and we show that it is blind to the first by construction: a standard negative keeps the real source and swaps the d...An interaction has two parts: someone decides to act, and then chooses whom to act on. Temporal link prediction has concentrated on the second, and we show that it is blind to the first by construction: a standard negative keeps the real source and swaps the destination, and we prove that this cancels the source's activity exactly from the optimal score, so no model trained and evaluated this way is ever rewarded for learning it. Under the harder historical and inductive negatives, whose sources differ, the same factor becomes the dominant signal. We model it with Source Node Activity Modeling (SNAM), a self-exciting event intensity fitted by an exact point-process likelihood to decayed interaction counts the history states already contain; it has fewer than 20 parameters. On their own, never looking at the destination, these parameters beat DyGFormer and TPNet on four of five datasets under historical negatives. Added to the frozen scores of TPNet, TGN, DyGFormer and DSRD, four models of different design, without retraining anything, they raise AP on almost every backbone-dataset pair in both settings, by up to 25 points. Our full model ranks first overall against eleven baselines on 13 datasets and three protocols, and on million-event streams trains an epoch 9-100x faster than TPNet and DyGFormer. We conclude that source activity is a blind spot of temporal link prediction, and a cheap, transferable one to close. Code is available at https://github.com/Erutaner/Your-Temporal-Link-Predictor-Is-Blind-to-Who-Is-Active.
|
| 951 |
ResOPD: Tail Residualization for Sparse On-Policy Distillation
2610.04882
|
cs.LGcs.AI
|
Penghui Yang, Long Xing, Xuanlang Dai, Ziyu Liu, Kai Chen |
On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse t...On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse teacher interfaces, typically transmitting either the sampled-token score or a small Top-$k$ distribution. However, this sparse setting faces a fundamental dilemma: sampled-token estimators are unbiased but suffer from severe gradient variance, whereas directly optimizing Top-$k$ objectives introduces systematic bias. To improve this trade-off, we propose ResOPD (On-Policy Distillation with Tail Residualization), which provides unbiased full-vocabulary reverse KL gradient estimation under on-policy sampling, with substantial variance reduction under the same sparse payload in the evaluated settings. ResOPD aggregates the unobserved vocabulary into an observable coarse tail event, computes its exact aggregate gradient, and samples only the fine-grained within-tail residual, which requires no additional teacher queries or forward passes. Extensive experiments demonstrate that ResOPD substantially reduces gradient variance, stabilizes online training dynamics, and improves downstream performance across the evaluated settings. These results establish ResOPD as an efficient, plug-and-play variance reduction primitive for sparse on-policy distillation.
|
| 952 |
SFlexRCA: Lightweight, Scalable, and Flexible Root Cause Analysis for IIoT Edge Systems
2610.04893
|
cs.LGcs.AI
|
Amr M. Zaki, Farhoud Jafari, Honggeun Ji, Komal Sarda, Marin Litoiu |
Industrial Internet of Things (IIoT) systems generate high-dimensional sensor telemetry from interconnected components, where faults can propagate across the system. To address these challenges, we propose SFlexRCA (Scalable and Flexible Root Cause Analysis), ...Industrial Internet of Things (IIoT) systems generate high-dimensional sensor telemetry from interconnected components, where faults can propagate across the system. To address these challenges, we propose SFlexRCA (Scalable and Flexible Root Cause Analysis), a topology-free RCA framework designed for resource-constrained IIoT environments. SFlexRCA transforms multivariate telemetry into compact orthogonal representations and applies shared lightweight linear modeling, avoiding explicit graph construction, message passing, and per-variable or lag-specific parameter growth. We evaluate SFlexRCA on three publicly available IIoT datasets, BATADAL, SWaT, and WADI, spanning different numbers of monitored variables, temporal characteristics, and training-data regimes. SFlexRCA is compared with 10 statistical, causal, and non-causal baselines in terms of RCA accuracy, training efficiency, inference latency, and memory consumption. In addition, inference efficiency and memory consumption are evaluated on Raspberry Pi 3 and Raspberry Pi 5, while energy consumption is additionally measured on Raspberry Pi 5. We further investigate temporalcontext sensitivity, architectural and loss components, and alternative representations. Notably, SFlexRCA maintains strong localization performance on BATADAL despite its limited normal-operation training data, while its compact shared architecture avoids the parameter growth associated with causal and graph-based approaches. Its lightweight shared architecture further enables efficient deployment on resource-constrained IIoT edge devices. The SFlexRCA code is available at https://github.com/theamrzaki/RootCause- Analysis-Correlation-Attentive-Modeling.
|
| 953 |
ScopeSAE: Model-Scope Feature Discovery with Interpretable Layer Selection
2610.04905
|
cs.LGcs.AI
|
Qingwen Zeng, Zehao Fu, Shuyu Meng, Linghan Huang, Jiayi Zhang |
Sparse autoencoders (SAEs) are a central tool in mechanistic interpretability. However, existing SAEs are primarily trained per layer. The modeling subspace is therefore fixed by layer identity, independent of which token-layer states actually drive each predi...Sparse autoencoders (SAEs) are a central tool in mechanistic interpretability. However, existing SAEs are primarily trained per layer. The modeling subspace is therefore fixed by layer identity, independent of which token-layer states actually drive each prediction. We argue that this constraint contributes to several limitations observed in layer-wise SAEs, including low feature utilization, high dictionary redundancy, and features that lack direct behavioral grounding. In this paper, we propose ScopeSAE, which selects the modeling subspace per token by attributing each prediction to its most influential token-layer state via normalized gradient-based attribution, and learns features over the resulting prediction-relevant subspace. Empirically, ScopeSAE yields an effect we term reconstruction-better-than-original. Written-back reconstructions of the SAE produce lower next-token cross-entropy than the original activations, an outcome that, to our knowledge, has not previously been reported for SAEs. Through interventional analyses and a KL fine-tuning counter-experiment, we show that this effect is attributable to ScopeSAE's prediction-relevant subspace itself rather than to architectural changes. ScopeSAE further improves effective feature count, interpretability, utilization, and dictionary redundancy over existing layer-based baselines, suggesting that choosing the SAE modeling subspace by predictive relevance leads to more useful and behaviorally meaningful features.
|
| 954 |
Why Subliminal Learning Needs So Much Data: A Noisy Inverse View through Steering Vector Recovery
2610.04907
|
cs.LGcs.AI
|
Luoyu Chen, Xiaoyu Ding, Weiqi Wang, Chenhan Zhang, Zhiyi Tian |
Subliminal learning lets a student inherit a teacher's behavioral trait from semantically unrelated data, yet published demonstrations typically require tens of thousands of carrier examples. We ask where this data requirement comes from. Our testbed is sublim...Subliminal learning lets a student inherit a teacher's behavioral trait from semantically unrelated data, yet published demonstrations typically require tens of thousands of carrier examples. We ask where this data requirement comes from. Our testbed is subliminal steering: the teacher trait is a known residual-stream vector $\Delta_T$, so transfer can be measured directly as parameter recovery. On identical carrier prefixes, we compare token-level (hard) NLL supervision with full-distribution (soft) KL supervision. At initialization the two objectives give nearly collinear gradients, and both align poorly with $\Delta_T$. Under iterative optimization, however, they diverge: soft supervision recovers $\Delta_T$ almost exactly from a few hundred carriers, while hard supervision stays well below it even with tens of thousands. We explain this gap by casting steering-vector distillation as a noisy linear inverse problem. Locally, the carrier task maps the trait through its Fisher matrix $F$, so gradients point toward $F\Delta_T$ rather than $\Delta_T$. Gradient descent then acts as a progressively less-damped inverse of $F$. With soft targets, this inverse restores low-curvature directions. With hard labels, it also amplifies the sampling noise in those same directions. The result is an optimal inversion depth that grows with the number of independent carriers. Experiments on Qwen2.5-7B and Gemma-2-9B confirm four predictions: the Fisher distortion of the initial gradient, recovery ordered from steep to flat directions, an optimal depth that shifts with data scale, and the finding that resampling completions from a fixed prompt pool works as well as adding new prompts. In this setting, large carrier datasets are needed less to reveal the trait than to suppress label noise amplified by Fisher inversion. Code is available at \url{https://github.com/luoyuchenmlcv/subliminal-data}.
|
| 955 |
TSAE: Structured Sparse Autoencoders for Interpreting Time-Series Forecasting Models
2610.04925
|
cs.LGcs.AI
|
Baoxi Liu, Yi Xie |
Time-series forecasting informs critical decisions in energy dispatch, industrial operations, and environmental monitoring; understanding the patterns models rely on is essential for assessing reliability and identifying failures. Input attribution identifies ...Time-series forecasting informs critical decisions in energy dispatch, industrial operations, and environmental monitoring; understanding the patterns models rely on is essential for assessing reliability and identifying failures. Input attribution identifies important variables and time segments but offers limited insight into internal features, while standard sparse autoencoder (SAE) objectives do not directly constrain cross-variable structure or temporal continuity. We introduce TSAE, a structured sparse autoencoder for forecasting representations that decomposes hidden states into individually inspectable features. TSAE organizes cross-variable structure through shared and variable-routed private dictionaries, separates feature detection from magnitude estimation with gated encoding, and constrains neighboring sparse-code changes according to raw-segment similarity. These mechanisms support analysis of variable context, activation strength, and temporal evolution. Forecast-consistency fine-tuning further improves preservation of the frozen forecaster's outputs. The accompanying TSEVAL protocol separately audits fidelity, feature coherence, and physical calibration to ground feature interpretation. In three-seed experiments with frozen PatchTST on ETTh1, ETTh2, and ETTm1, TSAE achieves the lowest hidden-state reconstruction error, normalized forecast-reconstruction error (NFRE), and feature transition rate among five SAEs at comparable per-token activity. NFRE decreases by 4.3-26.4% relative to the next-best mean. Dataset-dependent tradeoffs in selectivity, physical correlation, and calibration show that fidelity and temporal-stability gains require independent semantic validation to support feature interpretation.
|
| 956 |
Prompt Dominance and Asymmetric Verifier Costs: Empirical Ablations of GRPO at 1B Scale on GSM8K
2610.04928
|
cs.LG
|
Yi Hou |
This paper studies GRPO at 1B scale from both directions: what estimator choices do to the learning signal, and what a degraded reward signal does to what is learned. We train OLMo-2-0425-1B on GSM8K with a from-scratch implementation and measure both sides in...This paper studies GRPO at 1B scale from both directions: what estimator choices do to the learning signal, and what a degraded reward signal does to what is learned. We train OLMo-2-0425-1B on GSM8K with a from-scratch implementation and measure both sides in controlled sweeps, including a verifier-quality experiment that degrades the training reward and the test-time selector identically. Four results stand out. The prompt is the first-order decision: the zero-shot prompt leaves the base model at 0.08% (its outputs are degenerate continuations, not wrong answers), so almost no group carries a gradient, and training succeeds because the 3-shot prompt reaches 18.3%. At this scale the estimator variants sit within seed noise, with Dr. GRPO ahead on both seeds. In the off-policy regime, clipping is the whole story: training on data without a clipped ratio loses 4-6 points relative to the on-policy reference, while GRPO-style clipping and GSPO recover the loss entirely. Finally, the same weak verifier is far cheaper in RL than in test-time selection: a 10%-flip verifier leaves RL's attainable gain intact (91% and 106% retained across two seeds) where selection retains 57%, and a format-only verifier leaves RL with 16-30% of its gain and selection with essentially nothing. Flip noise acts as an affine transform on the expected reward, and the group-normalized advantage with Adam's rescaling removes it exactly; the residual is a second-order variance effect that the matched-step comparison at a 30% flip rate tests.
|
| 957 |
Billion-Scale Thumbnail Optimization for Uncurated Short-Form Videos via Multi-Armed Bandits
2610.04931
|
cs.LG
|
Ying Han, Ling Liu, Fabio Soldo, Vu Nguyen, Danio Wang |
This paper introduces a real-time thumbnail optimization system deployed at a global $O(B)$ scale on a major short-form video platform. Unlike traditional long-form content, where custom thumbnails are heavily curated by creators, a considerable fraction of sh...This paper introduces a real-time thumbnail optimization system deployed at a global $O(B)$ scale on a major short-form video platform. Unlike traditional long-form content, where custom thumbnails are heavily curated by creators, a considerable fraction of short-form videos are published without human-selected artwork. To address this uncurated corpus, we present a fully automated, end-to-end framework that replaces static default frames with dynamic, data-driven selections across billions of videos. To the best of our knowledge, this is the first published work demonstrating an online Multi-Armed Bandit framework successfully deployed at an $O(B)$ scale for uncurated short-form video discovery. Our solution pairs a multi-stage candidate generation pipeline with a low-latency serving infrastructure. By initializing the exploration framework with image-specific priors derived from a deep visual quality model, the system minimizes exploration costs and dynamically serves optimal thumbnails at serving time. Global deployment demonstrates statistically significant improvements in core user discovery and engagement metrics.
|
| 958 |
Learning under Localized Minority Imbalance
2610.04936
|
cs.LG
|
Amin Hosseininasab, Steven M. Shugan |
Class-imbalance methods implicitly assume that the minority class is uniformly undersampled relative to the majority class. However, in many real-world settings, minority instances may be disproportionately under-observed in certain regions of the feature spac...Class-imbalance methods implicitly assume that the minority class is uniformly undersampled relative to the majority class. However, in many real-world settings, minority instances may be disproportionately under-observed in certain regions of the feature space. For example, small businesses that go bankrupt may disappear from records, while those that survive remain visible, making bankruptcy appear less common among small firms than it actually is. This gives rise to localized minority imbalance (LMI), a challenge that is often overlooked and extends beyond general class-count imbalance. We show that under LMI, existing imbalance mitigation techniques can fit observation-induced biases in the training data and generalize poorly to under-observed regions of the true minority distribution. To address this, we propose a tree-based stratified approach that recursively partitions the feature space with the goal of reducing within-stratum LMI distortion. For each resulting stratum, we pair its majority instances with the full observed minority set and train a base classifier to create an ensemble. Extensive experiments over benchmark tabular datasets simulated with LMI show that our stratified ensembling approach outperforms popular and state-of-the-art imbalance mitigation techniques. We also introduce a gold-standard evaluation protocol that uses unbiased test sets, and demonstrate that conventional hold-out evaluation from the same LMI-biased data can substantially mislead performance. Overall, our results highlight that the cause of imbalance is as important as the correction method.
|
| 959 |
D-DOIT: Training-free Adaptation of Discrete Diffusion via Doob's h-Transform
2610.04938
|
cs.LGcs.AI
|
Jieke Wu, Qijie Zhu, Weimin Wu, Zeqi Ye, Minshuo Chen |
We propose D-DOIT (Discrete Doob-Oriented Inference-time Transformation), a training-free and efficient adaptation method for discrete diffusion models with generic rewards. D-DOIT formulates adaptation as sampling from a reward-tilted target distribution and ...We propose D-DOIT (Discrete Doob-Oriented Inference-time Transformation), a training-free and efficient adaptation method for discrete diffusion models with generic rewards. D-DOIT formulates adaptation as sampling from a reward-tilted target distribution and realizes this transport through Doob's h-transform of the discrete diffusion reverse kernel, using only reward values rather than reward gradients. Unlike continuous diffusion, masked discrete diffusion samples categorical token-reveal transitions rather than continuous state updates. D-DOIT derives the corresponding discrete Doob's h-transform, which guides sampling by reweighting reverse transition probabilities instead of adding a drift correction. To make this transformation practical, D-DOIT avoids expensive future rollouts. At each guided step, D-DOIT samples candidate next states, uses the model prediction head to complete each candidate into a clean sequence, evaluates each completion with the reward oracle, and resamples the next state with probabilities proportional to the rewards. An optional late-stage best-of-K refinement further improves sample quality by branching trajectories only near the end of denoising, avoiding the $K$-fold cost over the full trajectory. Empirically, across regulatory DNA design and protein inverse folding benchmarks, D-DOIT outperforms training-free guidance baselines. It improves enhancer activity and cell-type specificity while preserving sequence naturalness, and achieves the highest success rate in protein inverse folding.
|
| 960 |
A Graph-Based Inspection and Intervention Tool for Assessing Mechanistic Learning in PINNs
2610.04939
|
cs.LGcs.AI
|
Adwait Patkhedkar, Alifaraz Lakhani, Prathmesh Mohite, Abhijeet Salunke |
We ask whether physically meaningful correspondences discovered inside a trained scientific model remain meaningful outside the conditions under which they were discovered. We introduce GIIT (Graph-based Inspection and Intervention Tool), which represents gove...We ask whether physically meaningful correspondences discovered inside a trained scientific model remain meaningful outside the conditions under which they were discovered. We introduce GIIT (Graph-based Inspection and Intervention Tool), which represents governing physics as a computational physics dependency graph, maps graph nodes to internal network components via sensitivity- and trend-based discovery, and tests the resulting mapping under targeted intervention. Evaluating on temporal extrapolation out-of-distribution (OOD) regimes of increasing severity, we find that temporal extrapolation is associated with layer-wise correspondence drift and reduced functional correspondence. Specifically, while the model achieves low physics residual in-distribution (1.915 x 10-4 on full ID and 1.01 x 10-4 on an ID sub-window t in [0.5, 0.8]), the functional mapping changes even before leaving the training domain, and the shift continues as the evaluation window extends beyond the training domain: residual error increases from 2.81 x 10-2 on the boundary-crossing window (t in [0.5, 1.5]) to 3.30 x 10-1 on severe OOD (t in [1, 2]), accompanied by a systematic leftward shift of internal layer mappings, where the average winner layer index drops from 5.71 (ID) to 3.00 (ID sub-window) down to 2.29 (severe OOD). Only 2 of 7 physical nodes maintain stable layer assignments across the boundary-crossing window, and only 1 of 7 under severe extrapolation, revealing potential internal functional correspondence changes before severe degradation in conventional output metrics. Additional results for linear oscillatory systems are further detailed in the appendix.
|
| 961 |
Software World Models: From Consequence Prediction to Decision Value
2610.04940
|
cs.LGcs.AI
|
Tongli Su, Yuntong Hu, Liang Zhao, Bowen Zhu, JayaSai Somasundaram |
A coding agent may safely modify one repository while silently breaking downstream services, libraries, or datastores that depend on it. Exhaustively running integration tests after every agent action is impractical, so the agent must predict these failures be...A coding agent may safely modify one repository while silently breaking downstream services, libraries, or datastores that depend on it. Exhaustively running integration tests after every agent action is impractical, so the agent must predict these failures before executing them. Existing software world models predict the agent's own observations, while static change-impact analysis only identifies where a change may propagate. We instead introduce the Software World Model (SWM), which models the broader system affected by a code change and predicts its blast set: the components that the change will break. SWM follows three stages: explore, learn, and act. Explore executes candidate changes from restored system states, prioritizing regions where observed failures contradict the dependency graph. Learn fine-tunes a language model on these execution outcomes to predict downstream breakage. Act converts sampled predictions into per-consumer break probabilities for change ranking, proactive migration, and deciding when another execution is worth its cost. On held-out synthetic systems, SWM improves blast-set F1 from 0.431 for static reachability to $0.571\pm0.037$, more than halves ranking regret, and improves migration return at all nine evaluation checkpoints. The two methods are complementary: reachability is stronger on dependencies represented in the graph, while SWM recovers failures caused by couplings the graph misses. Experiments on held-out real libraries further show that predicting structured failure outcomes, rather than only scalar risk, is important for downstream decision quality.
|
| 962 |
TempoBridge: Source-Conditioned Flow Matching with Optimal Transport Couplings for Single-Cell Population Transitions
2610.04945
|
cs.LGcs.AI
|
Bowen Han, Lingbei Meng, Shihuan Luo, Yupeng Zang, Wenlin LI |
Destructive single-cell measurements provide unpaired population snapshots rather than observations of the same cells across conditions. Local cell states and transition requests may also be insufficient to distinguish responses across source populations. We i...Destructive single-cell measurements provide unpaired population snapshots rather than observations of the same cells across conditions. Local cell states and transition requests may also be insufficient to distinguish responses across source populations. We introduce TempoBridge, a common source-conditioned transport formulation for temporal, genetic, and chemical population transitions. Source cells initialize latent transport and provide a fixed empirical population summary. The velocity field receives this summary alongside the evolving cell state, flow time, and a structured transition descriptor. Minibatch optimal transport (OT) supplies couplings only for conditional flow-matching training paths; inference requires neither target expression nor OT computation. On held-out donors, TempoBridge achieves an Energy distance of 0.129 versus 0.144 for scGen. Genetic mean-expression $L_2$ error is 2.261 versus 3.156 for scGPT-scratch under Seen 2/2. On held-out compounds, condition-averaged drug-effect correlation is 0.598 versus 0.561 for the CellFlow adapter. Temporal ablations show higher mean distributional error after removing source context, replacing optimal transport with random pairing, or replacing flow matching with static residual regression. Together, these results demonstrate the predictive utility of a common source-conditioned transport formulation across held-out donors, gene combinations, and compounds.
|
| 963 |
Bridging the EHR Divide: Asymmetric Contrastive Learning for Cross-National Medical Representation Transfer
2610.04946
|
cs.LG
|
Qingyang Zhang |
Cross-system transfer of longitudinal Electronic Health Record (EHR) representations is challenging because clinical coding, patient populations, and healthcare workflows differ substantially across institutions and countries. We introduce Asymmetric Supervise...Cross-system transfer of longitudinal Electronic Health Record (EHR) representations is challenging because clinical coding, patient populations, and healthcare workflows differ substantially across institutions and countries. We introduce Asymmetric Supervised Contrastive Learning (Asymmetric SupCon), a task-specific pre-training objective motivated by the heterogeneity of negative clinical outcomes. The objective clusters patients sharing a target positive outcome without explicitly attracting negative trajectories toward one another. We pre-train temporal Transformer encoders on longitudinal records from 3.98 million patients in the Taiwanese National Health Insurance Research Database (NHIRD) and transfer them to two U.S. EHR datasets, MIMIC-IV and EHRSHOT. A hybrid semantic mapping pipeline combining direct mappings with embedding-based retrieval enables transfer across heterogeneous clinical vocabularies. On MIMIC-IV, NHIRD pre-training consistently improves over random initialization while substantially narrowing the performance gap to task-specific in-domain pre-training. On EHRSHOT, the transferred models show particularly strong few-shot performance for incident disease prediction. A controlled objective ablation shows that Asymmetric SupCon achieves the best AUPRC on three of four evaluated tasks and is 0.003 AUPRC below Standard SupCon on the fourth. These results support asymmetric contrastive pre-training as an effective approach for task-specific cross-national EHR representation transfer. Code is available at https://github.com/qingYzhang/Asymmetric_SupCon.
|
| 964 |
How Should Teachers Be Prepared? RL on Student-Induced States for On-Policy Distillation
2610.04950
|
cs.LG
|
Xiaoyu Ma, Haoyue Liu, Zhichao Wang, Jionghao Zhu, Xiaoying Tang |
On-policy distillation (OPD) improves the reasoning capabilities of small language models through token-level teacher supervision on student-generated trajectories. Yet can teachers that excel at solving problems independently also guide student reasoning effe...On-policy distillation (OPD) improves the reasoning capabilities of small language models through token-level teacher supervision on student-generated trajectories. Yet can teachers that excel at solving problems independently also guide student reasoning effectively? Prior work shows that when student prefixes follow reasoning paths that differ from the teacher's own or contain errors, teachers can be less accurate when continuing from these prefixes than when solving problems independently. To this end, we propose Prep-OPD, which uses reinforcement learning (RL) before distillation to train the teacher to adapt to the student's existing reasoning state and correct course when errors arise. Training optimizes teacher continuations from fixed student prefixes using final-answer correctness as the reward. The prepared teacher then trains the student through trajectory guidance and token-level supervision. We evaluate Prep-OPD on eight mathematical reasoning benchmarks, using Qwen3-4B-Instruct-2507 as the teacher and Qwen3-0.6B and Qwen3-1.7B as students. With the 4B teacher and 1.7B student, Prep-OPD improves average accuracy over standard OPD and the strongest baseline, Relay-OPD, by 8.28 and 2.30 percentage points, respectively. Controlled experiments further show that teacher RL conditioned on student-generated prefixes yields higher student accuracy than problem-start teacher RL with and without handoff on Qwen3-1.7B. Reusing the same prepared teacher also improves Qwen3-0.6B.
|
| 965 |
EDISCO: Equivariant DIScrete Diffusion for Euclidean Combinatorial Optimization
2610.04953
|
cs.LGcs.AI
|
Ruogu Chen, Jie Han |
Euclidean combinatorial optimization problems (ECOPs), such as the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP), possess inherent symmetries under the two-dimensional Euclidean group E(2), including rotations, reflections, an...Euclidean combinatorial optimization problems (ECOPs), such as the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP), possess inherent symmetries under the two-dimensional Euclidean group E(2), including rotations, reflections, and translations. Existing learning-based methods, including recent diffusion-based methods, rely on data augmentation or regularization to approximate E(2)-equivariance. This paper presents EDISCO, the first discrete diffusion model for ECOPs with exact E(2)-invariant generative distributions over node-index solutions. EDISCO introduces an E(2)-equivariant edge-score network coupled with a categorical continuous-time Markov chain over discrete edge variables, and exact posterior sampling provides efficient multi-step inference. This design gives EDISCO a local geometric inductive bias: edge neighborhoods with the same relative geometry and combinatorial context are represented consistently regardless of absolute position or orientation, making learning more efficient and inference more robust than non-equivariant methods. EDISCO outperforms previous learning-based state-of-the-art solvers on synthetic TSP from 100 to 10000 nodes and CVRP from 50 to 2000 customers, while using only 33-50% of the training instances. Trained only on uniform synthetic data, EDISCO also outperforms competing learning-based baselines under spatial distribution shift and CVRP constraint-tightness shift. Code is available at https://github.com/ValleyC/EDISCO.
|
| 966 |
Trinity: One Differentiable Physics for Training, Refining and Scoring Generative Floorplanners
2610.04957
|
cs.LG
|
Shih-Ying Yeh, Tzu-Sian Wang, Xuehai Wang, Jia-Hua Lee, Daniel Z. Kaplan |
Floorplanning arranges the blocks of a chip and decides their shapes under objectives that press blocks together, short wirelength and a small outline, and constraints that hold them apart, non-overlap, clusters, MIB shapes and boundary blocks. Recent diffusio...Floorplanning arranges the blocks of a chip and decides their shapes under objectives that press blocks together, short wirelength and a small outline, and constraints that hold them apart, non-overlap, clusters, MIB shapes and boundary blocks. Recent diffusion placers train on reference layouts alone and leave this coupled system to guidance, post-hoc loops and a legalizer, reporting only the endpoint, which hides what the generator contributes. We re-implement four of them under one recipe on FloorSet, score raw, refined and legalized layouts on one scale, and propose Trinity, a flow-matching floorplanner whose six differentiable functions for the constraints and objectives are its training loss term, the energy of a closed-form refiner after sampling and the base of a soft cost for every stage. The network thus learns the correction prior placers apply in their samplers, and sampling needs no guidance. Stage by stage, the training term lowers a plain transformer's raw soft cost by 26% and matters most at short budgets, the shared refiner decides more of the final cost than the generator and matches a ported placer's loop in 16 to 660 times fewer steps, Trinity's refined soft cost is 36% below the best ported pipeline, the soft cost ranks settings as the contest's hard cost does, and on the FloorSet val set the pipeline reaches a mean hard cost of 1.014 in 1.63 s per case.
|
| 967 |
MAGIC: Topology-Aware Analytic Graph Few-Shot Class-Incremental Learning
2610.04963
|
cs.LG
|
Junlin Chen, Yuhan Wang, Xuefei Wang, Xiao Wang, Ruijie Wang |
Graph few-shot class-incremental learning (GFSCIL) requires a model to continually recognize emerging classes from only a few labeled nodes while preserving previously acquired knowledge. Beyond the catastrophic forgetting inherited from conventional graph con...Graph few-shot class-incremental learning (GFSCIL) requires a model to continually recognize emerging classes from only a few labeled nodes while preserving previously acquired knowledge. Beyond the catastrophic forgetting inherited from conventional graph continual learning, GFSCIL presents two distinctive challenges: extremely limited novel-class supervision causes severe overfitting, while cross-session edges---edges connecting newly arriving nodes with historical nodes---alter historical propagation neighborhoods and thereby induce representation drift. We propose MAGIC, a replay-free GFSCIL framework that combines a frozen graph representation backbone (e.g., an intrinsically parameter-free backbone such as SGC or a pretrained graph foundation model) with closed-form analytic continual learning. To alleviate novel-session overfitting, MAGIC learns a topological prior from the base graph that can characterize both homophilous and heterophilous relations, and injects this prior through Potts Markov random field inference to refine supervision for novel classes. To mitigate representation drift, MAGIC transfers previous predictions from the old representations of affected historical nodes to their updated representations through drift-aware analytic distillation. Experiments across five datasets and eight baselines demonstrate the effectiveness of MAGIC. Under the 5-shot setting, MAGIC improves Mean Accuracy and Final Accuracy by 5.48 percentage points and 9.33 percentage points on average, and reduces Performance Drop by 10.78 percentage points on average compared with the best baselines. MAGIC also shows clear advantages under the 1- and 3-shot settings, with larger gains as the number of supports increases. Moreover, MAGIC requires substantially less training time.
|
| 968 |
Hamiltonian Metric Learning and Energy-Based Training: A Dissipative Geometric Framework for Optimization
2610.04969
|
cs.LG
|
Sparsho Chakraborty, Mohammad Alamgir, Nishanth M, Akshit Nanda, Ram Prasad Padhy |
Optimization in machine learning is usually expressed through iterative rules that update model parameters using information from the loss landscape. In this work, we study an alternative viewpoint in which parameter optimization is treated as the evolution of...Optimization in machine learning is usually expressed through iterative rules that update model parameters using information from the loss landscape. In this work, we study an alternative viewpoint in which parameter optimization is treated as the evolution of a dissipative dynamical system. The model parameters are regarded as generalized coordinates, the loss acts as a potential energy, and a positive-definite metric defines the local kinetic geometry of parameter space. Starting from a variational formulation, we derive the corresponding Hamiltonian dynamics, the geometric force generated by a position-dependent metric, and a metric-compatible Rayleigh dissipation law. The resulting continuous system satisfies a monotonic energy-dissipation relation, while its discrete form allows the influence of curvature on the optimization trajectory to be studied directly. We illustrate the framework using a controlled CIFAR-10 image-reconstruction problem for which the optimum is known analytically. With the image Hessian used as the metric, the curvature dependence of the quadratic modal dynamics is removed. We then reparameterize the same image as a matrix product state, producing a genuinely position-dependent metric and a nonzero geometric force. Poincar\'e return maps provide a complementary phase-space view of the resulting contraction dynamics. These examples establish HAMLET as a geometric, energy-based framework for studying optimization as dissipative motion in parameter space. On a five-seed MNIST MLP benchmark, HAMLET attains the highest mean test accuracy (97.95\%) and the lowest mean test negative log-likelihood (0.0696) among the three evaluated optimizers.
|
| 969 |
Priced Guidance: Can Language Models Generate Future Research Ideas?
2610.04976
|
cs.LG
|
Kaiyue Wen, Tengyu Ma, Percy Liang |
We evaluate language models' capability to generate novel research ideas through the lens of compression. We aim to lower-bound the potentially tiny probability that a language model generates the essence of a future research idea without any hints. Rather tha...We evaluate language models' capability to generate novel research ideas through the lens of compression. We aim to lower-bound the potentially tiny probability that a language model generates the essence of a future research idea without any hints. Rather than estimate this probability through expensive repeated sampling, our Priced Guidance framework measures the compression cost: how many additional bits of information are needed to guide the model to recover the target idea. We prove that if the model can recover the target idea with at most K bits of guidance in expectation, then it can generate the idea without any guidance with probability at least $2^{-K}$. In our framework, the language model, called the generator, can pose a sequence of multiple-choice questions and specify a probability distribution over possible answers. A guide, which is a language model with access to the target idea, selects answers. If a selected answer has prior probability p, the generator pays $-\log_2 p$ bits. The generator aims to produce an idea that matches the essence of the target idea with minimal cumulative cost. This cumulative cost equals, up to an additive constant, the number of bits of information sent by the guide. Using this methodology, we evaluate five generator LLMs (Opus 5, Fable 5.1, GPT-6 Astra, GPT-5.6 Sol, and GLM 5.3) on the core ideas in 87 recent high-quality deep learning papers and we use LLM as judge to determine whether the generated idea matches the target in terms of the central research object and defining mechanism. Fable 5.1 achieves the lowest median compression cost at 69.9 bits, substantially lower than gzip's median of 5,712 bits for losslessly compressing the summary of the target idea. A uniform ensemble of Fable 5.1, Opus 5, and Astra further reduces the median compression cost to 55.8 bits and improves the generation probability lower bound by 18,000 times.
|
| 970 |
What Will Post-Training Fix? Per-Problem Gains Are Shared Across Independent RL Runs, and Existing Checkpoints Predict Them Better Than A Priori Signals
2610.04978
|
cs.LG
|
Xiaoxian Duan |
Data selection, curricula and the evaluation of post-training recipes all assume that we can tell, before training, which problems a model will improve on. We test this assumption directly. For two base models, DeepSeek-R1-Distill-Qwen-1.5B and Qwen2.5-Math-1....Data selection, curricula and the evaluation of post-training recipes all assume that we can tell, before training, which problems a model will improve on. We test this assumption directly. For two base models, DeepSeek-R1-Distill-Qwen-1.5B and Qwen2.5-Math-1.5B, we evaluate eighteen post-training runs on up to 1532 competition math problems with many samples per problem, and compare signals available before training against a noise ceiling derived from the agreement between disjoint subsets of runs. Three findings hold on both base models. What post-training fixes is shared: independent runs agree on which rarely solved problems improve, with a noise ceiling of about 0.9, yet the two base models agree with each other only at rho=0.25 - the shared component belongs to the base model, not to the problem. A priori signals capture a minority of it: base pass rate, the likelihood of a correct solution, a larger model's pass rate and their combinations explain only 0.30 and 0.24 of the explainable variance in gains. Existing checkpoints are the better predictor: the per-problem gains of a single checkpoint from another family predict a new run better than every a priori signal, alone or combined (0.52 vs. 0.33 and 0.34 vs. 0.17 against the combined signals). The conclusions hold on problems from 2025-2026 competitions and when the baselines on the two sides of every comparison are estimated independently. We propose the noise ceiling as a standard companion to per-problem signals.
|
| 971 |
How Long, Not How Close: A Learned Temporal Metric for Planning in Latent World Models
2610.04988
|
cs.LG
|
Lama Moukheiber, Haotian Xue, Yongxin Chen |
Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans ...Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the demonstrations already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number of environment steps between them, and blends this distance into the planner's cost. It requires no rewards, policies or success labels and, being a cost rather than a model, applies to frozen world models with one latent vector per state that plan by a latent distance. We evaluate TEMPO on eleven simulated environments (e.g., maze navigation, tabletop pushing, robotic arm control and three-dimensional manipulation) with the LeWM and PLDM planners. With a small MLP that adds at most 0.3% to a plan's arithmetic, TEMPO improves both planners at every goal distance, including the one-plan setting of their evaluations, raises LeWM from 36% to 99% on TwoRoom three plans from the goal, and remains competitive on a broad range of 2D and 3D navigation, reaching and manipulation tasks.
|
| 972 |
Pessimistic Minimax Learning for Public-Private Information Games under Unilateral Coverage
2610.04997
|
cs.LG
|
Shuze Daniel Liu, Claire Chen, Jiuqi Wang, David Simchi-Levi |
We study offline learning in two-player zero-sum contextual games with public and private information, motivated by strategic settings such as auctions and negotiations with private valuations. We introduce unilateral prescriptive concentrability and show that...We study offline learning in two-player zero-sum contextual games with public and private information, motivated by strategic settings such as auctions and negotiations with private valuations. We introduce unilateral prescriptive concentrability and show that asymmetric information can change offline coverage through its effect on equilibrium behavior. For finite state-action spaces, we develop a pessimistic algorithm with an $\tilde{O}(1/\sqrt{n})$ exploitability rate, matching the standard sample-size dependence for fully observed minimax games. We further develop a pessimistic policy mirror descent framework, PPA-PMD, for general function approximation and obtain a unified $\tilde{O}(1/\sqrt{n} + 1/\sqrt{T})$ exploitability rate with no-regret actor updates. Together, these results provide the first theoretical framework for offline equilibrium learning under public-private information constraints.
|
| 973 |
Beyond Overparameterization: Provable Learning of Input-Convex Multi-Layer Polynomial Networks with Active Queries
2610.04999
|
cs.LG
|
Jinqi Tang, Qian Chen, Shihong Ding, Cong Fang |
The theoretical understanding of multi-layer neural networks is largely confined to overparameterized settings, which obscure parameter identifiability and incur high sample complexity. Neural tangent kernel (NTK) provides a general theory for wide networks, b...The theoretical understanding of multi-layer neural networks is largely confined to overparameterized settings, which obscure parameter identifiability and incur high sample complexity. Neural tangent kernel (NTK) provides a general theory for wide networks, but does not offer efficient sample-complexity guarantees. Recent feature-learning results go beyond kernel methods for single-neuron, multi-index, and hierarchical targets. However, the analysis is often restricted to shallow or specific architectures and to the overparameterized regime. We break this paradigm to achieve parameter-level recovery of deep target networks, albeit by using active data queries. Specifically, we study $L$-layer polynomial networks with even degree-$k$ monomial activations and nonnegative higher-layer weights. This structure makes the target network input-convex, while the optimization landscape remains highly nonconvex with respect to the parameters. Leveraging input convexity and active queries, we propose \textbf{ASPIRE} (\textbf{A}ctive \textbf{S}am\textbf{P}ling for \textbf{I}terative \textbf{R}ecovery via \textbf{E}igendirections), a layerwise sampling-based diagonalization algorithm that recovers all network parameters to $\delta$-accuracy with sample complexity $ \widetilde O_{k,L}\left(d^{L^2+O(L)}\delta^{-2e}\right) $ in polynomial time. To our knowledge, this is the \emph{first} parameter-recovery guarantee for deep target networks whose exponent grows only polynomially with depth, as well as the \emph{first} justification for the effectiveness of using high-quality data in neural network training, with a remarkably \emph{exponential} separation.
|
| 974 |
Physics-Augmented Graph Transformers for Patch-Antenna Forward and Inverse Design
2610.05004
|
cs.LG
|
Avi Epstein, Snir Nehemia, Haim Suchowski, Lior Wolf |
Full-wave electromagnetic (EM) simulation enables accurate patch-antenna analysis but is computationally expensive for large-scale forward prediction and inverse design. We present a mesh-native, physics-augmented graph-learning framework that treats radiation...Full-wave electromagnetic (EM) simulation enables accurate patch-antenna analysis but is computationally expensive for large-scale forward prediction and inverse design. We present a mesh-native, physics-augmented graph-learning framework that treats radiation-pattern prediction as signal reconstruction on an irregular surface mesh. For the forward problem, a GPS graph transformer is trained with Physics-Augmented Intermediate Supervision (PAIS), an auxiliary node-level objective that predicts complex surface currents, the physical intermediate linking geometry to radiation. PAIS improves multiple GNN backbones at no inference-time cost, while shuffled-current and non-physical controls show the gain comes from physical correspondence. Direction-conditioned decoding and a differentiable radiation-integral consistency loss further exploit this structure. On an 80,000-sample CST benchmark, GPS+PAIS reaches MSE 0.17 / PSNR 19.67, generalizes to a PCA split, and transfers zero-shot to canonical patches. For inverse design, surrogate-filtered diffusion beats nearest-neighbor retrieval by 32% relative MSE.
|
| 975 |
MetaKernelBench: Measuring GPU Kernel Knowledge Transfer Beyond Code
2610.05014
|
cs.LG
|
Xueyi Chen, Shiyu Liu, Xin Jin, Yuhua Zheng, Xin Li |
Recent GPU kernel optimization agents retain what they learn in knowledge bases or as distilled skills. Kernel benchmarks score each attempt's implementation for correctness and speed but leave the reuse value of retained experience unmeasured. We introduce Me...Recent GPU kernel optimization agents retain what they learn in knowledge bases or as distilled skills. Kernel benchmarks score each attempt's implementation for correctness and speed but leave the reuse value of retained experience unmeasured. We introduce MetaKernelBench, which measures whether experience distilled from an attempt in one kernel domain-specific language (DSL) improves a fresh attempt at the same problem in another. Its 74 problems are fused subgraphs in six families, each posed as a pair of CuTe DSL and TIRx variants that differ only in the DSL. The agent first attempts each variant solo and is instructed to distill what it learns into a natural-language skill, which is transferred whether or not the source attempt passes verification. The skill is the only extra input to a skill-conditioned attempt by the same model in the other DSL. We compare each skill-conditioned attempt with the solo attempt on the same variant under matched per-attempt budgets, scoring correctness and end-to-end runtime. Across six models and both directions, paired lift over solo attempts ranges from -19% to +29%. Four models gain in both directions, yet regressions occur on 16% to 45% of problems in every model and direction. Outcomes follow the source attempt's result relative to the target's solo attempt rather than source success alone, improving in 71% of comparisons when the source stands above and regressing in 54% when it stands below. MetaKernelBench complements implementation-quality metrics by measuring same-problem cross-DSL kernel knowledge transfer.
|
| 976 |
Operational Abstractions of Neural Network Concepts via Topological Representations
2610.05017
|
cs.LGcs.AI
|
Mathieu Pont, Christoph Garth |
Concept-based methods provide a semantic level for interpreting and manipulating learned representations, but existing editing approaches are typically specialized to particular interventions and do not provide a common and editable representation of concept o...Concept-based methods provide a semantic level for interpreting and manipulating learned representations, but existing editing approaches are typically specialized to particular interventions and do not provide a common and editable representation of concept organization. To achieve this, we introduce Topological Concept Representations (TCR), a post-hoc operational abstraction that jointly characterizes the concepts encoded in a learned representation and their relationships. TCR constructs an intermediate concept space from concept recoverability and interaction scores, and compactly encodes its organization through topology. Interventions are expressed through modifications of this abstraction that are propagated back to the underlying learned representation. This allows different concept-level operations to share the same optimization framework and separates the desired concept organization from the mechanism to achieve it. We establish stability and reparameterization-invariance properties of TCR and its connections to existing concept-editing formulations. We use TCR to disentangle concepts as a preprocessing step for existing erasure methods, improving worst-group accuracy by 21.89 on average at comparable concept leakage. We further use TCR to transfer concepts from teacher to student models, improving concept recoverability by up to 5.54 while also improving or maintaining competitive test top-1 accuracy.
|
| 977 |
Calibrated Weak Supervision for Post-Harvest Burned-Cropland Mapping Under Label Scarcity
2610.05040
|
cs.LG
|
Raunak Bhagate, Maitri Polisetty, R I Minu |
Mapping post-harvest burned cropland is difficult when fires are small and fragmented and reliable labels are scarce. We developed a calibrated weak-supervision framework for Punjab, India, using Sentinel-2 spectral change, VIIRS active-fire context, and MODIS...Mapping post-harvest burned cropland is difficult when fires are small and fragmented and reliable labels are scarce. We developed a calibrated weak-supervision framework for Punjab, India, using Sentinel-2 spectral change, VIIRS active-fire context, and MODIS MCD64A1 as a coarse external calibration and agreement reference. Three pseudo-label recipes, five feature representations, and linear, tree-based, boosted, and neural classifiers were evaluated using nested district-held-out cross-validation over three seeds and five folds. NBR and dNBR were excluded from classifier inputs. The best configuration used the very-strict recipe, a multilayer perceptron, and the full optical feature set (mean Cohen's kappa 0.395, AUROC 0.753, F1 0.747, balanced accuracy 0.703); Random Forest, XGBoost, and LightGBM were practically tied. Higher agreement with held-out pseudo-labels did not establish improved label correctness or independent burned-area accuracy. For deployment, a Random Forest with the very-strict recipe and full optical features was retained. An externally calibrated threshold of 0.60 yielded district-level MODIS agreement of R-squared 0.636, a mapped-to-MODIS burned-area ratio of 1.005, and Spearman correlation of 0.779 with district fire counts. Pixel-level MODIS agreement remained modest (F1 0.205, kappa 0.093). Zero-shot transfer to Haryana was promising (mean kappa 0.641), but Punjab cross-year stability was weak, and Sentinel-1/Sentinel-2 feature concatenation did not improve the optical baseline. Optical observations ended before the seasonal fire context, limiting coverage of late burns. The framework supports district-scale burden assessment and hotspot screening, with limited support for exact scar boundaries or temporally stable annual mapping.
|
| 978 |
E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
2610.05048
|
cs.LGcs.AI
|
Yifei Liu, Minghao Fang, Xinyu Gu, Chengkai Yao, Mengdi Liu |
On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training,...On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E$^2$-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E$^2$-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E$^2$-OPSD remains simple, requiring no additional forward passes or networks.
|
| 979 |
LogSig-SSM: Time-Series Modelling with Multi-Scale Log-Signature Compression for State-Space Models
2610.05051
|
cs.LGcs.AI
|
Felix Oury, Nicolas Calvo Peiro, Reiko J. Tanaka |
Time-series data are often sampled irregularly at high frequencies and exhibit long-range dependencies, which makes long-horizon modelling difficult. Continuous-time models such as neural controlled differential equations (NCDEs) and neural rough differential ...Time-series data are often sampled irregularly at high frequencies and exhibit long-range dependencies, which makes long-horizon modelling difficult. Continuous-time models such as neural controlled differential equations (NCDEs) and neural rough differential equations (NRDEs) can handle irregular sampling, but they scale poorly to long sequences. Selective state-space models (SSMs) such as Mamba scale linearly with sequence length, but they provide limited recurrent mixing across hidden dimensions within a single block. We propose LogSig-SSM (Log-Signature Compression for State-Space Models), which first compresses long multivariate time series into a shorter sequence of tokens using multi-scale windowed log-signatures, and then processes these tokens with a selective SSM backbone. LogSig-SSM is scalable and robust to irregular sampling, combining log-signature tokens that capture higher-order cross-channel interactions with a selective SSM that models long-range dependencies. The model also admits a continuous-time interpretation as an NCDE/NRDE-style system driven by a log-signature-based input, in which selectivity induces an input-dependent rescaling of the latent dynamics. Across four benchmarks, namely long-sequence classification on UEA, high-frequency physiological regression on PPG-DaLiA, multivariate weather forecasting, and irregularly sampled clinical prediction on PhysioNet Sepsis, LogSig-SSM outperforms or matches strong SSM and continuous-time baselines while training up to $30\times$ faster and using up to $37\times$ less GPU memory than Mamba on the longest sequences.
|
| 980 |
Revealing After Overwriting: An Exponential POMDP OPE Lower Bound under History-Dependent Logging
2610.05063
|
cs.LG
|
Youyu Luo, Pengzhan Zhou, Zhida Qin, Jia Wang, Zuotao Fu |
Multi-step revealing can make off-policy evaluation tractable under memoryless logging. With history-dependent logging, state decodability and target-relevant evidence can separate. For every horizon $H\ge3$, we construct two exactly realizable POMDPs with fou...Multi-step revealing can make off-policy evaluation tractable under memoryless logging. With history-dependent logging, state decodability and target-relevant evidence can separate. For every horizon $H\ge3$, we construct two exactly realizable POMDPs with four actions, at most four states per layer, a known logger, and a memoryless target. Action overlap, history coverage, and observation-only revealing remain bounded independently of $H$, yet the target values differ by $1/2$ and the KL divergence between the logged laws is $\Theta(4^{-(H-1)})$, forcing exponential sample complexity. Logger memory makes states distinguishable, while reset erases the model-distinguishing evidence preserved by the target. A separate construction retains this barrier with common, known observation-only revealing operators. Under action and history coverage, we give a finite-class OPE guarantee using common observable value representations that remain valid at every history. The sample bound depends polynomially on their second-moment cost. In the common-operator construction, the same value direction has constant marginal decoding cost but exponential history-conditioned cost. Finally, on a fixed four-action continuum, we derive matching passive and budgeted readout rates. With one known channel and unit read cost, early reads are optimal. With unknown sensor bias, early reads alone remain exponentially costly. Combining them with post-reset calibration gives sample complexity independent of $H$ when both read types receive fixed positive expected budgets per trajectory.
|
| 981 |
Outcome-Guided On-Policy Self-Distillation
2610.05070
|
cs.LG
|
ZheXu Wang, Mao-Lin Luo, Yankun Hong, Zi-Hao Zhou, Bo Ye |
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability...On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.
|
| 982 |
Discrete Action Matching: Learning Stochastic Dynamics from Samples via State Graphs
2610.05071
|
cs.LGcs.AI
|
Mikhail Persiianov, Alexander Korotin |
Learning population dynamics from unpaired temporal marginals is an ill-posed inverse problem that requires structural assumptions on the underlying dynamics. We introduce $\textit{Discrete Action Matching}$ (DAM), a finite-state counterpart of Action Matching...Learning population dynamics from unpaired temporal marginals is an ill-posed inverse problem that requires structural assumptions on the underlying dynamics. We introduce $\textit{Discrete Action Matching}$ (DAM), a finite-state counterpart of Action Matching based on discrete Wasserstein geometry. For a prescribed marginal path and transport geometry, we derive an action-minimization objective for its canonical minimum-kinetic-energy current. Our key observation is that the density dependence of the discrete action reduces to neighboring density ratios. Along an empirical interpolation of the snapshots, DAM first estimates these ratios and then learns an action potential. The learned fields also define a graph-supported Markov sampler. Experiments on controlled synthetic dynamics and real mouse gastrulation data evaluate marginal reconstruction and interpolation. Additional experiments approximate numerical surface-transport paths from paired samples.
|
| 983 |
ReDiffNet: Differential RGB-Infrared Learning for Low-Light UAV Oriented Vehicle Detection
2610.05074
|
cs.LG
|
Qifan Zhang, Ziran Zhou, Ruijie Li, Jincheng Tang, Hao Wang |
Low-light UAV-based RGB-infrared oriented small-vehicle detection is important for nighttime traffic monitoring, emergency response, and urban inspection. Illumination variations, headlight glare, local shadows, and thermal-response degradation cause spatially...Low-light UAV-based RGB-infrared oriented small-vehicle detection is important for nighttime traffic monitoring, emergency response, and urban inspection. Illumination variations, headlight glare, local shadows, and thermal-response degradation cause spatially varying modality reliability, while the small visual extent of vehicles further weakens boundaries, orientation cues, and thermal responses. Accordingly, selecting trustworthy observations based on local modality reliability while further exploiting complementary discriminative information in regions with ambiguous modality preference is key to constructing effective multimodal representations. Based on this insight, we propose ReDiffNet, a reliability-conditioned differential representation network in which modality reliability guides both evidence selection and complementary recovery. Specifically, degradation-aware reliability learning estimates relative spatial reliability, uncertainty-guided differential recovery exploits cross-modal differences to recover complementary cues in ambiguous regions, and reliability-conditioned reconstruction integrates retained and recovered evidence into a unified representation. ReDiffNet achieves 85.3% and 73.9% mAP50 on DroneVehicle and VEDAI, respectively, supporting its effectiveness.
|
| 984 |
How Execution Assumptions Change Short-Horizon Sharpe Rankings: Evidence from a Synthetic Trading Benchmark
2610.05077
|
cs.LG
|
Weicheng Xue |
Backtests of LLM trading agents often assume that every order fills at the closing price. We ask whether this choice changes only reported returns or also the order of the agents. Five prompted LLM signal policies and seven classical baselines trade the same s...Backtests of LLM trading agents often assume that every order fills at the closing price. We ask whether this choice changes only reported returns or also the order of the agents. Five prompted LLM signal policies and seven classical baselines trade the same synthetic price paths under six execution settings, from near-ideal fills to latency, spread, participation, and impact stresses. The main experiment contains $2{,}462$ runs with matched decision frequencies and paired market paths. On the compressed two-asset board, agreement between the near-ideal and default-stress rankings falls to Kendall $\tau_b=0.21$ in the high-volatility regime, compared with $0.82$ in the calm regime. The seed-bootstrap intervals, $[0.00,0.52]$ and $[0.48,0.94]$, are wide and overlap. On a fixed 11-policy board, agreement rises from 0.24 with two assets to 0.85 with ten; the two-asset point estimate differs substantially from the wider settings we tested. Rank changes are related to turnover, and comparisons with buy-and-hold also depend on how that anchor is initialized. The experiment does not compare LLM trading skill. It shows that, on a short horizon, an execution convention can become part of the benchmark's headline. Execution assumptions and rank stability should be reported alongside returns.
|
| 985 |
Direction-Conditioned Policies for Online Goal-Conditioned Reinforcement Learning
2610.05087
|
cs.LGcs.AI
|
S K Swaminathan, Damiya Gondha, Theyanesh Eswaramoorthy Rajahkrishnan, Aritra Hazra |
Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Poli...Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and distance in representation space. At deployment, DCP applies the same interface directly to the final goal, requiring neither waypoint selection nor planning. Across nine navigation and manipulation tasks, DCP attains higher final success rates than CRL on seven tasks and spends more time near the goal on seven. Controlled maze experiments further show that DCP captures shortest-path geometry more accurately and that the supplied direction causally influences the actor's behavior. We identify waypoint coverage and ranking as limits to exploration, and show that learned candidate generation improves goal reaching in two controlled mazes.
|
| 986 |
Advectra: Asymmetric Latent Transport for Non-Stationary Physics
2610.05098
|
cs.LG
|
Nodens Koren, Thomas Hofmann, Georgios Kissas |
Many latent neural operators represent input and output fields in a stationary latent chart. In particular, common latent routing mechanisms use fixed or shared assignment weights for feature projection and reconstruction, limiting their ability to model trans...Many latent neural operators represent input and output fields in a stationary latent chart. In particular, common latent routing mechanisms use fixed or shared assignment weights for feature projection and reconstruction, limiting their ability to model transport-dominated systems where coherent structures move relative to fixed coordinate frames. We propose Advectra, a transport-aware latent operator that introduces a regularized kinematic coordinate map to decouple source and target coordinate systems. This yields an approximately co-moving latent reference frame and enables asymmetric feature aggregation and reconstruction. Combined with a geometry-aware ordering mechanism for state-space models, Advectra captures advective dynamics while maintaining stable global interactions. Advectra achieves the best performance among evaluated geometry-constrained and form-free baselines on advection-dominated benchmarks, including passive scalar transport in Navier--Stokes flows and Rayleigh--Taylor instability, while demonstrating strong generalization on real-world engineering tasks. These results highlight the benefit of explicit moving-frame structure in neural operators for non-stationary physics.
|
| 987 |
METRO: Metric-Enhanced Token Routing Operator
2610.05100
|
cs.LG
|
Nodens Koren, Thomas Hofmann, Georgios Kissas |
State-of-the-art neural operators scale to complex meshes via slice-and-process architectures, yet many rely on linear compatibility scores for latent tokenization. Under common feature normalization, such scores are equivalent to isotropic Euclidean clusterin...State-of-the-art neural operators scale to complex meshes via slice-and-process architectures, yet many rely on linear compatibility scores for latent tokenization. Under common feature normalization, such scores are equivalent to isotropic Euclidean clustering, while without normalization they induce unbounded linear decision regions. In both cases, they lack slice-specific anisotropic locality, which can lead to redundant and entangled latent slices. To address this, we propose Metric-Enhanced Token Routing Operator (METRO), a geometry-aware routing mechanism that replaces linear projection with a learnable Mahalanobis metric. By enabling each latent slice to learn a local anisotropic tensor, METRO shapes receptive fields into exponentially localized, oriented ellipsoids that naturally align with flow features like boundary layers and wakes. As a drop-in replacement, METRO yields consistent improvements across both Transformer and Mamba backbones. Empirically, our method achieves substantial performance gains on irregular domains, outperforming baselines on both standard PDE benchmarks and complex industrial design tasks. Finally, METRO exhibits enhanced robustness in out-of-distribution regimes across varying Reynolds numbers and geometric configurations.
|
| 988 |
Private Component-by-Component Learning
2610.05102
|
cs.LG
|
Dvir Karni, Eliad Tsfadia |
We study differentially private learning problems in the realizable setting, where a hypothesis is specified by $k$ components. A direct iteration of private component learners is obstructed by a simple difficulty: an approximate choice of the next component m...We study differentially private learning problems in the realizable setting, where a hypothesis is specified by $k$ components. A direct iteration of private component learners is obstructed by a simple difficulty: an approximate choice of the next component may destroy exact realizability of the labeled sample, even when the next component is locally accurate. We restore realizability using the LabelBoost procedure of Beimel, Nissim, and Stemmer [SODA '15, Algorithmica '21] and recycle data through two alternating reservoirs. The resulting learner, for a target privacy $\varepsilon$, pays only $\widetilde O(\sqrt{k}/\varepsilon)$ overhead relative to the active sample requirement of a single component learning step at target accuracy $\Theta(\alpha/k)$. For learning $d$-dimensional halfspaces over a finite coordinate grid of size $L$, exact realizability makes the direct component-depth objective quasi-concave. Instantiating the framework with the IPConcave algorithm of Nissim, Tsfadia, and Yan [SODA '26] and with the quasi-concave optimizer of Cohen, Lyu, Nelson, Sarl'os, and Stemmer [STOC '23] yields a realizable sample complexity of $\widetilde{O}\left(\frac{1}{\varepsilon \alpha}\cdot \min\{d^{2.5} \log^*L, \:\: d^{2.5} + d^{1.5} 2^{\log^*L}\}\right),$ which improves on the previously known bound of $\widetilde{O}\left(\frac{1}{\varepsilon \alpha}\cdot\min\{\frac{1}{\alpha}\cdot d^{5.5}\log^*L,\:\: d^{2.5}2^{\log^*L}\}\right).$ We also apply the framework to Boolean compositions: given proper private learners for classes $H_1,\ldots,H_k$, we obtain a proper private learner for $G(H_1,\ldots,H_k)$ for any fixed Boolean function $G:\{0,1\}^k\to\{0,1\}$. Compared with the closure theorem of Alon, Beimel, Moran, and Stemmer [COLT '20], this reduces the overhead on a common component sample bound from $\widetilde O(k/\varepsilon)$ to $\widetilde O(\sqrt{k}/\varepsilon)$.
|
| 989 |
Best-of-$N$ Guidance for Test-time Diffusion Alignment
2610.05108
|
cs.LGcs.AI
|
Richard Lee Kim, Yeongmin Kim, Gyuwon Sim, Taekyu Kim, Minsang Park |
Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i....Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3$\times$ ImageReward score of the latest sample-based guidance method with a 1.6$\times$ speedup. We release the code at https://github.com/aailab-kaist/BoNG.
|
| 990 |
Bayesian Entropy-based Reordering for Calibrated Diffusion Language Models
2610.05125
|
cs.LGcs.AI
|
Zhejun Jiang, Mijung Park |
Masked Diffusion Language Models (MDLMs) generate sequences by iteratively replacing masked tokens with model predictions. At each denoising step, the decoder chooses which positions are sufficiently confident to commit. Existing decoding methods typically rel...Masked Diffusion Language Models (MDLMs) generate sequences by iteratively replacing masked tokens with model predictions. At each denoising step, the decoder chooses which positions are sufficiently confident to commit. Existing decoding methods typically rely on softmax confidence, which can be miscalibrated. We introduce BayesER (BAYESian Entropy-based Reordering), a post-hoc Bayesian decoding framework that uses predictive uncertainty to guide token commitment. In BayesER, we construct a lightweight approximate posterior centered at the pretrained checkpoint, similar to Laplace-LoRA but without training LoRA adapters. We average predictions over posterior samples and use predictive entropy to prioritize reliable positions. We examine how posterior predictions affect position ordering and token selection across benchmarks spanning code generation, mathematical reasoning, planning, and molecular generation. We show that BayesER reduces sequence-level calibration error while preserving or improving accuracy relative to common decoding schemes, including confidence-threshold decoding. Additionally, a posterior fitted on one code-generation dataset reduces calibration error on another without refitting, suggesting that Bayesian uncertainty may provide a transferable signal for more reliable MDLM decoding.
|
| 991 |
Component-Level Evaluation of Adaptive PINN Training for CFD-Oriented Crystal Growth Simulation
2610.05127
|
cs.LG
|
Niruta Chapagain, Rohit Raj, Bertwin Kurisinkal Shine, Aditya A S |
Physics-informed neural network (PINN) training minimizes a weighted combination of partial differential equation (PDE), boundary-condition, and initial-condition losses. Because adaptive methods modify these weights during training, their weighted total losse...Physics-informed neural network (PINN) training minimizes a weighted combination of partial differential equation (PDE), boundary-condition, and initial-condition losses. Because adaptive methods modify these weights during training, their weighted total losses are not always directly comparable. We compare fixed-weight PINN, gradient-normalized PINN (GNPINN), and a rule-based adaptive controller (AgenticPINN) under matched settings on a heat-equation benchmark and a simplified Czochralski-oriented thermal-fluid problem. In the crystal-growth MLP experiment, adaptive control reduced the PDE residual from the order of $10^{-5}$ to $10^{-6}$, while the boundary-condition loss increased from the order of $10^{-5}$ to $10^{-2}$. On the heat-equation benchmark, GNPINN achieved the lowest relative $L_2$ field error (0.054), whereas AgenticPINN obtained the smallest PDE residual but a relative $L_2$ error of 1.368. Gaussian-process surrogates were additionally evaluated using case-wise holdout tests on corrected Czochralski CFD parameter sweeps. The temperature-field error for the temperature sweep was approximately 6%, whereas the axial-velocity error for the crystal-rotation sweep was approximately 42%. These findings show that adaptive control can improve equation satisfaction while weakening other physical constraints. PINN training should therefore be evaluated using separate PDE, boundary-condition, and solution-error metrics rather than weighted total loss alone.
|
| 992 |
Hidden in the Comments: A Context-Injection Attack Surface in Code LLMs
2610.05139
|
cs.LG
|
Noor Munir, Francesco Quinzan, Stephen Roberts |
Code large language model (Code LLM) assistants generate code from heterogeneous development contexts, including open files, imported modules, pasted snippets, and comments, much of which may originate from untrusted sources. We investigate whether insecure in...Code large language model (Code LLM) assistants generate code from heterogeneous development contexts, including open files, imported modules, pasted snippets, and comments, much of which may originate from untrusted sources. We investigate whether insecure instructions embedded in such contexts can steer Code LLMs toward vulnerable code without access to model weights or training data. We evaluate ten open-weight Code LLMs spanning 3B--13B parameters, including four base and six instruction-tuned models, across ten web-application weakness classes. We compare completion tasks containing insecure instructions embedded as code comments with benign tasks without malicious instructions. Attack-condition completions contained a medium-or-higher weakness in {\bf 77.4--92.3}\% of cases, compared with {\bf 1.7--5.1}\% in the benign condition. Base and instruction-tuned models averaged 86.5\% and 84.5\% vulnerable outputs, respectively; equivalence testing and three matched model pairs indicated reductions of at most 8.1\% after instruction tuning. Susceptibility showed no clear association with model scale or specialization. Among vulnerable attack outputs, 86.2--91.0\% were rated high or critical, and the effect persisted without the pattern-based detector. Post-generation screening reduced but did not eliminate the risk, the strongest screen leaving roughly one-third undetected. These findings identify inference-time context injection as a substantial attack surface and motivate provenance-aware training objectives.
|
| 993 |
Arithmetic Actor Heads and Training Stabilization for Out-of-Distribution Reinforcement Learning
2610.05143
|
cs.LG
|
Yifan Zhang, Liang Zheng |
Reinforcement learning (RL) policies can deteriorate under out-of-distribution (OOD) magnitude shifts. Starting from soft actor-critic (SAC) and its Bayesian Amnesic Piecewise-Robust (BAPR) predecessor, we study the causal-symbolic BAPR (CS-BAPR) family. The p...Reinforcement learning (RL) policies can deteriorate under out-of-distribution (OOD) magnitude shifts. Starting from soft actor-critic (SAC) and its Bayesian Amnesic Piecewise-Robust (BAPR) predecessor, we study the causal-symbolic BAPR (CS-BAPR) family. The practical method combines six training-stabilization settings with alternative actor heads: a Neural Addition Unit (NAU) with a Neural Multiplication Unit (NMU)-inspired quadratic correction, a Kolmogorov-Arnold Network (KAN), or a multilayer perceptron (MLP) with rectified linear unit (ReLU) or hyperbolic-tangent activations.
|
| 994 |
Cross-Modal Contrastive Learning for the Retrieval of Immunotherapy-Associated Molecular Signatures from Histopathology
2610.05157
|
cs.LGcs.AI
|
Sigrid Vila-Bagaria, Mar Teixid\'o, Miquel Pi\~nol, Felip Vilardell, Robert Montal |
Gastric Adenocarcinoma is a leading cause of cancer mortality. Although "Inflamed/Non-Inflamed" subtypes have been proposed to predict immunotherapy response, their identification relies on a costly 10-gene RNA signature. We propose a Cross-modal Contrastive M...Gastric Adenocarcinoma is a leading cause of cancer mortality. Although "Inflamed/Non-Inflamed" subtypes have been proposed to predict immunotherapy response, their identification relies on a costly 10-gene RNA signature. We propose a Cross-modal Contrastive Multiple Instance Learning (CCMIL) framework for cross-modal retrieval, imputing these molecular signatures directly from standard Hematoxylin & Eosin (H&E) slides. By leveraging a supervised contrastive objective, CCMIL aligns visual morphological patterns with molecular phenotypes into a shared latent space. This establishes an interpretable search-by-case retrieval engine, enabling pathologists to query a whole slide image to surface transcriptomically coherent neighbors and approximate RNA signatures without genomic sequencing at inference. Our results demonstrate that this retrieval-first approach captures the continuous phenotypic spectrum of tumor inflammation and yields clinically interpretable attention heatmaps. Furthermore, the learned representation also supports competitive downstream classification, providing a practical molecular pre-screening strategy.
|
| 995 |
Measuring Learned Monotone Temporal Aggregation at Matched Admissibility
2610.05196
|
cs.LG
|
Yew Lee Tan |
Risk regulation imposes directional constraints on scores; we adopt their strict per-input form -- the score monotone non-decreasing in every exposure input -- as a normative commitment. Deployed pipelines -- monotone hand-crafted aggregates feeding sign-const...Risk regulation imposes directional constraints on scores; we adopt their strict per-input form -- the score monotone non-decreasing in every exposure input -- as a normative commitment. Deployed pipelines -- monotone hand-crafted aggregates feeding sign-constrained gradient boosting -- already satisfy it by composition, so constrained-versus-unconstrained comparisons price a guarantee the incumbent has for free. We instead hold admissibility fixed on both sides and measure what learning the aggregation is worth. Our instrument is a recurrent network whose state is classical risk statistics (an exponentially weighted moving average and a high-water mark with learned transforms), monotone by construction in every input and per MC-dropout sample. The central finding, by functional regression, is a subsumption boundary: a learned monotone channel reproduces the geometrically weighted separable family of hand-crafted statistics, one channel per member, to Spearman $\rho \ge 0.996$, approximates window statistics with measurable ceilings, and fails at consecutivity ($\rho = 0.924$) and time localization (0.628), both structural, and at the exposure floor (0.829), a learnability boundary. One explicit admissible basis repairs each failure (rank correlation 1.000). In or near the separable family, learned and engineered aggregation are substitutes, and the learned channel is never statistically behind at full sample size and specified capacity. Its advantages are incumbent-specific: a committed grid pays up to 0.019 AUC in decay regions it leaves uncovered (the learned channel stays within 0.004 of the strongest engineered consumer at every swept point); the highest-dimensional comparator degrades fastest with scarce data; and beyond the training support, grid-fed tree-ensemble scores go flat while a strictly increasing head keeps ranking. No single incumbent is dominated on all three axes.
|
| 996 |
Cross-Time Directional Selection in Diffusion Sampling
2610.05199
|
cs.LG
|
Dhia naouali |
How strongly do the remaining diffusion-sampling steps amplify a perturbation at a late latent state? Standard measurements answer this question with newly sampled isotropic noise, even though perturbations encountered during sampling have already been transfo...How strongly do the remaining diffusion-sampling steps amplify a perturbation at a late latent state? Standard measurements answer this question with newly sampled isotropic noise, even though perturbations encountered during sampling have already been transformed by earlier steps. We compare these two cases directly. For each trajectory, we transport a centered perturbation from an earlier step to a late state, then replay its direction at the same magnitude as a newly sampled isotropic perturbation; both then undergo the same remaining updates. Across the samplers we study, median paired shaped-to-fresh angular-gain ratios range from $1.14$ to $2.35$. Shaped angular gain exceeds its matched fresh counterpart in every trajectory in the original main cohorts. The same pattern appears when endpoint change is measured by latent RMS. The effect also appears in perceptual feature representations: earlier-sampling directions cause larger feature changes at the endpoint, even when the corresponding pixel-space change is comparable. Permuting shaped directions across trajectories weakens the effect, including within class, indicating that the advantage depends on alignment with the receiving trajectory as well as on shared directional structure.
|
| 997 |
Learning without Overwriting: A Theory of Self-Distillation and Supervised Fine-Tuning in Continual Reasoning
2610.05200
|
cs.LG
|
Shinichi Uemura, Taiji Suzuki |
On-policy self-distillation (OPSD) of large language models (LLMs) has demonstrated the ability to improve reasoning capabilities while preserving previously acquired knowledge. Despite substantial empirical success, the dynamics of OPSD in continual reasoning...On-policy self-distillation (OPSD) of large language models (LLMs) has demonstrated the ability to improve reasoning capabilities while preserving previously acquired knowledge. Despite substantial empirical success, the dynamics of OPSD in continual reasoning remain incompletely understood. Modeling LLM reasoning as search over a directed acyclic graph, we provide a unified theoretical analysis of both the dynamics of post-training---OPSD and supervised fine-tuning (SFT) in continual learning---and the impact of pre-training on subsequent performance. Our findings establish three key insights with an optimization guarantee: (i) OPSD with hints from correct outputs enables continual learning without forgetting by sparse yet effective gradient descent updates induced by the hint structure. (ii) SFT on correct reasoning paths can lead to catastrophic forgetting due to dense updates along the training paths, which overwrite the information previously acquired. (iii) Diversity in pre-training is crucial for enabling a post-trained model to reach a correct output when a rollout starts from an intermediate state. Our results, supported by theoretical analysis, show that reliable continual reasoning depends on how post-training updates interact with the reasoning structure established during pre-training.
|
| 998 |
Ranking Bandits for Carousel Interfaces with Observable Browsing Depth
2610.05220
|
cs.LG
|
Takuma Yasuda, Atsuyoshi Nakamura |
Carousel interfaces allow a recommender system to directly observe how far a user has browsed. This signal distinguishes displayed but unclicked items from items that were never displayed, whereas conventional ranking-bandit models, including cascade and posit...Carousel interfaces allow a recommender system to directly observe how far a user has browsed. This signal distinguishes displayed but unclicked items from items that were never displayed, whereas conventional ranking-bandit models, including cascade and position-based models, generally treat examination as latent. We formulate a ranking-bandit problem in which a learner presents a list of $L$ items, observes the user's maximum browsing depth, and receives click feedback only for positions up to that depth. The objective is to maximize the expected number of clicks under an unknown item-attractiveness vector and a browsing-depth distribution. We propose three algorithms based on UCB, Thompson Sampling, and DMED, all of which update item statistics only from observed exposures. We derive an instance-dependent logarithmic upper bound for our UCB-based algorithm and an asymptotic upper bound for our DMED-based algorithm that coincides with the lower bound as its parameter $\alpha\downarrow 0$, establishing asymptotic optimality in this limit. Simulations in synthetic shallow- and deep-browsing environments, together with experiments parameterized from RecGaze interaction logs, show that OD-TS attains final mean regret similar to PBM-TS, while the proposed methods achieve lower final mean regret than PBM-UCB.
|
| 999 |
EMG-GPT: Predictive Pretraining on Residual-Quantized EMG Tokens for Hand Pose Estimation
2610.05235
|
cs.LGcs.AI
|
Ettore Magni, Rolandos Alexandros Potamias, Stefanos Zafeiriou, Konstantinos Barmpas |
Surface electromyography (sEMG) is a low-power, cost-effective biosignal for hand-pose estimation and gesture classification. In this work, we examine whether self-supervised pretraining on sEMG can yield transferable representations for continuous hand-pose e...Surface electromyography (sEMG) is a low-power, cost-effective biosignal for hand-pose estimation and gesture classification. In this work, we examine whether self-supervised pretraining on sEMG can yield transferable representations for continuous hand-pose estimation. We introduce EMG-GPT, a causal transformer-based model that operates on discrete sEMG representations from a frozen residual vector quantization (RVQ) tokenizer and learns temporal dynamics through depth-autoregressive future-code prediction. The model combines within-frame integration with causal temporal modeling while preserving the geometry of the pretrained codebook. EMG-GPT shows competitive results in both Regression and Tracking tasks, supporting EMG-only pretraining as a viable approach for learning transferable sEMG representations.
|
| 1000 |
Pythia: Toward Foundation World Models for Multimodal Time Series
2610.05240
|
cs.LGcs.AI
|
Xilin Dai, Hongzhou Chen, Yifan Hu, Yiding Liu, Zewei Dong |
Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain ...Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model that learns context-conditioned latent dynamics across datasets through a joint-embedding predictive architecture. A stop-gradient numerical reference guides contextual corrections to predicted future states. A separate probabilistic decoder then adapts to the frozen predictive representation and observed history, decoupling world-model pretraining from observation-space forecasting. On MUSE, Pythia-Tiny's normalized mean absolute scaled error (MASE) and weighted sum quantile loss (WSQL) are 0.6879 and 0.4269, reducing errors by 6.26% and 5.00% relative to the strongest model evaluated in the published MUSE leaderboard. Through a series of controlled experiments, we investigate how to design a time-series world model through shared pretraining and how joint-embedding predictive learning can incorporate multimodal information. The results support separating predictive representation learning from probabilistic readout and show complementary contributions from entity descriptions, events, and covariates.
|
| 1001 |
Kolmogorov-Arnold Networks for Personal Context Recognition on ExtraSensory
2610.05250
|
cs.LG
|
Hoang-Thang Ta |
Kolmogorov--Arnold Networks (KANs) have attracted increasing attention in recent years, with applications across a wide range of AI tasks. In this paper, we evaluate several KAN variants on the ExtraSensory dataset for personal context recognition and compare ...Kolmogorov--Arnold Networks (KANs) have attracted increasing attention in recent years, with applications across a wide range of AI tasks. In this paper, we evaluate several KAN variants on the ExtraSensory dataset for personal context recognition and compare them with a multilayer perceptron (MLP) and TabM. We conduct the main experiments using five user folds and three random seeds per fold and report the average Macro-F1, Micro-F1, and training time. We also perform shallow ablation studies on grid size, the number of grids, and data normalization to examine their effects on KAN performance. The results show that all evaluated KAN variants significantly outperform MLP in terms of Macro-F1 and Micro-F1 and achieve performance comparable to TabM. However, KAN variants generally require more training time, while TabM provides a more favorable balance between predictive performance and training efficiency. These results suggest that KANs are promising for personal context recognition, while their computational efficiency remains an important challenge. Our source code is publicly available at: https://github.com/hoangthangta/ExtraSensory-KANs.
|
| 1002 |
Loopy: Low-Bit Quantization Framework for Looped Language Models
2610.05265
|
cs.LG
|
Zeyu LI, Yipu ZHANG, Jintao Chen, Xin LI, Wei ZHANG |
Looped language models provide a parameter-efficient way to scale iterative test-time computation by repeatedly executing a shared recurrent core. Post-training quantization (PTQ) can reduce the memory footprint and inference cost of looped language models, bu...Looped language models provide a parameter-efficient way to scale iterative test-time computation by repeatedly executing a shared recurrent core. Post-training quantization (PTQ) can reduce the memory footprint and inference cost of looped language models, but errors introduced by a quantized shared core affect subsequent cores. Among PTQ methods, channel scaling and orthogonal rotations preserve the floating-point computation while producing representations with different quantization quality. We find that quantization configuration candidate rankings can change with recurrent depth, motivating configuration selection at the target deployment depth. However, evaluating every candidate over the full calibration set at this depth is costly. We therefore propose Loopy, a PTQ framework that formulates shared-core quantization through a recurrent-depth-aware objective, selecting shared low-bit representations by their final prediction loss at the target deployment depth. Channel scaling and orthogonal rotations parameterize the candidate representations. To approximately solve this selection problem efficiently, Loopy progressively allocates calibration windows to promising candidates while preserving complete target-depth execution, using only forward evaluations. Across eight settings, Loopy achieves the state-of-the-art results among different baselines. On Ouro-1.4B under W4A4, Loopy reduces LAMBADA perplexity by 36.5% relative to SpinQuant. Our code is available at https://github.com/Shameless0817/Loopy-review.git.
|
| 1003 |
A Unified Scaling Law for Time Series Foundation Models
2610.05269
|
cs.LG
|
Xilin Dai, Yiding Liu, Zewei Dong, Jiang-Ming Yang, Qiang Xu |
We develop a Unified Scaling Law and a Unified Theory of Time Series Learning to understand how model capacity and historical information support forecasting. Across different lookback lengths and forecast horizons, we analyze 18,768 experimental cells from 21...We develop a Unified Scaling Law and a Unified Theory of Time Series Learning to understand how model capacity and historical information support forecasting. Across different lookback lengths and forecast horizons, we analyze 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks spanning six domains. Our empirical methodology integrates local resource relations into a parsimonious, fitted five-parameter law: capacity gains increase with history, context gains diminish toward saturation, and horizon effects enter as a common shift. Fitted without Toto 2.0, the law predicts its horizon-averaged capacity-scaling curves with mean absolute percentage errors of 1.09% and 1.50% at input lengths 2048 and 4096. To understand how history supports prediction, our learning theory uses Gaussian regression to analyze rule identification and predictive capability. We hypothesize that full-shot models learn by accumulating information in weights, while frozen time series foundation models (TSFMs) use history by extracting information through activations. Matched-history comparisons establish the predictive value of additional history. Controlled parameter exchanges and activation interventions provide evidence that history-derived rule information can be retained, reused across queries, and used to recover a contribution to long-context prediction. Together, these findings inform capacity scaling, context allocation, and the development of models that retain and apply historical rules. Code and main results are available at https://github.com/Fifthky/UniScale.
|
| 1004 |
Fast Convergence through Distributed Augmentation for Class-Imbalanced Federated Learning
2610.05279
|
cs.LG
|
Arathi Nair M, J. Harshan, Anwitaman Datta |
In federated learning, mitigating class imbalance is essential to improve minority-class performance. A common approach to address this problem is to augment minority-class samples to achieve local class balance. Existing approaches treat augmentation as a heu...In federated learning, mitigating class imbalance is essential to improve minority-class performance. A common approach to address this problem is to augment minority-class samples to achieve local class balance. Existing approaches treat augmentation as a heuristic and do not establish how the amount of augmentation influences the convergence of federated learning, leading to excessive augmentation and increased training time. To address this limitation, we first establish the relationship between augmentation and the convergence behavior of federated learning. Leveraging this insight, we propose DAFL, a distributed augmentation framework that determines the minimum augmentation required for each client-class pair by jointly minimizing augmentation and training time while constraining global class imbalance, thereby improving minority-class F1-score. Experimental results demonstrate that DAFL consistently improves minority-class F1-score while substantially reducing training time, particularly under severe global class imbalance and high label proportion imbalance.
|
| 1005 |
Smoothed Gradient Method for Nonconvex Federated Stochastic Bilevel Optimization
2610.05290
|
cs.LG
|
Xinwen Zhang, Peiran Yu, Zhaosong Lu, Hongchang Gao |
In recent years, federated stochastic bilevel optimization has attracted increasing attention due to its wide range of applications in machine learning. To reduce the computational overhead associated with second-order Hessian and Jacobian matrices, several fi...In recent years, federated stochastic bilevel optimization has attracted increasing attention due to its wide range of applications in machine learning. To reduce the computational overhead associated with second-order Hessian and Jacobian matrices, several first-order methods have been proposed. However, existing methods typically impose restrictive assumptions on the lower-level function, suffer from a strong dependence on the condition number in their convergence rates, and require different learning-rate scales for variables across the upper- and lower-level problems, limiting their practical applicability and complicating hyperparameter tuning. To address these challenges, we propose a stochastic doubly smoothed gradient method for nonconvex federated stochastic bilevel optimization problems, which decouples the learning rates of upper- and lower-level variables and does not require a strongly-convex lower-level loss function. We establish rigorous theoretical guarantees for the proposed algorithm, demonstrating an improved convergence rate of $O(\kappa^{15/2}/\epsilon^5)$ and a communication complexity of $O(\kappa^{4}/\epsilon^3)$, where $\kappa$ denotes the condition number and $\epsilon$ represents the solution accuracy. Notably, these bounds exhibit significantly better dependence on the condition number $\kappa$ than those of existing methods. Extensive experiments validate the effectiveness of our algorithm.
|
| 1006 |
Erased, Rerouted, or Rescaled? Post-Training and the Causal Quotient of a Language Model's Belief State
2610.05292
|
cs.LG
|
Weihan Li, Tianshi Zheng, Junhao Wu, Xinlei Chen |
What happens to information a pretrained model already encodes when post-training no longer rewards using it? The common language of representation compression conflates three fates: information may be erased, rerouted away from the decision while still repres...What happens to information a pretrained model already encodes when post-training no longer rewards using it? The common language of representation compression conflates three fates: information may be erased, rerouted away from the decision while still represented, or rescaled to occupy less variance while still represented and used. We make these fates identifiable in models whose pretraining recovers Bayesian belief states. A reward that reads only a coarse function of the hidden state defines an exact reward-null kernel. The kernel lets us separately measure whether the information remains recoverable, whether decisions causally depend on it, and how much activation variance it occupies. Theory says what is protected: KL-anchored reinforcement learning preserves the reference policy's log-odds among equally rewarded outputs, supervised and unanchored objectives carry no such constraint, and spectral compression implies neither erasure nor loss of use. In controlled worlds, post-training mostly reroutes or rescales reward-null information and leaves it decodable. Without an anchor decisions can stop using it although the representation survives, and with one they keep using it. Erasure appears only under prolonged weight decay, for distinctions that neither reward nor next-token prediction can see. Open language models show the same dissociation: in-context belief geometry stays decodable under late-layer spectral compression, and within-class behavior depends on the anchor. Post-training thus selects a causal quotient of the pretrained belief state: the reward defines decision-equivalence, the anchor and the state update protect part of what it ignores, and optimization decides whether the rest is erased, rerouted, or rescaled.
|
| 1007 |
FlexCast: Adaptive Weather Forecasting from Arbitrary Field Sets
2610.05296
|
cs.LG
|
Yuang Zhang, Chen Hui, Weisi Lin, Haiqi Zhu, Xiulai Wang |
Most deep learning weather models assign a fixed set of variables and pressure levels to predefined channels, limiting transfer across atmospheric field configurations. This dependence on a fixed field set limits the transferability of trained models across at...Most deep learning weather models assign a fixed set of variables and pressure levels to predefined channels, limiting transfer across atmospheric field configurations. This dependence on a fixed field set limits the transferability of trained models across atmospheric field configurations. We propose FlexCast, a field-adaptive weather forecasting model that uses a single set of parameters to produce identity-aligned forecasts for variable-cardinality subsets drawn from a 69-field ERA5 registry. Specifically, a metadata-conditioned adapter the first encodes variable identity, pressure level, and field type and combines them with spatial features. Then, shared rank-16 projec?tions are modulated by metadata-dependent gates to produce field?specific features, while masked set fusion aggregates the available fields into a fixed-width representation. Subsequently, a multiscale U-Transformer processes the fused atmospheric features, while an identity-aware query decoder produces forecasts for the requested fields. Finally, FlexCast learns a standardized six-hour increment and applies it recursively to generate forecasts at longer lead times. Experiments on the 2020 ERA5 test set demonstrate that FlexCast operates across varying field configurations. Compatible cross-field context is associated with lower forecast errors, whereas mismatched context increases them.
|
| 1008 |
Do Neural Networks Learn Structure-Preserving Maps? A Case Study in Latent-to-Hilbert Embeddings
2610.05297
|
cs.LG
|
Muhammad Adnan Shahzad |
We ask whether a neural network can learn a structure-preserving map from a compressed latent space to a Hilbert-space representation. Using an 8-dimensional autoencoder bottleneck on MNIST and $n$-qubit product-state targets from PCA-based angle encoding, we ...We ask whether a neural network can learn a structure-preserving map from a compressed latent space to a Hilbert-space representation. Using an 8-dimensional autoencoder bottleneck on MNIST and $n$-qubit product-state targets from PCA-based angle encoding, we report four findings. Although the target angles are generated by a nonlinear sigmoid transformation of the latent projections, the resulting mapping is well approximated by a linear function over the observed latent distribution: linear regression from $z$ to the true target angles achieves $R^{2} = 0.98$, while regression to the MLP's recovered angles achieves $R^{2} = 0.91$. The learned map's primary direction is strongly aligned with the target-induced direction, with cosine similarity $0.989$, while remaining nearly orthogonal to the input's principal direction, with cosine similarity $0.002$. The map is genuinely rank-4: removing any singular direction degrades inner-product preservation by $2.5$--$5.9\times$ despite a singular-value spectrum with two dominant and two small values. The learned subspace does not coincide with the PCA basis used to construct the target, and different random seeds recover the same primary direction but diverge in higher ranks. Finally, kernel ridge regression with an RBF kernel outperforms a tuned MLP (IP error $0.0144$ vs.\ $0.0197$), suggesting that for approximately linear structure-preserving mappings, classical kernel methods may be a simpler and more effective alternative.
|
| 1009 |
ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
2610.05303
|
cs.LG
|
Haodong Lu, Dong Gong |
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-conte...A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
|
| 1010 |
Robust Parameter-Efficient LLM Adaptation on Analog Hardware
2610.05318
|
cs.LG
|
Jindan Li, Zhaoxian Wu, Tianyi Chen |
Analog in-memory computing is a promising platform for on-device execution of large language models because it performs matrix--vector multiplications (MVMs) in memory and in parallel, reducing data movement. However, limited digital-to-analog converter precis...Analog in-memory computing is a promising platform for on-device execution of large language models because it performs matrix--vector multiplications (MVMs) in memory and in parallel, reducing data movement. However, limited digital-to-analog converter precision, input noise, and finite conductance states can degrade model accuracy, while full-model retraining to address these effects can be costly. We develop an optimizer-agnostic, parameter-efficient adaptation method based on Low-Rank Adaptation (LoRA), keeping the pretrained weights stored on analog arrays fixed while training the LoRA weights to adapt to downstream tasks and hardware non-idealities. Reliable adaptation requires handling errors in both forward and backward MVMs and physical weight updates. We use input reshaping to reduce input-induced MVM errors and update accumulation to retain small updates before programming them to finite-state analog devices. Across Llama-3.2-1B-Instruct and Llama-3-8B with both Muon and AdamW, input reshaping improves analog LoRA fine-tuning under noisy MVM computation. Update accumulation separately preserves sub-threshold updates and substantially improves adaptation under finite-resolution programming, including configurations with as few as 20 conductance states. Additional experiments show consistent held-out negative log-likelihood improvements across noisy analog settings.
|
| 1011 |
Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization
2610.05329
|
cs.LG
|
Hanzhang Wang, Tianqi Shen, Zonglin Liu, Junze He, Difan Zou |
Large language models (LLMs) are typically pretrained in high precision but increasingly deployed with low-precision post-training quantization (PTQ). Recent studies have shown that using weight averaging during pretraining can improve PTQ performance compared...Large language models (LLMs) are typically pretrained in high precision but increasingly deployed with low-precision post-training quantization (PTQ). Recent studies have shown that using weight averaging during pretraining can improve PTQ performance compared with learning-rate decay, suggesting that it might provide a simple way to improve the pretraining-to-quantization transition. But the mechanism behind weight averaging remains insufficiently explained. This leads to inconsistent and fragile performance gains, thereby preventing practitioners from applying such a technique confidently. As a response, we formulate weight averaging as a trade-off between retaining training progress and improving robustness under perturbation. We further derive a continuous family of averaging kernels that unifies conventional strategies and achieves the Pareto frontier between the two competing goals. Critically, a theoretical framework for performing weight averaging under PTQ is developed. It can be shown that coarser quantization is more susceptible to perturbations, whereas finer quantization could be less affected. Thus, our results could provide unified theoretical guidance for performing weight averaging under different PTQ conditions. Experiments validate both the predicted behavior and the proposed averaging strategy. Code is available at https://github.com/MOFA-LAB/weight-averaging-for-ptq.
|
| 1012 |
Compact set-valued deep ensembling in multi-class classification
2610.05332
|
cs.LG
|
Kim-Dung Tran, Dang-Man Nguyen, Vu-Linh Nguyen, Xuan-Truong Hoang, S\'ebastien Destercke |
This paper tackles visible challenges in deep ensemble learning, where deep neural networks serve as ensemble members: training and storage burdens, and robustness of cautious (set-valued) predictions targeting multiple utilities, which may involve reward-sens...This paper tackles visible challenges in deep ensemble learning, where deep neural networks serve as ensemble members: training and storage burdens, and robustness of cautious (set-valued) predictions targeting multiple utilities, which may involve reward-sensitivity. To mitigate the training and storage burdens, we propose to employ compact ensembles, such as Bayesian Neural Networks and Convolutional Neural Networks with the Monte-Carlo dropout prediction option, to produce probabilistic predictions. For each query instance, these probabilistic predictions are then used to define a representative distribution optimizing some statistical distance. The representative distribution is then employed to define the Bayes-optimal prediction (BOP) of any utility. To address the potential unrobustness of singleton prediction making, we propose a family of set-utilities satisfying some desirable properties and whose set-valued BOPs can be found efficiently. Empirical evidence is then given to illustrate the potential (dis)advantages of the proposed ensemble learning framework.
|
| 1013 |
Diffusion Transformers are Provably Optimal In-context Generators
2610.05333
|
cs.LG
|
Guoji Fu, Tomoya Wakayama, Ryotaro Kawata, Atsushi Nitanda, Wee Sun Lee |
Generative foundation models are attracting interest for their ability to produce desired outputs from demonstrations given at inference time, without updating parameters. However, since a few demonstrations cannot uniquely identify the intended task, the chal...Generative foundation models are attracting interest for their ability to produce desired outputs from demonstrations given at inference time, without updating parameters. However, since a few demonstrations cannot uniquely identify the intended task, the challenge is how to learn and sample from an output distribution that reflects this task uncertainty. In this work, we theoretically analyze how a Diffusion Transformer (DiT), pretrained across diverse tasks, learns and generates predictive distributions for a new query from demonstrations. We first show that the natural target to generate from finite demonstrations is not an output derived from estimating a single task, but rather a predictive distribution that captures the task uncertainty remaining after observing the demonstrations. We then prove that a DiT can learn this predictive distribution through score estimation, using attention to aggregate information from demonstrations and diffusion to generate samples. Owing to this property, with sufficient pretraining resources and diffusion sampling steps, the resulting DiT achieves the minimax optimal rate over a H\"older class of test-time tasks. These results imply that DiT acts as a statistically grounded in-context generator capable of generating distributions adapted to new tasks while retaining the uncertainty inherent in finite demonstrations.
|
| 1014 |
Green-Routed Neural Operators:\\Physics Determines Where the Network Reads
2610.05337
|
cs.LGcs.AI
|
Chenhao Si, Ming Yan |
We identify a mismatch between the physical role of transport fields in many PDEs and their usual role in neural operators: PDEs use them to select read coordinates, whereas neural operators typically treat them only as input values. We address this mismatch w...We identify a mismatch between the physical role of transport fields in many PDEs and their usual role in neural operators: PDEs use them to select read coordinates, whereas neural operators typically treat them only as input values. We address this mismatch with the Green-Routed Neural Operator (GRNO), which uses the governing equation to determine where latent features are sampled. A parameter-free equation adapter evaluates the diagnostic relation and constructs a departure map whose values are the read coordinates. A multiscale encoder-decoder combines centered and routed reads of latent features to learn the complete finite-time update. Across five two- and three-dimensional PDE systems, GRNO achieves the lowest mean final relative $L^2$ error on four under 40-step autoregressive evaluation and remains competitive on Keller-Segel. Fixed-weight route interventions reveal strong dependence on direction and spatial alignment in four systems, with weak dependence in Keller-Segel. In independently trained ablations, GRNO achieves lower mean errors than variants that supply the transport field only as an input feature, substitute a learned displacement for the equation-specified route, or apply the route with a spatial misalignment, across all five systems. It also substantially outperforms directly advecting the physical state and learning the remaining update, indicating that equation-specified read coordinates provide an effective structural prior for long-horizon PDE forecasting.
|
| 1015 |
On Semi-Markov Suboptimality in Hierarchical Reinforcement Learning
2610.05338
|
cs.LGcs.AI
|
Bingyun Liu, Yuheng Jing |
Hierarchical reinforcement learning uses temporally extended subtasks for exploration, yet committing to their execution can restrict both deployment and policy learning. We identify and separate the resulting execution and policy suboptimality. Task and execu...Hierarchical reinforcement learning uses temporally extended subtasks for exploration, yet committing to their execution can restrict both deployment and policy learning. We identify and separate the resulting execution and policy suboptimality. Task and execution trees distinguish reward objectives from policy choices and decision interruption. A Unified Value Function for HRL and a four-stage Generalized Hierarchical Bellman Equation then support a common analysis of both losses. Under bounded rewards and uniform termination, we establish hierarchical policy and execution improvement results. With the remaining node policies fixed, task-subtree compatibility and node-policy optimality under the original execution mode establish when Markov execution is optimal. The resulting decomposition leads to independent execution choices for behavior, targets, and deployment. We instantiate this principle through execution improvement and one-stage or two-stage policy improvement at arbitrary hierarchy depth. Option-based and goal-conditioned experiments demonstrate complementary gains from changing execution and changing the learning target. Controlled stochastic environments show how these gains depend on stochastic transition strength and spatial structure. This framework makes execution design an explicit component of hierarchical policy optimization.
|
| 1016 |
Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks
2610.05346
|
cs.LGcs.AI
|
Emanuele La Malfa, Saar Cohen, Gabriele La Malfa, Mickel Liu, Christian Schroeder de Witt |
Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isola...Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show that no fixed bounded window of recent prompts is sufficient in general: safety-relevant information may occur arbitrarily far back in the interaction. We formalize a watchman, an online state mechanism that carries this information forward, and show that under explicit assumptions it enables zero-failure defense with positive benign helpfulness. Under stronger conditions, it is also optimal among zero-failure defenders. An exact watchman may nevertheless require exponentially many states, while exact maliciousness detection can require exponentially many queries in an unstructured black-box model. These state and query lower bounds do not by themselves imply hard learning: the construction underlying the state lower bound is efficiently learnable from labeled examples, whereas certifying worst-case safety can require substantially more information under restricted access. We also show that self-play equilibrium alone does not certify usefulness, motivating a constrained formulation that maximizes worst-case benign helpfulness among zero-failure defenders. Empirically, training role-specific attacker and defender LoRA adapters over frozen LLMs via multi-turn self-play strengthens both roles: attackers become more effective at eliciting harmful responses, while defenders become more robust to attack, with improvements also observed on unseen attack objectives.
|
| 1017 |
Task Inference Beyond Least Squares in Behavioral Foundation Models
2610.05350
|
cs.LG
|
Kuan-Hsun Tu, Chien-Sheng Chiang, Hsin-Wei Chen, Ping-Chun Hsieh, Tsung-Wei Ke |
Behavioral Foundation Models (BFMs) aim to solve a wide range of downstream tasks without test-time policy learning by inferring a task vector from the reward function. While efficient, the retrieved policies are often suboptimal because of how this task vecto...Behavioral Foundation Models (BFMs) aim to solve a wide range of downstream tasks without test-time policy learning by inferring a task vector from the reward function. While efficient, the retrieved policies are often suboptimal because of how this task vector is inferred, typically with ordinary least squares (OLS). OLS minimizes reward reconstruction error but leaves the ordering of rewards unconstrained, which can bias the successor measure of the retrieved zero-shot policy away from that of the optimal policy. In this work, we propose BLS, an efficient test-time inference method that balances minimizing reward reconstruction error with reducing successor-measure mismatch. Theoretically, we provide a suboptimality gap upper bound characterized by both successor-measure and reward-function residuals. Empirically, we evaluate BLS on top of state-of-the-art BFMs across benchmarks for locomotion, manipulation, and humanoid control. BLS outperforms existing task inference baselines with negligible computational overhead. Project page: https://embodiedai-ntu.github.io/BLS
|
| 1018 |
Rethinking Tabular Foundation Models On Data Streams
2610.05352
|
cs.LGcs.AI
|
Nilesh Verma, Daniel Nowak-Assis, Afonso Louren\c{c}o, Albert Bifet, Bernhard Pfahringer |
Tabular foundation models (TFMs) outperform established machine learning models on tabular benchmarks through in-context learning. Building on this success, interest is growing in applying them to data streams, where data arrive continuously and evolve over ti...Tabular foundation models (TFMs) outperform established machine learning models on tabular benchmarks through in-context learning. Building on this success, interest is growing in applying them to data streams, where data arrive continuously and evolve over time. On a stream, a TFM adapts by updating its context rather than its parameters, so its accuracy and cost depend on which examples it keeps and how often it rebuilds its context. We therefore present a systematic study of TFMs on data streams, covering memory management, computational cost, and stream-specific challenges such as concept drift and delayed labels. We find that TFMs achieve the highest predictive performance and that simply retaining the most recent examples is as effective as existing memory management techniques. They also recover faster than streaming learners after drift and keep the highest accuracy under label delay. This accuracy, however, comes at a high serving cost, since a nearly unchanged context is re-encoded at every prediction. These results point to architectural efficiency as the way forward for in-context stream learning.
|
| 1019 |
The Effect of Missingness-Pattern Mismatch on Method Selection for Time-Series Classification: A Controlled Empirical Study
2610.05368
|
cs.LG
|
Ruiqi Zhao, Zishun Yuan, Zhentao Wang, Jiahao Quan, Kangzheng Li |
Classifiers for time-series classification are commonly selected on validation data, but the temporal pattern of missing observations at deployment may differ from the pattern seen during validation. We examine whether such a mismatch affects validation-based ...Classifiers for time-series classification are commonly selected on validation data, but the temporal pattern of missing observations at deployment may differ from the pattern seen during validation. We examine whether such a mismatch affects validation-based classifier selection. In a controlled $2 \times 2$ design, validation and test sets of 64 univariate UCR datasets were masked with either random point missingness or circular block missingness at six rates from 5% to 30%, imputed by linear interpolation, and used to select among three prespecified candidates: 1NN-DTW, MiniRocket with a Ridge classifier, and a statistical-feature Random Forest. Training data remained complete, and selections made under matched and mismatched validation patterns were compared on the same masked test sets. Mismatched validation reduced the test balanced accuracy of the selected classifier by 1.14 percentage points on average (95% CI 0.79 to 1.51), with losses on 49 of the 64 datasets. The loss was negligible at 5% missingness and increased to 2.46 percentage points at 30%. It was concentrated in point-masked deployment (1.84 percentage points), where block-masked validation shifted selection away from the usually best candidate, while the effect for block-masked deployment was small and not significant. Mismatch changed the selected classifier in 35.5% of paired comparisons, but a changed selection did not always reduce performance. A supplementary analysis with non-wrapping linear blocks reproduced these findings with a larger effect (1.67 percentage points). Matching the missingness pattern of validation data to the expected deployment pattern is therefore a simple safeguard for method selection, particularly at higher missingness rates.
|
| 1020 |
Robust Ensemble Guidance for Scientific Inverse Problems
2610.05371
|
cs.LG
|
Zixiang Li, Wei Wang, Yunchao Wei, Yao Zhao, Yue Song |
Ensemble guidance combines pretrained diffusion priors with black-box forward models to solve inverse problems without differentiating through the physical simulator. However, observation coordinates with large predictive spread or extreme residuals can domina...Ensemble guidance combines pretrained diffusion priors with black-box forward models to solve inverse problems without differentiating through the physical simulator. However, observation coordinates with large predictive spread or extreme residuals can dominate the ensemble correction, degrading reconstruction accuracy. We show that two simple modifications, weighting and clipping, substantially improve this correction. Our method, Robust Ensemble Guidance (REG), uses ensemble predictive spread to balance observation scales and adaptively clips standardized residuals to limit the influence of extreme discrepancies. Both operations reuse existing particles and forward predictions, requiring no additional denoiser or forward-model evaluations. Under a local linear Gaussian model, we derive conditions for reduced one-step estimation risk, bound the influence of individual observation coordinates, and characterize when these benefits persist with finite ensembles. Experiments on Navier-Stokes inversion, black-hole imaging, and acoustic full-waveform inversion demonstrate improved reconstruction over the underlying ensemble solver. In particular, REG increases black-hole reconstruction PSNR by 6.2-8.2 dB across three observation regimes and reduces Navier-Stokes reconstruction error by 26.4\% in a matched-budget comparison. These findings highlight the importance of observation heterogeneity and residual influence in designing reliable generative solvers for scientific inverse problems.
|
| 1021 |
Distributed Subliminal Learning: Replacing Model Updates with Random-Carrier Outputs
2610.05378
|
cs.LGcs.AI
|
Dario Fenoglio, Gabriele Dominici, Martin Gjoreski, Marc Langheinrich |
Collaborative learning typically exchanges model parameters: federated clients communicate updates, while independently adapted foundation models are combined by exchanging adapters or checkpoints. This makes communication scale with model size and requires lo...Collaborative learning typically exchanges model parameters: federated clients communicate updates, while independently adapted foundation models are combined by exchanging adapters or checkpoints. This makes communication scale with model size and requires local specializations to be reconciled in weight space, where interference is common. We ask whether knowledge can instead be shared through model behavior on task-unrelated inputs. We introduce Distributed Subliminal Learning (DSL), a collaborative learning primitive in which participants adapt a common model locally, probe it with task-unrelated inputs, and transmit only the resulting carrier outputs. A coordinator pools these outputs and distills them into a shared model. The primitive supports one-shot foundation-model composition through carrier completions and iterative federated learning through carrier logits, without transmitting model updates or requiring task-related proxy data. In LLM composition, compared with LoRA averaging, DSL achieves higher preference retention (94.56% vs. 87.76%) and a larger GSM8K gain over the base model (22.0 vs. 0.6 points), while reducing upload by 30.6-49.0$\times$. In federated classification, DSL reaches 96.83% on MNIST with 8.9$\times$ less uplink than FedAvg and provides lower-communication operating points on CIFAR-10 and Tiny ImageNet. These results establish random-carrier outputs as a practical communication primitive for knowledge sharing across distinct collaborative learning paradigms.
|
| 1022 |
Does Explainability Survive Data Drift?
2610.05379
|
cs.LGcs.AI
|
Samuel Ozechi |
Model performance monitoring is a standard practice in machine learning deployments. Detection performance is tracked continuously, and model decay is expected as the relationship between the feature and target variables degrades, a phenomenon known as concept...Model performance monitoring is a standard practice in machine learning deployments. Detection performance is tracked continuously, and model decay is expected as the relationship between the feature and target variables degrades, a phenomenon known as concept drift. Explanation fidelity, however, is rarely monitored with the same discipline, even in domains such as financial systems, healthcare, and other regulated environments where explanations are required for governance purposes. This paper investigates whether explanations can decay under data drift, even when the feature-target relationship remains stable, and whether explanations produced before drift occurs remain faithful to the decisions of the model that replaces them. Using the IEEE-CIS Transaction Fraud Detection dataset, we find statistically significant covariate shift but no statistically significant evidence of concept drift under the implemented conditional-drift tests, thereby providing an empirical setting in which input distributional change can be studied separately from detectable changes in the feature-target relationship. Local explanations are generated with ExIFFI and evaluated at three levels: path validity, structural behaviour, and fidelity under controlled intervention. Results show that while prior explanations retain substantial decision relevance to a retrained model, they are consistently less faithful than newly generated explanations, with no evidence of a systematically widening gap across the evaluated windows. The study shows that explanation fidelity requires its own monitoring, that structural stability of explanations does not guarantee functional fidelity, and that explanations should be treated as artifacts tied to the model that produced them.
|
| 1023 |
Efficient Graph Generation via Direct Prediction and Flow Matching
2610.05397
|
cs.LG
|
Susie Lu |
Generative modeling of graph-structured data is crucial for tasks ranging from drug discovery to social network simulation. Among these models, denoising diffusion models have achieved great success in graph generation by learning to progressively reverse a pr...Generative modeling of graph-structured data is crucial for tasks ranging from drug discovery to social network simulation. Among these models, denoising diffusion models have achieved great success in graph generation by learning to progressively reverse a process that adds noise to the original graph. However, the standard noise-prediction approach of diffusion models is suboptimal for graph data. The goal for a graph generative model is to learn the clean graphs' topological properties, such as connectivity and degree distribution. Because a diffusion model that predicts noise does not explicitly learn these topological properties, it is challenging for the model to output graphs with the desired structural statistics. To address this challenge, we introduce Direct Graph Flow Matching (DiGFM), a novel graph transformer model guided by two goals: predict clean graphs and improve sampling efficiency. Distinct from the prevailing diffusion approach, DiGFM employs a continuous flow-matching paradigm and integrates direct graph prediction. Specifically, DiGFM maps the prior noise distribution to the clean graph distribution via a multi-step process: the model repeatedly predicts the underlying clean graph, and a transformation is employed to convert the model output to the velocity vector that points in the direction toward the clean graph distribution. This design enables DiGFM to generate high-quality samples using only 2.5% to 15.6% of the steps required by diffusion-based models, which leads to a 5.3x to 257x speedup in wall-clock inference time. Experiments demonstrate that DiGFM outperforms or matches prior state-of-the-art models across general graph benchmarks and molecular datasets, generating graphs with strong adherence to ground-truth structural statistics at significantly faster inference speeds.
|
| 1024 |
No Concept Escapes the Audit: Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models
2610.05401
|
cs.LGcs.AI
|
Kaiyuan Deng, Yuchen Li, Gen Li, Yang Xiao, Geng Yuan |
Text-to-image diffusion models can generate prohibited content, which motivates concept erasure through machine unlearning. Most erasure methods intervene at the text interface, through prompt modification or localized updates to text-conditioning weights, and...Text-to-image diffusion models can generate prohibited content, which motivates concept erasure through machine unlearning. Most erasure methods intervene at the text interface, through prompt modification or localized updates to text-conditioning weights, and they are evaluated by what the model outputs for given prompts. Such evaluation cannot see what the network still encodes. Latent-space auditing, which bypasses text conditioning and probes the denoising network directly, shows that erased concepts remain recoverable from internal representations. We find that this also holds for methods built to be robust against adversarial prompts, and that the problem grows with the number of erased concepts. We propose Auditing-Aware Unlearning for Verifiable Concept Erasure in Diffusion Models (AVCE), a framework that grounds erasure in the model's latent representations. AVCE audits the embedding neighborhood of each concept and condenses the discovered vulnerable directions into an anchor at the weakest geometric point. It edits cross-attention and self-attention projections in closed form at this anchor, then fine-tunes the two pathways with pathway-level auditing losses, using orthogonal gradient projection to consolidate multiple concepts. Experiments on SD v1.5, SDXL, and Flux 1.0 across object, explicit-content, and artistic-style unlearning show that AVCE reduces attack success rates by 5.07x and improves auditing scores by 3.84x over the strongest baseline, while preserving competitive generation quality.
|
| 1025 |
BeliefGraph-JEPA: Structured Latent World Models for Action-Conditioned Time Series
2610.05409
|
cs.LGcs.AI
|
Yue Li, Kangqi Ni, Zhen Tan, Tianlong Chen |
Action-conditioned time-series forecasting requires accounting for how future actions and exogenous forcings influence multiple targets through partially observed effects with different delays and persistence. Direct conditioning leaves the evolution and targe...Action-conditioned time-series forecasting requires accounting for how future actions and exogenous forcings influence multiple targets through partially observed effects with different delays and persistence. Direct conditioning leaves the evolution and target-specific influence of these effects implicit in the predictor, while static relational graphs specify connections without tracking evolving effects. This motivates representing future-driver influence through structured latent states that evolve over the forecast horizon and route information to individual targets. We introduce BeliefGraph-JEPA, a structured latent world model that factorizes driver influence into typed latent-effect states. These states are rolled forward under future drivers and routed through a graph to target-specific nodes, forming the predictive base of a joint-embedding predictive architecture. A capacity-controlled residual supplements this base with direct driver information. On four multi-target clinical, agricultural, environmental, and industrial systems, the framework outperforms a range of pretrained and supervised known-future-covariate baselines. Matched controls isolate latent dynamics, future rollout, graph routing, and residual capacity; future rollout and graph-first residual routing improve forecasting across all four systems.
|
| 1026 |
Hierarchical Time-aware Bootstrapping for Off-Policy Subgoal Value Learning
2610.05446
|
cs.LGcs.AI
|
Bingyun Liu, Yuheng Jing |
Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal ins...Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data through subgoal relabeling, but after a label change, the value update targets the relabeled subgoal instead of the subgoal the high-level policy originally needed to update. We propose Hierarchical Time-aware Bootstrapping (HTB), which evaluates specified subgoals under the current low-level policy while retaining accumulated task rewards. Remaining execution time distinguishes subgoal continuation from a new high-level decision. Together with primitive-action conditioning, it enables off-policy Bellman updates based on the stationary environment transition law. HTB combines these one-step updates with multi-step suffix returns and truncated relabeling, reducing dependence on intermediate value estimates. A shared value component supports learning across actions, while nonnegative residuals constrain upward corrections relative to that component. At a fixed mixture weight of 0.95, HTB achieves 32.8% AntFall success versus 9.6% for matched local HIRO over five paired seeds at 10M environment steps. Ablations identify contributions from recursive continuation and mixed supervision; fixed-policy tests show more accurate predictions for actions whose returns were excluded from fitting.
|
| 1027 |
Groupwise Distortion Guarantees for Preference-Based Alignment
2610.05450
|
cs.LG
|
Jacob Brodkey, Roberto Tamez, Aaron Roth |
Preference-based alignment methods such as reinforcement learning from human feedback (RLHF) and Nash learning from human feedback (NLHF) aggregate pairwise preferences to learn an LLM policy, but a natural goal is maximizing social welfare (average cardinal u...Preference-based alignment methods such as reinforcement learning from human feedback (RLHF) and Nash learning from human feedback (NLHF) aggregate pairwise preferences to learn an LLM policy, but a natural goal is maximizing social welfare (average cardinal utility), which comparisons alone need not identify. G\"olz, Haghtalab, and Yang (GHY) measure the gap by distortion: the worst-case ratio between the welfare of the best fixed lottery (distribution over responses) and of the learned lottery. They show NLHF is optimal when every user receives the same lottery. Account-based LLMs, however, have information about their users and can serve different lotteries to different people. We give an efficient algorithm, GLHF, that learns a single group-conditioned policy from one comparison per user. Under individual Bradley--Terry comparisons, GLHF asymptotically matches GHY's optimal population distortion bound simultaneously on every group in a prespecified, possibly overlapping collection, with sample complexity growing logarithmically in the number of groups and inversely with the smallest group mass. A sharper guarantee for groups with similar preferences approaches distortion of one when members share a feasible favorite response. In experiments using human coffee ratings and synthetic LLM-generated ratings, GLHF lowers distortion in every evaluated group and substantially reduces worst-group distortion relative to NLHF and other group-agnostic baselines.
|
| 1028 |
VERA: Verdict-Conditioned Reliability for Adaptive LLM Judges
2610.05452
|
cs.LG
|
Qiushui Xu, Syamil Mohd Razak, Tao Yuan, Piotr Habas |
Accurately estimating judgment reliability is a central challenge in adapting LLM judges to newly verified feedback while preserving previously learned behavior. However, existing approaches often rely on output-level confidence, which can be overconfident and...Accurately estimating judgment reliability is a central challenge in adapting LLM judges to newly verified feedback while preserving previously learned behavior. However, existing approaches often rely on output-level confidence, which can be overconfident and poorly aligned with judgment correctness. We propose VERA, a VErdict-conditioned Reliability Axis that estimates reliability from hidden activations by distinguishing correct from incorrect judgments within each predicted-verdict group. Using VERA as a control signal, we develop a VERA-guided periodic adaptation framework that integrates reliability-ranked corrective updates, reliability-residual replay, and periodic refresh of the reliability directions. After VERA-guided adaptation on Chatbot Arena, 8B- and 14B-parameter judges outperform the strongest baseline on each of four held-out public benchmarks, with relative gains of up to 23.01%. The framework also improves focal-class recall by up to 16.1% relative to the strongest adaptive baselines on a separate proprietary temporal auditing task.
|
| 1029 |
Measuring and Reducing Cross-Vendor Mismatch in Language Models
2610.05458
|
cs.LG
|
Erland Hilman Fuadi, Chong Tian, Xiaosong Ma, Qirong Ho |
Running the same language model on different graphics processing unit (GPU) vendors can produce different logits, even when the model weights and inputs are the same. We analyze cross-vendor mismatch in two dense and two mixture-of-experts (MoE) models with fi...Running the same language model on different graphics processing unit (GPU) vendors can produce different logits, even when the model weights and inputs are the same. We analyze cross-vendor mismatch in two dense and two mixture-of-experts (MoE) models with five metric families, namely bitwise equality, logit differences, top-K consistency, token agreement, and task accuracy. We trace one source of the mismatch to accumulation order inside vendors' matrix instructions. Upcasting to FP32 reduces the dense model's logit error by 43% at three times the runtime, yet keeping only the MLPs in BF16 retains 94% of this gain at 1.3 times the runtime, so most of the cost of full upcasting buys little. In the MoE models, FP32 and FP16 both lower the probability error but raise the logit error and change expert selection, and FP16 fails in the dense model. An output-head low-rank adapter (LoRA) does not help either, since the final hidden state does not predict the mismatch. The mismatch also carries into training. With every seed fixed, a student distilled from a teacher running on AMD answers 431 MMLU questions differently from one distilled from the same teacher on NVIDIA. Under FP32 upcasting, bitwise equality barely changes while the output distributions move most of the way to the reference, so judging cross-vendor agreement by a single measure misreads both its cost and its gains. Code is available at https://github.com/crova-project/crova.
|
| 1030 |
Population Scaling or Data Dilution? Dynamics of Local Topology Evolution in Decentralized Learning
2610.05476
|
cs.LGcs.AI
|
Yin-Kuan Liang, Yan Gao, Yang Long |
Scaling decentralized learning changes not only the number of clients $N$, but also the dynamics of information propagation and consensus. We argue that the effect of increasing $N$ cannot be understood in isolation, because data allocation, topology-dependent...Scaling decentralized learning changes not only the number of clients $N$, but also the dynamics of information propagation and consensus. We argue that the effect of increasing $N$ cannot be understood in isolation, because data allocation, topology-dependent mixing, and communication capacity may change simultaneously. We study these coupled effects on CIFAR-10 with $N\in\{10,50,100,200\}$, comparing a degree-two Ring, a Static Random graph, and Local-First Heuristic Evolution (LFHE), a locally adaptive topology process based on friend-of-friend discovery. The Ring provides an analytically transparent failure mode: its Metropolis spectral gap decays as $\Theta(N^{-2})$, implying progressively slower contraction of model disagreement as the population grows. Experiments show that holding the nominal local dataset size fixed substantially reduces the apparent population penalty observed when a fixed total dataset is divided among more clients. The remaining degradation depends strongly on communication structure: Ring enters a high-disagreement regime, whereas Static Random and LFHE remain close to consensus. Increasing LFHE's degree threshold further improves accuracy and consensus, but at a substantially higher model-transmission cost. These results show that decentralized scaling is governed by coupled learning and communication dynamics, rather than by the number of clients alone.
|
| 1031 |
Underscoring the Problem: Why Softpick Fails at Initialization
2610.05488
|
cs.LG
|
Aryan Sood, Jaikaran Singh, Ishaan Bansal |
Softmax attention gives every token a nonzero weight, which in trained models concentrates into attention sinks and massive activations that widen the dynamic range low-precision inference must cover. Softpick removes this constraint by rectifying scores, elim...Softmax attention gives every token a nonzero weight, which in trained models concentrates into attention sinks and massive activations that widen the dynamic range low-precision inference must cover. Softpick removes this constraint by rectifying scores, eliminating sinks and lowering hidden-state kurtosis, but its advantage fades at scale. We reframe this failure as a normalization problem. Softpick's denominator splits into positive- and negative-shifted sums $D^+$ and $D^-$, used identically in the forward and backward pass, preventing their roles from being isolated. We separate them into a family of operators that independently choose each denominator. The failure originates at initialization: every layer contains rows where $D^+$ is exactly zero, while near-dead rows produce gradient norms above $10^{12}$ regardless of the backward denominator. Only Softpick and a stop-gradient variant, which keeps $D^+ + D^-$ forward but backpropagates through $D^+$ alone, train from scratch. At 230M parameters, the stop-gradient operator matches Softpick on quantization, has fewer dead heads, and retrieves passkeys more reliably, trailing only on peak attention-weight kurtosis.
|
| 1032 |
Universality and Convergence of Generative Flows
2610.05490
|
cs.LG
|
Leo Brunswic |
Generative flows sample from an unnormalized target by training a flow to be balanced, and the training loss is the signal a practitioner watches. We ask what that signal is worth: whether a small loss certifies an accurate sampler, whether the loss can be dri...Generative flows sample from an unnormalized target by training a flow to be balanced, and the training loss is the signal a practitioner watches. We ask what that signal is worth: whether a small loss certifies an accurate sampler, whether the loss can be driven to zero, and how fast gradient descent does so. The loss decides the first. Losses that compare the two sides of the balance by their difference bound, in total variation, the error of the sampler the flow implies, with explicit constants that do not involve the policy; flow-matching losses that compare them through a ratio admit no such bound, already on a single cycle, whenever their generator is continuous at balance. On graphs, the backward policy decides the other two. Once it is frozen, balance becomes invariance under the backward chain, so that existence is free on finite graphs, and one constant --- the norm of that chain's Green operator, which plays the role of an inverse spectral gap --- fixes the order of the curvature of the loss around the balanced flow, from above and below, and sets a floor under the rate at which training converges near it. The mechanism is that gradient descent diffuses the flow along the backward policy. For the squared-logarithm generator of detailed and trajectory balance, training the balance loss on states converges globally on every finite path-connected graph, from every positive initialization. The constant can be infinite while backward trajectories are short on average, and exact flow matching can then fail. The bounds and rates are tested by exact computation on enumerable state spaces, and every theorem carries a certification status computed from a Lean~4 development.
|
| 1033 |
When Does Retrieval Help? A Study of In-Context Adaptation in Vision-Language-Action Models
2610.05492
|
cs.LG
|
Zixuan Liu, Joris K\"oster, Zizhan Zheng, Siavash Khajavi |
Vision-language-action (VLA) models have shown strong potential as generalist robot policies, but adapting them to unseen tasks often requires costly parameter updates. Recent work such as RICL introduces in-context adaptability by retrieving expert demonstrat...Vision-language-action (VLA) models have shown strong potential as generalist robot policies, but adapting them to unseen tasks often requires costly parameter updates. Recent work such as RICL introduces in-context adaptability by retrieving expert demonstrations based on the current VLA observation and providing them as additional context at test time. The effectiveness of this adaptation therefore depends critically on the retrieval mechanism. In this work, we systematically study how different retrieval methods affect both retrieval quality and task performance within the RICL framework. Specifically, we compare four different methods: image-based retrieval, retrieval augmented with VLA's state, retrieval using features from the VLA backbone, and random retrieval. Our experiments yield three main findings. First, no retrieval method consistently dominates the others in task success, while surprisingly, random retrieval achieves a non-trivial success rate. Second, standard retrieval-quality diagnostics do not reliably reflect downstream VLA performance. Third, demonstrations from different but related tasks can provide useful transferable information. Together, these results provide an initial step toward understanding how retrieval mechanisms shape the in-context learning capability of VLA models and their downstream task performance, while highlighting the need for more careful design and evaluation of retrieval mechanisms for reliable test-time adaptation.
|
| 1034 |
LEON: Location Embeddings from OSM Neighborhoods via Hexagonal Graph Masked Autoencoders
2610.05497
|
cs.LG
|
Szymon Soltysiak, Radoslaw Malek, Jedrzej Kusnierz, Milosz Chojecki, Piotr Szymanski |
Geographic information systems increasingly rely on sophisticated spatial representation learning techniques to extract meaningful patterns from complex geospatial data. This paper introduces LEON, a novel self-supervised framework that adapts Graph Masked Aut...Geographic information systems increasingly rely on sophisticated spatial representation learning techniques to extract meaningful patterns from complex geospatial data. This paper introduces LEON, a novel self-supervised framework that adapts Graph Masked Autoencoders (GraphMAE) for geospatial region representation learning. Our method leverages the inherent spatial structure of geographic data by constructing hexagonal grid graphs using H3 indexing and applying masked autoencoding techniques to learn robust spatial embeddings from OpenStreetMap (OSM) amenity distribution patterns. We evaluate LEON on multiple real-world datasets including EuroSAT satellite imagery classification and various geographic prediction tasks (housing prices, crime prediction, and urban analytics). Experimental results demonstrate that LEON achieves significant improvements in spatial understanding, with up to 1.87% accuracy improvement on EuroSAT classification and consistent performance gains across geographic prediction benchmarks. The learned embeddings exhibit highly structured and distinct properties, making them particularly suitable for downstream spatial analysis tasks. Our findings suggest that self-supervised learning provides an effective paradigm for geospatial region representation learning using widely available crowdsourced data.
|
| 1035 |
Logic-Logit: A Logic-Based Approach to Choice Modeling
2610.05501
|
cs.LG
|
Shuhan Zhang, Wendi Ren, Shuang Li |
In this study, we propose a novel rule-based interpretable choice model, Logic-Logit, designed to effectively learn and explain human choices. Choice models have been widely applied across various domains---such as commercial demand forecasting, recommendation...In this study, we propose a novel rule-based interpretable choice model, Logic-Logit, designed to effectively learn and explain human choices. Choice models have been widely applied across various domains---such as commercial demand forecasting, recommendation systems, and consumer behavior analysis---typically categorized as parametric, nonparametric, or deep network-based. While recent innovations have favored neural network approaches for their computational power, these flexible models often involve large parameter sets and lack interpretability, limiting their effectiveness in contexts where transparency is essential. Previous empirical evidence shows that individuals usually use heuristic decision rules to form their consideration sets, from which they then choose. These rules are often represented as disjunctions of conjunctions (i.e., OR-of-ANDs). These rules-driven, consider-then-choose decision processes enable people to quickly screen numerous alternatives while reducing cognitive and search costs. Motivated by this insight, our approach leverages logic rules to elucidate human choices, providing a fresh perspective on preference modeling. We introduce a unique combination of column generation techniques and the Frank-Wolfe algorithm to facilitate efficient rule extraction for preference modeling---a process recognized as NP-hard. Our empirical evaluation, conducted on both synthetic datasets and real-world data from commercial and healthcare domains, demonstrates that Logic-Logit significantly outperforms baseline models in terms of interpretability and accuracy.
|
| 1036 |
Lightweight Semantic EEG Foundation Model for Frozen Cross-Disorder Transfer
2610.05503
|
cs.LG
|
Rita Huan-Ting Peng, Nhat Bui |
Large-scale EEG foundation models have demonstrated promising transferability across neurological disorders, but often require millions of parameters and substantial computational resources. In this paper, we present the Universal Semantic EEG Foundation Model...Large-scale EEG foundation models have demonstrated promising transferability across neurological disorders, but often require millions of parameters and substantial computational resources. In this paper, we present the Universal Semantic EEG Foundation Model (USE-FM), a lightweight EEG foundation model that learns transferable neural representations through self-supervised signal reconstruction on the Temple University Hospital EEG Corpus (TUEG). After pretraining, the encoder is frozen and evaluated on two clinically distinct downstream tasks, abnormal EEG detection (TUAB) and epileptic seizure recognition (TUEP), using a unified frozen-transfer protocol against recent EEG foundation models, including LUNA-Base and CBraMod. With only 1.46 million parameters, approximately one-fifth the size of existing models, USE-FM achieves competitive overall performance, including strong sensitivity and F1-score on TUEP (SEN $75.00 \pm 14.14$, F1 $70.37 \pm 4.01$), while maintaining competitive performance on TUAB (AUC $85.24 \pm 5.61$). Beyond downstream classification, latent representation analysis using $k$-means clustering together with PCA and t-SNE demonstrates that USE-FM learns organized semantic EEG representations comparable to substantially larger foundation models. These results suggest that large-scale self-supervised pretraining enables lightweight architectures to learn transferable semantic EEG representations, providing a computationally efficient foundation for cross-disorder analysis and future clinical decision support in neurological disorders.
|
| 1037 |
LiFT: Loop Flow Transformers
2610.05538
|
cs.LGcs.AI
|
Mohammad Mahdi Derakhshani, Pedro M. P. Curvo, Gertjan J. Burghouts, Jan-Willem van de Meent, Cees G. M. Snoek |
We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent ...We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target. Because we index these targets by a continuous depth coordinate, a trained model can loop far beyond its training depth with no retraining, early exits, or other modifications. In our experiments, these longer rollouts improve generation, so inference computation can grow without adding parameters. On ImageNet at 256x256, LiFT-L/2 achieves an FID 3.34 points lower than our dense DiT-XL/2 baseline while using approximately 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.
|
| 1038 |
Soft Strategy Selection for Batch-Mode Active Learning
2610.05544
|
cs.LGcs.AI
|
Rushil Gupta, Romain Lopez |
Real-world deployment of active learning typically forces practitioners to choose an acquisition strategy before any data is labeled. This is a daunting task: strategy performance varies widely across settings (e.g. datasets, surrogate models) and cannot be as...Real-world deployment of active learning typically forces practitioners to choose an acquisition strategy before any data is labeled. This is a daunting task: strategy performance varies widely across settings (e.g. datasets, surrogate models) and cannot be assessed without deployment. Existing strategy selection methods explore one strategy from a portfolio at each round and identify the optimal one using bandit feedback or model retraining. Many acquisition rounds are therefore spent exploring strategies rather than collecting the most informative data. Such overhead is a major barrier to AL-driven design of high-throughput experiments, such as genetic perturbation screens and directed evolution, where AL runs consist of only a few rounds with large batch sizes. This regime permits a natural alternative: acquiring data using multiple AL strategies within a single batch. We refer to this as soft strategy selection and introduce FractAL, a method specifically designed for this task. FractAL infers a per-strategy reward using influence-function-based data attribution, which requires no additional retraining, and then computes budget shares for each strategy in the portfolio using online mirror descent. We benchmark FractAL across 7 setups spanning classification, regression, and genetic perturbation effect prediction. The results highlight that strategy selection is a hard problem: every existing method performs worse than random sampling on at least one setup. FractAL, however, matches or outperforms every baseline, including random sampling, on all 7 setups. Its allocations concentrate budget on the strongest strategies in the portfolio while pruning the weakest. FractAL is therefore a reliable choice for real-world deployments, where the optimal strategy is unknown in advance, an important step towards making AL practical for high-throughput experiments and modern scientific discovery.
|
| 1039 |
Joint Estimation of Common-Slope Decay Rates and Spatial Amplitudes Using Parameterized Nonnegative Matrix Factorization
2610.05549
|
cs.LGeess.AS
|
Jeremy B. Bai, Filip Elvander, Sebastian J. Schlecht |
We formulate joint estimation of common-slope decay rates and amplitudes from room impulse responses (RIRs) as parameterized nonnegative matrix factorization with the Itakura--Saito divergence as the loss function (IS-NMF). Estimation at each short-time Fourie...We formulate joint estimation of common-slope decay rates and amplitudes from room impulse responses (RIRs) as parameterized nonnegative matrix factorization with the Itakura--Saito divergence as the loss function (IS-NMF). Estimation at each short-time Fourier transform frequency bin produces detailed reverberation time (RT) curves directly from RIR powers with no backward integration needed. Standard space-alternating generalized expectation-maximization (SAGE) algorithm yields closed-form amplitude updates and a convex subproblem for each decay rate update. To accelerate estimation, we introduce contribution-weighted SAGE, which emphasizes observations where each component contributes strongly to the modeled power. Experiments with synthetic data show accurate recovery of well-separated decays and faster loss reduction than standard SAGE. Application to measured coupled-room RIRs yields frequency-dependent RT curves and reveals complementary space-time contributions of the shared decay components.
|
| 1040 |
When Low Prediction Error Misleads Planning: Diagnosing Representation, Dynamics, and Decision Failures in Latent World Models
2610.05550
|
cs.LG
|
Rui Min, Xianyao Li, Fang Xu, Jing Du |
The component that dominates a latent world model's prediction error need not be the one whose repair most improves action selection. We show this by comparing action sequences from identical physical starts and separating endpoint error into a candidate-pool ...The component that dominates a latent world model's prediction error need not be the one whose repair most improves action selection. We show this by comparing action sequences from identical physical starts and separating endpoint error into a candidate-pool center and action-relative responses. Across four model families and four tasks, a confirmation pool of 256 new starts per task and 300 shared candidates per start shows that center error dominates MSE in 14/16 model-task cells. Yet in six of these cells, an oracle that corrects only the action-relative responses yields better physical rank correlation and top-30 elite quality than one that corrects only the center, while leaving more latent MSE (family-wise corrected intervals). The preference differs across the evaluated settings: a separate LeWorldModel (LeWM) study that executes oracle-selected actions favors center repair on PushT and on Reacher with a render-matched goal. Matched-candidate tests localize ordering loss: for LeWM, encoding realized endpoints raises physical Spearman from 0.464 to 0.975 on that Reacher setting and from 0.193 to 0.631 on PushT (64 starts per task), while Cube's encoded-goal cost remains uninformative. A 72-run objective study improves selected response diagnostics, while incremental closed-loop planning gains remain unconfirmed. These results separate error magnitude from the decision effects of oracle correction and motivate evaluating representation, prediction, and planning as separate stages.
|
| 1041 |
An equality condition for the Dobrushin bound on attention rollout and how often it holds in trained transformers
2610.05558
|
cs.LG
|
Przemys{\l}aw Rola |
The Dobrushin coefficient of each attention-rollout factor satisfies $\kappa(\frac12(I+A))\le\frac12(1+\kappa(A))$, and multiplying these inequalities over layers bounds the coefficient of the whole rollout. We characterise exactly when the layerwise bound is ...The Dobrushin coefficient of each attention-rollout factor satisfies $\kappa(\frac12(I+A))\le\frac12(1+\kappa(A))$, and multiplying these inequalities over layers bounds the coefficient of the whole rollout. We characterise exactly when the layerwise bound is tight: equality holds if and only if some token pair attaining $\kappa(A)$ is mutually self-dominant - each of the two attends to itself at least as strongly as the other attends to it. The condition is far from automatic: uniformly random stochastic matrices satisfy it only 24-30% of the time. When tested on the head-averaged attention of each individual input and restricted to content tokens - image patches, words or tabular features, excluding cls, register and separator tokens - the condition holds for essentially every input at every layer of DINOv2 (three model sizes), RoBERTa and DistilBERT. In the supervised models DeiT-B and ViT-B/16 it holds for 91% and 64% of input-layer pairs respectively, with all failures occurring late in depth. In FT-Transformer trained on two standard tabular benchmarks it holds for only 11-44% of input-layer pairs. The special tokens account for almost all failures in DINOv2 and the language models: when they are included, the condition holds for only 82-97% of input-layer pairs.
|
| 1042 |
Cut Binary Cross Entropy: Efficient Large-Vocabulary Loss and Gradient Kernels for Sequential Recommendation
2610.05559
|
cs.LGcs.AI
|
Yaoyiran Li, Haowen Ning, Mohamed Hammad |
Industrial sequential recommender systems operate over massive item catalogs (e.g., 10^5--10^7 items). Multi-label recommendation models are trained with Binary Cross-Entropy (BCE) loss over the full vocabulary, but standard BCE materializes a dense [B, N, V] ...Industrial sequential recommender systems operate over massive item catalogs (e.g., 10^5--10^7 items). Multi-label recommendation models are trained with Binary Cross-Entropy (BCE) loss over the full vocabulary, but standard BCE materializes a dense [B, N, V] logits tensor in High Bandwidth Memory (HBM), incurring prohibitive $O(BNV)$ memory and fatal Out-Of-Memory (OOM) errors. While chunked loss optimizations exist for Softmax Cross-Entropy in LLMs, large-scale multi-label BCE optimization remains unexplored across deep learning ecosystems. We propose CutBCE, an exact, hardware-accelerated BCE loss and gradient operator implemented in JAX and Pallas for large-vocabulary workloads. CutBCE introduces (1) an exact fused reformulation evaluating dense background loss and sparse target corrections; (2) a custom Vector-Jacobian Product (VJP) with a dedicated Pallas TPU backward kernel computing logit tiles on-chip in both passes so logits and their gradients never reside in HBM; (3) dynamic VMEM budgeting and sharding-aware collective hoisting for distributed meshes; and (4) count-based zero-overhead training metrics. On single-chip TPU v5e/v6e mini-benchmarks, CutBCE eliminates OOM errors with up to 91.9% speedup. On 8-chip TPU slice training for multi-label SASRec with 876k items (Yambda-50M), CutBCE reduces peak HBM by 65.7% (>14 GiB saved per chip) and increases training speed by 225.9% with comparable accuracy. CutBCE is open-sourced at https://github.com/AI-Hypercomputer/RecML/blob/main/recml/core/ops/binary_cross_entropy_ops.py.
|
| 1043 |
Poisson-GENERIC Neural Operators: Exact Metriplectic Structure in Function Space via Casimir Entropies
2610.05570
|
cs.LG
|
Jason Sulskis, Sathya Ravi |
Existing thermodynamically consistent neural operators impose the GENERIC degeneracy conditions by projecting the reversible operator onto the complement of the entropy gradient. This makes the operator state-dependent and forfeits the Jacobi identity, so the ...Existing thermodynamically consistent neural operators impose the GENERIC degeneracy conditions by projecting the reversible operator onto the complement of the entropy gradient. This makes the operator state-dependent and forfeits the Jacobi identity, so the result is metriplectic-degenerate rather than metriplectic. We instead obtain degeneracy the way GENERIC does. For nonlinear transport, the reversible operator $L$ is the compatible Lie-Poisson pencil $\alpha D+\lambda(uD+Du)$; otherwise it is a constant, trivially Poisson Fourier multiplier. On an augmented state $(u,s)$ with a latent entropy density, $S=\int s$ is a Casimir of $L$, so $L\,\delta S/\delta z=0$ holds identically without projection. The energy combines a fixed mechanical quadratic, a learned gauge-free potential, and a convex internal energy. The friction operator $M=AA^\top$ satisfies $M\,\delta E/\delta z=0$ pointwise, and its Onsager parity structure permits diffusion and damping while provably excluding transport. For any parameters, skewness, positivity, both degeneracies, and the Jacobi identity (on the resolved band for the Lie-Poisson term) hold to machine precision. Heat conduction and damped waves admit exact closed-form friction operators, the second law bounds physical energy under a checkable curvature condition, and a discrete-gradient integrator yields exact discrete first and second laws. On four PDEs in 1D and 2D with three backbones (FNO, Transolver, CNO), the model wins 61 of 72 seed-level comparisons against same-backbone unconstrained baselines, learns the exact transport and wave symbols, matches the true dissipation rate within 13% on heat and Burgers, and dissipates nothing on advection. A constant-$L$ ablation isolates the cost of exact Jacobi as the loss of Burgers, while a learned-entropy ablation injects energy on every reversible-irreversible problem.
|
| 1044 |
Disentangling Task Difficulty from Run-Level Failure in Agent Failure Prediction
2610.05572
|
cs.LG
|
Mohsen EsfandyariDoulabi, Lawrence Arkoh, Biruk Tadesse, Vaishvi Patel, Mehul Sharma |
Predicting whether an LLM agent will fail has emerged as a promising direction for supporting intervention during execution. Recent approaches report strong predictive performance, often with AUROC values between 0.85 and 0.94. However, predictors are typicall...Predicting whether an LLM agent will fail has emerged as a promising direction for supporting intervention during execution. Recent approaches report strong predictive performance, often with AUROC values between 0.85 and 0.94. However, predictors are typically trained by pooling runs from many tasks. We hypothesize that part of this performance comes from recognizing that some tasks are harder than others, rather than detecting whether a particular run is heading toward failure. This distinction matters because task-level difficulty supports decisions about where to allocate computation, while run-level prediction is needed to decide whether to intervene in an ongoing trajectory. We study benchmarks with repeated attempts of the same task by the same agent and separate cross-unit comparisons from comparisons between successful and failed runs of the same model-task unit. Across the evaluated corpora, more than 99.93% of the positive-negative pairs underlying pooled AUROC are cross-unit. Accordingly, predictors that never observe the current run can achieve high pooled performance, including a difficulty oracle with AUROC up to 0.945, while remaining at chance within task. Early run-level discrimination is consistently weak across trajectory predictors, released monitors, and hidden-state probes, although it improves later in execution and is stronger for weaker agents. Under fixed token budgets, task-level allocation outperforms abort-only strategies, while early stopping becomes beneficial only when within-task AUROC reaches about 0.84-0.93, far above the 0.50-0.55 range observed for early monitors. These results show that failure prediction should be evaluated not only by outcome accuracy, but by whether the captured signal supports the intended deployment decision.
|
| 1045 |
What Does an Observability Foundation Model Know?
2610.05577
|
cs.LGcs.AI
|
Dhyey Dharmendrakumar Mavani, Rian Atri, Tairan Ji |
A linear probe can show that a label is recoverable from a model's hidden states, but not whether that goes beyond what the input already reveals, or whether the model uses it. We audit Toto, an observability forecasting foundation model, on the Benchmark of O...A linear probe can show that a label is recoverable from a model's hidden states, but not whether that goes beyond what the input already reveals, or whether the model uses it. We audit Toto, an observability forecasting foundation model, on the Benchmark of Observability Metrics (BOOM) across five series-disjoint resplits, comparing linear probes on its frozen residual stream with models that read the raw input window and with Toto's architecture stripped of its trained configuration. Short-vs-medium cadence and metric type are more linearly recoverable from Toto's residuals than from the strongest raw-window model in every resplit (macro-F1 0.766 vs. 0.633 and 0.545 vs. 0.498). Domain is nearly tied, and series cardinality is recovered far better from the raw window. MOMENT-base shows related cadence, metric-type, and domain readouts. Recoverability is not use: exchanging Toto's residuals with those of high-burst donors moves a future-burstiness readout as intended but does not make forecasts consistently burstier than a randomized donor. A BOOM-trained coordination probe has negative zero-shot R^2 on the tested external benchmarks. We report each label against its strongest baseline.
|
| 1046 |
When the Cross-Silo Federation Goes Offline: Continual Learning for Site Onboarding with Limited Unlabeled Data
2610.05598
|
cs.LGcs.AI
|
Ahmadreza Eslaminia, Klara Nahrstedt, Chenhui Shao |
An organization often holds too little labeled data to train a model that generalizes, and the records that would supply the rest sit with organizations that cannot release them. Cross-silo federated learning offers a way through, since participants exchange m...An organization often holds too little labeled data to train a model that generalizes, and the records that would supply the rest sit with organizations that cannot release them. Cross-silo federated learning offers a way through, since participants exchange model parameters rather than records, but it ordinarily settles two aspects of the arrangement in advance, the participating sites and the classes the model can predict, and deployment can breach both. A new site joins after training, once the established sites have finished their engagement and gone offline, and its records arrive unlabeled, mixing conditions the model already recognizes with conditions no participant has observed. We present an autonomous three-stage procedure that expands the model entirely at the joining site: reconstruction experts screen for novelty, clustering separates the flagged records into candidate conditions, and class means describe the old classes, all inside one shared representation. Those classes were learned from records that never leave their owners, so the usual defenses against forgetting are unavailable, and the procedure supplies the evidence they would have carried from either of two dissimilar sources, prototypes held by the federation or records held by the joining site. On a real industrial condition-monitoring dataset, run end to end with no label consulted, either source holds old-class accuracy at 0.868 or above with forgetting at most 0.063, and the two differ by 0.021, so a configuration can be chosen by the disclosure it permits rather than the accuracy it delivers. Both keep old- and new-class accuracy in balance where every alternative we measure gives up one for the other, and both retain more of the old classes than distillation- and regularization-based baselines. The balance still holds with only 6 labeled records per arriving condition and 3 retained per old class.
|
| 1047 |
An LLM-in-the-loop RL Framework for Bioinformatics Feature Selection
2610.05600
|
cs.LGcs.AI
|
Xinyuan Wang, Deepti Agrawal, Yanjie Fu |
High-dimensional bioinformatics data, characterized by a large number of features relative to the number of samples, pose major challenges such as the ``curse of dimensionality,'' leading to overfitting, high computational cost, and poor generalization. Tradit...High-dimensional bioinformatics data, characterized by a large number of features relative to the number of samples, pose major challenges such as the ``curse of dimensionality,'' leading to overfitting, high computational cost, and poor generalization. Traditional feature selection methods often suffer from limited scalability and adaptability in such domains. We propose an LLM-in-the-loop reinforcement learning (RL) framework for bioinformatics feature selection, where the RL agent formulates feature selection as a sequential decision-making task, while the large language model (LLM) enhances the process in two ways: (1) guiding exploration through domain-informed advice, and (2) providing hybrid rewards that integrate data-driven performance with knowledge-driven evaluation. The LLM also produces explanations to improve interpretability for human experts without altering the RL policy update. Experiments on diverse bioinformatics datasets show that the LLM-in-the-loop framework outperforms baselines, achieves stable performance across downstream models, and converges faster than pure RL.
|
| 1048 |
Your Unlearning Gives You Away: Identifying Erased Concepts in Diffusion Models
2610.05601
|
cs.LGcs.AI
|
Kaiyuan Deng, Yuchen Li, Yang Xiao, Bo Hui, Geng Yuan |
Existing attacks on unlearned diffusion models assume that the erased concepts are known in advance and focus on recovering them. In practice, however, model providers may not disclose which concepts have been removed, and even with access to the original base...Existing attacks on unlearned diffusion models assume that the erased concepts are known in advance and focus on recovering them. In practice, however, model providers may not disclose which concepts have been removed, and even with access to the original base model, an adversary may still lack a clear target to attack. In this paper, we aim to answer the following critical but overlooked questions: which concepts have been erased from the model, and how many have been erased in total? To this end, we present Tracer, a framework that rapidly and accurately identifies erased concepts and estimates their number. Tracer efficiently identifies erased concepts without generating and classifying images. By combining lightweight spectral analysis of weight footprints, it enables efficient search over large candidate vocabularies. To distinguish multiple erased concepts, we introduce a footprint coverage objective that guides sequential discovery. Tracer estimates the number of erased concepts by detecting a sharp decline in candidate confidence as the selected concepts account for the erasure footprint, without requiring labeled examples for calibration. The framework requires only lightweight linear algebra and limited forward probes, with no prior knowledge of the unlearning algorithm. Experiments across text-to-image and text-to-video backbones and diverse unlearning methods demonstrate that Tracer identifies erased concepts and estimates their number in seconds, achieving 150 to 137,000 times and 133 to 20,000 times speedups over MIA and brute-force search on image and video models, respectively, with substantially higher identification accuracy.
|
| 1049 |
Delay-coordinate reconstruction and conditional-moment causal diagnostics in stochastic systems
2610.05632
|
cs.LG
|
Jun Ohkubo |
Partial observation and delay-coordinate reconstruction are rooted in deterministic dynamical-systems theory, whereas there are many systems with intrinsic stochasticity. We propose a conditional-moment interpretation of delay-coordinate reconstruction for sto...Partial observation and delay-coordinate reconstruction are rooted in deterministic dynamical-systems theory, whereas there are many systems with intrinsic stochasticity. We propose a conditional-moment interpretation of delay-coordinate reconstruction for stochastic systems, in which a delay vector is used to reconstruct conditional moments of a future distribution rather than a unique future sample path. Two complementary arguments motivate this viewpoint. First, the probability density of a stochastic differential equation obeys a deterministic Fokker-Planck equation and, under certain assumptions, is represented by an infinite deterministic hierarchy of moments. Hence, a finite-moment closure suggests a Takens-like finite-dimensional approximation. Second, a discussion based on the Koopman operator theory clarifies that the time evolution of an observable in the Mori-Zwanzig formalism yields a conditional expectation in stochastic systems. Then, the orthogonal "noise" term in the coefficient-space Mori-Zwanzig equation vanishes in the stochastic cases; this result is consistent with the moment-based argument. As an application of this stochastic delay-reconstruction viewpoint, we revisit convergent cross mapping (CCM) for diagnosing certain causal relationships. Although CCM based on the embedding theorem cannot generally be applied to stochastic systems, it is possible to examine certain types of causal relationships by using conditional moments. Using coupled logistic systems with additive and multiplicative coupling mechanisms, we discuss how causal relationships are embedded in stochastic systems.
|
| 1050 |
When Does a Diffusion Model Decide What to Draw ?
2610.05645
|
cs.LG
|
Snigdha Chandan Khilar |
A diffusion model starts from pure noise and removes it step by step. Somewhere along the way it stops being able to become "anything" and becomes committed to, say, a horse rather than a truck. We measure when this happens, and what a trained model gets wrong...A diffusion model starts from pure noise and removes it step by step. Somewhere along the way it stops being able to become "anything" and becomes committed to, say, a horse rather than a truck. We measure when this happens, and what a trained model gets wrong about it, on CIFAR-10. The most direct measurement is to freeze a half-finished image, restart the generation from that point many times, and count how often each class comes out. We call this probability the committor. Measured this way, the model settles coarse questions (vehicle or animal?) at roughly twice the noise level of fine ones (which animal?). A much cheaper measurement, the noise level at which a classifier's opinion about two classes splits into two distinct groups, gets the order of these decisions right (rank correlation 0.73-0.88) but not their exact timing. We then compare pretrained models with their
|
| 1051 |
Training and Scaling Compute-Optimal Physiological Waveform Foundation Models
2610.05649
|
cs.LG
|
Pingzhi Li, Jie Peng, Shuqing Luo, Zachary Plotkin, Tianlong Chen |
We investigate the scaling laws and compute-optimal training of physiological waveform foundation models (FMs). We train Aether, a family of over one hundred FMs ranging from 20M to 2.1B parameters, on up to 36.3M hours of physiological waveforms. We construct...We investigate the scaling laws and compute-optimal training of physiological waveform foundation models (FMs). We train Aether, a family of over one hundred FMs ranging from 20M to 2.1B parameters, on up to 36.3M hours of physiological waveforms. We construct eight clinical prediction tasks from MIMIC-III and evaluate the FMs through linear probing. The 720M FM outperforms all existing baseline FMs across all eight tasks. A scaling law of model size, pretraining hours, and labeled patients predicts downstream ranking error, i.e. $1-\mathrm{AUROC}$, effectively with $0.5\%$ prediction MAE at held-out resource scales and $0.9\%$ MAE when extrapolating to 2.1B parameters. We present three findings: (1) Compute-optimal training scales both FM size and pretraining hours. Under the fitted law, a $10.0\times$ increase in compute FLOPs scales model size by $1.2\times$ and pretraining hours by $8.2\times$. (2) Larger FMs use waveform data more efficiently, and greater pretraining exposure increases the benefit of model scaling. Starting from 25M parameters and 4.8M pretraining hours, doubling FM size reduces the predicted hours needed for the same performance by $51.8\%$. (3) Pretraining and clinical supervision reinforce each other: more labeled patients increase the return to pretraining, while larger FMs and longer pretraining reduce labeling requirements. For the example of the 720M FM, extending pretraining from 120K to 36.3M hours reduces the predicted patient requirement by $61\%$ at a target ranking error. These findings provide a quantitative training recipe and a promising and durable scaling path for physiological waveform modeling and downstream clinical prediction.
|
| 1052 |
Graph Data Augmentation via Contrastive Generator Inversion ($\texttt{DCBA}$)
2610.05653
|
cs.LG
|
Mateusz Stolarski, Micha{\l} Czuba, \L{}ukasz Krai\'{n}ski, Katarzyna Musial, Pawe\l{} Pra\l{}at |
Graphs provide a natural representation of many complex systems, ranging from social platforms to ecosystems. However, the development of graph-based machine learning methods is often constrained by the limited availability of large and diverse graph datasets....Graphs provide a natural representation of many complex systems, ranging from social platforms to ecosystems. However, the development of graph-based machine learning methods is often constrained by the limited availability of large and diverse graph datasets. In this paper, we introduce $\texttt{DCBA}$, a model-based approach to graph data augmentation that infers the configuration of a synthetic graph generator from an observed network. We instantiate the proposed framework using the $\texttt{ABCD}$ generator, which produces scale-free networks with community structure. Our model learns a joint representation of graphs and generator parametrisations using a multi-positive contrastive objective with soft negative weighting. The learned representation enables the prediction of an $\texttt{ABCD}$ configuration whose stochastic realisations preserve the macrostructural properties encoded by the generator. Experiments show that $\texttt{DCBA}$ recovers generator parameters more accurately and robustly than an algorithmic inverse-modelling baseline. Its downstream utility is further demonstrated in community detection, where inferred configurations used to fine-tune $\texttt{PRoCD}$ improve AMI on average by $161\%$ on synthetic and $273\%$ on real-world networks.
|
| 1053 |
Bellman-Centric Learning: Near-Optimal Regret for Linear Bandits with Memory
2610.05659
|
cs.LG
|
Jingyuan Liu, Huiwen Jia |
We study linear bandits with memory, where past actions induce endogenous nonstationarity through an arbitrary known, bounded matrix-valued memory map. To trade off exploration and exploitation while accounting for the memory dynamics, we develop RSM-LinUCB, a...We study linear bandits with memory, where past actions induce endogenous nonstationarity through an arbitrary known, bounded matrix-valued memory map. To trade off exploration and exploitation while accounting for the memory dynamics, we develop RSM-LinUCB, a Bellman-centric algorithm that learns as in linear bandits and plans as in reinforcement learning. This design admits a novel regret decomposition which separates the memory-induced error from the cumulative reward estimation error along the learner's trajectory. We prove a high-probability regret bound of $\widetilde O\big(dRS(M+1)+\sigma d\sqrt T\big)$, where $T$ is the learning horizon, $d$ is the parameter dimension, $M$ is the memory length, $R$ and $S$ bound the memory-map operator norm and reward-parameter norm, respectively, and $\sigma$ is the sub-Gaussian noise scale. Our results reveal that the multiplicative memory-horizon coupling in prior bounds is not intrinsic: memory only contributes an additive cost, up to logarithmic factors. We also prove a matching minimax lower bound, establishing near-optimality. We further extend the algorithm to generalized linear rewards, preserving this separation with near-optimal memory and leading statistical dependence. Our algorithms outperform the baselines in numerical experiments on synthetic instances and semi-synthetic KV- and semantic-cache tasks.
|
| 1054 |
Sharp Integrality Gaps in Calibration Distance
2610.05679
|
cs.LG
|
Zinan Wang, Xinhao Yang |
We study the offline gap between deterministic calibration distance C and its fractional relaxation L for binary unit-weight sequences under total absolute-change cost. We sharpen the offline comparison C <= L + O(sqrt(T)) (Qiao and Zheng, 2024, Theorem 2) ...We study the offline gap between deterministic calibration distance C and its fractional relaxation L for binary unit-weight sequences under total absolute-change cost. We sharpen the offline comparison C <= L + O(sqrt(T)) (Qiao and Zheng, 2024, Theorem 2) to the sharp worst-case order Theta(T^(1/3)). If Delta_T is the supremum of C - L over length-T inputs, then T^(1/3)/1000 <= Delta_T <= 41T^(1/3) for T >= 216. The upper bound holds for every input, while each T >= 216 has a rational lower-bound input. For every input with m distinct forecasts, C <= L + m, and the unrestricted-sample worst-case sparse order is Theta(m). For rational forecasts and accuracy, with binary-encoded multiplicities of separately assignable unit identities, a grid-free polynomial-bit-time procedure returns B <= L <= U, U - B < eta, and an exactly calibrated compact repair of cost at most U + m <= L + m + eta.
|
| 1055 |
Square-Root Regret for Adversarial Multiplayer Bandits without Collision Information or Shared Randomness
2610.05688
|
cs.LG
|
Chenyu Gan |
We study adversarial multiplayer bandits with $K$ arms and $2\le m<K$ labeled players, without collision information, shared randomness, or an external communication channel. We design a constructive communication and synchronization protocol with a Monte C...We study adversarial multiplayer bandits with $K$ arms and $2\le m<K$ labeled players, without collision information, shared randomness, or an external communication channel. We design a constructive communication and synchronization protocol with a Monte Carlo public constructor. With probability at least $1-CN^{-32}$ over preprocessing, where $N=2Km(T+1)$, its fixed published output satisfies \[ R_T\le C K^{5/2}\sqrt T\log^2(2Km(T+1)) \] simultaneously for every oblivious reward sequence chosen after preprocessing. Here $R_T$ is expected regret over the players' private execution randomness. Positive reward observations establish a common learning schedule and synchronize players before learning begins. The cost of delayed communication is charged to the support of positive rewards, ensuring that periods with little useful feedback incur only limited regret. A slow--fast learning procedure then maintains valid reward estimates while assignments and scores are exchanged.
|
| 1056 |
Planetary Geospatial Foundation Models: A New Paradigm for Global Public Health
2610.05699
|
cs.LG
|
Arbaaz Muslim, Aviv Slobodkin, Katherine Wheeler-Martin, Eric Zhou, John Brittain |
The efficacy of traditional disease prediction is limited by spatial gaps and temporal lags, which impact the timing and targets of resource deployments. Outbreaks escalate undetected, chronic disease burdens are quantified years later, and at-risk populations...The efficacy of traditional disease prediction is limited by spatial gaps and temporal lags, which impact the timing and targets of resource deployments. Outbreaks escalate undetected, chronic disease burdens are quantified years later, and at-risk populations in data-sparse regions remain unaddressed. Planetary geospatial foundation models complement existing epidemiological workflows to provide operational improvements, encoding multimodal search, mobility, and environmental signals into generalizable place representations. As illustrations of this complementarity, we present independent global health case studies of Google Earth AI's Population Dynamics Foundation Model (PDFM) -- a foundation model for geospatial inference -- across four domains (vaccine-preventable, communicable, noncommunicable, maternal mental health), five tasks (spatial extrapolation, interpolation/nowcasting, probabilistic forecasting, prospective forecasting, risk stratification), and four countries (USA, Canada, Mexico, and the Democratic Republic of the Congo). Across these case studies, PDFM addresses critical surveillance gaps across domains: improving US-Canada border MMR vaccination coverage predictions by capturing cross-border behavioral spillovers domestic models miss; nowcasting cardiovascular disease to accelerate data availability; enhancing short-term municipal Mexican dengue forecasts for timely outbreak vector control; improving forecasts of cholera hotspots; and adding a transferable signal to individual-level postpartum-depression risk prediction in US states the model had never seen, while not replacing individual socioeconomic data or closing demographic screening gaps. Together, these results showcase capabilities of geospatial foundation models for public health surveillance.
|
| 1057 |
ReMaD: Tuning-free Domain Adaptation for Classification and Out-of-Distribution Detection
2610.05718
|
cs.LG
|
Elijah Bolluyt, Cristina Comaniciu |
We introduce Reduced-rank Mahalanobis Distance (ReMaD), a novel prototypical distance-based refinement to classification and out-of-distribution (OOD) detection using pretrained models without finetuning. We use embeddings of the target dataset to fit closed-f...We introduce Reduced-rank Mahalanobis Distance (ReMaD), a novel prototypical distance-based refinement to classification and out-of-distribution (OOD) detection using pretrained models without finetuning. We use embeddings of the target dataset to fit closed-form distribution statistics in the model's latent space which can classify in-distribution samples and detect OOD samples, all without training or prior knowledge of the OOD data. Building on prototype classification and OOD detection, we analyze the distribution properties of large pretrained models when processing new datasets; based on this analysis, we formulate a simple modification to Mahalanobis Distance to adapt models' latent space distributions to new domains by removing unused features, without the finetuning or hyperparameter searches required by other adaptation procedures. We demonstrate the efficacy of this method to adapt existing large pretrained image embedding models to new classification domains outside their trained capabilities by testing across four target datasets, with competitive performance in both classification and OOD detection.
|
| 1058 |
Inferring physical fields in coupled systems with unknown parameters from incomplete observations using physics-constrained attentive neural operators
2610.05723
|
cs.LG
|
Shilun Wei, Xiaoqiang Sun, Wei Li, Kejun Tang |
Given incomplete measurements of a single physical field in a coupled system with unknown parameters, can we infer its full physical state and identify the underlying parameters? This problem is challenging because multiple coupled fields must be reconstructed...Given incomplete measurements of a single physical field in a coupled system with unknown parameters, can we infer its full physical state and identify the underlying parameters? This problem is challenging because multiple coupled fields must be reconstructed simultaneously from limited observations of only one, while the system parameters are unknown. In this work, we propose a machine learning framework for full-field reconstruction and parameter identification of unknown physical systems from sparse observations of a single physical field. Specifically, the cross-attention encoder propagates sparse sensor observations onto a regular grid to construct a sensor-conditioned latent representation, while a Fourier neural operator (FNO) decoder captures global spatial dependencies to reconstruct all coupled physical fields. The network parameters and unknown physical parameters are jointly optimized by minimizing observation losses, governing equation residuals, and boundary/initial condition constraints. The proposed approach is validated on two- and three-dimensional lid-driven cavity flows, a two-dimensional cylinder wake, and a two-dimensional non-ideal magnetohydrodynamics problem, demonstrating the recovery performance of unobserved fields and physical parameters from incomplete observations.
|
| 1059 |
PACMI: Provenance-Aware Cascading Memory Invalidation for Long-Term LLM Agents
2610.05732
|
cs.LG
|
Yiqi Wang, Jiaqi Liu, Jiaqi Zhang, Zhangkai Wu, Yiqun Duan |
LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memor...LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memories may remain semantically relevant, continue to affect dependent records, and retain value as historical evidence. This calls for two capabilities: dependency tracking to identify downstream effects and historical preservation to retain useful past records. We propose Provenance-Aware Cascading Memory Invalidation (PACMI), a framework that represents memories and new evidence in a provenance graph with typed dependency edges. PACMI assigns records to a four-state validity lattice, propagates validity changes to dependent memories, and uses the resulting states for retrieval and stale-premise detection. We also introduce a diagnostic benchmark with 100 cases and 300 queries across five domains. The evaluation separates node, context-, and answer-level performance. PACMI achieves the highest final-answer accuracy on this benchmark, and its paired difference from the strongest baseline is significant under an exact McNemar test. The premise checker achieves perfect precision, recall, and F 1 on the controlled query distribution. Cascading propagation primarily improves memorystate correctness: removing it increases final-answer errors from 3 to 11, but the paired difference does not reach the 0.05 significance threshold. Code and data will be made publicly available.
|
| 1060 |
CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training
2610.05744
|
cs.LGcs.AI
|
Jing Li, Jian Meng, Yingmeng Gao, Suming Qiu, Linyuan Qiu |
Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads...Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in degraded training efficiency and hardware utilization for underloaded experts, while hot experts require additional resources to accommodate excessive workloads. Recent studies address imbalanced MoE training through intricate parallelism strategies or resource reallocation. However, these system-level approaches often introduce additional resource requirements and considerable orchestration complexity, which become increasingly difficult to afford when training trillion-parameter LLMs under constrained computational resources. This work introduces CIPHER-MoE, which mitigates MoE workload imbalance while keeping the router's token-side Top-K selection unchanged. CIPHER-MoE applies affinity-aware Expert-to-Token filtering with explicit capacity control to reduce hotspot expert workloads without additional hardware resources or complex runtime design. The proposed method has been evaluated on large-scale MoE models, including DeepSeek-V4-Pro, showing up to 64.9 percentage points Top-1 expert workload reduction and 1.10$\times$-1.94$\times$ training acceleration, while preserving the training quality. The source code will be released soon.
|
| 1061 |
Beyond In-Distribution Preservation: Recovering Generalization in Quantized VLAs via Vulnerability-Oriented Tuning
2610.05745
|
cs.LG
|
Shen Ruan, Wenchang Gao, Jin Wang, Siao Liu, Zhoxizhuoma |
Post-training quantization has been shown to preserve VLA performance under standard evaluation conditions, but whether it preserves the full-precision model's robustness and generalization remains underexplored. In this study, we systematically study the robu...Post-training quantization has been shown to preserve VLA performance under standard evaluation conditions, but whether it preserves the full-precision model's robustness and generalization remains underexplored. In this study, we systematically study the robustness and generalization of post-quantized VLA policies under environmental disturbances. Empirical results show that quantized policies can become fragile to subtle environmental variations despite retaining comparable in-distribution performance. We further observe that action discrepancies are concentrated in a small subset of rollout states, while teacher guidance has opposite effects depending on discrepancy: it improves generalization at high-discrepancy states but can degrade it at low-discrepancy states. These findings reveal that effective post-quantization recovery requires selectively intervening on vulnerable states rather than globally distilling the student. We therefore propose Policy-Induced Vulnerability-Oriented Tuning (PIVOT-Q), a vulnerability-aware On-Policy Distillation (OPD) framework that selectively corrects vulnerable states encountered during quantized-student rollouts using the frozen full-precision policy as a teacher. PIVOT-Q identifies vulnerable states using discounted accumulated discrepancies over a short horizon, applies phase-balanced sparse supervision, and uses a Behavioral Anchor to prevent unnecessary changes. Experiments under seven LIBERO-Plus environmental variations demonstrate consistent recovery across multiple VLA backbones and quantization methods. Notably, PIVOT-Q consistently outperforms full-state distillation across all settings while using only 7.4% of its state-level distillation budget. Our code is available at https://github.com/ruanruan-andy/PIVOT-Q.
|
| 1062 |
Image resolution enhancement for advanced semiconductor nodes
2610.05809
|
cs.LG
|
Lucas Rencker, Omid Tajalizadehkhoob, Khalid Elsayed, Artem Khachaturiants, Helda Pahlavani |
Advanced semiconductor nodes are pushing the limits of feature sizes and require metrology with sub-nm resolution without compromising on the throughput as needed for in-line process control. Recently, high-throughput scanning probe microscopy (SPM) based metr...Advanced semiconductor nodes are pushing the limits of feature sizes and require metrology with sub-nm resolution without compromising on the throughput as needed for in-line process control. Recently, high-throughput scanning probe microscopy (SPM) based metrology and inspection tools capable of meeting these needs have been introduced to the market and qualified for use in HVM. While innovative measurement methods and tool architecture have allowed for a leap of improvement in throughput, the next step in further reducing imaging time can be obtained through the application of machine learning for enhancing the resolution of measured images for extraction of relevant parameters. In this work, we provide the general framework under which a neural network-based resolution enhancer is designed and used for SPM images. We showcase the effectiveness of this framework using measurements performed on Line/Space structures with a pitch of 200 nm. For the reusability of a pre-developed pre-trained model, we additionally leverage transfer learning and show that a new model for slightly differing structures can be re-trained and calibrated with a smaller data set of measurements performed on Line/Space structures with a pitch of 100 nm.
|
| 1063 |
Feature identification for parameter extraction and defect detection using machine learning
2610.05812
|
cs.LG
|
Yan Guo, Helda Pahlavani, Artem Khachaturiants, Khalid Elsayed, Jakob van de Laar |
Process control of advanced semiconductor nodes is not only pushing the limits of metrology equipment requirements in terms of resolution and throughput but also in terms of the richness of data to be extracted to enable engineers to finetune the process steps...Process control of advanced semiconductor nodes is not only pushing the limits of metrology equipment requirements in terms of resolution and throughput but also in terms of the richness of data to be extracted to enable engineers to finetune the process steps for increased yield. The move towards 3D structures requires extraction of critical dimension parameters from structures which can vary largely from layer to layer. For in-line process control, the necessary automation forces the development of layer and equipment-specific dedicated image processing algorithms. Similarly, with the increase in stochastic defects in the EUV era, detection of defects at the nm scale requires the identification of features captured in low resolution to meet the throughput requirements of HVM fabs, which can again lead to custom algorithm development. With the emergence of ML-based image processing methods, this process of algorithm development for both cases can be accelerated. In this work, we provide the general framework under which the images obtained from high-speed scanning probe microscopy-based systems can be used to train a network for either feature detection for parameter extraction or defect identification.
|
| 1064 |
Usefulness of Quantile-Aware Diffusion Modeling for Highly Imbalanced Tabular Data
2610.05825
|
cs.LG
|
Abu Talha, Peng Liu, Souradyuti Paul |
Classification problem in the context of highly imbalanced data is a major challenge in many real-world applications (e.g., FinTech, healthcare, etc.). In these cases, the vast majority of instances belong to a single class and a small fraction represent the m...Classification problem in the context of highly imbalanced data is a major challenge in many real-world applications (e.g., FinTech, healthcare, etc.). In these cases, the vast majority of instances belong to a single class and a small fraction represent the minority class (often the most critical class). Recently, diffusion models have emerged as powerful approaches to reduce the degree of ``imbalanced-ness'' in the dataset; they work by generating synthetic data by capturing complex data distributions using iterative transformations. However, standard diffusion models are not inherently suited to highly skewed or heavy-tailed data, due to inbuilt quadratic error loss, which lacks the structural sensitivity to capture rare, extreme values, and minority-class nuances. We propose a novel approach, namely, Quantile-TabDDPM, based on a quantile-regularized denoising objective that combines the standard quadratic error loss with a quantile loss term to explicitly capture rare events while preserving the theoretical grounding of the original denoising objective. We extensively evaluated our approach on a real-world credit card transaction dataset characterized by extreme class imbalance. The results demonstrate that the integration of diffusion-based synthetic data generation with a quantile-regularized denoising objective provides a robust and effective framework for fraud detection in highly imbalanced datasets.
|
| 1065 |
HiER-BLS: A Hierarchy-Guided and Error-Correcting Robust Incremental Broad Learning System
2610.05834
|
cs.LG
|
Gongli Zhang, C. L. Philip Chen, Zhulin Liu |
Broad Learning System (BLS) supports analytical training and incremental expansion, but its growth needs guidance on which inputs new blocks should learn from. Weight errors pose a further challenge by displacing learned outputs across class boundaries. We pro...Broad Learning System (BLS) supports analytical training and incremental expansion, but its growth needs guidance on which inputs new blocks should learn from. Weight errors pose a further challenge by displacing learned outputs across class boundaries. We propose HiER-BLS to couple hierarchy-guided representation growth with error-correcting learning. Successive blocks focus on inputs selected by feature importance and correlation while preserving earlier representations. The evolving branch guides encoded learners through subspace size and sample confidence, so its learning experience informs both their feature views and supervision. For finite broad readouts, we show how codeword correlations transform fitted class scores. Prediction preservation depends on the distance from the actual output to the nearest decoding boundary relative to the model's sensitivity to weight errors. Experiments on five image and five tabular datasets demonstrate improved classification performance over representative BLS variants. Component studies show that hierarchy guidance benefits the encoded branch even when the guiding branch has lower standalone accuracy, with further gains from combining their scores. Longer codes continue to improve accuracy under stronger Gaussian weight errors after clean accuracy has largely saturated.
|
| 1066 |
CACFG: Curvature-Aware Classifier-Free Guidance and Optimal Control
2610.05845
|
cs.LGcs.AI
|
Max Collins, Dasith de Silva Edirimuni, Jordan Vice, Tim French, Ajmal Mian |
Diffusion models generate samples by learning to reverse a fixed corruption process, and classifier-free guidance (CFG) is the standard mechanism for conditioning this process on a desired class or prompt. CFG can be applied at varying guidance strengths, and ...Diffusion models generate samples by learning to reverse a fixed corruption process, and classifier-free guidance (CFG) is the standard mechanism for conditioning this process on a desired class or prompt. CFG can be applied at varying guidance strengths, and while higher strengths improve image quality and conditional alignment, too high a guidance strength can degrade image quality and diversity. Furthermore, CFG violates principled diffusion sampling dynamics, and existing explanations for why it works despite the violation disagree on the underlying theory or do not extend to deterministic samplers used in practice. We address both these issues. We first frame CFG sampling as a continuous-time optimal control problem, treating the sampling trajectory as a sequence of controls chosen to maximise the probability of the desired condition. Solving the resulting Hamilton--Jacobi--Bellman equation shows that CFG is recovered under specific path costs when using an unconstrained control set. We argue this lack of constraint is responsible for CFG's failure at high guidance strengths, since it permits the sampling path to move arbitrarily far from the current image estimate. To fix this, we propose curvature-aware CFG (CACFG), which constrains the control set to a hypersphere informed by the Gaussian regularisation used when training variational autoencoders. We show that the control inputs produced by CFG sampling routinely violate this bound, and that across diffusion models, datasets, and guidance schedules, CACFG achieves superior generative quality at mid-to-high guidance strengths with a less severe quality-diversity tradeoff than regular CFG.
|
| 1067 |
RepICL: Reusable In-Context Prediction Across Heterogeneous Representation Spaces
2610.05852
|
cs.LG
|
Yu-Hsiang Liu, Kuan-Yu Chen, Chih-Sheng Chen, Meng-Hsuan Chang, Yu-Chen Den |
Frozen representations are widely reused for downstream classification, yet each new task typically requires fitting a new predictor. We ask whether the few-shot prediction procedure itself can instead be learned once and reused across datasets and representat...Frozen representations are widely reused for downstream classification, yet each new task typically requires fitting a new predictor. We ask whether the few-shot prediction procedure itself can instead be learned once and reused across datasets and representation spaces. To study this question, we introduce RepShiftBench, comprising 1,218 encoder--dataset tasks across text, image, and audio, with separate evaluation of generalization to unseen datasets, unseen encoders, jointly unseen datasets and encoders, and unseen modalities. The benchmark exposes a substantial gap: Logistic Regression fitted independently on each episode outperforms every evaluated in-context learner across all settings. We introduce RepICL, a meta-trained in-context learner that canonicalizes each episode through episodic whitening before prediction. Its inductive variant, RepICL-I, surpasses Logistic Regression in all 12 benchmark settings, while RepICL-T substantially outperforms existing transductive methods. Ablations identify episodic whitening as the primary source of these gains, while showing that it is not a universally beneficial preprocessing step. Across both variants, the gains concentrate on queries for which simple support prototypes favor the wrong class or provide little separation between the true class and competing classes. Transduction provides its largest additional gains when limited support coverage gives a misleading view of class separation. Together, these results demonstrate that a shared few-shot prediction procedure can generalize beyond the representation spaces observed during training.
|
| 1068 |
The Blind Spot Paradox: When Adaptive Classifiers Defeat Drift Detectors
2610.05853
|
cs.LG
|
Rapha\"el Minato, Fabrice Popineau, Arpad Rimmel, Bich-Li\^en Doan |
Monitoring concept drift from an adaptive classifier's error stream creates an operational conflict with the model's own update loop. When internal adaptation outpaces evidence accumulation, accuracy recovers before cumulative detectors (CUSUM, Page-Hinkley) c...Monitoring concept drift from an adaptive classifier's error stream creates an operational conflict with the model's own update loop. When internal adaptation outpaces evidence accumulation, accuracy recovers before cumulative detectors (CUSUM, Page-Hinkley) can reach threshold. Instrumenting an Adaptive Random Forest (ARF) shows that surviving trees absorb 98.6% of the post-drift error transient through incremental leaf updates alone. The first background tree swap accounts for just 0.71% of this erased error volume, but drops external detection rates by 31 percentage points. We derive the finite-horizon boundary where cumulative evidence fails to cross threshold and measure a critical magnitude floor ($\Delta e_c = 0.120$) below which false-alarm budgets preclude detection. This failure manifests as missed shifts on stationary streams and false-alarm flooding triggered by internal tree swaps on noisy baselines. We validate on synthetic shifts, ARMA-GARCH series (ProteuS), and tabular benchmarks (BAF, INSECTS); on the synthetic sweep at a standard threshold, the blind spot appears at $\Delta e \approx 0.25$, showing why classical benchmarks like SEA ($\Delta e \le 0.21$) failed to reach it.
|
| 1069 |
TurboPairFormer: Fast and Stable Protein Folding Model Training with an Optimized Triangle Attention Kernel
2610.05854
|
cs.LG
|
Yide Ran, Chelsea Lowman, Jan Doma\'nski, David Hartmann, Jenke Scheen |
Triangular attention is a core computation in AlphaFold3-style biomolecular models, with cubic cost in token count. Its shared pair bias adds a gradient reduction across attention slices to the usual reductions over queries and keys. The open-source backends w...Triangular attention is a core computation in AlphaFold3-style biomolecular models, with cubic cost in token count. Its shared pair bias adds a gradient reduction across attention slices to the usual reductions over queries and keys. The open-source backends we examine handle these reductions through repeated probability recomputation, floating-point atomics, or full score-gradient storage. Separately, computing the softmax backward correction from BF16-rounded forward outputs loses numerical precision. We present TurboPairFormer, a triangular attention implementation for NVIDIA Hopper GPUs that addresses both issues. Our key-tile-parallel backward algorithm recomputes each probability tile once for the query, key, value, and pair-bias gradients, using ordered partial reductions for deterministic accumulation without floating-point atomics or full score-gradient storage. Output-residual compensation retains a BF16 approximation of the output-rounding residual to compute the backward correction more accurately in FP32, without changing the BF16 output. With BF16 inputs at crop sizes 384, 640, and 768 and head dimensions 16 and 32, TurboPairFormer achieves the lowest mean query, key, and pair-bias gradient RMSE against an FP64 reference among the implementations compared in this paper. Residual compensation reduces these RMSE values by 28-47% in controlled ablations. All four gradients are bitwise identical across five repeated calls in all 600 input cases under fixed execution conditions. Integrated into OpenFold3 with our triangle multiplication kernels, TurboPairFormer achieves the lowest GPU computation time per optimizer step among the evaluated backend configurations on 16 H100 GPUs, with speedups of $1.73\times$ over OpenFold3's Triton backend and $1.13\times$ over cuEquivariance at crop size 768.
|
| 1070 |
CCQ: A Multi-State Child Care Quality Dataset to Support AI for Children's Health Research
2610.05863
|
cs.LGcs.AI
|
Victor Li, Yuzhang Xie, Ziwei Dong, Qingyang Zhu, Wenjing Ma |
High-quality child care in early life is a critical determinant of children's growth and development. Research on child care quality has been constrained by fragmented, non-research-friendly, and privacy-bound datasets. We present CCQ (Child Care Quality), a l...High-quality child care in early life is a critical determinant of children's growth and development. Research on child care quality has been constrained by fragmented, non-research-friendly, and privacy-bound datasets. We present CCQ (Child Care Quality), a large-scale, de-identified dataset for applied data science research at the intersection of AI and early childhood health. CCQ integrates 59,372 child care provider records across 12 U.S. states, covering diverse provider types as well as data schemas. To ensure research utility while protecting privacy, we implement an automated, LLM-based curation pipeline that anonymizes, cleans, and standardizes raw state records into two complementary releases: a cleaned textual release and a fully preprocessed tabular release. We also benchmark traditional machine learning models, tabular foundation models, and language models on quality rating prediction and important features analytics. Within a state, tabular classifiers on the preprocessed tables perform best. Across states, zero-shot transfer is near chance, but modest target-state supervision recovers most of the within-state performance, and pretraining on other states benefits finetuned language models. We release both datasets with all code to accelerate AI-driven research on child care quality and ultimately improve children's health and development.
|
| 1071 |
Global Communication or Graph-Specific Memory?
2610.05874
|
cs.LG
|
Hamed Shirzad, Danica J. Sutherland |
Scalable Graph Transformers are commonly trained and evaluated on static large graphs in a transductive setup. Many scalable Graph Transformer components can be formulated as a constant-size shared memory, similar to virtual nodes, providing compressed informa...Scalable Graph Transformers are commonly trained and evaluated on static large graphs in a transductive setup. Many scalable Graph Transformer components can be formulated as a constant-size shared memory, similar to virtual nodes, providing compressed information about the whole graph. The counterpart of these models in language models and other domains is justified as the input changes, and this mechanism learns to compress some useful information about the input. In transductive learning on a single fixed graph, however, any shared memory can be seen as a constant at test time. This raises the question of what exactly this shared memory does in this static setup. We give preliminary evidence that optimizing a shared memory directly performs similarly to global communication methods, and so normal local message-passing models can embed similar information in their weights. Thus, these settings may be a poor fit for evaluating global communication in graph neural networks.
|
| 1072 |
OGAM: Connecting Systematic Testing to Runtime Assurance through Object-Grounded Attention Monitoring for VLA Policies
2610.05878
|
cs.LGcs.AI
|
Haki Darwish, Xiangyu Yin, Changwen Li, Rongjie Yan, Francisco Gomes de Oliveira Neto |
Benchmarks expose vision-language-action (VLA) policies to few canonical instructions, while exhaustive deployment testing is impossible. We introduce Object-Grounded Attention Monitoring (OGAM), connecting systematic testing to runtime assurance: testing reve...Benchmarks expose vision-language-action (VLA) policies to few canonical instructions, while exhaustive deployment testing is impossible. We introduce Object-Grounded Attention Monitoring (OGAM), connecting systematic testing to runtime assurance: testing reveals attention divergence between successful and failed executions, and OGAM uses this signal to stop failures beyond the finite suite. We generate scene-grounded instructions through pairwise combinations of action templates and objects, and separately test meaning-preserving paraphrases. All 87 out-of-benchmark cases reveal problematic behavior across OpenVLA, OpenVLA-OFT, UniVLA, and $\pi_{0.5}$: none completes any of the 24 feasible instructions, while infeasible or hazardous requests also trigger behavior substitution. At each action query, we project gradient-weighted visual attention through object masks and group it by instruction role for comparison across tasks and policies. Dynamic time warping aligns this course with a successful reference despite speed differences; conformal calibration on successful episodes sets the early-stopping threshold for sustained deviations, with a nominal false-stop target of $\alpha=0.05$. Across four policies, OGAM stops 87-100% of failed episodes at median times of 5-12s within a 20s budget, with observed false-stop rates of 3-5%, without failure-labeled training. Finite testing thus identifies attention patterns that support online intervention before failure fully unfolds.
|
| 1073 |
The Optimization Landscape of Learning Compacted Context Models
2610.05885
|
cs.LG
|
Thomas Villeneuve, Alex Sandomirsky, Charles O'Neill, Max Kirkby, Michael Psenka |
Many works approach continual learning through the lens of infinite context windows. As an agent puts more observation into context (concretely the KV cache), compacting said context is akin to direct memory manipulation, without affecting the base model's wei...Many works approach continual learning through the lens of infinite context windows. As an agent puts more observation into context (concretely the KV cache), compacting said context is akin to direct memory manipulation, without affecting the base model's weights. Many works pose KV compaction as an optimization problem: learn a smaller set of KV vectors that matches the behavior of the full KV cache. While this preserves base model behavior, optimizing through a frozen base model results in a highly nontrivial optimization problem with a brittle and flat loss landscape. In this paper, we characterize what makes these optimization problems difficult and demonstrate that a heavily simplified Perceiver-based architecture not only matches performance of a full Perceiver transformer in continuous context compaction, but outperforms baselines on compaction utility. Results are presented on MCQ tasks across Finance, Legal, Gutenberg, and Code.
|
| 1074 |
Collaborative Personalized Preference Alignment for LLMs under Data Deficiency
2610.05898
|
cs.LG
|
Liyan Yang, Yige Yuan, Zhiqin Yang |
Real-world users often exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult. Lea...Real-world users often exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult. Learning shared initializations across users can support few-shot adaptation. However, heterogeneous preferences and competing objectives cause gradient conflicts across users and within each user, hindering effective initialization learning. This raises a central question: \textbf{how can we collaboratively learn aligner initializations that support few-shot adaptation to diverse user preferences?} To answer this question, we propose \textbf{A}pproximate \textbf{P}areto \textbf{O}ptimality (APO). We first group users whose updates are compatible, so that their information can be combined with less interference. Within each group, we combine gradient descent with controlled ascent to coordinate competing objectives and move towards preference-specific points on the Pareto front. This produces an initialization that is close to the optima of the users in the group. We then iteratively refine it using updates from few-shot local adaptation, making it more effective for personalization. Furthermore, we establish conditional suboptimality bounds for a one-local-step collaborative update and characterize how initialization error affects subsequent stochastic adaptation. Experiments on Fed-ChatbotPA and UltraFeedback show consistent improvements over existing methods using only 20 local examples.
|
| 1075 |
CoHyFuse: Condition-wise Hypergraph Fusion with Global Connectome in Task-fMRI
2610.05913
|
cs.LG
|
Boseong Kim, Haejun Chung, Ikbeom Jang |
Task-fMRI connectomes reveal state-dependent neural reconfigurations, yet conventional methods marginalize these signals by aggregating distinct conditions into static pairwise graphs, thereby obscuring condition-specific multi-ROI organization. We introduce C...Task-fMRI connectomes reveal state-dependent neural reconfigurations, yet conventional methods marginalize these signals by aggregating distinct conditions into static pairwise graphs, thereby obscuring condition-specific multi-ROI organization. We introduce CoHyFuse, a condition-aware ROI-centered hypergraph framework that constructs a task-state-specific incidence matrix from condition-wise functional connectivity (FC)-profile embeddings, allowing the same ROI to form different multi-ROI hyperedges across task phases. Condition-specific neighborhood sizes $K_q$ further adapt the hyperedge scale to each task state, and the resulting condition embeddings are fused with a complementary whole-session FC branch for prediction. In the AABC cohort (N=1,074), CoHyFuse achieved the best mean out-of-fold predictive performance among evaluated baselines on FACENAME Fluid Cognition Composite (FCC) prediction (7.83$\pm$0.10 MAE, 0.439$\pm$0.026 \(R^2\)) and VISMOTOR age prediction (7.52$\pm$0.37 MAE, 0.592$\pm$0.022 \(R^2\)). In an auxiliary CMI-HBN attention-deficit/hyperactivity disorder (ADHD) classification benchmark (N=223), CoHyFuse obtained 72.0$\pm$2.1\% macro-AUC and 74.2$\pm$2.9\% accuracy. Ablation studies support the contributions of condition-wise incidence construction and dual-view fusion, suggesting that state-resolved ROI-set structure provides complementary predictive information beyond whole-session FC alone. Occlusion analysis identifies the Distraction condition as the primary driver of model prediction, pointing toward the Salience/Ventral Attention Network (SAN)--FrontoParietal Network (FPN) and within-SAN hyperedge-defined ROI-set motifs as candidate model-relevant patterns. This framework provides an interpretable, state-resolved view of the connectome for downstream cohort analysis.
|
| 1076 |
Large Stepsizes Federated Learning on Logistic Regression with Linearly Separable Data: The Case of Heterogeneous Devices
2610.05915
|
cs.LG
|
Hok Fong Wong, Hoi-To Wai, Chung-Yiu Yau |
This paper revisits the distributed learning problem for training a multinomial logistic regression model with the Federated Averaging ($\texttt{FedAvg}$) algorithm. We concentrate on a scenario with arbitrarily large stepsizes and heterogeneous update rules w...This paper revisits the distributed learning problem for training a multinomial logistic regression model with the Federated Averaging ($\texttt{FedAvg}$) algorithm. We concentrate on a scenario with arbitrarily large stepsizes and heterogeneous update rules where the devices may perform a different number of local updates in each round. We show that, with linearly separable data, $\texttt{FedAvg}$ is stable with any stepsizes and the objective values converge to zero at the rate of ${\cal O}(1/R)$, where $R$ is the number of communication rounds. Our result also demonstrates that the effects of device heterogeneity vanish asymptotically. For sufficiently large $R$, the objective values decrease monotonically and is bounded by ${\cal O}( 1 / (R T_{\rm avg}))$, where $T_{\rm avg}$ is the average number of local update steps per communication round across devices. Numerical experiments support our findings.
|
| 1077 |
Non-invasive Seizure Detection Using Wearable Wrist-worn Accelerometry and Deep Learning
2610.05919
|
cs.LGcs.AI
|
Nilushika Udayangani Hewa Dehigahawattage, Kishor Nandakishor, Marimuthu Palaniswami |
Seizure monitoring and detection are crucial for reducing the morbidity and mortality associated with seizures. Current epilepsy care, often involving expensive video-electroencephalography (VEEG) monitoring, requires specialized expertise and is limited to in...Seizure monitoring and detection are crucial for reducing the morbidity and mortality associated with seizures. Current epilepsy care, often involving expensive video-electroencephalography (VEEG) monitoring, requires specialized expertise and is limited to in-hospital settings, and intrusive in nature. Seizure diaries, on the other hand, suffer from unreliability due to under-reporting, leading to incorrect therapeutic decisions. Wearable non-invasive seizure detection may offer a more tolerable and feasible solution for long-term ambulatory monitoring. This study explores a wearable remote monitoring system utilizing a single wrist-worn accelerometer device and capable of detecting multiple types of seizures, including shorter duration events. We enrolled 79 patients under video-electroencephalography monitoring to wear accelerometer devices and collect data. Concurrent VEEG recordings were reviewed by board-certified epileptologists to produce annotations, including seizure onset, offset, and seizure type. Using this data, we constructed a deep neural network based on the time-series ResNet architecture, which could discriminate among seizure and non-seizure events. Our proposed approach achieved a seizure detection sensitivity of 95.65% and an overall false alarm rate of 0.15/24 hours during the evaluation, which spanned 5576 hours of total recording. Additionally, it resulted in an area under the receiver operating characteristic curve (AUC-ROC) of 0.98 and an area under the precision-re call curve (AUC-PRC) of 0.67 when averaged over 20 patients who experienced 46 convulsive seizures. These promising results suggest that the proposed seizure detection system can be effectively used for long-term ambulatory seizure monitoring. Future steps include validating our findings in larger datasets and assessing the utility of detection for additional seizure types.
|
| 1078 |
The Arbitrary-Placement Problem in Entropy-Minimizing Selection, and a Residual-Entropy Formulation
2610.05925
|
cs.LG
|
Alyssa H. Shin, Claire H. Shin |
Entropy-based selection objectives suffer from a fundamental degeneracy: minimizing Shannon entropy $H(p_A)$ rewards confident selection regardless of whether the selected candidate is informative. We address this limitation with the residual entropy $D = H(p_...Entropy-based selection objectives suffer from a fundamental degeneracy: minimizing Shannon entropy $H(p_A)$ rewards confident selection regardless of whether the selected candidate is informative. We address this limitation with the residual entropy $D = H(p_A) - H(p_\beta)$, where $p_\beta$ is induced by candidate trust weights. We prove the exact identity $D = -\mathrm{KL}(p_A\Vert p_\beta) - \Delta$, where $\Delta$ measures whether the score-induced distribution and trust profile favor the same candidates. Boundary cases establish basic safety: under uniform trust, $D\leq0$ automatically, so an equal-trust, non-starving state is never penalized, while at any one-hot limit, $D\to0$ regardless of the selected candidate. For the intermediate regime where selection occurs, we prove that $D\leq0$ when candidate ordering by trust agrees pairwise with ordering by informativeness, and derive a tighter certificate based on the leading candidate's margin over its competitors. These results are independent of the candidate-scoring function and apply to both stationary and dynamically changing information. Experiments with a gradient-based mixture-of-experts router confirm that the ordering conditions can hold during real optimization and show that correct ordering improves downstream performance when candidates are non-interchangeable and selections are used directly rather than averaged. Beyond routing, margin-based reweighting matches or outperforms fixed-strength baselines in a class-imbalance task, while informative selection in a production video-prediction system reduces MSE by approximately 20$\%$ and transfers to a related species. Residual entropy, therefore, provides a safety criterion for selection and a usable signal for deciding when that selection is informative.
|
| 1079 |
ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning
2610.05935
|
cs.LGcs.AI
|
Seil Kang, Hangoo Kang, Tarun Suresh, Youngeun Kim, Shreyas Pimpalgaonkar |
Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verifica...Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes at the cost of policy staleness. We introduce ThunderSyncRL, which starts gradient computation as soon as all required inputs are fixed, without policy staleness. For group relative policy optimization (GRPO), ThunderSyncRL computes each trajectory's score gradient as soon as the reward for that trajectory arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. We prove that gradient streaming produces the same GRPO and OPD updates as batch-synchronous training, without changing either objective. On SWE-bench Verified and Terminal Bench 4.0, we train models to the same performance up to $1.9 \times$ faster than synchronous training. With zero policy staleness, ThunderSyncRL also outperforms asynchronous training at a fixed budget by up to $2.47$ percentage points.
|
| 1080 |
Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco
2610.05949
|
cs.LG
|
Abdelghani Belgaid, Zakaria Mahmoud, Fahd Chibani, Oumnia Ennaji, Younes Boudoul |
Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, ou...Site-specific fertilizer recommendation systems adapt nutrient advice to location, soil properties, crop type, and production targets, but scientific reuse is constrained when recommendation functions remain accessible mainly through interactive interfaces, outputs are not versioned, and trained approximations cannot be independently loaded or benchmarked. This technical report presents the Turba fertilizer machine learning stack, a three-layer open-source implementation for reproducible site-specific fertilizer recommendation in Morocco. \texttt{turba-client} provides programmatic access to publicly accessible site profiles, crop-specific target-yield spaces, and N, P$_2$O$_5$, and K$_2$O recommendation workflows; \texttt{turba-data} distributes analysis-ready snapshots; and \texttt{turba-models} packages crop-specific machine learning surrogates of recommendation outputs. The architecture links upstream retrieval, versioned analytical snapshots, reproducible cross-model benchmarking, and loadable offline surrogates while preserving the distinction between recommendation-system outputs, observed agricultural data, and model-generated predictions. The first dataset was constructed from 44,096 unique ESA WorldCereal locations. Scenario expansion across supported cereal workflows generated 132,017 crop-location recommendation requests under a medium target-yield setting. The resulting 22-variable dataset spans 10 regions, 66 provinces, and 1,149 communes. Nine regression families were evaluated under a fixed deterministic 80/20 protocol, and the current release packages five best-performing crop-specific models. The machine learning task is recommendation-function emulation rather than prediction of observed crop response. The stack provides a reproducible basis for spatial and temporal validation, uncertainty estimation, field-trial comparison, and future integration with additional data.
|
| 1081 |
Discovered, Not Designed: Population Evolution for Collaborative and Compute-Intensive Model Discovery
2610.05950
|
cs.LG
|
Bo Peng, Lizhu Zhang, Yuhang Zhou, Mingyi Wang, Yifan Wu |
LLM-driven evolution enables iterative model development, but two practical goals remain underexplored: finding model designs that transfer across related tasks and sustaining improvement when training is expensive. We introduce Population Evolution (PE), a co...LLM-driven evolution enables iterative model development, but two practical goals remain underexplored: finding model designs that transfer across related tasks and sustaining improvement when training is expensive. We introduce Population Evolution (PE), a collaborative, hierarchical framework that connects ongoing local searches through shared experimental evidence. PE evaluates code changes across related training instances and shares the results to guide subsequent proposals and promotion to larger training scales. For expensive targets, PE searches small training subsets and screens candidates through peer and intermediate evaluations before full-target training. We introduce RMD-Bench to evaluate both settings across ranking, watch-time prediction, RL algorithm discovery, and LLM/VLM pretraining. Compared with standalone evolution at matched source iterations, PE raises mean best local gains from 7.01% to 8.97% in ranking and from 2.84% to 3.85% in watch-time, while improving the best larger-scale outcome in all three joint-discovery families. In watch-time discovery, PE improves best larger-scale gains with four of five harnesses and all four proposers. On new recommendation datasets under shared target-side calibration, every evaluated PE design improves over the reference in mean performance. Under matched total GPU compute, completed LLM discovery runs yield a best relative accuracy gain of 2.48% and 13 successful candidates for PE, versus 0.92% and none for direct evolution. VLM loss reduction reaches 8.78% versus 5.05% under matched total GPU compute.
|
| 1082 |
Physics-Informed but Not Physics-Consistent: Error Geometry and Subspace Projection for Neural AC Power Flow
2610.05959
|
cs.LGcs.AI
|
Changhun Kim, Timon Conrad, Redwanul Karim, Karan Pahlajani, Julian Oelhaf |
Recent neural power-flow solvers, including emerging foundation models, achieve accurate voltage predictions, yet such accuracy does not necessarily imply physically consistent solutions. Even small complex voltage errors can yield large AC power-balance resid...Recent neural power-flow solvers, including emerging foundation models, achieve accurate voltage predictions, yet such accuracy does not necessarily imply physically consistent solutions. Even small complex voltage errors can yield large AC power-balance residuals. We study this accuracy-consistency gap across PIGNN-GC, GridSFM, gridfm-graphkit, and LUMINA on realistic 2224-bus Great Britain network (GBnetwork) scenarios, with cross-grid evaluation of GridSFM over 31 systems. Using a singular value decomposition (SVD) basis fitted to training AC power-flow solutions, we find that neural prediction errors contain substantial components outside the dominant solution subspace. To address this mismatch, calibrated solution-subspace projection (CSP) suppresses off-subspace prediction components after train-only bias calibration, reducing Mean PB by 67.0%, 37.8%, 40.5%, and 68.9% for PIGNN-GC, GridSFM, gridfm-graphkit, and LUMINA, respectively, relative to calibrated predictions, while improving voltage-magnitude accuracy in all four models. These results identify output-error geometry as an important factor in physics-consistent neural AC power flow. Code: https://github.com/Kimchangheon/neural-acpf-error-geometry
|
| 1083 |
Strategic Multi-Agent Learning for Interpretable Action Valuation of All Players in Football
2610.05961
|
cs.LG
|
Kenjiro Ide, Taiga Someya, Kohei Kawaguchi, Keisuke Fujii |
Valuing player actions in football requires accounting for strategic interactions among 22 players, including off-ball movements and defensive positioning. Existing reinforcement-learning-based methods commonly aggregate decisions at the team level or estimate...Valuing player actions in football requires accounting for strategic interactions among 22 players, including off-ball movements and defensive positioning. Existing reinforcement-learning-based methods commonly aggregate decisions at the team level or estimate player values independently, leaving strategic interdependence among players insufficiently represented. This study proposes an action valuation framework inspired by Markov perfect equilibrium (MPE) for all players. Each possession is modeled as a finite-horizon dynamic game, with each player represented as an autonomous agent whose policy depends on the current game state. MPE is used as a motivating solution concept rather than an exact equilibrium. To improve interpretability, we use Expandable Decision-Making States (EDMS) and decompose the Q-value into a successor-feature basis and a linear reward-weight vector. The value basis is estimated by linear TD initialization followed by nonlinear refinement. Using tracking and event data from 95 J1 League matches, we compare the proposed formulation with an independent reinforcement learning baseline. Because the two formulations define TD errors in different target spaces, TD MSE is used only for within-formulation consistency. With EDMS fixed, the independent baseline assigns the highest value to forward movement in 99.21% of evaluated off-ball states, whereas the most frequent direction under the proposed formulation accounts for 17.63%. Team-level average Q-values show a negative association with season-level expected goals for the baseline and a weakly positive association for the proposed formulation. Qualitative analyses illustrate context-dependent valuations of off-ball movements and defensive positioning. Overall, the proposed formulation produces more context-sensitive action rankings, although the comparison does not isolate the MPE-inspired component.
|
| 1084 |
Fast Last-Iterate Convergence in Zero-Sum Markov Games with Bandit Feedback
2610.05968
|
cs.LG
|
Yuheng Zhang |
We study last-iterate convergence in unknown two-player zero-sum discounted Markov games with bandit feedback. The players learn independently along a single trajectory without observing each other's actions. We develop Adaptive Regularized TD Learning (ARTD),...We study last-iterate convergence in unknown two-player zero-sum discounted Markov games with bandit feedback. The players learn independently along a single trajectory without observing each other's actions. We develop Adaptive Regularized TD Learning (ARTD), which achieves a $\widetilde{\mathcal{O}}(t^{-1/4})$ duality gap bound for the current policies under a uniform hitting time assumption, with high probability simultaneously over all rounds and starting states. This improves the $\widetilde{\mathcal{O}}(t^{-1/(9+\nu)})$ rate of Cai et al. (2023), for any fixed $\nu>0$, under the same feedback model and hitting time assumption. Our algorithm requires no knowledge of the hitting time bound, the time horizon, or the confidence level. To stabilize policy learning as value estimates change, we separate fast temporal difference averaging from bounded value updates. We adapt log-barrier regularization to the progress of value estimation, controlling both policy and value errors throughout learning. Together, these mechanisms enable fast convergence of the policies actually played, even when the players learn independently from bandit feedback.
|
| 1085 |
Reachability-Aware Diffusion Policy Optimization
2610.05969
|
cs.LG
|
Hikmet Simsir, Kutay Demiray, Ozgur S. Oguz |
Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information withou...Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.
|
| 1086 |
Transfer-Stratified On-Policy Distillation for RL-Improved Reasoning Teachers
2610.05974
|
cs.LGcs.AI
|
Xiaoyu Chen, Bo Shao, Tiangang Zhu, Bintao Wu, Linjun Shou |
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages ...Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.
|
| 1087 |
EpicWorldModel: Exploration-driven Planning with Latent World Models
2610.05996
|
cs.LGcs.AI
|
Bowen Feng, Julian Ost, May Mei, Anirudha Majumdar, Felix Heide |
Latent world models based on Joint-Embedding Predictive Architecture (JEPA) are deterministic by design. While successful in fully observable scenarios, this paradigm breaks down when past observations and actions lead to multiple plausible future possibilitie...Latent world models based on Joint-Embedding Predictive Architecture (JEPA) are deterministic by design. While successful in fully observable scenarios, this paradigm breaks down when past observations and actions lead to multiple plausible future possibilities, e.g., due to occlusion. We introduce EpicWorldModel, a framework to train stochastic JEPAs for environments and tasks with inherent uncertainty under partially observability. We jointly train the EpicWorldModel predictor with its latent representation space to directly predict multiple potential future states using a flow-matching objective, when the goal-relevant scene content is absent from the conditioning history. We show that flow predictive variance, motivated by its relation to an upper bound on predictive entropy, serves as a useful exploration guidance for planning. By incorporating this uncertainty signal into Cross-Entropy Method (CEM)-based planning, our approach balances goal-reaching with exploration of uncertain regions where occluded goals are most likely to be located. We demonstrate the effectiveness of EpicWorldModel through a series of latent planning experiments with the best or on-par performance across tasks, showing up to 22% empirical improvement in success rate over LeWorldModel.
|
| 1088 |
Langevin Flow Maps: Efficient Molecular Dynamics and Transition Path Sampling
2610.05998
|
cs.LG
|
Sam McCallum, Niklas Rindtorff, Alexander Tong, James Foster |
Molecular dynamics simulations proceed by integrating the Langevin equations over many small femtosecond timesteps. This poses a challenge for estimating ensemble properties and transition dynamics that occur on much longer timescales. We introduce Langevin Fl...Molecular dynamics simulations proceed by integrating the Langevin equations over many small femtosecond timesteps. This poses a challenge for estimating ensemble properties and transition dynamics that occur on much longer timescales. We introduce Langevin Flow Maps, which extend machine-learned force-fields to additionally learn the stochastic Langevin integrator. We show that Langevin Flow Maps enable large-timestep molecular dynamics and recover accurate dynamical properties of the system, while running an order of magnitude faster than current machine-learned force fields. Further, by training on a diverse molecular dataset, we demonstrate a path towards transferable Langevin Flow Maps.
|
| 1089 |
Learning While Scheduling Jobs under Context-Dependent Service Rates: An Anytime Rate-Optimal Algorithm
2610.06006
|
cs.LG
|
Seoungbin Bae, Dabeen Lee |
We study contextual queueing bandits, where a learner schedules jobs while learning unknown service rates modeled by logistic functions of job-server features. Performance is measured by queue length regret, the expected excess queue length at round $t$ relati...We study contextual queueing bandits, where a learner schedules jobs while learning unknown service rates modeled by logistic functions of job-server features. Performance is measured by queue length regret, the expected excess queue length at round $t$ relative to an oracle that knows the service rates. Existing decaying-regret guarantees either have a suboptimal decay rate or require a known fixed horizon. They also assume context-wise slack and a strictly positive minimum eigenvalue of the feature covariance. In this paper, we propose WISE (Widest Interval Selection with Elimination), achieving rate-optimal $\widetilde{\mathcal O}(t^{-1/2})$ queue length regret at every sufficiently large time without knowing the horizon. We assume capacity slack, meaning that expected incoming workload under best-server service is below service capacity, and impose no covariance lower bound. Our analysis uses a workload potential measuring the expected service attempts needed by waiting jobs on their best servers. Its drift on nonempty rounds combines a negative term ensured by capacity slack with errors from suboptimal service choices. Then an elliptical potential count bounds how often WISE selects wide confidence intervals, thereby limiting the number of rounds with large service errors. We also sharpen the arrival-rate dependence of an existing lower bound and make its dependence on feature dimension and server count explicit. We prove another lower bound that quantifies the increase in regret as the normalized capacity slack decreases; to our knowledge, this is the first such lower bound for CQB. Simulations show small regret even when context-wise slack fails.
|
| 1090 |
Spectral Geometry of Attention: From Information Routing to Uncertainty
2610.06012
|
cs.LG
|
Giulio Vigan\`o, Simone Melzi, Maks Ovsjanikov |
In this work, we study transformer attention through the lens of spectral geometry and operator theory. We view each attention head as a functional map between Hilbert spaces of functions on the token sequence and derive a Token Difference Operator, whose spec...In this work, we study transformer attention through the lens of spectral geometry and operator theory. We view each attention head as a functional map between Hilbert spaces of functions on the token sequence and derive a Token Difference Operator, whose spectral structure controls how token-space information is routed to the output. We show that standard Euclidean spectra are structurally biased by sinks, conflating mass concentration with genuine routing capacity. By recasting token space in the intrinsic probability geometry induced by attention, the token difference spectrum disentangles sink effects from routing capacity and provides a spectral description of the dimensionality of the head output. This yields a unified framework for analyzing attention maps, explaining sinks, routing collapse, and output dimensionality within a single operator-theoretic framework. In practice, by grounding attention heuristics in spectral geometry, we develop a novel attention-based uncertainty estimator that complements probability-based scores, with the largest gains on long-context inputs.
|
| 1091 |
Adaptive Expert Guidance for Efficient On-Policy Reinforcement Learning
2610.06019
|
cs.LGcs.AI
|
Daniele Affinita, Ming Xu, Rudolf Reiter, Davide Scaramuzza, Pascal Fua |
With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a h...With massively parallel simulation, on-policy Reinforcement Learning methods such as PPO have become standard in many domains. However, learning from scratch is sample-inefficient and fails to exploit the potential existence of a suboptimal expert, such as a heuristic, a model-based controller, or a policy trained on a related task. Such an expert is often available and can guide early training, but its sub-optimality limits final performance. The challenge then becomes balancing expert guidance against learning from rewards. Existing methods set the expert's influence through a blending weight, a schedule, or an evaluation-driven curriculum. Alternatively, they adapt it with additional learned components such as critics over expert actions or auxiliary agents. However, none optimizes it using the same on-policy objective as the policy itself. We propose a method in which the learner and the expert alternate control within each training episode, and the expert's share of control is a single learnable parameter optimized jointly with the policy. The learner benefits from the expert early in training, but its share of control declines as the learner becomes more competent, until eventually vanishing completely. This leaves the learner acting alone and better than the suboptimal expert. We evaluate our method on 34 tasks across two benchmarks, spanning discrete and continuous action spaces, using both learned and model-based experts. Our method improves sample efficiency over guided and unguided baselines while requiring minimal hyperparameter variation. The expert's share decays to zero as the learner improves, vanishing when the expert is no longer useful.
|
| 1092 |
Joint Precision Neural Networks: Task-Aware Dependency and Predictive Learning
2610.06023
|
cs.LG
|
Andrea Cavallo, Samuel Rey, Antonio G. Marques, Elvin Isufi |
Exploiting meaningful latent structures from data to solve downstream tasks is a fundamental challenge in signal processing and machine learning. While Principal Component Analysis (PCA) and coVariance Neural Networks (VNNs) successfully leverage the covarianc...Exploiting meaningful latent structures from data to solve downstream tasks is a fundamental challenge in signal processing and machine learning. While Principal Component Analysis (PCA) and coVariance Neural Networks (VNNs) successfully leverage the covariance matrix to process data, they inherently capture both direct and indirect correlations. The precision matrix (inverse covariance) overcomes this by explicitly encoding conditional independencies, making it largely studied in graphical lasso and graph topology identification. However, finite-sample precision estimates are notoriously unstable, and regularized estimators remain task-agnostic. In this work, our principal contribution is tackling the challenging problem of task-aware graph inference. We propose Precision Neural Networks-Joint (PNN-Joint), a framework that jointly estimates a sparse, statistically grounded precision matrix alongside graph neural network weights via an alternating optimization scheme. As a foundational framework to support this, we introduce Precision Neural Networks (PNNs), a broader class of graph convolutional networks operating on precision estimators, and establish their spectral connections to PCA and VNNs alongside their stability to finite-sample errors. Extensive empirical evaluations on synthetic data, as well as real-world neuroimaging and motion sensor datasets, demonstrate that PNN-Joint yields highly interpretable task-aware graphs, exhibits remarkable robustness in low-data regimes, and consistently achieves the best or second-best performance among competitors on real-world tasks.
|
| 1093 |
MercerFlow: Flow Matching in a Kernel-Induced Latent Space for Probabilistic Forecasting
2610.06039
|
cs.LGcs.AI
|
Ilya Kuleshov, Egor Serov, Alexey Zaytsev |
Recent work has shown that probabilistic flow matching for time series forecasting benefits from a data-matched prior. The resulting prior introduces local correlations, which a sequential architecture usually absorbs: a recurrent neural network (RNN), a struc...Recent work has shown that probabilistic flow matching for time series forecasting benefits from a data-matched prior. The resulting prior introduces local correlations, which a sequential architecture usually absorbs: a recurrent neural network (RNN), a structured state-space model (S4), or a Transformer. However, such a backbone costs GPU memory and time per epoch. A cheaper alternative is MLP-based latent-space flow matching: embed the time series via an invertible map to a single latent vector and learn the flow there, so a tabular MLP can treat the series as a set of features. The relationship between the prior and the choice of linear latent map is understudied in conditional flow matching (CFM) forecasting, yet we found it strongly affects performance. Fixed transforms such as Fourier or discrete cosine (DCT) are only well-conditioned for Ornstein--Uhlenbeck priors, while a principal-component (PCA) map fit to the data is a strong but training-set-dependent reference sensitive to train--test shift. Instead, we propose to use the Mercer eigenbasis of the prior kernel: it diagonalises the centred covariance exactly, decouples from training data, and adapts to non-stationary and periodic priors. On five GluonTS benchmarks (ETTh1, ETTh2, Weather, Electricity, Traffic) under a shared protocol with TSFlow, the resulting MLP matches or beats it on CRPS at about $4.7\times$ less training memory and $3.5\times$--$4.4\times$ less time per epoch.
|
| 1094 |
GO-Based Clustering for Learning Cluster-Level Causal Gene Regulatory Networks
2610.06042
|
cs.LG
|
Azlaan Mustafa Samad, Wei Zhang, Ad\`ele H Ribeiro |
Discovery of causal relationships in high-dimensional Gene Regulatory Networks (GRN) is computationally challenging and often difficult to interpret due to dense connections. Therefore, grouping genes together into functional modules can improve tractability a...Discovery of causal relationships in high-dimensional Gene Regulatory Networks (GRN) is computationally challenging and often difficult to interpret due to dense connections. Therefore, grouping genes together into functional modules can improve tractability and biological interpretability. However, existing cluster level causal discovery methods assume access to a predefined admissible partitions, requiring the graph over clusters to be acyclic. Constructing such partitions is therefore challenging. In this work, we introduce GO-based Clustering for Causal Discovery (GO4CD), an algorithm that uses Gene Ontology (GO) to construct biologically meaningful gene partitions at multiple levels of granularity, while favoring those more likely to be admissible for causal discovery. GO4CD groups together genes participating in a shared biological process, and propagates gene annotations through the ontology hierarchy to achieve different granularity of partitions. Furthermore, we integrate GO4CD with Causal Learning over Clusters (CLOC) algorithm and evaluate recovery of true Markov equivalence class both with an oracle of conditional independencies and on simulated gene expression data using multivariate conditional independence tests. We evaluate GO4CD on multiple E.coli regulatory subnetworks and find that it is inadmissible in 18.1% of the cases, compared with 65.3-82.3% for the semantic-similarity baselines. Our results indicate that GO4CD is substantially better suited to learning causal GRNs defined over biologically meaningful gene clusters.
|
| 1095 |
Pay to Learn, Share to Earn: Incentivized Federated Multi-Player Bandits
2610.06062
|
cs.LG
|
Pavamana K J, Chandramani Singh |
Federated multi-player multi-armed bandit problems model collaborative sequential decision-making where multiple players interact with a common bandit environment and share information through a central server to accelerate learning. Existing federated bandit ...Federated multi-player multi-armed bandit problems model collaborative sequential decision-making where multiple players interact with a common bandit environment and share information through a central server to accelerate learning. Existing federated bandit frameworks typically assume that all players willingly share their local observations with the server. However, this assumption is often unrealistic in practical settings where players are self-interested and may not participate in collaboration without explicit incentives. To address this challenge, we propose an incentive-aware federated bandit framework in which players receive rewards for sharing information with the server and incur costs when buying information from the server. We develop a UCB-based algorithm, termed Buying-UCB, that balances individual exploration and collaborative learning by incorporating both sharing incentives and information acquisition costs into the learning process. We theoretically analyze the proposed algorithm and derive upper bounds on the group regret and buying cost. Our analysis further characterizes the trade-off between fully collaborative federated learning and completely independent learning. Extensive numerical experiments validate the theoretical findings and demonstrate the effectiveness of the proposed framework under different collaboration and pricing regimes.
|
| 1096 |
Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning
2610.06083
|
cs.LGcs.AI
|
Zexin Li, Ruili Yao, Yiming Zeng, Xiaoxue Gao |
Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box...Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image-classification attacks (FGSM, MI-FGSM, and NI-FGSM) with a per-step objective yields perturbations that transfer but are no stronger than random noise of the same budget. We then propose a trajectory-level attack that optimizes a sequence of perturbations over a receding horizon through a differentiable model of the environment and a temperature-smoothed surrogate policy, with the same optimizers. On CartPole-v1 with ten DQN and DDQN agents and 100 surrogate--victim pairs, the trajectory-level attack outperforms per-step attacks and random noise in the white-box, cross-model, and cross-algorithm settings.
|
| 1097 |
Rethinking Least-Core Computation in Contextual-Distractor Games
2610.06087
|
cs.LG
|
Hiroshi Kera, Toshinori Yamauchi, Sai Ganesh Nagarajan |
Game-theoretic attribution explains a model by assigning credit to its features or training examples. The least core has attracted interest as an alternative to Shapley-style averaging because it can expose players that cause substantial harm in rare, high-val...Game-theoretic attribution explains a model by assigning credit to its features or training examples. The least core has attracted interest as an alternative to Shapley-style averaging because it can expose players that cause substantial harm in rare, high-value contexts. However, least-core allocations are generally nonunique, and the choice of allocation can affect the resulting explanation. In this study, we investigate how payoff selection and coalition sampling affect least-core attribution. Our experiments show that selector choice matters for distinguishing useful and harmful contributions, and that sampling can degrade harmful-player identification across the tested selectors even when useful players remain well identified. These observations motivate efficient computation with all coalition constraints and a well-defined selector. We introduce entropic least core (ELC), a smooth approximation whose unique minimizer follows a continuous path along the temperature to the nucleolus, a classical refinement of the least core. Our experiments show that ELC approximates the nucleolus faster than an LP-based nucleolus solver while retaining small payoff errors, with further GPU acceleration at larger problem sizes. In the tested full-coalition contextual-distractor games, ELC matches the minimum-norm selector in identification accuracy and more accurately ranks distractors by harm.
|
| 1098 |
Lossy Compression of PDE Training Inputs: Field Reconstruction Error Does Not Order the Cost to a Trained Operator
2610.06095
|
cs.LG
|
Huy Hoang Le |
Operator-learning benchmarks are stored at full precision and have grown to terabyte scale. Rate-distortion theory says how many bits the stored field needs, while a practitioner needs to know how accurate an operator trained on the compressed data will be. We...Operator-learning benchmarks are stored at full precision and have grown to terabyte scale. Rate-distortion theory says how many bits the stored field needs, while a practitioner needs to know how accurate an operator trained on the compressed data will be. We show that the first does not determine the second, and measure why, compressing the input fields while targets and test inputs stay at full precision. A solution operator attenuates a perturbation of its input. Pushing a compressed field through a surrogate already trained at full precision measures how much of the perturbation that surrogate transmits. The fraction is consistent with the smoothing behaviour of the underlying equation, and it spans more than two orders of magnitude across PDE families. Field reconstruction error is computed before the attenuation and cannot see it. For operators trained with mean squared error it inverts 36 of 104 cost comparisons across datasets, where a probe built from the same forward passes inverts 12. Two families that PDEBench stores with identical initial conditions differ threefold downstream at identical field error. Under the relative-L2 objective of the reference recipe the separation narrows, while the ordering of the family-level median transmission factors is unchanged. After one full-precision training run, the probe evaluates an entire rate curve by forward passes alone. It ranks datasets and rates consistently across the codecs and architectures we test, while its magnitude does not transfer between them.
|
| 1099 |
On the Geometry of Multimodal Saturation: Riemannian VICReg
2610.06096
|
cs.LGcs.AI
|
Nessim Ben Abbes, Duc Han Le, Sabri Mtibaa, Van-Tam Nguyen |
In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired...In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired runs under VICReg. The same failure occurs in 51.1% of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Riemannian VICReg (R-VICReg) generalizes classical VICReg: it aligns views by squared geodesic distance on learnable negative-curvature product factors and recovers VICReg exactly as curvature vanishes. Over the same 45 paired runs, R-VICReg raises the probability that the third modality helps from 44.4% to 64.4%, with gains concentrated where VICReg saturates.
|
| 1100 |
Flash-OPD: Fast On-Policy Distillation
2610.06105
|
cs.LG
|
Wei Chen, Junle Chen, Yitong Yang, Zhaoyang Xu, Jiaxin Lin |
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules ...On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision compatibility can vary substantially across trajectories, making a single rollout horizon difficult to match their heterogeneous reliable lengths: an overly short horizon may truncate useful supervision, while an overly long one wastes computation beyond reliable regions. Our key insight is that the trajectory-specific reliability boundary need not be predicted before generation. By viewing reliability as the first-passage of accumulated low teacher--student compatibility events, the boundary is inherently unknown before sampling, yet whether it has been reached can be determined exactly from the observed prefix. Building on this insight, we propose *Flash-OPD*, which shifts from rollout-horizon control to adaptive trajectory-level boundary verification. *Flash-OPD* interleaves cached generation with teacher verification and independently stops each trajectory according to its observed compatibility events. To reduce verification overhead, the recent event rate is used only to schedule the next verification point, while the actual stopping decision always relies on the exact cumulative count. This separation prevents estimation errors from causing premature termination while enabling efficient verification during generation. Extensive experiments across diverse datasets and teacher--student settings show that *Flash-OPD* achieves $2.2\times$--$7.5\times$ speedups over standard OPD while maintaining or improving accuracy.
|
| 1101 |
ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training
2610.06116
|
cs.LG
|
Yuanshi Liu, Boyuan Jiang, Liang Hou, Xin Tao, Pengfei Wan |
Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions...Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.
|
| 1102 |
Mind the Drift: Diagonal Linear Networks Under Large Learning Rates
2610.06120
|
cs.LG
|
Aniket Sanyal, Tom Jacobs, Rebekka Burkholz |
Large learning rates can qualitatively change the trajectory of neural network training, often pushing optimization into regimes far from classical gradient-flow behavior. The Edge of Stability (EoS) offers a valuable lens on the dynamics such learning rates i...Large learning rates can qualitatively change the trajectory of neural network training, often pushing optimization into regimes far from classical gradient-flow behavior. The Edge of Stability (EoS) offers a valuable lens on the dynamics such learning rates induce. We study corresponding dynamics in diagonal linear networks, where we uncover a competition between two distinct implicit biases that jointly determine the sparsity of the recovered solution in regression settings. Complementary to the Gain, which captures the average discretization error accumulated by Gradient Descent relative to Gradient Flow, we derive a closely associated but overlooked quantity: the Drift. Under large learning rates, it describes an imbalance between different discretization errors and represents a systematic shift in the optimization trajectory. While the Gain grows monotonically in certain regimes, and can bias towards denser, flatter interpolators, the impact of the Drift depends on its alignment with potential solutions, which can either counteract or reinforce the effect of the Gain. Consequently, its behavior drives model selection, particularly during early training epochs. To validate our theoretical insights, we introduce an intervention that actively steers the Gain to recover sharper, sparser solutions. Thus, our analysis reveals that large learning rates do not universally hinder the recovery of sparse solutions. On the contrary, they can be harnessed to control the implicit bias of training.
|
| 1103 |
Integrating Survival-Based Aging Models with Data-Driven RUL Prognostics
2610.06128
|
cs.LG
|
Abhishek Srinivasan, Juan Carlos Andresen, Sepideh Pashami, Anders Holst |
Predictive maintenance requires reliable remaining useful life (RUL) estimation. Existing methods mainly follow two paradigms: wear-based aging models that capture cumulative degradation and sensor-driven data models that reflect instantaneous health condition...Predictive maintenance requires reliable remaining useful life (RUL) estimation. Existing methods mainly follow two paradigms: wear-based aging models that capture cumulative degradation and sensor-driven data models that reflect instantaneous health conditions, each providing only partial information. In this work, we propose a probabilistic fusion framework that integrates wear-based and sensor-based prognostic components through failure probability distributions. Based on explicit structural assumptions linking wear, latent health, sensor observations, and failure, we derive a principled combination rule that enables uncertainty-aware integration with adaptive weighting of the components. Experimentally, we assess this combination rule by learning the wear-based component using a parametric survival model and the sensor-based component using a 1D convolutional neural network (1D-CNN) with a post-hoc uncertainty model. Evaluation on multiple N-CMAPSS datasets demonstrates that the fused model improves point accuracy, preserves the C-index, and produces narrower yet well-calibrated prediction intervals compared to either component alone. The results highlight the complementary roles of wear-based survival model and sensor-based deep learning model, and show that their probabilistic integration provides a structured pathway toward more robust and consistent prognostics over the life-time.
|
| 1104 |
Machine learning for journal entry testing: A type-aware evaluation of anomaly detectors under a review budget
2610.06133
|
cs.LGcs.AI
|
Jan Gronewald, Michel Scherer, Nijat Mehdiyev |
Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall...Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, type coverage and first-hit rank. We evaluate nine unsupervised detectors, a supervised reference and feedback-driven Deep Semi-Supervised Anomaly Detection (DeepSAD) on four real client ledgers with injected typed anomalies and a public synthetic ledger. On the largest client ledger, principal component analysis (PCA), an autoencoder (AE) and a variational autoencoder (VAE) each place on average 98 anomalies among the first 100 postings, but at least 95.8 belong to one type. FSR instead favours a nearest-neighbour (kNN) detector and changes the top-ranked detector on three of four client ledgers. Representation also matters: one-hot encoding exposes unseen accounts, whereas frequency encoding leaves unseen contra accounts largely undetected. On the public ledger, the Histogram-Based Outlier Score (HBOS) and Empirical Cumulative Distribution-Based Outlier Detection (ECOD) reach all eight markings within 1,386 entries, whereas kNN, the hit leader at 1,000 entries, first reaches cross-linked clearing at rank 4,641, and the supervised row-level reference misses this marking within 1,000 entries. There, the adaptive DeepSAD review protocol raises mean hits per 100 reviews from 40.0 to 68.3 but type coverage only from 2.7 to 3.0. These findings show that high hit rates can conceal systematic blind spots and suggest that feedback can reinforce existing detection patterns without broadening anomaly coverage.
|
| 1105 |
Co-Optimizing Graph Sparsification and Approximate Computing for Energy-Efficient FPGA-Based GCN Inference
2610.06138
|
cs.LG
|
Nathaniel Kaye Mellor, Shreejith Shanker, George Floros |
Graph Convolutional Networks (GCNs) have emerged as a powerful framework for learning from graph-structured data, yet their deployment on resource-constrained edge platforms remains challenging due to the computational and memory demands of sparse graph aggreg...Graph Convolutional Networks (GCNs) have emerged as a powerful framework for learning from graph-structured data, yet their deployment on resource-constrained edge platforms remains challenging due to the computational and memory demands of sparse graph aggregation. This work presents an FPGA-based GCN accelerator that combines DSpar graph sparsification, 8-bit quantization, and approximate multipliers on the AMD Kria KV260. Evaluated on Cora, LastFM Asia, and Amazon Photo, the design explores the interaction between sparsification and approximation across graphs with widely varying densities. Results show that the effectiveness of approximate arithmetic is governed by accumulation depth within GCN computations. Approximate multipliers are most effective when applied to sparse aggregation operations, while graph sparsification further improves their viability by reducing aggregation depth. The combined approach achieves up to 9.88$\times$ speedup while maintaining 86.6\% classification accuracy on Amazon Photo, and 1.52$\times$ speedup with 77.0\% accuracy on Cora, with total power consumption below 1 W. These results demonstrate that graph sparsification and approximate computing are complementary techniques whose co-optimization enables efficient low-power GCN inference on edge FPGA platforms.
|
| 1106 |
TIGER: Time-Series Classification with In-Context-Learning Gated Ensemble of Representations
2610.06156
|
cs.LG
|
Johann Faouzi |
A representation family is a distinct way of extracting features from time series. Ensemble algorithms that combine several representation families remain the most accurate approach to time series classification. Current state-of-the-art ensembles, most notabl...A representation family is a distinct way of extracting features from time series. Ensemble algorithms that combine several representation families remain the most accurate approach to time series classification. Current state-of-the-art ensembles, most notably HIVE-COTE~2.0, pair a bespoke classification algorithm with each representation family and combine their predictions using a fixed, non-adaptive rule. We present TIGER (Time-series classification with In-context-learning Gated Ensemble of Representations), which instead applies the same small portfolio of three general-purpose classifiers (Ridge, Extra Trees, and Naive Bayes) to four representations from four distinct families, stacking the resulting twelve base learners' predictions into a meta-feature matrix. The final prediction is produced by an adaptive meta-classification rule that chooses, independently for each data set, between a weighted hard majority vote and TabICLv2, a pretrained tabular foundation model used in-context as a meta-classifier, based on the mean number of training samples available per class. On a 142-data-set benchmark drawn from the UCR time series classification archive, TIGER obtains the best mean accuracy, balanced accuracy, and F1-score among six compared algorithms, including HIVE-COTE~2.0, and significantly outperforms each of the other five individually. TIGER's adaptive rule also meaningfully outperforms either of its two constituent meta-classification methods used alone, and its single hyperparameter, tuned using only a twenty-data-set development subset, is shown to generalize to the full evaluation benchmark. We further characterize TIGER's design through an extensive set of ablation experiments and report the design alternatives that we investigated and ultimately discarded.
|
| 1107 |
Anlu: Enabling In-Context Time Series Anomaly Detection in Foundation Models via Counterfactual Supervision
2610.06180
|
cs.LGcs.AI
|
Tian Lan, Yifei Gao, Yimeng Lu, Xuming An, Meng Wang |
Whether a time-series pattern is anomalous often depends on the operating regime of the monitored process. A missing event can signal a fault in one regime and be routine in another, and the query alone may not reveal which regime applies. We study in-context ...Whether a time-series pattern is anomalous often depends on the operating regime of the monitored process. A missing event can signal a fault in one regime and be routine in another, and the query alone may not reveal which regime applies. We study in-context learning (ICL) for time series anomaly detection (TSAD) through reference-conditioned detection, where a reference record provides evidence about expected behavior and model parameters remain fixed at inference. Supplying the reference is not enough: when training anomalies are recognizable from the query alone, the detector can fit its targets while ignoring the reference. We therefore introduce counterfactual supervision, which pairs one query with two references that support different normal rules and labels the query under each. At positions where the two labels disagree, no detector that ignores the reference can fit both targets. Anlu learns from this supervision by adding a reference memory and zero-initialized gated adapters to a frozen time-series foundation model (TSFM) pretrained for anomaly detection. On the 350 TSB-AD-U evaluation sequences, Anlu raises the mean VUS-PR of the frozen TSFM from 0.542 to 0.607. Replacing the reference with zeros lowers Anlu's score to 0.499.
|
| 1108 |
Two-Point Local Optimality in $k$-Means via Boundary-Point Screening
2610.06182
|
cs.LG
|
Wenlong Lyu, Xujie Xiao, Yuheng Jia |
Lloyd's algorithm and the discrete local (D-local) optimization method (Li et al., 2025) for $k$-means provide only weak local-optimality guarantees, and their solution quality remains sensitive to initialization. In this paper, we introduce $r$-point local op...Lloyd's algorithm and the discrete local (D-local) optimization method (Li et al., 2025) for $k$-means provide only weak local-optimality guarantees, and their solution quality remains sensitive to initialization. In this paper, we introduce $r$-point local optimality, under which no reassignment of at most $r$ samples decreases the objective function, and focus on $r=2$. The main computational obstacle is the $\mathcal{O}(n^2(k^2+d))$ cost of exhaustive two-point certification for $n$ samples in $d$ dimensions and $k$ clusters. To address this challenge, we prove that (i) every improving two-point move of a D-local optimum must involve a cluster shared by both reassignments, and (ii) only certificate-defined boundary points can participate in an improving pair. Exploiting this structure, we propose Boundary-Point-Screened Two-Point Local Search (BPS-2PLS), which terminates at a two-point local optimum. For fixed $k,d$ and nonvanishing cluster occupancy, the number $m$ of retained candidates satisfies $m=\mathcal{O}_{\mathbb{P}}(\log n)$ under i.i.d. sampling from a bounded-support distribution with bounded density or from a Gaussian mixture. Across twelve benchmarks, BPS-2PLS attains the lowest available mean WCSS on ten. In a subsampling study, screening retains 0.10% to 2.81% of samples on average at the largest tested sizes. The code is available at https://github.com/lwl-learning/BPS-2PLS.
|
| 1109 |
Structured Representation Learning for Behavior Cloning: How can we learn to safely control a nuclear power plant?
2610.06211
|
cs.LG
|
Perceval Beja-Battais (CB), Alain Grosset{\^e}te (CB), Nicolas Vayatis (CB) |
Learned models for industrial control are usually judged by aggregate accuracy, but accuracy at the component level does not guarantee safety once it is embedded in the system it is meant to serve. We study this gap on a behavior-cloning task: imitating an exp...Learned models for industrial control are usually judged by aggregate accuracy, but accuracy at the component level does not guarantee safety once it is embedded in the system it is meant to serve. We study this gap on a behavior-cloning task: imitating an expert Nonlinear Model Predictive Control (NMPC) policy for load-following of a Pressurized Water Reactor (PWR), an industrial system with tight safety constraints. We propose a structured architecture encoding variables from each timescale into separate latent spaces, reflecting the physical decomposition of the system, before training a controller to imitate the expert on the product latent space. On long-horizon rollouts, separated embeddings improve both accuracy and feasibility compared with a shared-embedding baseline. Sensitivity analysis further shows that our model yields interpretable representations aligned with the system's physics. However, standalone deployment still leaves several percent of trajectories infeasible regardless of the architecture. Using our method to warmstart the NMPC optimizer rather than acting standalone, we recover full feasibility and near-optimal cost while still cutting computation time by $\sim$15% relative to the expert controller, and even more for abrupt operating changes.
|
| 1110 |
Sampling Allocation of LinUCB: Optimal Design Limits in the Small-Gap Regime
2610.06213
|
cs.LG
|
Yujie Liu, Vincent Y. F. Tan, Yunbei Xu |
We study the sampling allocation of LinUCB in the small-gap regime, where the reward gaps are of order at most $n^{-1/2}$ over the decision horizon $n$. This scaling captures the hard instances underlying worst-case regret lower bounds, for which LinUCB is kno...We study the sampling allocation of LinUCB in the small-gap regime, where the reward gaps are of order at most $n^{-1/2}$ over the decision horizon $n$. This scaling captures the hard instances underlying worst-case regret lower bounds, for which LinUCB is known to be near optimal up to logarithmic factors in $n$. Using a mean-field perspective, we characterize this allocation through the empirical sampling distribution, a macroscopic object that averages the effect of adaptive decisions over the horizon, and identify its limit as $n\to\infty$. We establish that in this regime, the empirical sampling distribution induced by LinUCB converges to the set of D-optimal designs. This central result reveals that, in the small-gap regime, LinUCB not only achieves near optimal minimax regret but also allocates samples in a way that is asymptotically efficient for learning the reward parameter, thereby connecting regret-driven online learning with information-efficient experimental design. Building on the optimal design limit, we obtain two useful consequences. First, we refine the asymptotic regret analysis of LinUCB in the small-gap regime by characterizing its leading-order constant in the limit. Second, we show that, despite LinUCB's adaptive sampling strategy, the regularized least-squares estimator satisfies a central-limit-type theorem in the small-gap regime, thereby enabling valid statistical inference for the reward parameter.
|
| 1111 |
SO(3)-RoPE for Spherical Transformers
2610.06229
|
cs.LGcs.AI
|
Christian Libner, Chase van de Geijn, Alexander S. Ecker, Maurice Weiler |
Spherical data arise in many scientific applications. Often spherical transformers disregard the geometry of the underlying spherical domain, causing distortions and coordinate singularities near the poles. We introduce SO(3)-RoPE, a relative positional embedd...Spherical data arise in many scientific applications. Often spherical transformers disregard the geometry of the underlying spherical domain, causing distortions and coordinate singularities near the poles. We introduce SO(3)-RoPE, a relative positional embedding that incorporates spherical geometry into transformer attention through unitary SO(3) representations. Our formulation is SO(3)-equivariant and compatible with FlashAttention, retaining efficiency of vanilla transformers. On shallow water dynamics prediction over a rotating sphere, our SO3ViT outperforms an S2Transformer baseline with lower errors and reduced runtime.
|
| 1112 |
Parameter Estimation in Machining Dynamics with Regenerative Delay and Nonsmooth Friction using Physics-Informed Neural Networks
2610.06230
|
cs.LG
|
Meiyazhagan Jaganathan, Vikram Pakrashi, Aasifa Rounak |
A multi-domain eXtended Physics-Informed Neural Network (XPINN) framework is developed for nonsmooth Delay Differential Equations (DDEs). This is the first implementation to demonstrate the efficacy of partitioning the temporal domain into subdomains of intege...A multi-domain eXtended Physics-Informed Neural Network (XPINN) framework is developed for nonsmooth Delay Differential Equations (DDEs). This is the first implementation to demonstrate the efficacy of partitioning the temporal domain into subdomains of integer multiples of the characteristic time delay and progressively training the associated subnetworks while freezing previously learned parameters. The efficacy of the proposed framework is demonstrated using a machining dynamics model that incorporates both regenerative and nonsmooth frictional effects. Results demonstrate that the proposed multi-domain XPINN framework leads to better solution reconstruction in DDEs and improved parameter estimation compared to a generic PINN (SPINN) formulation. The proposed method works particularly well for extended temporal domains and non-constant history functions. The robustness of inverse XPINN (I-XPINN) is also assessed using reference data contaminated with Gaussian measurement noise. Results indicate that I-XPINN remains resilient to measurement noise and the physics-informed constraints guide the network toward accurately recovering the underlying dynamics. This demonstrates, for the first time, the potential of the proposed framework for reliable parameter identification in DDEs characterised by nonsmoothness and large time delays.
|
| 1113 |
Certification-Enhanced Generalization Bounds
2610.06238
|
cs.LG
|
Leo Elmecker-Plakolm, Matthew Wicker |
We investigate the use of formal methods to provide tight and sound generalization bounds for learning algorithms. By casting the traditional notion of algorithmic stability as a specification to be verified, we demonstrate that recent advances in reachability...We investigate the use of formal methods to provide tight and sound generalization bounds for learning algorithms. By casting the traditional notion of algorithmic stability as a specification to be verified, we demonstrate that recent advances in reachability analysis can yield provable bounds on the generalization of a given model and algorithm on a sample dataset. As sample-specific algorithmic stability is insufficient to bound the usual distributional notion of generalization, we develop a novel concentration inequality to connect the sample-specific results of formal certification algorithms to the required distributional analysis for bounding the expected generalization gap. The resulting framework enables the analysis of prior generalization bounds to extend far beyond their original restrictive assumptions. Our approach computes sound bounds on the expected generalization gap in a constant number of algorithm runs without making any analytical assumptions on the algorithm; to achieve non-vacuous bounds we only require that the certified reachable parameter set is bounded --- a condition that we do not assume but formally verify. In practice, we demonstrate that our framework provides formal generalization guarantees that are orders of magnitude tighter than alternative sound computational approaches at scales ranging from toy datasets to fine-tuning classification heads on top of modern large language models. While we implement certification-enhanced versions of several well-known stability results, future extensions of our approach will enable tighter bounds and enhanced practical adoption across the spectrum of modern generalization bounds.
|
| 1114 |
Few-Shot Prototype Head Adaptation for On-Device ECG Personalization on PSoC~6
2610.06241
|
cs.LGcs.AI
|
Guilherme Silva, Pedro Silva, Gladston Moreira, Eduardo Luz |
Wearable and bedside electrocardiogram (ECG) monitors must adapt to patient-specific morphology to maintain arrhythmia detection accuracy across users, yet personalization is typically performed offline and cannot account for individual physiology, electrode p...Wearable and bedside electrocardiogram (ECG) monitors must adapt to patient-specific morphology to maintain arrhythmia detection accuracy across users, yet personalization is typically performed offline and cannot account for individual physiology, electrode placement, or recording drift. On-device adaptation by backpropagation is expensive for microcontroller-class medical devices because it requires an optimizer state, repeated backward passes through convolutional layers, and labeled arrhythmic beats that may not be available at deployment time. This letter proposes prototype-only head adaptation as a compact personalization primitive for TinyML ECG systems. A one-dimensional convolutional neural network (1-D CNN; 1,314 parameters and 72.6k multiply-accumulate operations per beat) is trained offline on the MIT-BIH Arrhythmia Database under an inter-patient protocol, frozen as a feature extractor, and exported to a PSoC 6 microcontroller. Patient-specific adaptation then reduces to computing closed-form class means in a 32-dimensional embedding space, requiring no convolutional backward pass, no iterative optimization, and only one forward pass per support beat. Prototype adaptation improves inter-patient macro-F1 from 0.635/0.639/0.646 to 0.731/0.771/0.797 at 1/5/10-shot, outperforming linear stochastic-gradient-descent (SGD) head fine-tuning at every shot count for the target tiny backbone. On-device replay over 18 one-shot episodes on a PSoC 6 Cortex-M4F matches the host macro-F1 for the prototype head (0.798), with 11.39 ms per beat, 5.2 KB flash, and 22.2 KB SRAM. A restricted variant that updates only the normal-class prototype from passively buffered sinus beats yields a consistent +0.05 macro-F1 gain, reducing the annotation burden during initial
|
| 1115 |
RoSA: Rotational Sparse Adaptation for Memory-Efficient Fine-Tuning
2610.06243
|
cs.LG
|
Muhammad Azeem Lodhi, Chao Zhou, Rebekka Burkholz |
Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers...Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers within each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
|
| 1116 |
Constrained Goal-directed Planar Graph Generation with Grammar-based Reinforcement Learning
2610.06244
|
cs.LG
|
Nicolas Hochuli, Lorenzo Miele, Kristina Shea, Tino Stankovic |
Planar graphs are central to applications across science and engineering, yet existing generators provide limited support for goal-directed generation under hard structural and geometric feasibility constraints. We propose a dataset-free method for generating ...Planar graphs are central to applications across science and engineering, yet existing generators provide limited support for goal-directed generation under hard structural and geometric feasibility constraints. We propose a dataset-free method for generating planar graph embeddings by combining parametric graph grammars with safe reinforcement learning to optimize generic task-specific objectives while satisfying constraints during construction. We formulate the generation process as a constrained Markov decision process, where the graph grammar defines the state and action spaces. We further introduce an action projection that maps sampled actions toward state-dependent safe sets, improving constraint satisfaction during training. In contrast to classical graph generators and deep generative models, which typically offer limited goal-directed control or rely on weak constraint satisfaction, our method constructs feasible planar graph embeddings directly during generation. We also introduce a benchmark suite for constrained and goal-directed planar graph generation, together with classical and deep generative baselines. Across all benchmark tasks, our method consistently outperforms baselines while satisfying the formulated constraints.
|
| 1117 |
Generative World Models Enable Predictive Control of Laser Melt Pool Dynamics
2610.06250
|
cs.LG
|
Yiyang Yan, Markus Bambach, Mohamadreza Afrasiabi |
World models, which learn how environments respond to actions, are emerging as a powerful paradigm for planning through imagined futures, transforming decision-making across games, robotics and autonomous driving. Bringing this capability to manufacturing coul...World models, which learn how environments respond to actions, are emerging as a powerful paradigm for planning through imagined futures, transforming decision-making across games, robotics and autonomous driving. Bringing this capability to manufacturing could enable process decisions on timescales inaccessible to high-fidelity simulation. Here we introduce a generative world model for localized highly dynamic laser melt pool that predicts evolution from histories of temperature and phase morphology under candidate actions. Its generative latent dynamics capture the effects of unresolved melt flow, enabling more accurate recursive rollouts than deterministic regressors under transient laser inputs. Because the learned dynamics are differentiable, the model can serve directly as a predictive control plant. Gradients through imagined futures optimize laser schedules that regulate melt-pool depth over previously unseen geometry, path, initialization. We further distil this optimization into an amortized policy that produces control actions in a single forward pass, providing a proof of concept for real deployment on machines.
|
| 1118 |
Lipschitz Thinking: Ten Years of Certifiable-by-Design Robust Neural Networks
2610.06252
|
cs.LG
|
Fabio Brau, Giorgio Piras, Maura Pintor, Battista Biggio |
The Lipschitz property of a deep neural network provides a direct measure of its sensitivity to input perturbations and, when explicitly controlled, offers a principled way to limit the propagation of errors and improve robustness. Over the past decade, Lipsch...The Lipschitz property of a deep neural network provides a direct measure of its sensitivity to input perturbations and, when explicitly controlled, offers a principled way to limit the propagation of errors and improve robustness. Over the past decade, Lipschitz-bounded layers have been incorporated into increasingly expressive and high-performing deep models, narrowing the gap between empirical robustness and formal, by-design guarantees of stability. This article introduces the fundamental concepts underlying Lipschitz-bounded neural networks, explaining the principles behind Lipschitz-constrained layers, the mechanisms used to enforce their bounds, and how they yield robustness certificates at the cost of a single forward pass. The tutorial concludes by discussing emerging and open directions, highlighting Lipschitz control as a general framework for offering guaranteed, by-design stability.
|
| 1119 |
SimAuthor: Harnessing Foundation Models for Persistent Scientific Simulator Authoring
2610.06257
|
cs.LG
|
Yishan Wang, Ran Piao, Mathias Funk, Aaqib Saeed |
Foundation models can generate scientific code, but authoring a scientific simulator (an executable program encoding hypotheses about how mechanisms generate observable signals) requires iterative refinement. Scientific adequacy rarely admits a unique implemen...Foundation models can generate scientific code, but authoring a scientific simulator (an executable program encoding hypotheses about how mechanisms generate observable signals) requires iterative refinement. Scientific adequacy rarely admits a unique implementation or exact test, so simulators must instead be judged against limited real observations. We study this setting as scientific simulator authoring under weak empirical feedback, where distributional comparisons between simulated and real signals guide revision, and the target is the simulator itself rather than only its generated samples. We introduce SimAuthor, a persistent authoring harness that retains and revises executable simulators, separates scalar search scores from structured discrepancy feedback, and accumulates reusable implementation mechanisms. We evaluate SimAuthor on six biomedical tasks spanning cardiac and respiratory audio, photoplethysmography (PPG), and electrocardiography (ECG). Under a fixed 100-attempt budget, SimAuthor outperforms PUCT score search on all six tasks, generally outperforms textual-strategy optimization, and achieves the highest endpoint score on five of six. The authored simulators also improve on unseen recordings, transfer to independent pretrained representations, and yield substantial out-of-distribution gains in downstream ECG classification. Finally, 111 of 138 audited revisions alter program structure and account for 86.1% of the signed score improvement. These results suggest that persistent revision can progressively convert foundation-model knowledge into better executable scientific simulators from limited empirical evidence.
|
| 1120 |
OCL-PDE: A Generative Framework for PDE Inverse Problems with Observation-Complementary Latents
2610.06259
|
cs.LG
|
Ding Yang, Chuqi Chen, Chang Ma, Yang Xiang |
Partial differential equation (PDE) inverse problems are often ill-posed, making fine-scale details difficult to recover. We address this problem by introducing a learned observation-complementary latent representation that preserves reconstruction-relevant in...Partial differential equation (PDE) inverse problems are often ill-posed, making fine-scale details difficult to recover. We address this problem by introducing a learned observation-complementary latent representation that preserves reconstruction-relevant information and is combined with the observation to reconstruct the unknown field. Building on this representation, we propose OCL-PDE, a generative framework that encourages the observation to guide large-scale structure and the latent to supply complementary fine-scale details. OCL-PDE is built on a physics-aware autoencoder (AE) and conditional Flow Matching, supporting inverse reconstruction as well as forward PDE prediction. Experiments demonstrate improved reconstruction accuracy and fine-detail recovery compared with the evaluated baselines.
|
| 1121 |
Ramp Metering Control via Hybrid State Deep Reinforcement Learning in Partially Observable Connected Vehicle Environments
2610.06266
|
cs.LGcs.AI
|
Youcef Mehamlia, Nadir Farhi, Meriem Bouali |
Freeway on-ramp merges are major sources of congestion, causing significant economic and environmental costs. While Deep Reinforcement Learning (DRL) offers a promising solution for ramp metering, existing approaches rely primarily on aggregated macroscopic da...Freeway on-ramp merges are major sources of congestion, causing significant economic and environmental costs. While Deep Reinforcement Learning (DRL) offers a promising solution for ramp metering, existing approaches rely primarily on aggregated macroscopic data. Connected vehicles (CVs) provide vehicle-level observations that can complement aggregate traffic measurements, but their limited penetration produces incomplete microscopic information. This paper proposes a hybrid observation representation combining macroscopic traffic measurements with a two-channel grid encoding observed CV presence and speed. A Dueling Double Deep Q-Network processes these inputs to select ramp-metering green durations. The controller is trained under varying traffic demands and CV penetration rates and evaluated against ALINEA and macroscopic-only DRL variants in SUMO. Across 50 matched evaluation scenarios, the hybrid controller under partial CV visibility reduces the reported total travel time by 11.4 % and mean spillback duration by 84.9 % relative to ALINEA. Evaluating the same trained policy with full CV visibility yields a further travel-time reduction of approximately 1.6 %. Analysis across penetration rates suggests that the performance gap decreases as microscopic observations become more complete. These results support the use of complementary macroscopic and sparse microscopic observations for learning-based ramp metering. The source code implementation of the model is available at: https://github.com/youcefMehamlia/Multimodal-DRL-RMC
|
| 1122 |
What May an Agent Change About Itself? A Containment Floor for Self-Configuring Agent Runtimes
2610.06274
|
cs.LG
|
Sajib Hossain, Moeen Uddin Mahmud |
Many agent runtimes give the agent a tool for editing its own configuration. Some of that configuration grants abilities, such as enabling a tool. Other parts set the agent's limits: which directories it may write to, who may send it messages, which network ad...Many agent runtimes give the agent a tool for editing its own configuration. Some of that configuration grants abilities, such as enabling a tool. Other parts set the agent's limits: which directories it may write to, who may send it messages, which network address it listens on, how callers authenticate, and the gate that blocks risky writes. If the agent can edit those limits, a single ordinary request can widen them. We study this in a deployed, model-agnostic runtime. We propose a rule: the agent may change fields that grant abilities, and may never change fields that set its limits. We enforce the rule as a containment floor inside the configuration tool and measure what happens with and without it. Without the floor, a frontier model wrote a protected value on 25 of 72 ordinary requests that gave it permission to change settings, often when the request never named the field. Prohibitions written in the system prompt failed in a predictable way. A prompt that listed the protected field names stopped every request that used those names (0 of 36 saved, against 17 of 36 with no prompt) and did not stop the requests that only described the goal (10 of 36 saved, against 8 of 36). A prompt that described the forbidden effects did the reverse. With the floor, 0 of 167 protected writes were saved, although the models attempted a protected write in 65 of those cases. A search for other routes through the tool found only one, a pinned shell, which the floor's scope statement already excludes. The study covers two models and a single agent. We state what that does and does not support.
|
| 1123 |
dIon: Fragmentation-Based Invariance for Self-Supervised Learning of Tandem Mass Spectra
2610.06282
|
cs.LG
|
Alfred Nilsson, Joel Lapin, Samuel H. Payne, Mathias Wilhelm, Lukas K\"all |
We introduce a novel invariance for peptide tandem mass spectrometry data, unlocking self-supervised representation learning that improves de novo sequencing of peptides. This invariance exploits the physical relationship between precursor properties (mass and...We introduce a novel invariance for peptide tandem mass spectrometry data, unlocking self-supervised representation learning that improves de novo sequencing of peptides. This invariance exploits the physical relationship between precursor properties (mass and charge) and fragment-ion evidence, without requiring peptide sequence labels. We introduce dIon, which adapts the DINO framework with two latent prediction tasks, both recovering a clean teacher representation: one from a spectrum mixture, using the precursor as a selection query, and one from a partial spectrum with the precursor withheld. The first associates precursor information with fragment-ion evidence; the second prevents representational collapse onto that information alone. Mechanistic probes support both effects, and ablations show that the full objective performs best. Under identical end-to-end training, dIon initialization improves de novo peptide precision over training from scratch by 5.5 and 8.4 percentage points on the held-out MassIVE-KB and Kingdoms test sets, and by 2.3 and 4.8 percentage points with a larger supervised training corpus. The resulting models surpass fully supervised state-of-the-art de novo sequencing models on the diverse, multi-species Kingdoms corpus under the same greedy-decoding protocol. Without peptide labels, dIon learns strong native peptide-similarity geometry compared with other learned models; with limited peptide-supervised adaptation, it achieves the best retrieval and pair-discrimination performance across all representation benchmarks.
|
| 1124 |
When Are Concept Bottleneck Model Explanations Faithful and Compact?
2610.06285
|
cs.LGcs.AI
|
Stefano Teso, Emanuele Marconato, Steve Azzolin, Antonio Vergari |
Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal e...Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based variants, faithful explanations must include all concepts in the bottleneck, compromising interpretability when this is large. This result applies to both heuristic and faithful-by-construction formal explanations. To encourage the existence of compact faithful explanations, we suggest i) modeling concepts probabilistically as binary or categorical random variables (rather than logits), and ii) employing per-concept training-time sparsification via group lasso (rather than regular elastic net). We also extend algorithms from formal explainability to CBMs, and show they outperform natural heuristics in terms of guarantees and explanation size. Overall, our work warns against naive interpretability claims and provides formal conditions and practical strategies for ensuring CBMs are as interpretable as advertised.
|
| 1125 |
Trajectory-Guided Tokenization of Complex CSI for Wi-Fi Sensing
2610.06288
|
cs.LG
|
Ziyi Wang, Kenuo Xu, Jichu Jiang, Yumeng Yang, Zheng Chen |
Wi-Fi channel state information (CSI) enables contactless presence detection and gesture recognition. Its high-dimensional complex-valued time series require input representations that preserve informative temporal variations during compression. We propose Tra...Wi-Fi channel state information (CSI) enables contactless presence detection and gesture recognition. Its high-dimensional complex-valued time series require input representations that preserve informative temporal variations during compression. We propose Trajectory-Guided Tokenization (TGT), which combines complex trajectory decomposition with asymmetric attention to construct compact continuous tokens. For each antenna link and subcarrier, an orthonormal Helmert transform decomposes short, ordered temporal patches into local-center and centered-trajectory coordinates. Keys are learned from the centered-trajectory coordinates, while values retain both components. Learnable queries aggregate subcarriers into frequency slots, which are fused into temporal tokens. Trained jointly from scratch, TGT with TokenMLP achieves the highest mean accuracy of 92.83% among all evaluated frontend-backend combinations on the self-collected dataset. Experiments on EHUNAM and Widar further support the applicability of TGT to cross-domain presence detection and gesture recognition.
|
| 1126 |
SPDAlign: Interpretable Riemannian Alignment for EEG Forward Modeling Shifts
2610.06315
|
cs.LG
|
Shanglin Li, Shiwen Chu, Okan Ko\c{c}, Chenyu Liu, Qibin Zhao |
Electroencephalography (EEG) based brain-computer interfaces enable direct brain-to-device communication for applications such as rehabilitation and communication. However, their practical utility is often limited as the non-stationary nature of the EEG data i...Electroencephalography (EEG) based brain-computer interfaces enable direct brain-to-device communication for applications such as rehabilitation and communication. However, their practical utility is often limited as the non-stationary nature of the EEG data introduces distribution shifts across domains (e.g., sessions and subjects). Adapting machine learning models to be invariant to these shifts in an unsupervised way, without using costly labeled calibration data, would drastically improve the utility of EEG data. In this work, we use a classic generative model of EEG to study distribution shifts introduced by the domain-specific forward process, which is associated with factors such as head geometry. We theoretically show that such distribution shifts can be recovered solely through linear transformations on the Symmetric Positive Definite manifold. Building on this insight, we propose SPDAlign, an interpretable framework for promoting domain-invariant EEG learning. SPDAlign first aligns the domain-specific means and corrects global rotations across domains using a recent optimal transport technique called Wasserstein Procrustes. We systematically study the proposed approach through simulations and demonstrate its competitive performance on extensive public EEG datasets. Additionally, SPDAlign is a globally linear framework and is intrinsically interpretable, so that the framework can identify frequency ranges of interest, determine the spatial patterns reflecting source-sensor relationships, and address cross-subject variability.
|
| 1127 |
Dynamic Minimax Regret Optimization for Robust LLM Post-Training
2610.06329
|
cs.LG
|
Chengbo Zang, Haoyu Dong, Mehmet Kerem Turkcan, Gil Zussman, Zoran Kostic |
Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mi...Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mini-batch-only bandit feedback. The framework views the training as a two-player sampler-optimizer process: a sampler adaptively selects among data sources using bandit feedback, while an optimizer updates the model parameters using stochastic gradients from the selected source. We focus on the practically restrictive setting where source losses evolve with model training but historical data are not re-evaluated, requiring the sampler to track instantaneous worst-sources from stale partial feedback. We propose DUCB-OGD, a simple and scalable algorithm that couples a Discounted Upper-Confidence-Bound sampler with an Online Gradient Descent optimizer. The sampler maintains exponential moving average loss estimates and confidence radii based on discounted effective sample sizes, avoiding costly re-evaluation of past data or intrusive changes to standard training pipelines. For $K$ data sources and $T$ training steps, we prove that DUCB-OGD achieves a dynamic minimax regret of $\tilde{O}(K^{1/4}T^{3/4})$, which is optimal up to logarithmic factors for the undiscounted objective under our feedback model. Extensive experiments across supervised fine-tuning, preference optimization, and reinforcement learning show that DUCB-OGD integrates seamlessly into modern LLM training pipelines and improves worst-group robustness with negligible computational overhead compared with standard sampling baselines.
|
| 1128 |
Stability-Shaped Deep Graph Learning
2610.06344
|
cs.LG
|
Junyou Zhu, Langzhou He, Fenying Cai, Christian Nauck, Ping Xiong |
In deep graph neural networks, increasing depth enlarges the receptive field but often leads to over-smoothing, where node representations tend to align. We develop a unified, mode-wise stability framework for deep GNN propagation that provides a principled ch...In deep graph neural networks, increasing depth enlarges the receptive field but often leads to over-smoothing, where node representations tend to align. We develop a unified, mode-wise stability framework for deep GNN propagation that provides a principled characterization of over-smoothing. By interpreting layer depth as time and layer updates as graph-coupled dynamics, over-smoothing can be understood as an undesirable dynamical synchronization of features, for which the master stability curve provides a theoretical tool to assess the stability of synchrony. Guided by this theory, we further propose Stability-Shaped Deep Graph Learning (SDGL) to mitigate over-smoothing in deep GNNs. SDGL has two complementary instantiations: one induces controlled Turing instability to replace synchronization with spatial pattern formation, and the other maintains stable near-critical propagation. Experiments on diverse node- and graph-level benchmarks demonstrate the improved depth scaling and consistent accuracy gains over strong baselines, including graphs exhibiting long-range dependencies.
|
| 1129 |
Multimodal Deep Survival Analysis for Sinkhole Susceptibility
2610.06365
|
cs.LG
|
Lucas Yuan, Minhee Kim, Zihan Li, Chunli Dai, Sanduni S. Disanayaka Mudiyanselage |
Sinkholes are a widespread geohazard in karst terrain. In Florida, soluble carbonate bedrock, shallow groundwater, and intense rainfall combine to make subsidence both common and spatially heterogeneous. Predicting where and when sinkholes will occur is diffic...Sinkholes are a widespread geohazard in karst terrain. In Florida, soluble carbonate bedrock, shallow groundwater, and intense rainfall combine to make subsidence both common and spatially heterogeneous. Predicting where and when sinkholes will occur is difficult for two reasons. First, locations without reported sinkholes cannot be directly labeled or sampled as true negative locations. Second, the potential factors governing sinkhole risk span heterogeneous data modalities and therefore require careful integration within a unified modeling framework. We address both problems with our proposed model, a multimodal Cox proportional hazards framework for sinkhole susceptibility. Our contributions are threefold. First, we extend the proportional-hazards formulation to heterogeneous multimodal input through modality-specific encoders and a cross-modal fusion layer. Second, we treat unreported locations as right-censored rather than negative, avoiding hard-negative labeling and yielding continuous, time-aware susceptibility from the predicted survival function. Third, a statewide Florida case study with spatially blocked validation and ablation studies quantifies the benefit of multimodal integration. A Florida case study demonstrates that the proposed method effectively ranks sinkhole risk and produces a statewide susceptibility map that captures spatial variations in sinkhole occurrence.
|
| 1130 |
Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models
2610.06387
|
cs.LGcs.AI
|
Joe Dwyer |
Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs,...Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs, and barriers to participation for researchers operating outside large industrial laboratories. This review examines the evolution of LLM scaling theory from empirical scaling laws to compute-optimal training, with particular emphasis on parameter efficiency, token utilization, data efficiency, and resource-constrained environments. Foundational work on scaling laws is synthesized alongside later research on compute-optimal training, data pruning, efficient architectures, quantization, low-rank adaptation, and edge-oriented optimization. The literature indicates a shift from scale maximization toward more deliberate allocation of parameters, tokens, compute, and hardware resources. At the same time, important empirical, theoretical, and methodological gaps remain regarding whether scaling principles established on enterprise-grade infrastructure generalize to smaller models and constrained computing environments. This review organizes these developments into a unified framework for resource-efficient LLM training and argues that future progress should evaluate efficiency not solely through model performance, but through the relationship among performance, parameter count, computational cost, token allocation, and hardware constraints.
|
| 1131 |
Learning Pareto Stationary Fronts via Single-Pass Backpropagation
2610.06397
|
cs.LG
|
Elina Rojin Celik, Marcos Medeiros Raimundo, Isabel Valera |
We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-ob...We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference alignment. Casting the bilevel optimization problem as a Stackelberg game enables solving the original a posteriori MOO problem in a single forward-backward pass. As a result, MOSEL matches the time and memory efficiency of standard single-objective training while enabling scalable Pareto stationary front learning. Empirically, MOSEL uncovers diverse and optimal Pareto frontiers in strongly conflicting settings (e.g., fairness-accuracy). Remarkably, even in weakly conflicting regimes such as multi-task learning, it consistently converges to solutions closer to the utopia point, outperforming both standard single-objective training and specialized multi-task learning methods. These results highlight the broader potential of a posteriori MOO learning as a pathway to efficiently learn more diverse and robust representations, ultimately improving generalization.
|
| 1132 |
Quantifying the Stability of Multi-Step Reasoning via Error Amplification
2610.06404
|
cs.LGcs.AI
|
Dongyue Li, Ziniu Zhang, Minxuan Duan, Hongyang R. Zhang |
We consider the stability of multi-step reasoning processes, which have extensive applications in language models, including chain-of-thought and algorithmic reasoning. While longer sequences of reasoning can improve a model's generation capability at test tim...We consider the stability of multi-step reasoning processes, which have extensive applications in language models, including chain-of-thought and algorithmic reasoning. While longer sequences of reasoning can improve a model's generation capability at test time, the errors due to intermediate reasoning steps can accumulate in autoregressive generation, and thus grow substantially at the end. In this paper, we ask: What are the key factors determining the stability of multi-step reasoning? First, we show an inference error bound governed by the product of spectral norms of the Jacobians taken through the input space across generation steps. This product can be viewed as an error amplification factor, which could scale exponentially with the number of reasoning steps, serving as a quantitative measure of reasoning stability. Second, we analyze this measure in transformer models trained to predict simple tasks like linear and quadratic functions. We theoretically prove that the transformer model converges to a solution where the stability measure decays, thus yielding nearly zero inference loss over (arbitrarily) long steps. Finally, the stability analysis leads to several algorithmic implications for controlling the stability, through (i) chain-of-thought length compression that reduces the sensitivity of each step, and (ii) quantization-aware training that regularizes the input Jacobian norms. We validate the proposed algorithms by fine-tuning language models on graph-algorithmic reasoning tasks and symbolic state-tracking tasks. Across seven evaluations, our algorithms improve over baseline comparisons by 3.5% on average, and by 8.2% for longer-length inputs. Ablation analysis validates that the stability measure is drastically reduced by 3-8$\times$, confirming the regularization effect on the spectral norms of the (input space) Jacobians.
|
| 1133 |
FlashCart: Fast Cartesian Tensor Products for Equivariant Interatomic Potentials
2610.06409
|
cs.LG
|
Viktor Zaverkin, Payman Goodarzi, Sergey V. Sukhomlinov, Davit Hovhannisyan, Roland Aydin |
Machine-learned interatomic potentials extend atomistic simulations beyond the length- and timescales accessible to electronic-structure methods. However, the computational cost of equivariant architectures limits the local correlations they can represent in p...Machine-learned interatomic potentials extend atomistic simulations beyond the length- and timescales accessible to electronic-structure methods. However, the computational cost of equivariant architectures limits the local correlations they can represent in practice and therefore their achievable accuracy. Here we introduce FlashCart, which makes higher-order correlations affordable by combining generated GPU kernels with an architecture that recursively builds equivariant features and compresses them to a fixed width at each step. We express tensor products in independent Cartesian components and symbolically simplify them and their derivatives, producing fused kernels that often outperform optimized spherical counterparts. We then show that increasing correlation order improves accuracy more efficiently than increasing width, depth, or tensor rank. On SPICE-MACE-OFF, FlashCart models advance the measured accuracy-efficiency frontier: a model with $5.6$ million parameters achieves lower energy and force errors and $10\times$ faster inference than a transformer with $189$ million parameters.
|
| 1134 |
Training-Free Transformer Merging via Sequential Local Operator Alignment
2610.06415
|
cs.LG
|
Akansh Maurya, Ya-Wei Eileen Lin, Stefanie Jegelka, Sebastian U Stich, Rotem Mulayoff |
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-k...Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, overlooking the functional structure. Moreover, when merging earlier components, downstream components receive different activations than they do in the original model, thus, the merged and original execution paths no longer match. In this paper, we introduce Sequential Local Operator Alignment, a training-free method that merges transformers along the execution path of the partially merged model. Our method uses calibration data to estimate the local behavior of each functional component, aligns operators sequentially under the intermediate activation of the partially merged model, and subsequently factorizes the merged operators back into valid transformer parameters. We empirically show that this sequential step reduces error accumulation across layers. Furthermore, the proposed operator factorization step enables rank expansion, providing a principled mechanism for increasing multi-task capacity. We demonstrate that our approach generalizes across modalities, model scales, and varying numbers of tasks, from CLIP and RoBERTa to billion-parameter LLMs, and further extends naturally to the merging of LoRA-fine-tuned models. The results indicate improvements over strong merging baselines without requiring rank expansion, while optional expansion provides a further accuracy-inference-cost trade-off. Project link: https://akansh12.github.io/SLOA-Merge/
|
| 1135 |
Efficient Secure Federated Learning via Information-Theoretically Secure Key Distribution: A Medical Imaging Case Study
2610.06420
|
cs.LG
|
Ivan Don\`a, Hans H. Brunner, \'Alvaro Troyano Olivas, Chi-Hang Fred Fung, Momtchil Peev |
Federated Learning (FL) enables collaborative training of models across institutions without centralizing sensitive data, making it well-suited for privacy-concerned applications, such as medical imaging. To protect FL model updates during secure aggregation, ...Federated Learning (FL) enables collaborative training of models across institutions without centralizing sensitive data, making it well-suited for privacy-concerned applications, such as medical imaging. To protect FL model updates during secure aggregation, additive masking is commonly employed. However, its underlying classical key establishment is only computationally secure. On the other hand, physics-based Information-Theoretically Secure (ITS) key exchange introduces practical constraints: finite key generation rates and time-limited storage severely limit throughput and sustained training of uncompressed models. In this work, we address this bottleneck by developing an FL framework that integrates frozen backbones, knowledge distillation, and quantization. These techniques reduce communication payload and, consequently, key material consumption. Moving beyond simulation, we benchmark this framework on a real physics-based key distribution testbed involving a chest X-ray classification application. Our results show that key usage can be reduced by $\sim$35$\times$ while maintaining predictive accuracy. This prevents buffer depletion and key expiration, enabling sustainable FL training under physical key generation constraints.
|
| 1136 |
ARO: Aligned Representation learning for multi-Omics data
2610.06443
|
cs.LGcs.AI
|
Amogh Singh, Yash Shah, Chiara D'Ercoli, Arash Mehrjou, Patrick Schwab |
The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, t...The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities. Contrary to increasingly complex, larger models, e.g. Foundation Models (FMs), ARO prioritizes practical applicability in limited or incomplete data settings. ARO optimally reconstructs missing modalities (MSE of $0.15$ on the validation and test data in the Unmasked settings), with its learned latent embeddings enabling a downstream cancer classification task. Our findings indicate that analyzing diverse molecular layers as a single integrated system offers a reliable and cost-efficient approach, reducing dependence on large-scale experimental testing, while still supporting multi-omic exploration in limited data settings.
|
| 1137 |
Time-series Foundation Models for Predictive Control: The Role of Excitation
2610.06447
|
cs.LG
|
Mazen Amria, Jasper Hoffmann, Philipp Bordne, Anna Rothenh\"ausler, Lilli Frison |
Deploying model predictive control (MPC) requires constructing or identifying a predictive model for each target system. Time-series foundation models (TSFMs) offer an attractive option thanks to strong zero-shot forecasting capabilities across systems. Howeve...Deploying model predictive control (MPC) requires constructing or identifying a predictive model for each target system. Time-series foundation models (TSFMs) offer an attractive option thanks to strong zero-shot forecasting capabilities across systems. However, low forecast error does not guarantee that a TSFM captures the system's response to the alternative actions considered by the controller. We study this gap using residential heat-pump control as a test bed, measuring the agreement between predicted and ground-truth effects of control interventions. Importantly, we find that TSFMs can recover the system's input-response relationship when the context contains sufficient independent control excitation. Common fine-tuning pipelines and feature smoothing reduce, but do not eliminate, the need for in-context excitation. Our results indicate that current TSFMs used for predictive control require sufficiently informative control variation in the inference context. Initial closed-loop results show promise for shorter context windows.
|
| 1138 |
EMG-FM-Bench: A Comprehensive Benchmark for Foundation Model Transfer and Adaptation on Electromyography
2610.06450
|
cs.LG
|
Tianhao Wu, Xu Wu, Amirmohammad Radmehr, Jiawei Yu, Yi Wu |
Foundation models (FMs) are increasingly being developed for general time series and physiological signals, yet their transferability to downstream physiological tasks remains poorly understood. This question is particularly challenging for electromyography (E...Foundation models (FMs) are increasingly being developed for general time series and physiological signals, yet their transferability to downstream physiological tasks remains poorly understood. This question is particularly challenging for electromyography (EMG), where signal distributions vary substantially across users, sensing configurations, acquisition hardware, and downstream tasks. We introduce EMG-FM-Bench, a systematic benchmark for studying foundation-model transfer and adaptation on EMG. EMG-FM-Bench unifies 20 public datasets with over 1 million EMG segments and evaluates nine pretrained foundation models across four questions: how pretrained models perform when frozen or fully fine-tuned, how much pretraining helps compared with training the same model from scratch, how well models generalize to new users with limited labeled data, and how performance changes across different EMG tasks. Across the benchmark, linear probing provides useful information about pretrained representations, but full fine-tuning can substantially change downstream EMG performance. Comparing each pretrained model with the same model trained from scratch shows that the benefit of pretraining varies substantially across models and is not universal. Performance decreases when models are evaluated on new users, while five-shot adaptation improves macro-F1 in 70.2% of evaluated model-dataset combinations but recovers only part of the lost performance. Model performance is highly consistent between upper- and lower-limb classification and remains strongly correlated with continuous EMG-to-text decoding. Together, these results provide a systematic view of when pretrained time-series models transfer effectively to EMG and how their performance depends on fine-tuning, user variation, and downstream task.
|
| 1139 |
A Physics-Guided Transformer Framework for Electromigration Analysis in Multi-Segment Interconnects
2610.06464
|
cs.LGcs.AI
|
Pavlos Stoikos, Anuj Pathania, George Floros |
As technology scales to smaller nodes, increasing current densities make electromigration (EM) one of the dominant reliability challenges in on-chip interconnects. Accurate transient stress analysis is needed to identify wires susceptible to EM degradation, bu...As technology scales to smaller nodes, increasing current densities make electromigration (EM) one of the dominant reliability challenges in on-chip interconnects. Accurate transient stress analysis is needed to identify wires susceptible to EM degradation, but applying physics-based solvers across many interconnects remains computationally expensive. This paper proposes a physics-guided transformer framework for fast EM stress prediction in multi-segment interconnect lines. The framework converts each line into geometry- and DC-aware segment tokens and uses transformer attention to capture line-level context. A lightweight query decoder then predicts stress at selected locations and time instants. The model is trained with an objective that combines normalized supervised regression, linewise relative-$L_2$ loss, and physics-guided continuity and terminal-flux terms. Experiments on IBM power grid benchmarks show that the proposed model achieves relative-$L_2$ error below 8\% and reaches up to 2459.68$\times$ speedup compared with the matrix exponential~solver.
|
| 1140 |
Latent Flow Matching for Molecular Graph Generation
2610.06468
|
cs.LGcs.AI
|
Mathis Goupillon, Roman Bresson, Konstantinos Divriotis, Michalis Vazirgiannis |
Modern graph generative models typically operate directly in the discrete graph space, explicitly generating node and edge variables, which can become costly as graphs grow. In this paper, we perform generation explicitly on latent representations of entire gr...Modern graph generative models typically operate directly in the discrete graph space, explicitly generating node and edge variables, which can become costly as graphs grow. In this paper, we perform generation explicitly on latent representations of entire graphs obtained from a pretrained Variational Autoencoder with high reconstruction fidelity. The generated representations, obtained through flow matching, are then decoded only at the final step. Across molecular benchmarks of increasing size, our approach achieves strong validity and FCD while offering a favorable quality-efficiency trade-off compared with state-of-the-art explicit graph generative models. One of the main advantages of this formulation is that the graph representation only needs to be learned once, after which the same one can be reused across multiple generative objectives without retraining. We demonstrate generation guided by molecular properties and further introduce validity-aware generation though a classifier learned directly in latent space. All code will be made available upon acceptance.
|
| 1141 |
FairProp: Fair Node Representation Learning via Differentiable Propagation Layers
2610.06484
|
cs.LG
|
Emmanouil Kariotakis, Aritra Konar |
Graph neural networks (GNNs) are the standard tool for node representation learning and are increasingly used in high-stakes settings. Their message-passing backbone, however, can amplify topological bias, raising fairness concerns. We study group fairness at ...Graph neural networks (GNNs) are the standard tool for node representation learning and are increasingly used in high-stakes settings. Their message-passing backbone, however, can amplify topological bias, raising fairness concerns. We study group fairness at the level of downstream predictions for node classification, link prediction, and node regression, and bound the demographic parity gap for an arbitrary number of sensitive groups. Our node classification bound is provably no looser than the closest prior result. For link prediction, ours is the first bound on the parity gap of the deployed sigmoid-activated prediction rather than a pre-activation proxy, and for node regression we provide the first such bound. Across all three tasks, the analysis identifies two distinct sources of bias: the separation between group means and the within-group covariance of the final representations. Building on this insight, we embed fairness into propagation itself by augmenting the convex smoothing problem underlying APPNP with a convex group-mean constraint and a within-group covariance regularizer. Unfolding projected gradient descent on this problem yields FairProp, whose layers pair a propagation step with a closed-form projection and which provably converges linearly to the unique fair optimum. Experiments on three tasks show that FairProp, even with exact group-mean equalization alone, provides a strong inductive bias that achieves excellent fairness-utility trade-offs against strong baselines.
|
| 1142 |
Xaurora: Generative Weather Forecasting with Denoising Stochastic Interpolants from a Foundation Model Prior
2610.06509
|
cs.LG
|
Eliot Walt, Miltiadis Kofinas, Nikolaj M\"ucke, Efstratios Gavves, Dim Coumou |
Deep learning has revolutionised weather forecasting in recent years, especially through atmospheric foundation models, which offer competitive skill for a fraction of the computational costs of classic physics-based models. However, most existing foundation m...Deep learning has revolutionised weather forecasting in recent years, especially through atmospheric foundation models, which offer competitive skill for a fraction of the computational costs of classic physics-based models. However, most existing foundation models are deterministic, limiting the generation of large ensembles for accurate uncertainty quantification, extreme weather risk assessment, and long-range weather forecasting. Furthermore, these models incur a large, often prohibitive, computational overhead to train from scratch. To address these shortcomings, we turn a pretrained deterministic prior model, namely the Aurora foundation model, into a generative ensemble-prediction model. To that end, we introduce a novel generative method, Denoising Stochastic Interpolants, combined with a replay buffer for Stochastic Differential Equation (SDE) rollout, enabling probabilistic training of SDE trajectories. Our stochastic foundation model, Xaurora, is finetuned from the small Aurora version, yet it approaches the state-of-the-art on global ensemble metrics and is competitive with the large version of Aurora. Our method is parameter and sample efficient, and generates skilful 15-day forecasts in 13 minutes. Our results demonstrate that deterministic foundation models can be efficiently extended into even stronger stochastic models.
|
| 1143 |
Conditional Flow Matching for Single-Neuron Electrophysiology: Capturing Multimodal Responses Across Stimuli
2610.06520
|
cs.LG
|
Cameron Schofield, Luca Ghafourpour, Philip H. Wong, Costas A. Anastassiou, Richard E. Turner |
Neurons of the brain exhibit a rich repertoire of electrophysiology dynamics with the same repeated stimulus eliciting very different voltage responses from the same cell. One common approach in biophysically detailed models is to capture this variability thro...Neurons of the brain exhibit a rich repertoire of electrophysiology dynamics with the same repeated stimulus eliciting very different voltage responses from the same cell. One common approach in biophysically detailed models is to capture this variability through ensembles of deterministic parametrizations, at a cost of hundreds of thousands of CPU hours. Existing machine learning surrogates inherit the same limitation, where a stimulus is mapped to a single voltage response. We address this challenge by learning a conditional generative model for single-neuron electrophysiology, using flow matching with a velocity field conditioned on the input current. On biophysically detailed models of two human cortical interneuron types, the generated responses closely reproduce the electrophysiological feature distributions, spike-time structure, and excitability profiles, even matching the experimental recordings from the corresponding human cortical neurons. Near the firing threshold, firing and non-firing responses coexist at the same stimulus amplitude, and at high amplitudes, ensembles may split into low- and high-firing modes near depolarization block. We show that our model recovers both modes in each case, while a neural operator baseline suppresses spiking near threshold and blurs the gap between modes at depolarization block.
|
| 1144 |
WaveGSSM: Graph Wave State Space Models for Propagating Spatio-Temporal Patterns
2610.06540
|
cs.LG
|
Junyou Zhu, Fenying Cai, Ping Xiong, Christian Nauck, Langzhou He |
Spatio-temporal graph models typically encode each snapshot with a GNN and then connect the resulting representations through a temporal module. This space-then-time design is effective, yet it does not explicitly represent how a pattern moves across the graph...Spatio-temporal graph models typically encode each snapshot with a GNN and then connect the resulting representations through a temporal module. This space-then-time design is effective, yet it does not explicitly represent how a pattern moves across the graph. We show empirically that, for a propagating process, the same present field can lead to different futures when its recent rate of change differs, motivating an explicit representation of motion in the predictive state. We introduce WaveGSSM, a second-order graph state-space model that maintains two coupled latent states at each node, one for the current pattern and one for its temporal rate of change. A graph-wave transition updates the motion state through graph interactions and uses it to advance the pattern state, coupling spatial propagation and temporal evolution within a single rollout. We evaluate WaveGSSM on four temporal-graph benchmarks and global weather forecasting. It consistently achieves the best mean performance across the temporal-graph benchmarks and reduces the geopotential RMSE by 20.2% on average for 1- to 5-day weather forecasts relative to a backbone-matched snapshot model, while better preserving large-scale atmospheric patterns.
|
| 1145 |
A Fine-Grained Analysis of the LoRA Fine-Tuning Landscape with Implications for Data Selection
2610.06542
|
cs.LG
|
Bowen Zhang, Changrui Fang, Xinsong Ma, Jiaye Teng, Ziye Ma |
Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization la...Low-Rank Adaptation (LoRA) has become a standard approach for parameter-efficient fine-tuning, yet a fundamental practical question remains unresolved: how should the adapter rank be chosen? An overly small rank may lead to a poorly conditioned optimization landscape, whereas an unnecessarily large rank sacrifices the efficiency that motivates LoRA in the first place. Existing theoretical analyses provide only limited guidance on this trade-off, and their guarantees are typically established under restrictive theoretical settings. We address this gap by developing a substantially sharper landscape theory for LoRA, building on modern results from nonconvex low-rank matrix sensing. Our central insight is that the appropriate adapter rank should depend on the quality of the data-induced optimization geometry, rather than on the model alone. To formalize this connection, we introduce LoRA-RIP, a data-dependent restricted-isometry metric that characterizes the conditioning of the cross-entropy (CE) objective along LoRA-relevant low-rank directions. We prove that sufficient rank over-parameterization, with the required rank explicitly determined by the LoRA-RIP constant, eliminates spurious local minima, thereby extending existing RIP-based guarantees beyond the classical 1/3 regime. This characterization further enables principled data selection under a fixed rank budget. Experiments across language and vision tasks support these theoretical predictions, showing that rank and data quality are two coupled resources that should be jointly considered for more efficient and reliable LoRA fine-tuning.
|
| 1146 |
Empirical Variational Autoencoder
2610.06545
|
cs.LG
|
Kaede Shiohara |
We present Empirical Variational Autoencoder, a general generative framework for continuous-valued (i.e., non-vector-quantized) sequences. EVA is based on the evidence lower bound of the Variational Autoencoder (VAE) but learns autoregressive latent priors emp...We present Empirical Variational Autoencoder, a general generative framework for continuous-valued (i.e., non-vector-quantized) sequences. EVA is based on the evidence lower bound of the Variational Autoencoder (VAE) but learns autoregressive latent priors empirically from training data, which can be implemented only by an additional single linear layer on top of VAEs. By replacing the conventional standard-Gaussian constraint with the self-predicted priors, EVA significantly alleviates the latent distribution gap between prior and posterior which is typically observed in conventional VAEs, and leads to high-fidelity ancestral sampling for sequential data generation. Extensive experiments on image and sound synthesis demonstrate that EVA achieves competitive generation quality with autoregressive diffusion baselines despite its much faster inference time.
|
| 1147 |
LinearPFN: Amortized Variable Selection for Linear Models with Interactions
2610.06580
|
cs.LG
|
Louis Schiekiera, Max Zimmer, Christophe Roux, Manuel Arnold, Sebastian Pokutta |
Spike-and-slab regression is a standard Bayesian formulation of variable selection: it returns a posterior distribution over which candidate effects are active rather than a single selected subset, so that every candidate effect carries an inclusion probabilit...Spike-and-slab regression is a standard Bayesian formulation of variable selection: it returns a posterior distribution over which candidate effects are active rather than a single selected subset, so that every candidate effect carries an inclusion probability. Its cost grows exponentially with the number of candidate effects, so the posterior can be enumerated exactly only when the number of predictors is small. Beyond that reach, the posterior has to be approximated, typically by Markov chain Monte Carlo over the model space, which requires a fresh run for every dataset and, within a fixed budget of steps, may fail to converge. We present LinearPFN, a prior-data fitted transformer network that amortizes spike-and-slab inference for linear models with main effects and pairwise interactions. The network is pretrained once on synthetic datasets, drawn from an explicitly specified prior, and a single forward pass over a new dataset returns posterior inclusion probabilities, posterior-mean coefficients and posterior predictive distributions with no per-dataset fitting. The prior is conjugate by design, so that the posterior for each fixed set of active effects has a closed form, and wherever the exact posterior is still computable by enumeration we verify the network's outputs against it. On real predictor matrices from published social-science datasets, with outcomes drawn from the prior so that the true active set is known, LinearPFN attains a higher per-dataset selection AUC and a higher F1 under the median probability model rule than five classical baselines. The lead holds when the coefficients, the interactions or the noise depart from the prior. Code: https://github.com/schiekiera/LinearPFN. Trained model: https://huggingface.co/schiekiera/LinearPFN.
|
| 1148 |
Separators Make Carry Propagation Learnable:The Geometry of Latent Carry in a Multiplication Transformer
2610.06605
|
cs.LG
|
Sama Satariyan (SCAI), Raphael Cousin (SCAI), G{\'e}rard Biau (LPSM, IUF, MEGAVOLT) |
Transformers asked to multiply multi-digit numbers in a single forward pass often fail, and interpretability studies of pretrained language models find arithmetic solved by input-range heuristics rather than by an explicit carry. We train small Llama-style tra...Transformers asked to multiply multi-digit numbers in a single forward pass often fail, and interpretability studies of pretrained language models find arithmetic solved by input-range heuristics rather than by an explicit carry. We train small Llama-style transformers from scratch on 4x4 multiplication without chain of thought and find that the input format is decisive: inserting a space token between digits raises exact-match accuracy from 1% to 89%. Output positions are learned in carry-chain order, with the middle digits, which have the longest-range dependencies, learned last. Inside the model, the separator token that predicts each digit (its prediction slot) encodes the carry-in as an angle on a ring in the residual stream; examples with more distinct carry values fill more of the ring. Activation patching between examples matched on the column sum shows that this state is causally used before the last layer: patching the prediction slot alone transfers the source carry in up to 84% of cases after block 4 for one middle column of our best model, while for other columns the carry is first assembled at the neighboring answer slot before reaching its own. Remaining errors are almost always off by one, consistent with a small error on the carry or on the circular digit code.
|
| 1149 |
Differentially Private Mixing of Public Datasets Improves Private Learning
2610.06636
|
cs.LGcs.AI
|
Yufei Chen, Tejumade Afonja, Anvith Thudi, Nicolas Papernot |
Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on "public" data before finetuning with DP ...Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on "public" data before finetuning with DP on the sensitive data can reduce the drop in utility. However, the success of this depends on how relevant the selected public dataset is to the sensitive data. We introduce the first pipeline that privately learns the mixture of several public datasets to pretrain on for a given sensitive downstream task. Our key insight is that we can privately find the best mixture of multiple public datasets by privately learning a low-dimensional linear model. We tested our method on the NIH dataset for X-ray classification and the ENRON email dataset for language modeling. Applying our method to find tailored mixtures of X-ray datasets to pretrain on for diseases in the NIH ChestX-ray14 dataset, we improved macro AUC by up to 0.037 across privacy budgets compared to the baselines, with gains as large as +22.8% relative AUC on Cardiomegaly at $\epsilon=1$. For DP training on the ENRON dataset, pre-training on our mixture of The Common Pile (a collection of public-domain text datasets) decreased test perplexity by 16% relative to the baseline mixtures.
|
| 1150 |
Beyond the Model: The Critical Role of Data Filtering in Clinical Machine Learning
2610.06640
|
cs.LG
|
Noah Subedar, Colin Campbell, Wenjing Zhang, Dan Perri, Sarah Culgin |
Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increas...Machine learning (ML) studies using clinical data often rely on preprocessing and filtering pipelines before model development. The filtering decisions made in these pipelines can alter the dataset's statistical structure and may artificially reduce or increase the complexity of the prediction task. We argue that filtering choices should be treated as part of the scientific method rather than as a routine preprocessing step. We further discuss the need for explainable and transparent preprocessing pipelines that allow researchers to understand why specific filtering choices are made and how these choices affect the resulting data distribution and model performance. All of the source code for this work is available on GitHub.
|
| 1151 |
Considering Context: When World Models Need Context Encoders
2610.06651
|
cs.LG
|
Oleg Smirnov, Sofiane Ennadir, John Pertoft, Bjartur Hjaltason, Sara Karimi |
Methods for generalization in model-based reinforcement learning typically assume that an agent cannot recover the latent context governing the environment dynamics from its own experience, and therefore supplies it externally. We formalize and test this assum...Methods for generalization in model-based reinforcement learning typically assume that an agent cannot recover the latent context governing the environment dynamics from its own experience, and therefore supplies it externally. We formalize and test this assumption with \emph{predictive sufficiency}, which quantifies what access to the context adds to next-step prediction under the visitation distribution an agent induces, and separates that quantity into a history-recoverable part, a residual requiring the true context, and the deficit added by a finite model. We classify context-aware algorithms by the predictive risk their conditioning set can target and demonstrate across environments of increasing identification difficulty that the headroom does not follow the MDP class. The same task under different priors leaves predictive headroom in one setting and nothing distinguishable from zero in another, where the agent's behavior implicitly identifies the context and any benefit of such a mechanism cannot be attributed to missing information. Where headroom persists, the learned state exposes it only partially, and adding the true context still lowers the risk. Our contribution is a practical criterion for matching contextual mechanisms to the information available to them, estimated from the ordinary trained agent without a reference policy.
|
| 1152 |
The Birkhoff Geometry of Manifold-Constrained Hyper-Connections: Two Channels, Vertex Viscosity, and Sinkhorn as a Retraction
2610.06653
|
cs.LG
|
Xiaoyu Li, Zhizhou Sha, Chiwun Yang |
Hyper-connections widen the residual stream of a Transformer to $n$ parallel streams. Their manifold-constrained version (mHC) mixes the streams at each layer with a doubly stochastic matrix, which it computes by Sinkhorn normalization of exponentiated logits....Hyper-connections widen the residual stream of a Transformer to $n$ parallel streams. Their manifold-constrained version (mHC) mixes the streams at each layer with a doubly stochastic matrix, which it computes by Sinkhorn normalization of exponentiated logits. We give a geometric theory of this design on the Birkhoff polytope. First, a doubly stochastic mixer splits the stream into a mean channel, on which mHC is exactly a residual network, and a difference channel, which each layer contracts by its second singular value $\sigma_2 \le 1 - n \min_{ij} H_{ij}$. Thus the extra width is a fading memory with a horizon of $1/(1-\sigma_2)$ layers, and among nonnegative mixers only the permutations do not collapse. Second, the Sinkhorn-logit map is a global chart, and its logit gradient is exactly the Fisher-Rao gradient. Thus logit gradient flow follows a squared Fisher-Rao metric, and the straight-through update is exactly entropic mirror descent. Third, under logit gradient flow the logarithm of each entry moves at a rate of at most $4n^3\|\nabla f\|_\infty \varepsilon$, where $\varepsilon$ is the distance to the nearest permutation. Thus gradient flow approaches and leaves the vertices only at rate $1/t$, but mirror descent moves at an exponential rate. Fourth, the local convergence factor of Sinkhorn is $\sigma_2^2$, so a fixed iteration budget limits the horizon. Experiments confirm the predicted rates.
|
| 1153 |
TrustmeWatcher: An Application for Workplace Micro-Sensing and Explainable Well-Being Feedback
2610.06657
|
cs.LG
|
Chengyu Yu, Leon Jacopo Costa, Zoja An\v{z}ur, Mohan Li, Ga\v{s}per Slapni\v{c}ar |
Workplace sensing studies combine long-running behaviour traces with self-reports, yet the tools that collect those data often sit apart from the interface that returns results. We present TrustmeWatcher, the application built for the TRUST-ME project to conne...Workplace sensing studies combine long-running behaviour traces with self-reports, yet the tools that collect those data often sit apart from the interface that returns results. We present TrustmeWatcher, the application built for the TRUST-ME project to connect this work. TrustmeWatcher reuses ActivityWatch's OS-level watchers for computer-activity collection and adds its own application layer. It turns the collected traces into an interactive screen-time dashboard, synchronizes responses from short questionnaires completed on the StreamDeck, and presents questionnaires alongside video highlights. Activity records and self-reports are aligned into labelled records for model development. The scope of this paper is limited to ActivityWatch data as model input. Artificial intelligence (AI) uses these activity records to predict six normalized state scores and an overall well-being score. The trained model runs locally, and the dashboard presents its predictions in semantic bands. Explainable artificial intelligence (XAI) helps users understand how recorded activity contributed to a prediction. Privacy Control lets users pause or resume the camera and eye tracker used by the study. We describe the workflow, its user-device and sensing-setup boundaries, and its use with records from 17 participants. The result is a deployed application and study workflow that integrates activity review, study data collection, privacy control, local prediction, and a participant-facing interface for XAI evaluation.
|
| 1154 |
Learning What to Imitate: Entropy-Aware Distribution Mixing
2610.06671
|
cs.LG
|
Juan Garcia Giraldo, Matteo Santelmo, Eduard Durech, Imanol Schlag, Valentina Pyatkin |
Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{...Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imitate a continuation that it deems unlikely (i.e., low-probability) despite being confident in a different continuation (i.e., in a low-entropy state). To mitigate the degradation in generalisation and catastrophic forgetting caused by these conflicts, we propose \textbf{Entropy-Aware Mixing}: a dynamic per-token interpolation of the student and teacher distributions, gated by the student's predictive entropy. We implement both convex and geometric interpolations for both offline trace generation (via speculative decoding, then SFT) and on-policy forward-KL distillation. Our results show that entropy-aware mixing stabilises distillation, improving in-distribution and out-of-distribution math reasoning while better preserving general capabilities than fixed-teacher supervision. Nonetheless, the optimal entropy schedule depends on the training source, with offline-generated traces favouring concave schedules (greater overall teacher influence) and on-policy training favouring linear or convex schedules (teacher concentrated in high-entropy states).
|
| 1155 |
Improved Convergence of Large Stepsize Gradient Descent for Logistic Regression
2610.06675
|
cs.LG
|
Xiaochuan Gong, Ang Li |
We study gradient descent (GD) with a large constant stepsize for logistic regression on linearly separable data. Existing analysis shows an accelerated rate of $\widetilde{O}(1/\sqrt{\epsilon})$ to reach loss $\epsilon$ with an aggressive stepsize, although t...We study gradient descent (GD) with a large constant stepsize for logistic regression on linearly separable data. Existing analysis shows an accelerated rate of $\widetilde{O}(1/\sqrt{\epsilon})$ to reach loss $\epsilon$ with an aggressive stepsize, although the loss may initially oscillate. Tighter control of the oscillatory dynamics has been available only for two-dimensional data. We prove a substantially faster rate in arbitrary dimension: GD with a large stepsize $\eta=1/\epsilon$ reaches loss $\epsilon$ within $O(\ln^{p}(1/\epsilon))$ steps, where $p$ depends only on the margin and the rank of the data. Our proof improves the bound on the transition time of GD from the oscillatory to the stable phase, after which the loss decreases monotonically. We split the oscillatory phase into recursively nested intervals. The margin and the rank bound the nesting depth, and a counting argument bounds the number of intervals at each depth, together yielding the polylogarithmic step complexity.
|
| 1156 |
Closing the Context Gap: Activation Alignment for Tabular In-Context Learning
2610.06679
|
cs.LGcs.AI
|
Yoel Zeldes |
Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every ...Tabular foundation models perform in-context learning (ICL) by conditioning predictions on labeled training examples provided as context. Unlike traditional models that separate training from inference, these models must process all training examples in every forward pass, making each prediction expensive. Restricting the number of training examples reduces this cost but substantially degrades performance. Instead of discarding context, we propose activation alignment, a method that leverages the full context to teach a model how to behave when seeing only a subset. This is achieved by training a lightweight linear transformation on synthetic unlabeled data to map the intermediate activations of a data-constrained "student" (using partial context) toward those of a full-context "teacher" (using all data). Training the aligner requires no GPU and converges in seconds to minutes on commodity hardware. We evaluate on 38 classification datasets from the TabArena benchmark using the leading two tabular foundation models, TabPFN-3 and TabFM. Across all context budgets, the aligned student yields broad, statistically significant improvements over the unaligned baseline for both models. In low-data regimes, alignment recovers nearly half of the teacher's predictive advantage. The method provides a practical, low-overhead approach to achieving the inference speed of compact contexts while closing a significant fraction of the performance gap to the full-context teacher.
|
| 1157 |
OVAL: Output-Aware Local Page Bases for KV Cache Retrieval
2610.06686
|
cs.LG
|
Ashkan Shahbazi, Chayne Thrash, Soheil Kolouri |
Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query...Long context inference with large language models becomes increasingly expensive as attention must operate over an ever growing KV cache. Page sparse attention reduces this cost by representing each KV page compactly and retrieving only a subset for each query. Existing retrieval methods are designed to estimate attention scores or page relevance, but their objectives do not directly account for how approximation errors affect the resulting value weighted attention output. We introduce \method{}, an output aware page encoding derived from the joint structure of keys and values while preserving the key information needed for accurate retrieval. \method{} is training free and requires no additional value dependent statistics at inference time. Once constructed, its stored representation has the same size and decode time scoring cost as a key only spectral representation. Across long reasoning, long context understanding, and long generation benchmarks, \method{} consistently improves over the key only spectral baseline and performs competitively with recent KV cache compression and retrieval methods. On long reasoning benchmarks, it achieves strong avg@\(k\) performance across model benchmark pairs, while matching or surpassing leading baselines on several long context understanding and generation settings with modest decoding overhead. Code is available at \url{https://github.com/Ashkan13776/oval-kv}.
|
| 1158 |
Adapting prior-data fitted networks for tabular anomaly detection
2610.06693
|
cs.LG
|
Maximilian Bershtman, Niv Cohen |
While deep features have transformed anomaly detection in images and video, their impact on tabular data has been less substantial, partly due to the limited availability of strong deep representations. Recently, prior-data fitted networks (PFNs) have emerged ...While deep features have transformed anomaly detection in images and video, their impact on tabular data has been less substantial, partly due to the limited availability of strong deep representations. Recently, prior-data fitted networks (PFNs) have emerged as a promising source of such representations for tabular data. In this work, we investigate how PFN representations can be adapted and leveraged for anomaly detection. The question is harder than it looks. No anomalies are available before deploy- ment, so model parameters cannot be tuned with supervision, and the reference set that defines normal behavior may itself contain the very anomalies it is supposed to reveal. We begin our study using frozen TabPFN features. Scoring each sam- ple by its distance to its nearest neighbors in feature space already gives strong results. We identify which layers to use and a feature-extraction procedure suited to the task. Next, to further improve performance, we use the reference set to fine- tune the model, so that the resulting features better separate normal samples from anomalies. On the ADBench benchmark, our fine-tuning free approach (ZEN) reaches a higher mean AUROC than every baseline, and our fine-tuned method (FOCUS) improves on it further. Our approach also generalizes across PFN models.
|
| 1159 |
To Learn is to Wander: Learning Across Graphs and Tasks with Random Walks
2610.06694
|
cs.LG
|
Louis Tichelman (TU Wien, AITHYRA), Xingyue Huang (University of Oxford), Jinwoo Kim (KAIST), \.Ismail \.Ilkan Ceylan (TU Wien |
Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed...Graph foundation models aim to transfer across graphs, feature spaces, relational schemas, and prediction tasks, yet existing approaches typically generalize only within particular graph modalities or tasks. We propose Wander, a graph foundation model designed to operate across these settings within a single pretrained checkpoint. Following the prior-predictive perspective, we formulate graph learning as completion of a partially observed graph. We realize this task-general view through a common interface based on random walks, allowing the same model to operate across homogeneous and multi-relational graphs with varying features, labels, and relational schemas. Wander can increase its structural context at inference time without changing its learned parameters and, under suitable assumptions, universally approximates the corresponding Bayes-optimal predictor on bounded connected graphs. Empirically, a single pretrained checkpoint achieves state-of-the-art or highly competitive results across node classification, homogeneous link prediction, and knowledge-graph link prediction. Moreover, joint pretraining across graph modalities and tasks preserves performance in specialized settings while enabling positive transfer and the composition of separately learned capabilities at inference time.
|
| 1160 |
BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
2610.06725
|
cs.LGcs.AI
|
Gang Fu, Adel Javanmard, MohammadHossein Bateni, Vahab Mirrokni |
Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry ...Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal node the branching probability is centered on the arrival-weighted mean score of the traffic reaching that node. This mean is estimated using an exponential moving average, which promotes utilization of both child subtrees without an auxiliary load-balancing loss. We show that this moving-average estimate admits an explicit noise-lag trade-off. We prove that for linear node maps and log-concave arrival distributions, this mechanism prevents routing-mass collapse. We further establish that, under a frozen router, an expert's execution frequency controls its stochastic-gradient convergence rate, and that confident decisions near the root bound cross-device communication when experts are assigned to devices by tree prefix. We evaluate BRANCH-MoE against Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing on Criteo click-through-rate prediction, Forest Covertype, HIGGS, and YearPredictionMSD, using \(E=16\), top-\(4\) routing, and five random seeds. Our results show that hierarchical routing can preserve task quality and balanced utilization while inducing a topology that supports localized expert co-activation and reduced communication.
|
| 1161 |
Decoupling Time and Space: A Temporally Conditioned Refinement for EEG Source Imaging
2610.06726
|
cs.LG
|
Marco Morik, Jesse Palarus, Carmen Vidaurre, Klaus-Robert M\"uller, Shinichi Nakajima |
Electroencephalography (EEG) offers millisecond temporal resolution, but inferring underlying neural sources is a severely ill-posed spatial inverse problem. While deep learning has advanced spatial reconstruction, current architectures face a critical dilemma...Electroencephalography (EEG) offers millisecond temporal resolution, but inferring underlying neural sources is a severely ill-posed spatial inverse problem. While deep learning has advanced spatial reconstruction, current architectures face a critical dilemma: frame-by-frame models discard vital temporal context, whereas full 4D spatiotemporal networks introduce an architectural trade-off between reconstruction accuracy and inference cost. We propose a novel two-stream framework that explicitly decouples global temporal representation learning from per-time-point spatial refinement. A Transformer-based Temporal Condition Encoder processes the entire EEG sequence via factorized spatiotemporal attention, retaining sensor-resolved features. A fixed inverse then maps these features into source-indexed conditioning for a per-timestep Source-Space Transformer or volumetric convolutional refiner. Extensive evaluations on realistic synthetic data demonstrate that this temporal prior dramatically improves spatial localization, outperforming classical and spatiotemporal baselines, particularly in high-noise and multi-source regimes. Training across diverse leadfields and explicit operator mismatches improves transfer to unseen head geometries and brings template-based reconstruction closer to subject-specific inversion. Furthermore, we apply the model trained only on synthetic EEG data to real-world EEG. A logistic regressor fit on source power differences in eyes-open, eyes-closed conditions successfully decodes age groups.
|
| 1162 |
Hyperbolic Graph Representation Learning: Embed in One Metric, Optimize with Another
2610.06745
|
cs.LG
|
Federico Larroca, Paola Bermolen, Marcelo Fiori, Bernardo Marenco |
Hierarchical graphs embed in hyperbolic space with lower distortion than in Euclidean space owing to its negative curvature. However, their gradient-based learning is hampered at large radii, where the Poincar\'e ball and the Lorentz hyperboloid models fail nu...Hierarchical graphs embed in hyperbolic space with lower distortion than in Euclidean space owing to its negative curvature. However, their gradient-based learning is hampered at large radii, where the Poincar\'e ball and the Lorentz hyperboloid models fail numerically. Polar coordinates avoid this problem, but the hyperbolic metric scales the angular step by the hyperbolic sine of the radius, freezing angular motion. We observe that this factor is a choice, silently fixed by existing implementations: the Euclidean tangent parametrization, for instance, uses the radius itself. We show that other choices are not only possible but preferable. They are endpoints of a one-parameter family of optimization preconditioners with curvatures from $-1$ to $0$, while the embedding remains at curvature $-1$. We show that since the Euclidean preconditioner rearranges a layout but refines it poorly, while an intermediate one refines far better once a layout is in place, combining them in two stages reduces the loss on real-world trees by 46-74% over the best single curvature.
|
| 1163 |
MatrixFormer: A Foundation Model for Matrix Completion
2610.06751
|
cs.LGcs.AI
|
Dwaipayan Saha, Jacob Feitelberg, Kyuseong Choi, Raaz Dwivedi, Anish Agarwal |
Matrix completion underlies problems from tabular imputation to causal inference, yet existing tabular foundation models treat it as entry-by-entry prediction, repeating context for every target and discarding the matrix's two-dimensional structure. We introdu...Matrix completion underlies problems from tabular imputation to causal inference, yet existing tabular foundation models treat it as entry-by-entry prediction, repeating context for every target and discarding the matrix's two-dimensional structure. We introduce MatrixFormer, a pre-trained matrix-native transformer that predicts a full distribution for every missing entry in a single forward pass. MatrixFormer is trained entirely on synthetic low-rank and latent-factor matrices under diverse missingness patterns. Applied zero-shot and with the same model weights, MatrixFormer achieves competitive performance on causal inference panel-data tasks, language-model benchmark-score completion, tabular imputation, and recommendation systems matrix completion. These results position MatrixFormer as a general-purpose foundation model for matrix completion.
|
| 1164 |
Round-Trip KNN Clustering: multiscale hierarchical cluster detection on directed nearest-neighbour graphs
2610.06795
|
cs.LG
|
Eraldo Pereira Marinho, Caetano Mazzoni Ranieri, Fabricio Aparecido Breve |
We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undir...We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relation: which points a given point selects and which points select it. Incoming selections are treated as weighted votes that help decide which local connections remain visible during a recursive forward-and-reverse traversal. Repeating the procedure for increasing $K$ reveals how groups persist or merge as the neighbourhood scale grows; for the reference inverse-square model before structural refinement, clusters can merge but do not split. Because graph connectivity can occasionally join distinct groups through a sparse bridge or a small region of overlap, we add an optional label-free refinement. It first tests whether an already formed component is better described by two or three Gaussian subpopulations, and accepts a subdivision only when the proposed groups are large enough and consistent with the visible KNN graph. Across eight synthetic datasets and $K=2,\ldots,16$, independent C and Python implementations produced identical partitions in all 120 reference runs. Refinement increased adjusted Rand index from $0.7817$ to $0.9627$ on a variable-density benchmark and from $0.8083$ to $0.9853$ on a sparse-bridge benchmark. Comparisons with seven external clustering methods show competitive performance while preserving a label-free cluster-construction process.
|
| 1165 |
Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
2610.06804
|
cs.LGcs.AI
|
Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa, Mehdi Kamal |
A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probabili...A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters, but needs many scored candidates per query. We show that a model can instead be trained to produce such answers in one generation. On-policy power distillation (OPPD) runs a sequential Monte Carlo sampler in which the model being trained generates candidates and a frozen teacher's power distribution weights them; the same probabilities weight each answer in a maximum-likelihood update. Training raises single-generation accuracy by up to 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature, and one generation scores 2.4 and 3.5 points above published power sampling with 64 candidates, recovering 94 percent of the gain that 16 candidates give the untrained model. For context, against GRPO trained with verified rewards from the same checkpoint and budget, OPPD scores 3.8, 4.0 and 5.4 points higher on MATH500, GSM8K and AIME using no reference answers; the two are complementary, and OPPD applied after GRPO adds up to 9.3 points. Trained only on mathematics, OPPD raises HumanEval accuracy by up to 5.3 points. One loss coefficient moves the sharpening exponent the model absorbs between 1.19 and 2.02, against 1.14 for ordinary on-policy distillation, and it rises mostly on the model's own answers. Gains hold across model families and sizes, including a model already trained with verified rewards, where lowering the temperature gives nothing and OPPD adds 4.4 points on MATH500. Code: https://github.com/ArminAzizi98/OPPD.
|
| 1166 |
H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning
2610.06805
|
cs.LG
|
Wancong Zhang, Basile Terver, Michael Rabbat, Yann LeCun, Randall Balestriero |
Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-...Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down: the top level optimizes progress toward the goal, and each level's predictions become subgoals for the planner below it. When factors in the data evolve at separated timescales, higher levels discard fast, unpredictable detail and retain slower task-relevant state. Across four simulated navigation and manipulation environments, hierarchical planning improves over a flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% using less planner compute. Ablations attribute these gains to both temporal decomposition and higher-level goal representations. With inverse-dynamics supervision, the approach extends to diverse real-robot videos from DROID, where hierarchy improves offline planning fidelity at lower planner compute.
|
| 1167 |
Block Disentanglement in CRL: Bridging Identifiability and Visual State Estimation
2610.06809
|
cs.LG
|
Emre Acart\"urk, Pranamya Kulkarni, Puranjay Datta, Karthikeyan Shanmugam, Burak Var{\i}c{\i} |
Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impra...Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impractical to obtain. While there has been significant progress in understanding the identifiability guarantees of CRL, such guarantees often hold under highly stylized assumptions, which temper the direct application to real-world problems. This paper has a two-fold objective for interventional CRL. First, it establishes identifiability guarantees for substantially weaker interventional assumptions, resulting in block disentanglement of the causal variables, where the block structure depends on the realistically available intervention mechanisms. Secondly, the block disentanglement framework is used for embodied visual state estimation, in which the objective is to recover the latent physical variables of a robotic system directly from visual data (images and videos) without labeled data. These two components are critically complementary. The block disentanglement theory delineates identifiability guarantees under weakened assumptions, and the application demonstrates that the resulting objective remains effective in a controlled embodied setting despite further assumption violations, providing a theory-to-practice bridge needed to translate the promise of label-free CRL into practical problems.
|
| 1168 |
Private online learning and prediction for Littlestone classes
2610.06822
|
cs.LG
|
Amartya Sanyal |
We study mistake bounds for differentially private online learning and online prediction under oblivious realisable adversaries. Online learning requires the learner to release a hypothesis at each time step whereas in online prediction, the learner only needs...We study mistake bounds for differentially private online learning and online prediction under oblivious realisable adversaries. Online learning requires the learner to release a hypothesis at each time step whereas in online prediction, the learner only needs to make predictions without releasing a hypothesis. Using a novel lower bound for private online learning and an upper bound for private prediction, we show that the sample complexity of these two problems are separated by a factor that grows with the time horizon for every class of finite Littlestone dimension $d$. First, we prove that every $\br{\epsilon,\delta}$-private online learner has a deterministic realisable stream of length $T$ on which the mistake bound is at least $\bE\bs{M_T}=\Om{\frac d\epsilon \log\br{ T}^{2/3}}$. In particular, this is the first non-trivial lower in the range $1/T<\delta<1/\log T)$ left open in earlier works[SR22,DSS24,LWY24]. Second, we prove that for every class of of Littlestone dimension $d$, there exists an $(\epsilon,\delta)$-jointly private predictor with at most $2^{2^{cd^2}}\epsilon^{-2}\log^2\br{2/\br{\epsilon\delta}}$ expected mistakes, independently of $T$, for some absolute constant $c>0$. Thus, for every fixed class of finite Littlestone dimension when $\delta=\Theta\br{1/\log T}$, private learning requires $\Om{\br{\log T}^{2/3}}$ expected mistakes, whereas private prediction admits $\bigO{\br{\log\log T}^2}$.
|
| 1169 |
Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors
2610.06823
|
cs.LGcs.AI
|
Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi, Nazmus Sakib |
Large longitudinal cohorts often contain wrist accelerometry without optical heart-rate sensing, motivating recovery of cardiac information from motion signals already collected during sleep. We present SeqSmoother, a transformer-based temporal corrector for s...Large longitudinal cohorts often contain wrist accelerometry without optical heart-rate sensing, motivating recovery of cardiac information from motion signals already collected during sleep. We present SeqSmoother, a transformer-based temporal corrector for sleep heart rate (HR) estimation from wrist accelerometry. SeqSmoother combines spectral descriptors with an intermediate Nightbeat-derived frequency anchor and a physics-motivated sub-harmonic feature designed to identify harmonic frequency lock-on. All inference-time features are derived from wrist accelerometry, while ECG is used only to construct reference HR labels and training-label quality weights. We evaluate SeqSmoother using 13 participant-disjoint held-out folds and compare it with the official Nightbeat implementation under a matched 60-s window and 15-s step protocol. Across all out-of-fold predictions, SeqSmoother achieved a participant-macro MAE of 1.60 bpm. On Nightbeat-retained matched intervals, Nightbeat achieved lower absolute error than SeqSmoother (0.615 versus 1.091 bpm), while SeqSmoother provided estimates over a larger portion of the eligible recording; Nightbeat produced final estimates for 72.85% of the SeqSmoother-eligible out-of-fold grid. Separately, the proposed sub-harmonic ratio achieved an AUROC of 0.972 for identifying reference-defined harmonic lock-on candidates. These findings reveal an accuracy-availability trade-off between learned temporal modeling and quality-gated signal processing while providing empirical support for a physics-informed approach to identifying frequency-tracking failures in accelerometer-based sleep HR estimation.
|
| 1170 |
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points
2610.06833
|
cs.LG
|
Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu |
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; termi...Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.
|
| 1171 |
Game Plan: What AI can do for Football, and What Football can do for AI
2011.09192
|
cs.LG
|
Karl Tuyls, Shayegan Omidshafiei, Paul Muller, Zhe Wang, Jerome Connor |
The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to footba...The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players' and coordinated teams' behaviors. The research challenges associated with predictive and prescriptive football analytics require new developments and progress at the intersection of statistical learning, game theory, and computer vision. In this paper, we provide an overarching perspective highlighting how the combination of these fields, in particular, forms a unique microcosm for AI research, while offering mutual benefits for professional teams, spectators, and broadcasters in the years to come. We illustrate that this duality makes football analytics a game changer of tremendous value, in terms of not only changing the game of football itself, but also in terms of what this domain can mean for the field of AI. We review the state-of-the-art and exemplify the types of analysis enabled by combining the aforementioned fields, including illustrative examples of counterfactual analysis using predictive models, and the combination of game-theoretic analysis of penalty kicks with statistical learning of player attributes. We conclude by highlighting envisioned downstream impacts, including possibilities for extensions to other sports (real and virtual).
|
| 1172 |
On the Approximation Relationship between Optimizing Ratio of Submodular (RS) and Difference of Submodular (DS) Functions
2101.01631
|
cs.LG
|
Pierre Perrault, Jennifer Healey, Zheng Wen, Michal Valko |
We demonstrate that from an algorithm guaranteeing an approximation factor for the ratio of submodular (RS) optimization problem, we can build another algorithm having a different kind of approximation guarantee -- weaker than the classical one -- for the diff...We demonstrate that from an algorithm guaranteeing an approximation factor for the ratio of submodular (RS) optimization problem, we can build another algorithm having a different kind of approximation guarantee -- weaker than the classical one -- for the difference of submodular (DS) optimization problem, and vice versa. We also illustrate the link between these two problems by analyzing a \textsc{Greedy} algorithm which approximately maximizes objective functions of the form $\Psi(f,g)$, where $f,g$ are two non-negative, monotone, submodular functions and $\Psi$ is a {quasiconvex} 2-variables function, which is non decreasing with respect to the first variable. For the choice $\Psi(f,g)\triangleq f/g$, we recover RS, and for the choice $\Psi(f,g)\triangleq f-g$, we recover DS. To the best of our knowledge, this greedy approach is new for DS optimization. For RS optimization, it reduces to the standard \textsc{GreedRatio} algorithm that has already been analyzed previously. However, our analysis is novel for this case.
|
| 1173 |
Broaden Your Views for Self-Supervised Video Learning
2103.16559
|
cs.LG
|
Adri\`a Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Ross Hemsley |
Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by image techniques, where these two views are similarly extracted by cropping and ...Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by image techniques, where these two views are similarly extracted by cropping and augmenting the resulting crop. However, these methods miss a crucial element in the video domain: time. We introduce BraVe, a self-supervised learning framework for video. In BraVe, one of the views has access to a narrow temporal window of the video while the other view has a broad access to the video content. Our models learn to generalise from the narrow view to the general content of the video. Furthermore, BraVe processes the views with different backbones, enabling the use of alternative augmentations or modalities into the broad view such as optical flow, randomly convolved RGB frames, audio or their combinations. We demonstrate that BraVe achieves state-of-the-art results in self-supervised representation learning on standard video and audio classification benchmarks including UCF101, HMDB51, Kinetics, ESC-50 and AudioSet.
|
| 1174 |
Sharp Deviations Bounds for Dirichlet Weighted Sums with Application to analysis of Bayesian algorithms
2304.03056
|
cs.LG
|
Denis Belomestny, Pierre Menard, Alexey Naumov, Daniil Tiapkin, Michal Valko |
In this work, we derive sharp non-asymptotic deviation bounds for weighted sums of Dirichlet random variables. These bounds are based on a novel integral representation of the density of a weighted Dirichlet sum. This representation allows us to obtain a Gauss...In this work, we derive sharp non-asymptotic deviation bounds for weighted sums of Dirichlet random variables. These bounds are based on a novel integral representation of the density of a weighted Dirichlet sum. This representation allows us to obtain a Gaussian-like approximation for the sum distribution using geometry and complex analysis methods. Our results generalize similar bounds for the Beta distribution obtained in the seminal paper Alfers and Dinges [1984]. Additionally, our results can be considered a sharp non-asymptotic version of the inverse of Sanov's theorem studied by Ganesh and O'Connell [1999] in the Bayesian setting. Based on these results, we derive new deviation bounds for the Dirichlet process posterior means with application to Bayesian bootstrap. Finally, we apply our estimates to the analysis of the Multinomial Thompson Sampling (TS) algorithm in multi-armed bandits and significantly sharpen the existing regret bounds by making them independent of the size of the arms distribution support.
|
| 1175 |
The Llama 3 Herd of Models
2407.21783
|
cs.LG
|
Aaron Grattafiori (Jack), Abhimanyu Dubey (Jack), Abhinav Jauhri (Jack), Abhinav Pandey (Jack), Abhishek Kadian (Jack) |
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our larg...Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.
|
| 1176 |
Evolutionary feature selection for spiking neural network pattern classifiers
2604.26654
|
cs.LG
|
Michal Valko, Nuno C. Marques, Marco Castelani |
This paper presents an application of the biologically realistic JASTAP neural network model to classification tasks. The JASTAP neural network model is presented as an alternative to the basic multi-layer perceptron model. An evolutionary procedure previously...This paper presents an application of the biologically realistic JASTAP neural network model to classification tasks. The JASTAP neural network model is presented as an alternative to the basic multi-layer perceptron model. An evolutionary procedure previously applied to the simultaneous solution of feature selection and neural network training on standard multi-layer perceptrons is extended with JASTAP model. Preliminary results on IRIS standard data set give evidence that this extension allows the use of smaller neural networks that can handle noisier data without any degradation in classification accuracy.
|
| 1177 |
Online Control via Counterfactual Tracking
2607.13029
|
cs.LG
|
Yunzong Xu |
We study online control of a known linear dynamical system with adversarial costs and bounded disturbances, measuring regret against a general class of benchmark policies. We introduce counterfactual tracking, which separates the challenge of learning from the...We study online control of a known linear dynamical system with adversarial costs and bounded disturbances, measuring regret against a general class of benchmark policies. We introduce counterfactual tracking, which separates the challenge of learning from the challenge of controlling the system. An online learner builds a reference trajectory by selecting or averaging the trajectories that the benchmark policies would have generated under the realized costs and disturbances, and a corrective law steers the system toward that reference. Charging each change in the reference its recovery cost (the cost of steering the system onto the new reference) reduces the problem to online learning with switching costs. Conversely, under additional natural assumptions, we show that this reduction is tight: the two problems have the same minimax regret up to a system-dependent factor, uniformly over horizons and policy classes. The reduction gives sharp regret guarantees for policy classes beyond standard finite-memory parameterizations. For a class of $N$ possibly nonlinear or history-dependent policies, it achieves $O(\sqrt{T\log N})$ regret over $T$ rounds, provided their trajectories remain within a bounded distance of one another and recovery costs are bounded. For the full $\ell_1$ ball of disturbance-response controllers, it achieves $O(\sqrt{T\log T})$ regret, which is minimax optimal in $T$, without assuming a common decay rate for disturbance effects. The framework also improves the best known regret bounds for linear state-feedback policies.
|
| 1178 |
Shapley-based Structural Analysis of Neural Calibration for Stochastic Volatility Models
2610.03076
|
cs.LG
|
Sha\"in Afzali, Serena Della Corte, Antonis Papapantoleon |
Neural network-based approaches have emerged as efficient alternatives to traditional optimization-based procedures for the calibration of stochastic volatility models. However, existing work has focused primarily on predictive accuracy, with comparatively lit...Neural network-based approaches have emerged as efficient alternatives to traditional optimization-based procedures for the calibration of stochastic volatility models. However, existing work has focused primarily on predictive accuracy, with comparatively little attention devoted to understanding the structure of the learned inverse calibration mappings. In this work, we analyze neural calibration mappings for the Heston and rough Heston models across multilayer perceptron, highway, and softmax-parametrized highway architectures, using complementary Shapley-based methods from explainable AI. Specifically, we consider SHAP and $\nu$SHAP explanations, which capture distinct, complementary notions of feature relevance, corresponding to sensitivity and sufficiency of feature subsets, respectively. Short maturities and smile wings consistently dominate parameter inference, and the dominant attribution structure remains qualitatively stable across architectures despite differences in predictive accuracy and parameter count. Parameter-specific differences between SHAP and $\nu$SHAP further reveal how distinct regions of the implied volatility surface contribute to parameter recovery and expose substantial redundancy in the calibration input. Building on this redundancy, we show that $\nu$SHAP explanations can guide a significant reduction in input dimensionality for the rough Heston model while matching calibration accuracy relative to the full implied volatility surface. These findings demonstrate that complementary Shapley-based methods provide structural insight into learned inverse calibration mappings beyond predictive error metrics, and offer a practical route to feature selection in neural calibration problems.
|
| 1179 |
Efficient Analog-Initialized Latent Transport for Spatially Coherent Probabilistic Downscaling from Global to Kilometer Scales
2610.03736
|
cs.LGcs.AI
|
Oph\'elia Miralles |
Probabilistic downscaling must represent kilometer-scale structure left unresolved by a deterministic regional prediction. We introduce an analog-initialized latent transport method that uses historical regional residuals as a meteorologically informed empiric...Probabilistic downscaling must represent kilometer-scale structure left unresolved by a deterministic regional prediction. We introduce an analog-initialized latent transport method that uses historical regional residuals as a meteorologically informed empirical source. A frozen graph neural network predicts the deterministic state from the Integrated Forecasting System (IFS). For each input, residuals from 20 training dates with similar IFS conditions initialize the ensemble. A conditional flow-matching model transports these fields in a carefully designed 65-factor, 100,000-node latent space, and a grid decoder reconstructs 81 atmospheric channels including precipitation at 2.5km spacing. On a season- and cycle-balanced 2025 test, ensemble CRPS improves over deterministic mean absolute error for all 81 variables. Matched analogs improve point and spatial errors over Gaussian initialization and reduce spectral error compared with covariance matching. Comparisons with released CorrDiff, a local CorrDiff-mini model and full grid EDM provide context on the 2023 validation sample. Total training of the downscaling model takes approximately 12h on one NVIDIA H200, and inference takes about 12s per 81-field member. The frozen networks also generate fields over France when supplied with local static data and a high-resolution IFS-derived analog bank. The method combines empirical initialization with whole-domain latent transport for computationally practical regional ensembles.
|
| 1180 |
Deep Learning Denoising of Real SWOT Sea Surface Height Observations
2610.03739
|
cs.LG
|
Ga\'etan Meis, Ana\"elle Tr\'eboutte, Maxime Ballarotta, Marie-Isabelle Pujol, G\'erald Dibarboure |
The SWOT (Surface Water Ocean Topography) mission is currently providing unpreceded high-resolution measurements of Sea Surface Height (SSH), revealing ocean features at finer scales. Nevertheless, the two-dimensional observations of KaRIn altimeter of SWOT su...The SWOT (Surface Water Ocean Topography) mission is currently providing unpreceded high-resolution measurements of Sea Surface Height (SSH), revealing ocean features at finer scales. Nevertheless, the two-dimensional observations of KaRIn altimeter of SWOT suffer from instrumental errors. This noise degradation is altering the high frequencies of SWOT signal, and the small to sub-mesoscale dynamics of interest for some oceanographers. For this reason, Tr\'eboutte et al. (2023) have developed a convolutional neural network (CNN) based on U-Net architecture to separate the noise from the physical signals contained in the SSH. Their approach has demonstrated great potential on simulated SWOT measurements, and Dibarboure et al. (2024) report a positive influence on actual flight data from SWOT. However, degraded denoising performance, and occasional negative side-effects have been observed in atypical conditions (e.g. very high surface waves, internal tides solitons). In this study, we illustrate some of these limitations and we present an improved approach of the CNN-based denoising: we modified the training procedure to obtain a more robust version of the algorithm, to avoid biases and artifacts in the denoised SSH. This paper also presents a more complete validation process with a robust and standardized evaluation benchmark: these metrics could be of interest to assess other SWOT filtering and denoising algorithms.
|
| 1181 |
Logit-Aware MIMO AirComp for Distributed Mixture-of-Experts LLM Inference over Wireless Edge Networks
2610.03741
|
cs.LGcs.AI
|
Lyutianyang Zhang, Yunjian Jia, Liu Cao, Dengke Wang, Jinke Ren |
Distributed mixture-of-experts (MoE) inference is a promising architecture for deploying large language models (LLMs) at wireless edge networks because sparse experts can be placed across coordinated base stations (BSs), while the anchor node and user equipmen...Distributed mixture-of-experts (MoE) inference is a promising architecture for deploying large language models (LLMs) at wireless edge networks because sparse experts can be placed across coordinated base stations (BSs), while the anchor node and user equipment (UE) can offload LLM inference tasks to BSs. The communication bottleneck is the MoE aggregation, where the anchor BS must recover a weighted sum of selected expert outputs before each decoding step. Over-the-air computation (AirComp) is well matched to this operation because the wireless multiple-access channel naturally superposes simultaneous transmissions. However, conventional AirComp minimizes communication distortion, whereas MoE aggregation errors have unequal impact on LLM outputs. We propose a logit-aware MIMO AirComp framework that estimates local logit sensitivity as block-level weights and jointly optimizes receive combiners and BS precoders under per-BS power constraints. We also develop an alternating algorithm that protects decoding decisions from aggregation perturbations. Using OpenCompass, we evaluate Qwen3-30B-A3B-Instruct-2507-FP8 on GSM8K and ARC-Challenge. At 30 dB aggregation SNR, perturbed Qwen3 retains 99.3% of clean GSM8K accuracy and 94.8% of clean ARC-Challenge accuracy. In a direct ARC-Challenge closed-loop audit over 295 examples, SW-AirComp achieves 43.73% accuracy at 30 dB and 26.78% at 20 dB, outperforming unweighted, matched-filter, and zero-forcing AirComp. In the wireless simulator, SW-AirComp reduces decision-relevant aggregation distortion; at 20 dB, its final weighted-sum mean-squared error is 53.5% lower than unweighted AirComp under the same channels, samples, and sensitivity weights. Logit-RMSE and logit-gap diagnostics further show that the gain comes from reducing output-logit perturbation and lowering the risk of top-token changes.
|
| 1182 |
Structured Neural Modeling of Daily Arctic Sea-Ice Concentration Evolution: Physical-Trajectory-Driven Learning and Forecast-Domain Adaptation
2610.03743
|
cs.LG
|
Maqun Zhang, Feng Gao, Wankun Chen, Hui Yu, Yanhai Gan |
Accurate modeling of the daily evolution of sea ice concentration (SIC) is central to improving the credibility and operational forecasting capability of deep learning-based sea ice prediction. However, existing deep learning methods often couple the underlyin...Accurate modeling of the daily evolution of sea ice concentration (SIC) is central to improving the credibility and operational forecasting capability of deep learning-based sea ice prediction. However, existing deep learning methods often couple the underlying sea ice evolution relationships and data errors within high-dimensional nonlinear mappings, making it difficult to construct a stable and verifiable evolution core and apply it reliably to practical forecasting. To address this issue, this study proposes a reanalysis-forecast dual-domain decoupled framework for learning sea ice evolution operators. The framework builds upon a lightweight physical baseline to generate daily evolution trajectories, employs a temporally constrained joint multi-lead compensation network to com?pensate for unresolved processes, and introduces an ice-mass?aware transport mechanism to suppress numerical dissipation. In the forecasting stage, the parameters of the base evolution core are fixed, while a lightweight variable-semantic adaptation mechanism calibrates inter-domain distributions and evolution responses, thereby separating forecast-domain errors from base evolution errors. Experiments show that the constructed base evolution core can accurately and stably simulate daily sea ice evolution at both short-term and annual scales under reanal?ysis forcing, and can be effectively transferred to the forecast domain through lightweight adaptation, achieving stable prac?tical forecasting capability while preserving the base evolution structure. The source code will be made publicly available at https://github.com/zhangmaqun65535/SNM upon acceptance of this manuscript.
|
| 1183 |
Generalizable Neural Downscaling of Earth System Model Wind Fields via Continuous Dynamics Modeling
2610.03757
|
cs.LG
|
Chenxi Yu, Jianan Wei, Hanlin Kong, Hao Sun, Bian He |
Accurate high-resolution wind field simulations are critical for resolving fine-scale atmospheric dynamics, yet the simulation of wind fields in Earth System Models (ESMs) remains limited by coarse spatial resolution and systematic biases. To address this, dat...Accurate high-resolution wind field simulations are critical for resolving fine-scale atmospheric dynamics, yet the simulation of wind fields in Earth System Models (ESMs) remains limited by coarse spatial resolution and systematic biases. To address this, data-driven down scaling techniques have been widely used to enhance coarse-resolution ESM outputs. However, existing methods are typically tied to fixed discretizations, limiting generalization across models with different native resolutions. Here we formulate global near-surface wind downscaling as an operator-learning problem on continuous atmospheric state fields and develop a downscaling neural operator that maps coarse-scale fields to fine-scale counterparts across heterogeneous discretizations. The operator learning-based neural downscaling framework outperforms dominant baselines, recovers fine-scale physical structures, and generalizes to previously unseen ESMs and future climate scenarios without retraining, while preserving long-term wind projection trends. These findings establish a generalizable paradigm for high-resolution climate downscaling across diverse simulation outputs and future scenarios.
|
| 1184 |
BridgeCast: Bridging Ocean Wave Forecasts to Reanalysis via Flow Matching with Exogenous Variables
2610.03759
|
cs.LGcs.AI
|
Siyu Gan, Dongsheng Luo, Kunxiaojia Yuan, Dongjin Song, Jingchao Ni |
Ocean wave forecasting is essential for maritime safety, offshore operations, and coastal resilience, yet remains challenging due to systematic biases in physics-based models. Physical models, while widely used, rely on approximations and parameterizations tha...Ocean wave forecasting is essential for maritime safety, offshore operations, and coastal resilience, yet remains challenging due to systematic biases in physics-based models. Physical models, while widely used, rely on approximations and parameterizations that limit their accuracy under complex ocean-atmosphere conditions. To enhance ocean wave forecasting, we propose BridgeCast, within a physics-AI hybrid framework for bias correction. BridgeCast is a probabilistic model based on conditional flow matching (CFM) that learns to transform physical model forecasts into reanalysis-like fields. It treats physical forecasts as corrupted observations and employs a continuous-time generative process to bridge their distribution toward that of reanalysis data. BridgeCast is parameterized by a Transformer-based architecture that enables spatiotemporal modeling, incorporation of exogenous atmospheric variables, and flexible inference via both ordinary and stochastic differential equation formulations. Extensive experiments on real-world datasets demonstrate that BridgeCast consistently outperforms state-of-the-art baselines across regions and forecast lead times.
|
| 1185 |
Targeted Active Learning for Preference-Based Treatment Effects on Multivariate Outcomes
2610.03824
|
cs.LG
|
Lola Giordani, Mathieu Even, Chlo\'e Geoffroy, Jean-Christophe Corvol, Rapha\"el Porcher |
Treatment efficacy is traditionally demonstrated on the basis of a single primary outcome. However, clinical decision-making usually requires consideration of multiple outcomes, balancing expected benefits against potential risks. The relative value assigned t...Treatment efficacy is traditionally demonstrated on the basis of a single primary outcome. However, clinical decision-making usually requires consideration of multiple outcomes, balancing expected benefits against potential risks. The relative value assigned to these outcomes varies substantially from one patient to another. Given a preference rule over outcome profiles, treatment effects and optimal policies can be defined and estimated. Such a rule is rarely available in practice: it must itself be estimated from pairwise comparisons of outcome profiles, which are costly to collect from clinical experts. We propose an active learning framework that selects which comparisons to query. Standard criteria maximize the information gained on the preference rule itself. We instead target the quantities of interest, and select the query that most reduces uncertainty on the treatment effect and on the optimal policy induced by the learned rule. Under a Gaussian process model of the preference rule, we derive a closed-form approximation of this criterion. On semi-synthetic data built from a Parkinson's disease cohort with 13 clinical outcomes, our criterion achieves lower treatment effect estimation error and lower policy regret than existing criteria at equal query budget.
|
| 1186 |
Distribution Matching Evolutionary Algorithms for Rare Event Sampling
2610.03833
|
cs.LGcs.AI
|
Yonatan Gideoni, Yarin Gal |
A novel discovery is one which is both useful and surprising: a generative model's output is a useful discovery if it has a low probability of being generated (it's surprising) and a high reward (it's useful). Global optimization can directly increase the prob...A novel discovery is one which is both useful and surprising: a generative model's output is a useful discovery if it has a low probability of being generated (it's surprising) and a high reward (it's useful). Global optimization can directly increase the probability of sampling high rewards but typically requires updating model weights. Such gradient based optimization is expensive and bars using capable closed-source models. Instead, modern search methods for discovery sacrifice the global target, and use evolutionary algorithms with local reward maximizing objectives, permitting the search to focus only on high probability samples. In this paper, we interpret various evolutionary algorithms as approximate Markov Chain Monte Carlo, an optimization-free method to sample from complex distributions. This interpretation allows developing Distribution Matching Evolutionary Algorithms (DME), a class of search methods which sample from a global target distribution without updating weights. Empirically, DME has a higher sample efficiency than existing methods on problems requiring many samples to find a solution.
|
| 1187 |
Probability flow ODEs in score-based and reflected diffusion models
2610.03846
|
cs.LG
|
Rama CONT |
Probability-flow ordinary differential equations (PF-ODEs) are widely used as deterministic samplers for score-based diffusion models. Their usual justification is that the Fokker--Planck equation of a diffusion can be rewritten as a continuity equation driven...Probability-flow ordinary differential equations (PF-ODEs) are widely used as deterministic samplers for score-based diffusion models. Their usual justification is that the Fokker--Planck equation of a diffusion can be rewritten as a continuity equation driven by the score function of the forward diffusion. This identity does not, however, guarantee that the resulting velocity field generates a well-posed flow. We provide theoretical insights into the design of such deterministic samplers for generative models based on diffusions and reflected diffusions. We identify sufficient conditions for a regular Lagrangian PF-ODE flow to exist; reverse sampling and invertibility require two-sided divergence control. For learned scores, sampler stability is controlled by an unweighted velocity error, exposing a mismatch with density-weighted score matching and motivating architectural control of Jacobians, divergence, growth, and compression. Under the manifold hypothesis, positive-time regularization justifies an early-stopped PF-ODE while constants deteriorate near the data endpoint; an explicit sphere example shows that the exact deterministic flow becomes singular as the noise level vanishes even though the diffusion marginals remain well defined. These theoretical insights translate into concrete design principles for stable, invertible, and constraint-preserving diffusion samplers. We illustrate the practical relevance of these design principles using controlled numerical experiments.
|
| 1188 |
REACT: Physically and Chemically Consistent Reconstruction of Marine Active Tracers
2610.03888
|
cs.LGcs.AI
|
Wenbin Dai, Hao Zheng, Shiyu Liang, Chaofan Sun, Xueying Zhang |
Reconstructing global sea surface pH from sparse observations is critical for monitoring ocean acidification and understanding marine carbon cycling. Traditional assimilation and inverse models are physically grounded but costly for large-scale reconstruction....Reconstructing global sea surface pH from sparse observations is critical for monitoring ocean acidification and understanding marine carbon cycling. Traditional assimilation and inverse models are physically grounded but costly for large-scale reconstruction. Recent black-box and physics-guided AI models improve efficiency, but are mainly designed for passive tracers, where the reconstructed variable is also the transported inventory. In contrast, pH is an active carbonate tracer: it is the prediction target, while dissolved inorganic carbon (DIC) is the conserved carbon inventory. This mismatch can produce low pH error while violating carbonate closure and source-free carbon conservation. To address this, we introduce \textbf{REACT}, a carbon-first reconstruction framework that decouples transport, active correction, and chemical decoding. REACT transports a latent carbonate state with a conservative advection--diffusion solver, captures non-conservative carbon-cycle variations with a source module, decodes the corrected state into pH, and constrains the output through carbonate equilibrium. This design keeps pH as the target while enforcing consistency on the underlying carbon state. On simulation data, REACT reduces pH NRMSE by (14.7%) and chemical consistency error by (24.0%) over the best baseline. Cross-temporal-scale evaluations show robustness against error accumulation from coarse to fine temporal scales, and ablation studies validate the effectiveness of each component.
|
| 1189 |
Latent Score-Based Bayesian Cram\'er-Rao Bound Estimation for High-Dimensional Imaging Systems
2610.03956
|
cs.LG
|
Evan Scope Crafts, Thomas Wynn, Seonyeong Park, Mark Anastasio, Umberto Villa |
We propose a data-driven framework for estimating the Bayesian Cram\'er-Rao bound (CRB) in high-dimensional imaging systems with complex, analytically intractable priors. Direct CRB computation is challenging in this setting due to the need to model the prior ...We propose a data-driven framework for estimating the Bayesian Cram\'er-Rao bound (CRB) in high-dimensional imaging systems with complex, analytically intractable priors. Direct CRB computation is challenging in this setting due to the need to model the prior score and to form and invert the Bayesian Fisher information matrix in very high dimensions. To address these issues, we first reformulate the inverse problem in the latent space of a pre-trained variational autoencoder, thereby dramatically reducing the dimensionality of the bound estimation problem while preserving the spatial structure of the images. The Bayesian CRB is formed in this latent space and mapped back to the native parameter space using a change-of-variables formula. Second, to learn the latent prior score, we introduce a new sliced score matching objective defined in a Bochner space endowed with an $H^1(\Omega)$ Sobolev norm in space. This "Bochner-space sliced score matching" objective is consistent with standard sliced score matching, but penalizes errors in the spatial gradients of the score, suppressing high-frequency artifacts that otherwise contaminate the resulting CRB estimates. We validate the approach on a stylized quantitative photoacoustic computed tomography (qPACT) breast imaging problem with over one million unknown parameters, using a foundation-model autoencoder derived from Stable Diffusion. The proposed method yields stable, artifact-free Bayesian CRB estimates that reflect the highly non-Gaussian structure of the learned prior and reveal the substantial impact of the prior on the relative performance of competing qPACT design schemes.
|
| 1190 |
Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models
2610.03960
|
cs.LG
|
Luze Sun, Cristina Nita-Rotaru, Alina Oprea |
Instruction fine-tuning adapts a pretrained language model to follow instructions by training it on instruction--response pairs from a collection of tasks, such as summarization and question answering. Task-level poisoning exploits this task structure to manip...Instruction fine-tuning adapts a pretrained language model to follow instructions by training it on instruction--response pairs from a collection of tasks, such as summarization and question answering. Task-level poisoning exploits this task structure to manipulate the fine-tuned model into producing attacker-specified biased content on a particular target task, without requiring an explicit input trigger. Detecting such attacks is challenging because there is no explicit trigger to identify, the target task and biased content are unknown, and benign fine-tuning itself changes model behavior. We introduce Localize-and-Detect, a two-stage black-box auditing method for task-level poisoning that requires only outputs from both the base and fine-tuned models. In the first stage, we localize the target task by identifying candidate tasks on which the fine-tuned and base models have the largest differences in their next-token distributions. In the second stage, we search the shortlisted tasks for biased content that repeatedly appears in the fine-tuned model's responses but not in those of the base model. We evaluate Localize-and-Detect on 216 poisoned models across two model families, varying the target task, poisoning mode, poison budget, and type of biased content. Our evaluation demonstrates that Localize-and-Detect can effectively localize target tasks and detect biased content across a range of poisoning settings and models, with detection that tracks attack success and few false positives on clean models.
|
| 1191 |
Teaching Agents to Code Reliably
2610.03984
|
cs.LGcs.AI
|
Muhammad Ahmed Mohsin, Myeongsoo Kim, Kangrui Ruan, Shweta Garg, Varun Kumar |
Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three be...Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather than byproducts of scale, so a policy can carry them instead of a scaffold. Location diversity remains narrow, since attempts return to the same site and extra samples add no coverage. Edit diversity is left unexploited, since methodologies that differ resolve complementary issues no single run reaches. Verification misleads, since a test the agent writes for its own patch accepts many incorrect ones. Directing search by execution feedback and scoring each patch against its own reverted tree resolves 52.8% of SWE-bench Verified using 48.1% of the agent-steps an eight-sample baseline spends. Training moves these behaviors into the policy. On the 270 issues held out from SFT and RL training, weighted supervised fine-tuning raises pass@1 from 31.9% to 35.2% and pass@8 from 46.7% to 51.1%. A reinforcement objective then trains the verifier against gold-labeled repairs and incorrect variants, crediting the assertions that detect them. It raises pass@1 to 43.0% and pass@8 to 60.7%, lifts verifier precision from 26.8% to 41.7%, and more than halves false acceptance. Resolution improves on two of three out-of-distribution suites and verifier precision on all three, and the gains hold at 7B, 14B, and 30B against published coder baselines.
|
| 1192 |
SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown
2610.04008
|
cs.LGcs.AI
|
Yuxuan Liu, Haoran Li, Yuhao Zhang, Jiahe Guo, Hongyu Luo |
Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation rep...Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce SkillScriptBench, a 350-task benchmark designed to evaluate these capabilities separately. From a survey of over 35,000 GitHub-hosted Skill roots, we select 100 packages and construct 150 repair tasks. Each task pairs a package containing injected script faults with a maintenance request and executable checks of the required behavior. A complementary controlled track contains 200 tasks from 50 packages, each evaluated under the same maintenance request in four states: clean, documentation faults, script faults, and faults in both. Across four LLMs, methods that edit both documentation and scripts can repair script faults but do not consistently outperform Markdown-only revision on documentation repair or preservation. We therefore introduce AST-Guided Skill Revision, which uses abstract syntax trees and calling relationships to link maintenance requirements to relevant code locations. It restricts script edits to these locations and updates the documentation to match the revised scripts. Averaged across models, this revision stage yields absolute gains in repair success of 21.9% for Raw Package and 27.7% for CoEvoSkills on faulty packages. Absolute gains in the proportion of tasks solved in all three runs reach 20.8% and 31.5%, respectively, indicating more consistent repair success across repeated runs.
|
| 1193 |
Gated Graph Neural Networks for Learning Hidden Independent Cascade Dynamics
2610.04033
|
cs.LG
|
Anubha Goel, Illia Oleksiienko, Mateusz Wilinski, Alexandros Iosifidis, Juho Kanniainen |
Information and infectious diseases spread through social networks, but the spreading probabilities driving them are hard to estimate without per-node activation times. Applications seldom supply these, and inference must instead proceed through indirect and n...Information and infectious diseases spread through social networks, but the spreading probabilities driving them are hard to estimate without per-node activation times. Applications seldom supply these, and inference must instead proceed through indirect and noisy proxies for the terminal infection states. We study this inverse problem on a fixed, known graph under the Hidden Independent Cascade (HIC) model, with one spreading probability per node rather than a single global rate, so the number of unknowns scales with the number of nodes. Seed sets and observation parameters are known, while activation times and terminal infection states are latent, and the observed-data likelihood requires marginalizing over every spreading outcome. We propose a simulation-based amortized estimator that recovers the full node-level parameter vector without reconstructing individual latent cascades. Repeated seed-conditioned symptom observations are summarized as Symptom-Aware Cascade Features (SACF), which combine empirical symptom statistics with neighborhood and structural information. SACF are mapped to parameters by SAGE-HC, a permutation-equivariant gated graph neural network whose learned gates attenuate neighborhood messages corrupted by false positives and false negatives, and training on simulated HIC realizations yields a reusable inverse map. As a benchmark under the same hidden observations, we extend the Dynamic Message Passing learning framework to the HIC emission model. On synthetic and empirical graphs the two methods separate by topology. DMP is highly accurate on trees, whereas SAGE-HC is substantially better on heterogeneous, loopy graphs under noisy terminal symptoms.
|
| 1194 |
Adaptive Partitioning Schemes for Optimistic Optimization
2610.04039
|
cs.LG
|
Raja Sunkara, Ardhendu Tripathy |
Applications such as engineering design often require us to optimize a black-box function, i.e., a system whose inner processing is not analytically known and whose gradients are not available. Practitioners often have a fixed budget for the number of function...Applications such as engineering design often require us to optimize a black-box function, i.e., a system whose inner processing is not analytically known and whose gradients are not available. Practitioners often have a fixed budget for the number of function evaluations and the performance of an optimization algorithm is measured by its simple regret. In this paper, we study the class of "Optimistic Optimization" algorithms for black-box optimization that use a partitioning scheme for the domain. We develop algorithms that learn a good partitioning scheme and use flexible surrogate models such as neural networks in the optimization procedure. For multi-index functions on an $m$-dimensional subspace within $d$ dimensions, our algorithm attains $\tilde{O}(n^{-\beta / d})$ regret, where $\beta = 1 + \frac{d-m}{2m-1}$, as opposed to $\tilde{O}(n^{-1/d})$ for SequOOL, a state-of-the-art optimistic optimization algorithm. We use our approach to improve the quality of Activation-aware Weight Quantization (AWQ) of the OPT-1.3B model, achieving $\sim10\%$ improvement in performance relative to the best possible unquantized model.
|
| 1195 |
Geometry-Dependent Approximation for Non-Monotone $k$-Submodular Maximization
2610.04049
|
cs.LG
|
Vaneet Aggarwal |
We study nonnegative, non-monotone $k$-submodular maximization with $k\ge2$ labels under support constraints, and show how the certified approximation coefficient improves as the support region permits more uniform selection. For a compact convex down-closed s...We study nonnegative, non-monotone $k$-submodular maximization with $k\ge2$ labels under support constraints, and show how the certified approximation coefficient improves as the support region permits more uniform selection. For a compact convex down-closed support region $P\subseteq[0,1]^n$, the diagonal level $\zeta(P)=\max\{t\in[0,1]:t {\bf 1} \in P\}$ ranges from $\zeta=0$, which carries no geometric promise, to $\zeta=1$, which is unrestricted support. Our main structural result is a comparator-uniform linearization of the multilinear extension, built from an objective-independent action and a comparator-independent update field. For $k\ge3$, its validity reduces, independently of the number of elements, to four polynomial inequalities of degree at most three in one or two variables, only one of which depends on $k$. Explicit parameter choices give a nondecreasing certified profile $\underline{\alpha}_k(\zeta)$, in closed form on all of $[0,1]$ when $k=2$. At $\zeta=0$ we certify $0.4456\ldots$ for $k=2$ and $0.4541\ldots$ for every $k\ge3$, improving the recent $\sqrt2-1$ guarantee for one matroid or one knapsack, as well as the $1/3$-type guarantees for a fixed number of budgets; at $\zeta=1$ we certify $1/2$ for $k=2$, $(\sqrt{17}-3)/2$ for $k=3,4$, and $k/(2k-1)$ for $k\ge5$, whose excess over $1/2$ is of order $1/k$ rather than the previous $1/k^2$. Value-retaining rounding transfers these guarantees to matroid and knapsack constraints, and the same field yields $O(\sqrt T)$ approximate regret online under gradient or post-decision value feedback.
|
| 1196 |
Application of sequence learning for predicting radiation damage of the CMS electromagnetic calorimeter
2610.04058
|
cs.LG
|
Mario Ivan Gallegos Torres, Leonid Serkin, Guy Paic |
In this paper we use machine learning methods to predict radiation damage in the lead tungstate crystals of the CMS electromagnetic calorimeter at the Large Hadron Collider. We analyze LHC Open Data collected from 2016 to 2018 and study the time evolution of t...In this paper we use machine learning methods to predict radiation damage in the lead tungstate crystals of the CMS electromagnetic calorimeter at the Large Hadron Collider. We analyze LHC Open Data collected from 2016 to 2018 and study the time evolution of the crystal optical transparency. We apply deep neural network models to predict its behavior over different future time intervals, and find that encoder-decoder sequence-to-sequence architectures can effectively describe crystal aging.
|
| 1197 |
Exact Optimal Transport by Matching
2610.04085
|
cs.LG
|
Dmitry Kamenetsky |
Balanced discrete optimal transport between n sources and n targets of unit mass is exactly the minimum-cost assignment problem-a bipartite perfect matching-and is therefore solvable exactly by industrial matching engines in milliseconds to seconds. We ask whe...Balanced discrete optimal transport between n sources and n targets of unit mass is exactly the minimum-cost assignment problem-a bipartite perfect matching-and is therefore solvable exactly by industrial matching engines in milliseconds to seconds. We ask when the exact approach beats the standard approximate alternatives, entropic Sinkhorn and its accelerated variant Greenkhorn, and make the sparse-exact side certified by a textbook LP dual-feasibility clip. Three contributions. (i) Measurement: on dense 2-D instances, exact matching (Jonker-Volgenant) is faster and strictly more accurate than either approximate method throughout the moderate-n regime (0.01 s at n=500 to 11.5 s at n=8000); reaching a 1% quality target on the same hardware requires roughly 10-80 min for Greenkhorn (factors 4e2-6e4 over exact; plain Sinkhorn is 20-650x slower still), a rough power-law projection beyond the measured range. Greenkhorn's measured speedup over plain Sinkhorn is only 1.0-1.5x on most converged cells. (ii) A simple kNN-pool gap certificate: given a pool matching and its Blossom dual, a one-pass O(n^2) clip produces a dense-feasible lower bound; combined with the Sinkhorn dual potential (valid at every iterate, not just at convergence), the bound is valid on all 45 measured configurations and tightens monotonically with k. (iii) A multi-robot task-allocation sanity check where the discrete plan is the deliverable: per-round exact assignment costs 0.1-68 ms, while a Sinkhorn-plus-hardening pipeline costs 0.12-15.9 s and accumulates 6-27% extra travel over 15 rounds. The Sinkhorn family's large-n dense regime is acknowledged and left untouched. Code, data, and results under MIT: https://github.com/dimkadimon/OT-Blossom.
|
| 1198 |
Shared Geometry Is Not Shared Physics: A Layerwise Test of the Platonic Representation Hypothesis in Astronomy
2610.04130
|
cs.LG
|
Kshitij Duraphe, Aravind Kannappan, Dun Li Chan, Okiki Famutimi, Yaswant Sai Ejjagiri |
In this paper we investigate whether geometrically aligned astronomical representations are also scientifically interchangeable. We analyze every layer's representation from 36 pretrained models using cross-matched optical images, infrared images, and spectra....In this paper we investigate whether geometrically aligned astronomical representations are also scientifically interchangeable. We analyze every layer's representation from 36 pretrained models using cross-matched optical images, infrared images, and spectra. All 139 analyzed model-survey comparisons show significant local-neighborhood alignment somewhere in the network after accounting for the search over depth. However, alignment does not generally increase with network depth and can be stronger for the changes between consecutive computational blocks than for the block outputs themselves. We find that geometry-selected stages under-perform label-selected stages in every image-transfer comparison; moreover, every final HSC-COSMOS-Web model pair is aligned while every redshift transfer has negative R^2. Models can therefore recover a similar cross-survey geometry among astronomical objects without recovering a survey-invariant linear encoding of the physical properties tested here.
|
| 1199 |
Agentic Resource Allocation for Batch Multi-Objective Bayesian Optimization in Autonomous Materials Discovery
2610.04134
|
cs.LGcs.AI
|
Robert Robinson, Shakti Prasad Padhy, Sushant Sinha, Sk Md Ahnaf Akif Alvi, Juan Florez Coronel |
The discovery and development of advanced materials is a challenging process constrained by the high time and monetary costs of synthesis, processing, and characterization. The underlying design spaces can be enormous, often with multiple competing objectives....The discovery and development of advanced materials is a challenging process constrained by the high time and monetary costs of synthesis, processing, and characterization. The underlying design spaces can be enormous, often with multiple competing objectives. Bayesian optimization (BO) provides a principled approach for efficiently navigating such spaces, but most workflows rely on fixed exploration-exploitation policies that lack the capacity to adapt to shifting constraints in dynamic campaigns typical of self-driving laboratories. In this work, we develop a multi-objective BO framework for alloy design under resource constraints, benchmarking strategies for adaptive policy tuning at each iteration. Our evaluation covers a septenary refractory high-entropy alloy (RHEA) system focused on maximizing melting temperature and minimizing density, and an Fe-Co-Ni-based soft magnetic alloy system targeting saturation magnetization, coercivity, and hardness. We compare an exploitation-focused strategy, a fixed mixed exploratory/exploitative policy, and two distinct LLM-based adaptive strategies with different approaches to batch allocation and campaign signal interpretation, evaluated across baseline and mid-campaign resource event conditions including budget reductions, timeline cuts, and combined disruptions. Our results show that mixed allocation strategies accumulate substantially more mutual information than the exploitation-focused baseline at a proportionally smaller cost to hypervolume and optimization speed, with adaptive strategies outperforming a fixed-mixed allocation policy by adjusting their allocation in response to both evolving campaign statistics and resource constraints. These findings suggest that adaptive resource allocation offers a favorable tradeoff for materials discovery campaigns in reducing predictive uncertainty on Pareto-optimal compositions.
|
| 1200 |
One-Cycle Fault Classification and Faulted-Line Identification on the PROTECT-90 Dataset: An Initial Application Benchmark
2610.04155
|
cs.LG
|
Emad Abukhousa, Abdulaziz Qwbaiban, Saman Zonouz, A. P. Sakis Meliopoulos |
Open electromagnetic-transient datasets are beginning to make reproducible learning-based protection studies possible, but the practical use of these datasets still requires application-level benchmarks that define timing, sensing, and validation assumptions. ...Open electromagnetic-transient datasets are beginning to make reproducible learning-based protection studies possible, but the practical use of these datasets still requires application-level benchmarks that define timing, sensing, and validation assumptions. This paper presents an initial application benchmark on the recently released PROTECT-90 dataset for two protection-oriented tasks: fault-type classification and discrete faulted-line identification. A compact one-dimensional convolutional neural network (CNN) is evaluated using post-inception windows of 0.25, 0.5, 1, and 2 cycles under strict episode-wise splitting. A non-convolutional multilayer perceptron (MLP) is also trained as an architecture-control baseline. The results show that both tasks are nearly saturated under full observability, with one-cycle test accuracies of 99.84% for fault type and 100.00% for line identification. The main performance variation appears under reduced observability: current-only inputs preserve line identification accuracy at 100.00%, whereas voltage-only inputs reduce line identification accuracy to 53.09% with the CNN and 50.57% with the MLP. This indicates that the limiting factor is measurement information rather than neural architecture. Additional stratified checks show stable performance across topology states and fault-resistance bins, while CPU inference contributes only 0.528 ms to the one-cycle total decision time of 20.53 ms.
|
| 1201 |
Mitigating Over-squashing without Rewiring: A Sheaf Effective Resistance Perspective
2610.04157
|
cs.LG
|
Andr\'e Ribeiro, Germano Barcelos, Amauri H. Souza, Diego Mesquita, Ana Luiza Ten\'orio |
Graph Neural Networks (GNNs) often struggle to capture long-range dependencies due to over-squashing -- a phenomenon in which the repeated compression of node embeddings into finite-size messages causes representations to collapse. Over-squashing is most often...Graph Neural Networks (GNNs) often struggle to capture long-range dependencies due to over-squashing -- a phenomenon in which the repeated compression of node embeddings into finite-size messages causes representations to collapse. Over-squashing is most often diagnosed as a property of the graph topology, with effective resistance serving as a principled measure of the bottleneck. We provide a complementary view on the matter: building on cellular sheaves, we introduce sheaf effective resistance, a generalization of effective resistance that depends on the sheaf attached to the graph, and we prove that for flat vector bundles, the over-squashing sensitivity in the Jacobian sense is upper bounded by a quantity related to the sheaf effective resistance between the nodes. The bottleneck thus need not lie in the graph itself: it can be relocated, and reduced, by adjusting the sheaf. We instantiate this idea in FlatNSD, a simple message-passing variant of Neural Sheaf Diffusion, and show that it implicitly learns to modulate total sheaf effective resistance, performing well on benchmarks designed to stress over-squashing without altering the original graph topology.
|
| 1202 |
MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memory
2610.04195
|
cs.LGcs.AI
|
Priyanka Mudgal, Kai Zhao, Guilin Zhang, Andy Olsen, Ezekiel Miller |
Personal AI agents in enterprise multi-tenant deployments share a common vector store for long-term memory. Shared embedding spaces create a surface for cross-user memory leakage: a user's query can retrieve semantically adjacent memories belonging to another ...Personal AI agents in enterprise multi-tenant deployments share a common vector store for long-term memory. Shared embedding spaces create a surface for cross-user memory leakage: a user's query can retrieve semantically adjacent memories belonging to another user through ordinary cosine-similarity retrieval, without any exploit. We formalize this as cross-user admissibility failure and evaluate it across six experiments, plus follow-up ablations, under both sparse (TF-IDF) and production-faithful (MiniLM-L6-v2) retrieval. Non-adversarial, incidental leakage reaches 70--100\% under pooled {same-team} retrieval; adversarially crafted memories achieve 90--100\% top-$k$ placement, exceeding weaker keyword-based attacker baselines, with score lifts of $+0.416$ to $+0.511$ under production-faithful dense retrieval (Config B); and end-to-end response contamination reaches 5.00/5 under a production retrieval path and 4.67/5 with Claude Sonnet~4.5, with contaminated responses often scoring as helpful or more helpful than clean ones, a gap validated against human judgment. Among three architectural mitigations, only hard post-retrieval ownership gating consistently restores the clean baseline (1.00/5) across {two generation models, at a measured latency overhead of roughly 1.4~ms per query.
|
| 1203 |
ALoDLM: Adaptively Looped Diffusion Language Models
2610.04198
|
cs.LGcs.AI
|
Liancheng Fang, Zhuowei Li, Youngeun Kim, Tianchen Zhao, Rajat Koner |
Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a comp...Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a partially observed sequence, some unknown tokens are easy to predict, while others require substantially more computation. Existing DLMs nevertheless apply uniform computational depth to all unknown positions at each denoising step. We introduce ALoDLM, which replaces uniform computation with token-adaptive latent recurrence. At each denoising step, ALoDLM iteratively refines latent representations and allocates computation according to token difficulty. Tokens ready to commit are fed back as discrete context, while unresolved tokens retain and further refine their latent states through additional recurrent passes. To learn token prediction and computation allocation jointly, we formulate token-wise computation schedules as latent variables and derive a conditional negative evidence lower bound (NELBO). We train ALoDLM at 1.7B and 8B parameter scales. Across eleven benchmarks, ALoDLM outperforms all evaluated DLMs and the corresponding AR baselines in average benchmark score at both scales. ALoDLM also retains fast parallel decoding, yielding a strong quality-efficiency trade-off among evaluated autoregressive and diffusion models under optimized inference engines.
|
| 1204 |
Humanoid Rickshaw Pulling: Whole-Body Locomotion under Coupled Wheeled Loads
2610.04238
|
cs.LG
|
Yangzhi Yang, Xiansheng Lin, Zhaoming Xie, Xiaobin Xiong |
Humanoid robots could transport payloads substantially heavier than themselves by pulling passive wheeled vehicles instead of carrying the load. This capability, however, creates a coupled locomotion problem: the robot must maintain persistent upper-body conta...Humanoid robots could transport payloads substantially heavier than themselves by pulling passive wheeled vehicles instead of carrying the load. This capability, however, creates a coupled locomotion problem: the robot must maintain persistent upper-body contact while adapting to unknown, configuration-dependent forces arising from the payload, vehicle, and terrain. We present a whole-body control framework for humanoid rickshaw pulling that tracks commanded vehicle motion while preserving balance and stable grasps under uncertain load dynamics. During training, a privileged teacher exploits vehicle states, interaction forces, and load properties. Its actions and latent are distilled into a history-conditioned student that implicitly infers coupled dynamics from proprioceptive responses, followed by reinforcement-learning fine-tuning. Comparisons with \emph{No History} and \emph{Only History} baselines show that the resulting policy achieves accurate vehicle tracking while reducing vehicle oscillation, torso tilt, and actuation cost. Behavioral analysis shows that Unitree G1 propels the rickshaw and generates gait-synchronized whole-body reactions that stabilize its lateral and roll motions. Moreover, pulling redistributes joint effort and yields a lower robot-normalized cost-of-transport proxy than unloaded walking over most tested load--speed conditions. On hardware, a single policy performs starting, sustained pulling, turning, and stopping with both rigid payloads and human passengers, handling a loaded rickshaw mass of up to 115~kg without load-specific retuning. These results demonstrate robust heavy-load transportation through coordinated and persistent humanoid--vehicle interaction.
|
| 1205 |
Attention-Based Surface Representation Learning for Robot State Prediction and Open-Ended Surface Classification
2610.04240
|
cs.LGcs.AI
|
Oleg Kushnarev, Alexander Belyaev |
For ground robots operating in outdoor environments, understanding the properties of the underlying terrain is essential for ensuring reliable operation. In most perception-based studies, this problem is formulated as categorical classification with a fixed nu...For ground robots operating in outdoor environments, understanding the properties of the underlying terrain is essential for ensuring reliable operation. In most perception-based studies, this problem is formulated as categorical classification with a fixed number of classes defined during training. We propose an approach that enables new surface classes to be added as trainable vectors, which can subsequently be used to address higher-level tasks. By employing a learning paradigm based on predicting the robot's next state in time and using attention blocks, we improved classification accuracy to 98.56% on the Belyaev-Kushnarev dataset and 94.8% on BorealTC.
|
| 1206 |
Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems
2610.04285
|
cs.LG
|
Qi Feng, Gu Wang |
We develop an amortized, grid-free implementation of continuous Langevin dynamics based policy-value iteration for entropy-regularized, infinite-horizon relaxed stochastic control problems. The improvement rate of the exact iteration is a discounted aggregate ...We develop an amortized, grid-free implementation of continuous Langevin dynamics based policy-value iteration for entropy-regularized, infinite-horizon relaxed stochastic control problems. The improvement rate of the exact iteration is a discounted aggregate of relative Fisher information between the policy and the Gibbs law of its Hamiltonian. The associated score residual is the velocity with which the control's Langevin dynamics transport its law. We project this velocity onto a conditional sampler shared across states, instead of one Langevin dynamics per state, and the value dynamics onto a parametric critic, estimating both projections at sampled states to obtain coupled actor--critic flows. The score loss measures the actor's agreement with the current critic, while the policy-evaluation residual measures the critic's agreement with the actor. We also derive gradient and Hessian residuals, including a Feynman--Kac representation for the gradient equation, to control errors not detected by the projected value iteration. An exact decomposition of the HJB residual combines these errors into a policy-suboptimality bound under verification and logarithmic Sobolev assumptions. In the linear-quadratic class, both projections are exact and recover the pointwise iteration, and we provide numerical experiments on general models to demonstrate the coupled actor--critic learning in high-dimensions.
|
| 1207 |
CyTReX: Explainable AI-Based Cybersecurity Threat Reasoning Framework for DER Networks
2610.04286
|
cs.LGcs.AI
|
Damilola Popoola, Souradeep Bhattacharya, Manimaran Govindarasu |
Distributed Energy Resource (DER) environments rely on network communication protocols to coordinate control commands, measurements, and device states across edge assets and cloud systems. Edge anomaly detection systems (ADS) monitor this traffic to identify d...Distributed Energy Resource (DER) environments rely on network communication protocols to coordinate control commands, measurements, and device states across edge assets and cloud systems. Edge anomaly detection systems (ADS) monitor this traffic to identify deviations from normal communication behavior, flagging suspicious flows for further investigation. When the ADS flags abnormal network traffic, a single attack label is often insufficient for operational response: the label reports the detector's selected class but does not expose alternative threat interpretations that may warrant investigation. This paper presents Cybersecurity Threat Reasoning with Explainable Artificial Intelligence (CyTReX), an evidence-grounded threat reasoning framework for DER security that transforms network-level anomaly alerts into ranked, analyst-facing threat hypotheses designed to support Security Operations Center (SOC) triage and investigation. CyTReX constrains large language model (LLM) reasoning through a structured evidence packet, defined as a consolidated record of detection outputs, model explanations, and cyber threat intelligence (CTI) context. The evidence packet integrates edge-layer anomaly detection evidence, cloud reasoning layer attack interpretation, Shapley Additive Explanations (SHAP) network-feature attributions, surrogate decision rules, and Model Context Protocol (MCP)-enabled CTI enrichment. This ensures that every ranked hypothesis and attack-tree branch is traceable to explicit evidence rather than free-form LLM inference, and that incomplete or conflicting evidence is communicated rather than suppressed. Evaluation across five configurations shows that additional reasoning components improve hypothesis specificity, evidence traceability, and analytical grounding, with the complete pipeline providing the richest evidence-grounded reasoning context.
|
| 1208 |
Local Fisher Information Enables Sparse Causal Discovery
2610.04291
|
cs.LG
|
Byeongguk Kang, Donghyeon Lee, Euijong Song, Gunwoong Park |
Sparse causal discovery calls for methods that exploit graph structure without estimating high-dimensional densities. We introduce Fisher Information Completion Search (FiCS), a source-first algorithm for additive noise models that uses one local Fisher score ...Sparse causal discovery calls for methods that exploit graph structure without estimating high-dimensional densities. We introduce Fisher Information Completion Search (FiCS), a source-first algorithm for additive noise models that uses one local Fisher score for both ordering and parent selection. Under regularity and nonconstant-parent conditions, we prove that a node's local Fisher information equals the noise Fisher information exactly when the conditioning set contains all parents, provided that it contains no descendants. This Fisher parent completion identifies the parent set as the unique minimal Fisher completion. With a maximum conditioning set size $q$ at least the maximum indegree $d$, population FiCS queries marginals of at most $q+1$ variables and recovers the true directed acyclic graph under a positive ordering margin. Bounded conditioning also has a population advantage: reducing $q$ toward $d$ cannot decrease, and can strictly increase, the ordering margin. A growing non-Gaussian family separates local Fisher selection from conditional-variance and leaf-first Fisher ordering. For the regularized kernel Stein estimator, we establish high-dimensional DAG consistency under $q\{1+\log(p/q)\}+\log p=o(n)$, uniform Fisher separation, local approximation, and compatible ridge and parent penalty parameters. Experiments show the strongest gains when $n$ is small relative to $p$, quantify the effect of the conditioning size, and demonstrate competitive reference-graph recovery on three real-data benchmarks.
|
| 1209 |
Latent Safety Filters: When a Lossy Encoder Admits a Transferable Certificate
2610.04297
|
cs.LG
|
Johannes Mootz, Zahra Nili Ahmadabadi, Reza Akhavian |
Latent safety filters certify safety on a learned low-dimensional representation of the state, enabling constraints that resist analytic description. Because the encoder is lossy, a filter can report safe while the physical state is unsafe, with no detectable ...Latent safety filters certify safety on a learned low-dimensional representation of the state, enabling constraints that resist analytic description. Because the encoder is lossy, a filter can report safe while the physical state is unsafe, with no detectable model error. Existing transfer conditions leave the effect of discarded safety information implicit. We ask when a lossy encoder admits a safety certificate that transfers to the physical system, and show the answer is governed by the detectability of the discarded safety-relevant dynamics. We construct a system whose latent model is exact and whose latent signals always report safe, while the physical state becomes arbitrarily unsafe. For this system no certificate exists and no monitor downstream of the encoder can detect the failure. When the discarded dynamics contract, a latent barrier certifies true safety up to two explicit margins, one for the latent-model error and one for the variation of safety across states the encoder cannot distinguish. In the linear case and under boundedness and non-degeneracy conditions, every calibrated barrier transfers with a finite margin when the safety-relevant subspace is detectable, and none does otherwise. On learned cartpole encoders, the model error does not indicate for which representations the estimated bound is non-vacuous, while the second margin does.
|
| 1210 |
ML-OPF-Bench: Benchmarking Machine Learning for Optimal Power Flow
2610.04307
|
cs.LG
|
Xinyi Liu, Xuan He, Danny H. K. Tsang, Yize Chen |
Machine Learning (ML) methods promise a fast solution process for Optimal Power Flow (OPF). While inconsistent test cases, implementations, and evaluation metrics across existing studies make it challenging to determine which algorithmic advances are most crit...Machine Learning (ML) methods promise a fast solution process for Optimal Power Flow (OPF). While inconsistent test cases, implementations, and evaluation metrics across existing studies make it challenging to determine which algorithmic advances are most critical for real-world deployment. To this end, we propose ML-OPF-Bench, a unified benchmark for AC- and DC-OPF that evaluates representative ML algorithms under a consistent pipeline, stress-tests them across system sizes, distribution shifts, and resource budgets, and ranks them with a multi-objective framework. We find that prediction accuracy alone is not a reliable indicator of operational feasibility. Under heavily loaded, congested conditions, even the strongest in-distribution performers lose their advantage, while feasibility is maintained largely by post-processing that enforces the target constraints rather than by the underlying pure ML predictor. Data scaling shows that prediction accuracy and constraint violations follow different trajectories, whereas compute scaling shows that returns diminish and that larger models do not consistently perform better. These results expose critical trade-offs among ML methods' speed, accuracy, and feasibility, and offer practical guidance for future ML-OPF design. We open-source the benchmark as an extensible Python package for integrating new learning-based OPF algorithms and evaluating them under the same standard as the existing baselines.
|
| 1211 |
FLAT: Smoothing the Rugged Landscape for Learnable, Sample-Efficient Traffic Calibration
2610.04337
|
cs.LG
|
Haopeng Deng, Shuo He, Dayuan Wang |
Calibrating microscopic traffic models for digital twins is an expensive black-box optimization problem: tuning car-following and lane-changing parameters requires a full simulation run, affording only a tight budget per recalibration window. Matching raw traj...Calibrating microscopic traffic models for digital twins is an expensive black-box optimization problem: tuning car-following and lane-changing parameters requires a full simulation run, affording only a tight budget per recalibration window. Matching raw trajectories yields a rugged objective that sparse surrogates cannot learn, reducing sequential acquisition to near-random probing. We present FLAT, which couples what to optimize with where to sample next. An eight-dimensional behavioral fingerprint smooths the parameter-error landscape, making the objective learnable from a few dozen samples; annealed lower-confidence-bound (LCB) acquisition then spends each remaining run where it most reduces error. The surrogate, interchangeable among a Gaussian process (GP), random forest (RF), or multi-layer-perceptron (MLP) ensemble, plugs into the same LCB loop. Across six heterogeneous real-world scenes, FLAT-GP achieves the lowest scene-averaged behavioral error, winning 6/6 scenes against SPSA, GA, and CMA-ES and 5/6 against TPE under the matched budget. Some baselines need up to 4.4 times more simulations to match. Ablations show objective choice shifts final behavioral error by 81% on average, removing sequential LCB raises the six-scene mean by 20%, and surrogate choice shifts it by at most 4.2%, confirming gains trace to objective geometry and sequential allocation rather than surrogate capacity.
|
| 1212 |
A KKL Observer Perspective on Reservoir Computing
2610.04343
|
cs.LG
|
Anastasia Bizyaeva, Fernando Casta\~nos, Jaime A. Moreno |
Reservoir computing (RC) is a machine learning technique for data-driven modeling of dynamics for forecasting and control, primarily studied in computer science and physics literature with promising applications in neural network learning, physical computing, ...Reservoir computing (RC) is a machine learning technique for data-driven modeling of dynamics for forecasting and control, primarily studied in computer science and physics literature with promising applications in neural network learning, physical computing, and neuroscience. Why reservoirs learn and how to choose good reservoir architectures are considered important open questions. We show that the RC problem is mathematically an extension of a classic problem in systems and control theory, the Kazantzis-Kravaris-Luenberger (KKL) observer design problem. As a consequence, many of the questions considered open for RC stand to benefit from a large body of theory in the mature KKL literature, non-exhaustively including on questions of embedding, transverse stability, local and global uniqueness guarantees, and effective data-driven solution constructions. Elaborating on this connection, we show that the surprising forecasting ability of reservoirs is in fact a direct consequence of the well-known observer internal model principle, derive an upper bound on the prediction error over a fixed forecast horizon, and provide a partial explanation for why linear readout training in RC works reasonably well. This work illustrates how classical ideas from systems and control can provide strong theoretical backing and open new questions for modern machine learning methods.
|
| 1213 |
Frame-Level Temporal Alignment for Human-to-Robot Visual Adaptation
2610.04372
|
cs.LG
|
Xizhe Zhang, Jingfeng Zhang, Zirun Zhou, Hong Jia |
Transferring visual representations pretrained on human videos to robot manipulation requires learning reliable correspondences between human and robot demonstrations. However, paired demonstrations can differ in execution rate and in the proportion of non-key...Transferring visual representations pretrained on human videos to robot manipulation requires learning reliable correspondences between human and robot demonstrations. However, paired demonstrations can differ in execution rate and in the proportion of non-key frames that do not directly reflect task progress. Frames at the same relative timestamp may therefore represent different task stages, which can cause correspondence learning to fail. To address these issues, we propose Frame-Level Temporal Alignment (FLTA), a framework that uses two temporal priors to adapt visual encoders pretrained on human videos for robot manipulation. It learns shared task-progress representations without frame-level correspondence annotations, allowing frames at different relative temporal positions to match. A global progress prior combines normalized temporal positions with visual similarity to construct a soft correspondence target. A local temporal order prior penalizes backward transitions while allowing stays and varying forward rates to accommodate execution-rate differences. With ResNet-50 and ViT encoders, our method achieves relative improvements of 46.93% and 65.96%, respectively, over the best baselines in average simulation success rates and achieves higher task success rates on real-world manipulation tasks. These results also suggest that effective human-robot adaptation depends less on the number of parameters updated than on which parameters are selected. Our project page is available at https://rtx5090ultra.github.io/FLTA-Project-Page/.
|
| 1214 |
Largest Rashomon sets of decision trees for robust contextual optimization
2610.04385
|
cs.LG
|
Lorenzo Bonasera, David Pisinger |
Many decision trees fit the same data almost equally well, yet they can route a query point to different leaves and induce different local empirical distributions. We study decisions that meet prescribed cost, shortage or risk targets despite this predictive m...Many decision trees fit the same data almost equally well, yet they can route a query point to different leaves and induce different local empirical distributions. We study decisions that meet prescribed cost, shortage or risk targets despite this predictive multiplicity. We propose the joint Rashomon and robustness optimization framework for optimal decision trees. It jointly selects an operational decision and the largest Rashomon set of trees, so that the targets hold under the local empirical distribution that every tree in this set induces at the query point. We specialize the framework to the regression setting, and we show that a tree affects the decision only through the training observations sharing the query leaf, which we call its query neighborhood. As a result, the robust problem involves only finitely many distinct constraints, which can be examined in order of increasing estimation loss. We develop a constraint generation algorithm that combines query-path pricing with dynamic programming to identify violating neighborhoods without enumerating trees. On synthetic newsvendor instances, the algorithm typically needs few neighborhoods and runs substantially faster than full neighborhood enumeration. On restaurant demand data, the robust orders increase the mean tolerated excess estimation loss by 17.6% and reduce the empirical conditional value-at-risk of the worst 10% of realized costs by 8.3% relative to the sample average approximation orders of the optimal tree, while the mean cost difference is not statistically significant. An interpretability analysis further shows how the retained neighborhoods explain the decision and its robustness limit.
|
| 1215 |
Gaussian Flow Dynamics: Simulation-Free Neural SDE Learning Beyond One-Time Marginals
2610.04390
|
cs.LG
|
Grigory Bartosh, Christian A. Naesseth |
Simulation-free training of latent Stochastic Differential Equations (SDEs) relies on a variational posterior process whose one-time marginals are tractable, typically Gaussian. Such marginals, however, do not determine the underlying dynamics: many processes ...Simulation-free training of latent Stochastic Differential Equations (SDEs) relies on a variational posterior process whose one-time marginals are tractable, typically Gaussian. Such marginals, however, do not determine the underlying dynamics: many processes share the same marginals while differing in their temporal structure, and existing parameterizations fix this structure implicitly, which restricts the posterior family and biases the learned model. We introduce Gaussian flow dynamics, which construct stochastic processes directly from smoothly evolving Gaussian marginals while making the marginal-preserving, or gauge, degrees of freedom explicit and parameterizable. The construction admits state-dependent diffusion coefficients and recovers every linear SDE with additive noise and a non-degenerate Gaussian initial distribution. Building on it, we propose Gauge Matching, a simulation-free method for latent SDE learning that combines Gaussian flow dynamics with the SDE Matching objective. Gauge Matching costs at most quadratically in the latent dimension per step, like SDE Matching, but learns the temporal structure of the posterior beyond its one-time marginals. It comes within a nat of Helmholtz-SDE, which computes the gauge from the prior Jacobian at cubic cost, on the linear benchmark where the exact posterior is known, matches it on nonlinear systems, and applies where Helmholtz-SDE does not, to state-dependent noise.
|
| 1216 |
JASPER: Special Session on Joint Reliability And Security Assessment of SPlit Computing for Edge Robustness
2610.04396
|
cs.LG
|
Enrico Magliano, Giuseppe Esposito, Amir Hossein Shahdadian, Rama Mounika Kodamanchili, Juan David Guerrero Balaguera |
Split Computing (SC) enables efficient deployment of Deep Neural Networks (DNNs) by partitioning inference between edge devices and cloud servers. However, intermediate feature representations are simultaneously exposed to hardware faults and adversarial attac...Split Computing (SC) enables efficient deployment of Deep Neural Networks (DNNs) by partitioning inference between edge devices and cloud servers. However, intermediate feature representations are simultaneously exposed to hardware faults and adversarial attacks, which are traditionally evaluated independently. This paper presents a unified framework for the joint assessment of reliability and security in Split Computing. First, reliability is characterized through neuron-level fault injection using the Mean Relative Accuracy Degradation (MRAD) while security through feature-map-aware adversarial attacks simulations using the Attack Success Rate (ASR). Based on these complementary analyses, the Joint Vulnerability Score (JVS) is introduced, along with a confidence-aware extension that jointly captures prediction errors and confidence degradation. The framework is evaluated on ten Split Computing configurations based on ResNet-50 trained on ILSVRC-2012. Experimental results show substantial differences across compression strategies, with MRAD ranging from 44.3% to 61.2% under fault injection, while adversarial attacks achieve up to 98.8% ASR. Furthermore, the proposed joint metrics reveal vulnerability trends that remain hidden when reliability and security are analyzed independently, providing a more comprehensive methodology for designing dependable Split Computing systems.
|
| 1217 |
TimeNet: An Extensible Unified Data Infrastructure for Next-Generation Temporal Foundation Models
2610.04407
|
cs.LGcs.AI
|
Martin Maritsch, Timo Stoffregen, Thomas Kaar, Behsad Riemer, Maxwell A. Xu |
Temporal Foundation Models (TFMs) aim to generalize across domains, datasets, and tasks. Yet, their development remains constrained by fragmented, task-specific data formats, annotations, and processing pipelines. We introduce TimeNet, an open-source data stan...Temporal Foundation Models (TFMs) aim to generalize across domains, datasets, and tasks. Yet, their development remains constrained by fragmented, task-specific data formats, annotations, and processing pipelines. We introduce TimeNet, an open-source data standard and scalable infrastructure that decouples temporal data from task definitions and represents signals, metadata, annotations, and supervision in a shared, extensible data model. TimeNet supports multimodal signals with regular, irregular, or ordinal time axes and expresses different task families (including classification, forecasting, temporal localization, question answering, generation, and editing) as reusable views over the same recordings. This shared representation enables heterogeneous time-series datasets to be combined for large-scale model training across domains, modalities, and tasks. We demonstrate TimeNet by transcoding datasets with 1.5M task instances spanning diverse domains, modalities, temporal scales, and forms of supervision, while retaining practical I/O performance relative to native formats. TimeNet enables an existing TFN training pipeline to support joint training on a configurable number of heterogeneous datasets through configuration changes alone. We show this capability by training TFM across multiple datasets and tasks, obtaining a 14% F1 score improvement compared with models trained on individual datasets. These results show that TimeNet provides the data and systems foundation needed to move beyond task- and dataset-specific TFMs toward models that can learn jointly across heterogeneous domains, modalities, temporal scales, and forms of supervision from a common data model.
|
| 1218 |
What Does a Harness Buy? Tokens, Mostly
2610.04433
|
cs.LGcs.AI
|
Yangze Liu, Zhongyi Han |
A coding agent is a language model wrapped in a harness: the system prompt, the tool set, and the context management that turn a chat model into something that can work inside a repository. Production harnesses ship releases daily, vendors advertise pass-rate ...A coding agent is a language model wrapped in a harness: the system prompt, the tool set, and the context management that turn a chat model into something that can work inside a repository. Production harnesses ship releases daily, vendors advertise pass-rate gains from harness changes, and leaderboards mix harnesses freely. What is rarely measured is how much the harness itself moves the score when the model is held fixed. We run five models through three production harnesses, Claude Code, mini-SWE-agent, and OpenCode, on SWE-bench Verified, and rerun the same configurations to calibrate how much a score moves when nothing changes but the run. On 447 tasks and the two models we ran there, Claude Code and mini-SWE-agent, the heaviest and the lightest harness, are equivalent within five points. On a 45-task hard subset and five models, swapping the harness flips as many tasks as rerunning the same harness, 13% in both cases, and the tasks a harness wins in one run are not the tasks it wins in the next. The one harness effect that clears the noise is a loss, not a gain: OpenCode trails by up to 9 points on the large pool, and on one model half of that gap sits in runs its output cap cut short. What the harness does decide is the bill. With the same model, the same tasks, and one price list, cost per task differs by up to 3x across harnesses. The gap is set at the first call, by the preamble of system prompt and tool schemas each harness sends with every step, and scaled by the number of steps; per-step growth and per-call tool output differ far less. The provider's price for cached input scales the bill and does not reorder it. The rerun data also give the resolution a harness comparison needs: at the discordance we observe, 45 tasks catch a 13-point gap only half the time and no gap with 80% power, and 447 tasks resolve 5 points, still coarser than the gains many harness changes claim.
|
| 1219 |
Causally Fair Generation with Large Language Models
2610.04444
|
cs.LGcs.AI
|
Patrik Okanovic, Torsten Hoefler, Drago Plecko |
Large language models (LLMs) are increasingly used to generate, complete, and transform information in settings where their outputs can shape consequential decisions, raising concerns about their impact on demographic disparities. In this context, causal infer...Large language models (LLMs) are increasingly used to generate, complete, and transform information in settings where their outputs can shape consequential decisions, raising concerns about their impact on demographic disparities. In this context, causal inference provides a principled basis for assessing fairness, because it attributes observed disparities to the mechanisms that generated them, which a purely statistical approach cannot do even with infinite data. In LLM generation, a query may request several causally related variables, each of which is both an outcome of interest and a possible cause of other outputs, and the information supplied in the prompt need not follow a topological or a temporal order. This calls for methods that can analyze and selectively remove disparities from such a flexible generation process. In this paper we introduce Causally Fair Generation with LLMs (CFG, for short). CFG extracts relevant concepts, grounds generation in a reference population and causal diagram, and removes user-selected causal effects. CFG also allows pathways deemed justifiable for the task's utility to be retained, which is known in legal literature as business necessity. Further, we provide formal guarantees for our method when eliminating all discriminatory causal effects in the adapted population model, under appropriate causal assumptions. We evaluate CFG with four LLMs in three real-world settings based on population data and on a synthetic dataset with a known causal ground truth.
|
| 1220 |
DreamTest: World-Model Surrogates for Search-Based Testing of Deep Reinforcement Learning Agents
2610.04494
|
cs.LG
|
Qinghua Xu, Guancheng Wang, Boxi Yu, Liting Lin, Lionel Briand |
Testing deep reinforcement learning (DRL) agents in cyber-physical systems aims to uncover diverse failures before deployment, but each execution can be expensive. Surrogate-assisted testing reduces this cost by learning to predict which test configurations ar...Testing deep reinforcement learning (DRL) agents in cyber-physical systems aims to uncover diverse failures before deployment, but each execution can be expensive. Surrogate-assisted testing reduces this cost by learning to predict which test configurations are likely to fail. Prior surrogates treat the system as a black box and predict pass or fail outcomes directly; we instead model how a test unfolds and estimate failure from an imagined episode. We introduce DreamTest, a world-model surrogate for testing DRL agents. DreamTest adapts a recurrent state-space model to learn agent behaviour and environment dynamics from the agent's training log. Given a candidate configuration, imagined rollouts produce a failure score that guides search without executing every candidate in a simulator or real system. We evaluate DreamTest for failure prediction, test generation, and failure diversity on Parking, Humanoid, and DonkeyCar. Mean area under the precision-recall curve (AUPRC) exceeds the strongest baseline by 97%, 12%, and 39%, respectively, and gains on five out-of-distribution test sets reach 145%, 29%, and 44%. Under the same simulator-validation budget, the best "DreamTest + search" combinations find 29%, 22%, and 79% more novel failures on average. Across clusterings with k = 2-40, failures generated with DreamTest cover the most behavioural clusters for almost all k, indicating that DreamTest consistently discovers behaviourally diverse failures.
|
| 1221 |
ManifoldCache: Training-Free Diffusion Acceleration via Constraint Manifold Caching
2610.04510
|
cs.LGcs.AI
|
Prashant Pandey, Devineni Sri Venkatraya Chowdary, Brejesh Lall |
Diffusion models for structured scientific generation must produce samples satisfying hard geometric constraints imposed by physics, chemistry, or biology, yet inference in these settings is prohibitively slow, demanding hundreds to thousands of neural-functio...Diffusion models for structured scientific generation must produce samples satisfying hard geometric constraints imposed by physics, chemistry, or biology, yet inference in these settings is prohibitively slow, demanding hundreds to thousands of neural-function evaluations per sample. We unify eight state-of-the-art models spanning medical volumetrics, molecular conformations, protein backbone design, crystal structure prediction, and multi-view 3D scenes under a single abstraction, Constraint-Manifold Diffusion Models (CMDMs), in which the target distribution is supported on a manifold defined by an externally specified constraint map. All existing acceleration families fail on this class: quantization exhausts memory on high-dimensional volumetric operators; pruning breaks constraint fidelity; fast ODE solvers allow trajectories to drift off the constraint manifold; and feature-caching heuristics are blind to constraint geometry, inducing mode confusion in the high-noise regime. We introduce ManifoldCache, the first training-free, data-free accelerator designed from first principles for CMDMs. The key insight is that the conditional score decomposes orthogonally into a normal component, which enforces constraint satisfaction, and a tangential component, which navigates within the manifold. Exploiting this structure, we prove that the noise-schedule midpoint is a sharp safe-caching boundary: caching before it incurs provably bounded error, while caching after it guarantees a strictly positive fraction of trajectories suffer mode confusion, a gap that persists up to the boundary. We further prove that deeper network blocks admit provably larger certified cache strides within the safe phase, as a consequence of the score decomposition propagating through block Jacobians. The resulting schedule requires no calibration data, along with zero training overhead.
|
| 1222 |
EvoCast: Reliable Autonomous Research Agents for Iterative Forecasting Architecture Evolution
2610.04517
|
cs.LGcs.AI
|
Kaipeng Xu, Xianli Yan, Yan Wang, Xiang Liu, Shan Liu |
Deep time-series forecasting models have rapidly diversified, yet adapting them to a specific task still requires extensive expert effort in model selection, mechanism diagnosis, architecture design, implementation, and evaluation. Existing AutoML methods are ...Deep time-series forecasting models have rapidly diversified, yet adapting them to a specific task still requires extensive expert effort in model selection, mechanism diagnosis, architecture design, implementation, and evaluation. Existing AutoML methods are constrained by predefined search spaces, while general-purpose LLM research agents lack reliable control over experimental protocols and model promotion. We introduce EvoCast, a fully autonomous research-agent system for iterative forecasting architecture evolution. EvoCast first establishes and diagnoses a task-specific baseline through executed mechanism ablations, then generates evidence-grounded research directions from dataset characteristics, diagnostic results, prior rounds, and failure records. Its central design, cognition-authority separation, assigns open-ended hypothesis generation and code implementation to LLM agents, while deterministic program authorities control source-edit boundaries, canonical evaluation, and promotion decisions. Experimental outcomes are accumulated as evidence to guide subsequent rounds. Results show that EvoCast completes complex architecture modifications with higher implementation success and lower agent-side token/time cost, and develops task-specific architectures that outperform selected baselines, strong forecasting models, and agent baselines in three real-world forecasting cases. The code is available at https://github.com/18e0-x/EvoCast.
|
| 1223 |
All against the machine: the Solo score for rating skill in variable environments
2610.04523
|
cs.LG
|
David Reguera, Xavier R. Hoffmann, Irene P\'erez, Pol Colomer-de-Sim\'on, Miquel Masoliver |
We propose a distribution-free metric to rate individual skill in ``player-versus-environment'' settings, where participants face heterogeneous tasks without direct opponents. Such settings are common in digital platforms, games, education, finance, and the be...We propose a distribution-free metric to rate individual skill in ``player-versus-environment'' settings, where participants face heterogeneous tasks without direct opponents. Such settings are common in digital platforms, games, education, finance, and the benchmark evaluation of AI agents. They combine high randomness, tasks of widely varying difficulty, and unknown heterogeneity across individuals. Our metric maps each task outcome to a bounded performance score with zero population mean and variance bounded by $1/3$, regardless of the outcome distribution, so that scores are directly comparable across tasks. Aggregating these scores over tasks yields an interpretable skill score for each individual (the ``Solo'' rating), and a reshuffling null model tests whether that score exceeds what chance alone would produce. We validate the approach on the progression of $2\times10^5$ players in the mobile games Candy Crush Saga and Bubble Witch 3 Saga. The metric identifies high- and low-skill players with high statistical confidence, their classification persists over hundreds of subsequent levels, and a windowed version of the score tracks changes in performance along progression.
|
| 1224 |
Diffusion-Based Stress Testing of Overload Monitoring for Resilient Emergency Cellular Networks Using Internet CDR Proxies
2610.04526
|
cs.LGcs.AI
|
Bilal Hussain, Xiao Tang, Tan Li, Muhammad Azhar, Danista Khan |
Disasters can overload cellular control-plane signaling within minutes, yet fine-grained Radio Resource Control (RRC) or Next Generation (NG) Application Protocol (NGAP) telemetry is privacy-sensitive and costly to collect for analytics. Many emergency monitor...Disasters can overload cellular control-plane signaling within minutes, yet fine-grained Radio Resource Control (RRC) or Next Generation (NG) Application Protocol (NGAP) telemetry is privacy-sensitive and costly to collect for analytics. Many emergency monitoring pipelines therefore rely on coarse Call Detail Record (CDR) aggregates. We treat Internet activity in CDR grids as a practical proxy for hidden signaling stress under that constraint. We train a lightweight convolutional neural network (CNN) on stylized overload injections, stress-test it with diffusion-synthesized surges that preserve normal traffic structure, and adapt the detector by retraining on hard synthetic samples. Under stress-test conditions, the default alert threshold fails even though receiver operating characteristic (ROC) curves stay strong: the detector still assigns overloaded cells a larger overload probability than normal cells, but those probabilities fall below the default cutoff 0.5 and are labeled normal, so the F1-maximizing threshold -- selected post hoc on the same stress-test grids (oracle $\tau^*$) -- shifts by $0.32 \pm 0.03$ (operating-point drift). Across three random seeds, hard-sample adaptation raises thresholded performance (F1) from 0% (no alerts at the default cutoff 0.5 on any seed) to $85.67 \pm 14.37$% and ranking from ROC-AUC $0.886 \pm 0.040$ to $0.99996 \pm 0.00007$. Diffusion-synthesized surges expose threshold fragility that matched-condition training -- training and testing on the same stylized injections -- hides, and hard-sample adaptation restores usable alerts at the default cutoff. Together, these steps define a reusable pre-deployment stress test for emergency monitors. Internet-only CDR input further supports lightweight AI-native workflows that combine monitoring, recalibration, and adaptation.
|
| 1225 |
Quantum Machine Learning Protection of Military Quantum Key Distribution Against Cryptographically Camouflaged Attacks
2610.04543
|
cs.LG
|
Muhammad Shaheer Bin Junaid |
Quantum key distribution proves its protocol secure and says nothing about the hardware beneath it, so military and government operators fielding it for command-and-control keys monitor the channel for implementation attacks, and that monitoring has a blind sp...Quantum key distribution proves its protocol secure and says nothing about the hardware beneath it, so military and government operators fielding it for command-and-control keys monitor the channel for implementation attacks, and that monitoring has a blind spot. An adversary with a kleptographic foothold in the generator of a public per-block value \(x=g^v \pmod p\) can hide attacked blocks in honest noise, gating them on a predicate of its discrete logarithm, making detection a discrete logarithm problem that defeats every efficient classical monitor yet yields to a quantum kernel recovering \(v\) through Shor's algorithm. I formalise these cryptographically camouflaged attacks, reduce their hardness to an established learning separation, prove a single-frequency fidelity kernel cannot represent an interval predicate, and test them on Ghillie, a decoy-state BB84 simulator with a positive key rate to 142 km. From 10- to 14-bit groups over two seeds, a classical monitor reads 0.458 to 0.516 on camouflaged attacks while the quantum kernel reads 1.000, and both catch overt attacks above 0.99. Finite-precision recovery under depolarising noise and a hardened predicate lower the quantum result to 0.916 through 0.983 with the classical monitor at chance, and a feasibility probe on IBM Heron processors tracks the exact kernel within 0.034. A defender can therefore discard precisely the compromised key material, although the advantage is asymptotic, awaits fault tolerance, and holds only when the feature map matches the adversary's predicate, since a low-frequency map reads 0.545 on a residue pattern and 0.982 once aligned.
|
| 1226 |
Label Agreement Does Not Measure Authorization
2610.04544
|
cs.LGcs.AI
|
Amir Sabbaghziarani, Bradley Thomas Baker, Theodore J. LaGrow, Sergey Plis |
Many groups now delegate label ontology and metadata harmonization to agentic LLM pipelines. We built one and audited it. Our aggregate scores looked healthy, but the pipeline kept failing in ways they did not show, so we set out to find what they hid. Label a...Many groups now delegate label ontology and metadata harmonization to agentic LLM pipelines. We built one and audited it. Our aggregate scores looked healthy, but the pipeline kept failing in ways they did not show, so we set out to find what they hid. Label agreement asks whether a proposed label matches a reference. It does not ask whether the agent was entitled to propose it, whether the output was complete enough to act on, or whether the label moved when the evidence moved. We measured those three separately on COBRE and FBIRN, two schizophrenia and control neuroimaging cohorts from different consortia, and they come apart, from label agreement and from each other. Showing the agent an upstream proposal barely moves label agreement, 0.857 to 0.870, while agreement on the chosen action doubles, 0.409 to 0.830. Output that parses as JSON still drops a required field on 10% of one model's cases and 33% of the other's. And an agent that replays its first answer scores perfectly on original cases and zero once we change the evidence that decides them. Downstream, a row-order error that none of these metrics reports erases most of the diagnostic signal. So we measure these properties apart, pair each with a control, and gate commitment on the result, which makes failures visible and easy to route to a person. None of this prevents failure. Our reference labels are rule-derived, so agreement with them means consistency, not correctness. Code is available at https://github.com/amir-sbg/Label-Agreement-Does-Not-Measure-Authorization.
|
| 1227 |
ClimateBench v2.0: Probabilistic Climate Model Benchmarking
2610.04558
|
cs.LG
|
Duncan Watson-Parris, Willa Tobin, Ayta\c{c} Pa\c{c}al, Manuel Schlund, V. Balaji |
We present ClimateBench v2, a standardized protocol for evaluating climate models on diagnostics expected to be informative for their skill in projecting mid-century regional temperature and precipitation changes. The protocol is designed to evaluate any physi...We present ClimateBench v2, a standardized protocol for evaluating climate models on diagnostics expected to be informative for their skill in projecting mid-century regional temperature and precipitation changes. The protocol is designed to evaluate any physics-based, data-driven, or hybrid climate model on equal footing using a common set of observational and out-of-distribution tests. We define three tiers of evaluation. Tier I establishes physical credibility through entry-ticket tests of energy conservation, coupled (co-)variability, and basic forced responses. Tier II scores models against post-2015 observations of surface temperature, precipitation, radiative fluxes, sea ice, and key modes of variability using fair CRPS as the primary probabilistic score, complemented by distributional and ensemble-consistency diagnostics. Tier III tests out-of-distribution generalization through paleoclimate simulations spanning the Last Interglacial, Last Glacial Maximum, and Mid-Holocene, and through perfect-model experiments in which data-driven models must predict the future climate of existing Earth system models from historical data alone. We reserve all observational data after 2015 for testing, and submissions must include multiple ensemble members to enable probabilistic evaluation. This reservation exploits a new opportunity provided by the decade of observations accumulated since the end of the CMIP6 historical experiment, which constitutes an out-of-sample record of forced climate change (and internal variability) for the current generation of models, and we quantify, in an idealized setting, the information it carries about mid-century warming. We provide the evaluation code, observational reference datasets, and perfect-model training data as an open benchmark to drive measurable progress in climate projection across all modeling approaches.
|
| 1228 |
Gradient-Free Sampling from Generative Models via Stochastic Bounded Extremum Seeking
2610.04568
|
cs.LG
|
Alexander Scheinker |
We introduce a sampling approach for energy- and score-based generative models that requires no gradient evaluations of the model. Replacing the drift term that would normally contain the score $\nabla_\mathbf{x} \log p_\theta(\bf{x})$ with a high-frequency di...We introduce a sampling approach for energy- and score-based generative models that requires no gradient evaluations of the model. Replacing the drift term that would normally contain the score $\nabla_\mathbf{x} \log p_\theta(\bf{x})$ with a high-frequency dithered cosine of the model's \textit{value}, $\sqrt{\alpha\omega}\,\cos(\omega t + k \log p_\theta(\bf{x}))$, produces, in the high-frequency averaging limit, Langevin Markov chain Monte Carlo for energy-based models and the reverse-time SDE of score-based diffusion. We prove that trajectories of the dithered It\^{o} SDE converge to those of the target SDE, driven by the same Brownian motion, uniformly on compact time intervals in probability, by an averaging argument that extends bounded extremum seeking (ES) to It\^{o} processes, with an explicit $O(\omega^{-1/2})$ mean-square rate under global bounds. The approach is not confined to smooth targets: it extends to $C^{1,1}$ energies with discontinuous curvature (without ellipticity requirement) and to Sobolev energies whose Hessians exist only off measure zero sets; for Lipschitz energies with gradient kinks the averaged limit remains well posed; the Krylov-R\"ockner integrability class is the boundary of provability. The approach provides a hard \textit{a priori} bound on the per-step update rate and applies to explicitly time-varying targets on finite horizons. Gradient-free pixel-space sampling is not competitive with well-tuned backpropagation-based samplers at practical evaluation budgets; the regime where the approach offers an advantage is latent-space sampling when the model is a black box and the target drifts in time. We demonstrate latent-space tracking for time-varying images on CelebA-HQ ($256{\times}256$) from limited 1D projection measurements and latent-space EBM sampling on CIFAR-10.
|
| 1229 |
From Transformers to Weighted Automata: Towards the Verification of Large Language Models
2610.04569
|
cs.LG
|
Smayan Agarwal, Aslah Ahmad Faizi, Shobhit Singh, Aalok Thakkar |
Large language models (LLMs) are increasingly deployed in safety-critical settings, yet their black-box nature makes it difficult to provide formal guaranties about their behavior. Existing verification approaches rely primarily on empirical probing and testin...Large language models (LLMs) are increasingly deployed in safety-critical settings, yet their black-box nature makes it difficult to provide formal guaranties about their behavior. Existing verification approaches rely primarily on empirical probing and testing, leaving open the question of how to reason rigorously about general-purpose trans- former architectures. In this work, we establish a principled bridge between transformers and weighted automata, a classical model from formal language theory. This connection enables us to transfer verification tools from automata the- ory to the analysis of LLMs. Our contributions are twofold: First, we develop a formal correspondence between transformer architectures and weighted automata over reals, showing how distributional properties of LLMs can be captured within this framework. Second, we introduce an identity testing algorithm for weighted automata that provides a statis- tical method for distinguishing whether two stochastic models define the same distribution up to a tolerance threshold. This work provides the first formal bridge between modern neural se- quence models and classical automata theory, clarifying both the poten- tial and the computational challenges for rigorous LLM verification.
|
| 1230 |
Variational Quantum Attention for Molecular Graph Learning
2610.04588
|
cs.LGcs.AI
|
Yu-Cheng Lin, Yu-Chao Hsu, Tai-Yue Li, Nan-Yow Chen, Samuel Yen-Chi Chen |
Molecular property prediction is central to computational drug discovery, where graph neural networks learn to weight neighboring atomic environments during message passing. Yet it remains unclear how variational quantum circuits alter learned attention behavi...Molecular property prediction is central to computational drug discovery, where graph neural networks learn to weight neighboring atomic environments during message passing. Yet it remains unclear how variational quantum circuits alter learned attention behavior in molecular graphs. We introduce an edge-aware variational quantum attention mechanism for molecular graph learning, in which the receiving atom, neighboring atom, and connecting bond jointly determine the quantum attention state. Across five molecular property and bioactivity prediction tasks, QGAT achieves competitive performance relative to GATv2, with a consistent improvement on BBBP across all six evaluated circuit ansatzes. We further compare how the quantum and classical attention scores weight molecular structure beyond accuracy. In the Verubecestat BACE1 inhibitor series, QGAT achieves a higher Spearman correlation than GATv2 and assigns positive attributions to several structural changes consistent with reported structure-activity relationships (SARs). This case study shows that the two attention mechanisms can exhibit different prediction and attribution behavior across structurally related BACE1 analogues, while broader validation is required to determine how consistently these differences generalize across chemical series and targets. Circuit ablations further show that performance depends on the circuit design. Together, these results show that variational quantum attention can serve as a viable alternative molecular attention parameterization while inducing circuit- and chemistry-dependent behavior distinct from a matched classical scorer.
|
| 1231 |
Exploiting Hierarchical Controller Structure in Contextual Parameter Learning for Humanoid Loco-Manipulation
2610.04609
|
cs.LG
|
Sebastian Hirt, Lukas Theiner, Jan Peters, Rolf Findeisen |
Hierarchical control architectures are widely used to decompose complex control problems into interacting control levels and are particularly important in robotics, where planning, whole-body motion, and lower-level control must be coordinated across different...Hierarchical control architectures are widely used to decompose complex control problems into interacting control levels and are particularly important in robotics, where planning, whole-body motion, and lower-level control must be coordinated across different levels of abstraction and time scales. Their overall closed-loop performance, however, depends strongly on parameters distributed across the hierarchy, such that tuning controllers on different levels independently may neglect relevant cross-layer interactions. We propose a contextual Bayesian optimization framework for joint parameter learning in hierarchical control systems. Rather than modeling closed-loop performance only as a scalar black-box function, we retain separate observations of task performance, realization quality, and control effort. A correlated multi-output Gaussian process models these performance components, while their known aggregation into the overall closed-loop objective is evaluated analytically. The formulation exploits three complementary consequences of hierarchical control: informative performance quantities exposed by the hierarchy, coupling between parameters of different controller levels, and variations of these relations with operating conditions. We evaluate the approach for humanoid loco-manipulation, jointly tuning a centroidal predictive controller and a whole-body controller for physical box pushing under varying box mass. The proposed method achieves the lowest mean empirical regret during both training and adaptation among the considered baselines.
|
| 1232 |
Hypergraph Representation Learning with Hyperlink Random Effects
2610.04640
|
cs.LG
|
Zimeng Li, Shihao Wu, Gongjun Xu, Ji Zhu |
Hypergraphs record multi-way interactions among entities. Extracting information from the combinatorial structure underlying observed multi-way interactions is a central task in many real-world problems. Existing methods face several limitations. First, many d...Hypergraphs record multi-way interactions among entities. Extracting information from the combinatorial structure underlying observed multi-way interactions is a central task in many real-world problems. Existing methods face several limitations. First, many deep architectures for hypergraphs do not explicitly exploit the potential low-rank structure, which can sacrifice parsimony and interpretability in the learned representations. Second, many low-rank-based methods operate on tensor representations, which typically require hyperlinks to have uniform sizes and thus limit their applicability to general hypergraphs with non-uniform hyperlink sizes. Third, many methods ignore the fact that hyperlinks often arise from heterogeneous mechanisms. For example, medical symptoms may co-occur in the profiles of patients with very different conditions, and such heterogeneity should be incorporated into the learning process. In this work, we develop a general framework for hypergraph representation learning using hyperlink random effects while exploiting the low-rank structure in hypergraphs. The proposed framework accommodates latent heterogeneity in hyperlink formation while preserving entity interaction patterns. We establish identifiability of the model parameters and theoretical guarantees of representation-level recovery under this framework. The framework allows flexible specifications for the hyperlink random effects; in this paper, we study three choices: categorical, Gaussian mixture, and score-based effects, and develop corresponding estimation algorithms. Through simulation studies, we demonstrate the effectiveness of the proposed method in recovering latent structure and capturing heterogeneous interaction patterns. Empirical studies on real-world hypergraph datasets further illustrate the practical utility of our approach.
|
| 1233 |
GS-Codec: A Gaussian-Splatting Bottleneck for Neural Audio Coding
2610.04651
|
cs.LGcs.AIcs.SDeess.AS
|
Ron Aluf, Alon Canfi, Eliya Nachmani |
Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, ...Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER). Code and audio samples are available at https://ronaluf.github.io/gs-codec/
|
| 1234 |
Asking the Crowd the Right Question: Bias-Cancelling Weights for Federated Learning
2610.04671
|
cs.LG
|
Ilya Kuruzov, Dmitrii Vishovan, Kirill Novoselov, Yuriy Dorn, Darina Dvinskikh |
A federated objective is a weighted sum of client risks, and the weights are almost always fixed in advance. We treat them instead as the only instrument of a wisdom-of-crowds mechanism: clients are noisy views of one truth, each seeing it through an independe...A federated objective is a weighted sum of client risks, and the weights are almost always fixed in advance. We treat them instead as the only instrument of a wisdom-of-crowds mechanism: clients are noisy views of one truth, each seeing it through an independent distortion that is unbiased across the crowd. That the optimal weights are inversely proportional to the clients' error energies is classical; we begin at the question that answer presupposes, which energies belong there and whether a crowd can recover them from itself. Excess risk on the truth is of the exact order of the aggregate bias energy, so no optimizer can repair a bad weight vector; the truth itself is identifiable only up to a linear tilt, so subtracting estimated client biases provably reproduces uniform weighting. The expected per-client second moments, however, are exactly identified from the law of the crowd's disagreement by a well-conditioned linear inversion, a step random-effects meta-analysis cannot take because a source reports once; their realized counterparts are estimable up to an incoherence floor the algorithm can measure. This yields CROWD, which reads the disagreement off the optimization trajectory at no extra cost and matches a Bayesian minimax lower bound in the same constant: per instance as the horizon grows, and unconditionally as the prior becomes diffuse. For arbitrary distortions it stays competitive with the optimal weights, at a ratio governed by a geometric incoherence the algorithm can measure. On real scans split into sites with their own miscalibrated detectors it attains the oracle excess risk; on a companion federation that pulls bias and noise apart, weighting by noise variance is worse than not weighting at all, and CROWD is not.
|
| 1235 |
GPU-Accelerated Bregman Douglas-Rachford Splitting for Discrete Optimal Transport
2610.04715
|
cs.LG
|
Yifan Xu, Shiqian Ma |
We present GPU-accelerated Bregman Douglas--Rachford splitting algorithm (BDRS) for discrete optimal transport problem in three input formats: an explicit cost matrix, a point cloud with a ground cost between them, and a separable cost on a regular grid. For e...We present GPU-accelerated Bregman Douglas--Rachford splitting algorithm (BDRS) for discrete optimal transport problem in three input formats: an explicit cost matrix, a point cloud with a ground cost between them, and a separable cost on a regular grid. For each input format, we propose hardware-aware designs of mathematically equivalent representations for the BDRS iterations to enhance numerical stability and empirical runtime. We benchmark the three proposed implementations against eight GPU baseline solvers from the literature on the same device. We demonstrate that our implementations of BDRS achieve state-of-the-art performance on their respective input formats. To the best of our knowledge, this is the first cross-solver study of GPU DOT solvers with a unified measure of optimality.
|
| 1236 |
Latent-Lagrangian Neural Networks for Reduced Order Modeling of Non-autonomous Nonlinear Dynamical Systems
2610.04723
|
cs.LG
|
Anand Kumar Agrawal, Anders Thorin |
This work proposes a latent Lagrangian-based framework for reduced-order modelling of forced nonlinear dynamical systems. In contrast with conventional Lagrangian or Hamiltonian neural networks, our approach learns a set of latent coordinates sufficient to cap...This work proposes a latent Lagrangian-based framework for reduced-order modelling of forced nonlinear dynamical systems. In contrast with conventional Lagrangian or Hamiltonian neural networks, our approach learns a set of latent coordinates sufficient to capture the dynamics conjointly with two neural networks for the latent kinetic and latent potential energies, and leverages force supervision to eliminate the need for an ODE solver during training. Consistency of physical laws in the latent space is ensured through the principle of virtual work. Results show that the model effectively learns the subtle dynamics induced by the system's nonlinearity and non-convex potential energy, and generalizes well to unseen forces and initial conditions. These observations confirm the physical relevance of the proposed approach, and its interest for model reduction.
|
| 1237 |
PyINE: A Framework for Scalable Elicitation and Oversight via Code Execution
2610.04737
|
cs.LGcs.AI
|
Pierre-Luc St-Charles, Alessandro Palmas, Damiano Fornasiere, Storm Lei, Mirko Bronzi |
Reasoning models can remain capable of solving a task while still defaulting to cheaper but misleading shortcuts. This creates a central oversight problem: when a model gives an answer with plausible but incomplete reasoning, can an overseer determine whether ...Reasoning models can remain capable of solving a task while still defaulting to cheaper but misleading shortcuts. This creates a central oversight problem: when a model gives an answer with plausible but incomplete reasoning, can an overseer determine whether that output should be trusted? To study this problem, we introduce PyINE, a framework for scalable elicitation and oversight using instrumented Python programs as a verifiable execution substrate. In PyINE, programs define task environments, execution traces provide authoritative labels for outcomes and intermediate facts, and task variants can be generated mechanically rather than through static human annotation. We instantiate the framework in PyINE-v1, a first release built from nearly one million deterministic execution traces and over 500,000 matched LLM-generated code variants used for counterfactual evaluation. Using standard RL with verifiable rewards on cue-varied tasks, we train a shortcut-following model that improves substantially at predicting execution outcomes while still making systematic errors when misleading human-facing cues conflict with the program's realized behavior. We then evaluate activation probes, trained text classifiers, prompted judges, and a lightweight debate protocol as overseers of this model. We find that performance pooled at the dataset level can hide weak coverage of the failures that matter most: cheap learned overseers often miss rare shortcut-driven errors, while stronger model-based checks are more balanced but substantially costlier and harder to turn into reliable thresholded decisions. PyINE-v1 turns this failure-mode coverage problem into a reusable experimental setting for developing oversight methods that are verifiable, failure-mode-aware, and cost-sensitive.
|
| 1238 |
Exact Fast Batch Simulation for Tabular Reinforcement Learning
2610.04746
|
cs.LG
|
Haochen Zhang, Lingzhou Xue, Zhong Zheng |
Simulation is a fundamental computational primitive in reinforcement learning (RL), yet conventional simulation explicitly generates individual trajectories even when downstream procedures use only aggregate statistics. To address this, we develop an exact fas...Simulation is a fundamental computational primitive in reinforcement learning (RL), yet conventional simulation explicitly generates individual trajectories even when downstream procedures use only aggregate statistics. To address this, we develop an exact fast-simulation framework for finite-horizon tabular Markov decision processes. Our framework has two complementary modes. In direct batch simulation, a batch is represented by its aggregate Markov flow. With sufficient parallel simulation resources, this flow can be obtained by trajectory aggregation; when such simulation is unavailable or costly but the initial state and transition distributions are directly accessible, we instead generate an identically distributed flow through forward Markov-flow sampling without materializing individual trajectories. The latter reduces the simulator-side computational dependence on batch size $m$ from $O(m)$ to $O(1)$. In adaptive batch simulation, when batch length is determined by a data-dependent condition, exact multivariate-hypergeometric splitting recursively refines a candidate Markov flow while preserving the conditional law, reducing the cost dependence on $m$ from $O(m)$ to $O(\log m)$. Together, these modes accelerate simulation by keeping trajectories aggregated whenever possible and refining flows only when required to locate data-dependent boundaries. The framework applies broadly across simulator-based, offline, and online batch or stage-based RL, as illustrated with representative algorithms from each setting.
|
| 1239 |
Variance-Aware Fine-Grained Gap-Dependent Bounds for Online Reinforcement Learning
2610.04752
|
cs.LG
|
Haochen Zhang, Lingzhou Xue, Zhong Zheng |
We study model-free online reinforcement learning (RL) for episodic tabular Markov decision processes, focusing on both gap-dependent regret and policy switching cost. While fine-grained gap-dependent analysis has been established for model-free RL algorithms ...We study model-free online reinforcement learning (RL) for episodic tabular Markov decision processes, focusing on both gap-dependent regret and policy switching cost. While fine-grained gap-dependent analysis has been established for model-free RL algorithms using Hoeffding-type exploration bonuses, such results for model-free algorithms with variance-based exploration bonuses remain unknown, despite their superior worst-case and coarse-grained gap-dependent guarantees. In this paper, we resolve this open problem by establishing the first fine-grained gap-dependent regret upper bound for UCB-Bernstein+, a refined UCB-Bernstein algorithm, in variance-aware model-free online RL. Moreover, by integrating a stage-wise policy update design into our fine-grained framework and using refined variance-based bonuses, we achieve the best-known gap-dependent local switching cost to date. In addition, our analysis yields improved worst-case guarantees for both regret and local switching cost over the original UCB-Bernstein algorithm. Numerical experiments further demonstrate that UCB-Bernstein+ achieves favorable empirical performance in both regret and local switching cost.
|
| 1240 |
When Is Enough Enough in Self-Evolving LLM Systems?
2610.04756
|
cs.LGcs.AI
|
Enoch Yin, Bin Liu, Zhengling Qi |
Self-evolving large language model (LLM) systems repeatedly propose, evaluate, and incorporate updates to prompts, skills, or other persistent artifacts. Despite their growing effectiveness, these systems typically operate under a predetermined iteration or co...Self-evolving large language model (LLM) systems repeatedly propose, evaluate, and incorporate updates to prompts, skills, or other persistent artifacts. Despite their growing effectiveness, these systems typically operate under a predetermined iteration or compute budget, without a principled criterion to determine when further evolution is no longer worthwhile. This can lead to two undesirable consequences: unnecessary computation after performance has saturated and the risk of returning late updates that overfit or exploit the evaluation signal. These issues motivate us to study two fundamental questions: when should a self-evolving system stop, and what should it output once it stops? We address the first by formulating an online sequential testing problem and constructing an anytime-valid restart detector using the per-item paired evaluation outcomes already produced by self-evolving LLM systems. We address the second by formulating a change-point estimation problem and using the estimated transition to select an earlier artifact for output. The resulting procedure is plug-and-play and requires no modification of the underlying self-evolving algorithms. Across two self-evolving frameworks, three LLM model families, and five benchmarks, our method substantially reduces computation costs while maintaining comparable unseen-test performance. For example, on SearchQA with SkillOpt and DeepSeek V4 Flash, our method stops at round 4 rather than the full budget of 40, reducing token usage by 91.6% while achieving 82.43% unseen-test accuracy versus 82.00% under the full-budget run.
|
| 1241 |
Flow Policies as Actions of Skill-Level World Models: Learned and Symbolic Abstractions for Long-Horizon Planning
2610.04767
|
cs.LG
|
Andreu Matoses Gimenez, Andrei-Carlo Papuc, Chris Pek, Javier Alonso-Mora |
Latent world models enable robots to plan by predicting the consequences of actions. Planning long tasks with control-rate actions requires many prediction steps, which enlarges the search space and accumulates error. Skill-level actions shorten these sequence...Latent world models enable robots to plan by predicting the consequences of actions. Planning long tasks with control-rate actions requires many prediction steps, which enlarges the search space and accumulates error. Skill-level actions shorten these sequences, but a symbolic skill vocabulary requires domain knowledge and labeled demonstrations. We construct skill-level actions from the inputs of a flow-matching policy trained on demonstrations segmented into complete skills. The policy maps a noise seed and an observation, optionally with a code or label, to a complete skill execution, so one execution is one world-model transition. On this mechanism we propose four action abstractions with increasing task knowledge: a compressed seed, two discrete codes learned from the demonstrations, and a symbolic label. We evaluate them with a common world-model training procedure and planning framework on simulated block rearrangement tasks that require up to 14 sequential skills. The symbolic label succeeds in over 90% of the tasks that require up to six skills and degrades beyond. Without any label, an object-centric learned code matches it on single-skill tasks and retains half to three quarters of its success on tasks of two to five skills. Ablations attribute much of the label's advantage to its planner knowing which actions are applicable, rather than to the label itself. Beyond six skills the search, not the world model, limits success. Project page: https://andreumatoses.github.io/research/flow-skill-wm
|
| 1242 |
Learning Task and Motion Plans from Real Demonstrations with Hybrid Flow Matching
2610.04771
|
cs.LG
|
Zuleika Redondo Garcia, Andreu Matoses Gimenez, Javier Alonso-Mora |
Long-horizon mobile manipulation requires a task plan and the motion that executes it. Generative planners trained on demonstration produce both in one pass, requiring neither a symbolic domain nor search. To date, however, they have relied on thousands of scr...Long-horizon mobile manipulation requires a task plan and the motion that executes it. Generative planners trained on demonstration produce both in one pass, requiring neither a symbolic domain nor search. To date, however, they have relied on thousands of scripted demonstrations of fixed-base arms and executed open loop. This paper presents a hybrid flow matching planner: a single network generates the symbolic plan with masked discrete flow matching and the motion trajectory with continuous flow matching. Unlike prior generative planners, we aim to learn from a much smaller set of demonstrations and to execute the plan in closed loop. Two properties of the data compensate for the small dataset. A demonstration resumed from any of its intermediate actions is itself a demonstration, which multiplies the training samples and enables replanning after every action. Objects of the same kind are interchangeable, which turns demonstrations of one goal into demonstrations of every permuted goal. Our base implementation produces valid plans on 68% of held-out scenes; a training and generation scheme for the discrete plan raises this to 76%, and replanning after every action raises the task completion rate from 40% to 53% in a kinematic simulation. The planner matches the task completion rate of motion-only flow matching policies while additionally providing the symbolic plan, and it outperforms previous hybrid diffusion formulations on both task completion and plan validity. We validate the planner on a real mobile manipulator. Videos and project page: https://andreumatoses.github.io/research/hybrid-flow-planning
|
| 1243 |
RepTC: Representation-Aware Optimization for Efficient Traffic Classification on Edge IoT Devices
2610.04784
|
cs.LG
|
Adel Chehade, Edoardo Ragusa, Paolo Gastaldo, Rodolfo Zunino |
Traffic classification (TC) is crucial to secure Internet of Things (IoT) networks, whose edge nodes often operate under privacy, bandwidth, and energy constraints. Yet, encrypted payloads and limited computing power make accurate, real-time TC a challenging t...Traffic classification (TC) is crucial to secure Internet of Things (IoT) networks, whose edge nodes often operate under privacy, bandwidth, and energy constraints. Yet, encrypted payloads and limited computing power make accurate, real-time TC a challenging task. Existing learning-based TC approaches often fix the input configuration a priori, even though it directly influences both predictive performance and computational cost. This paper presents RepTC, a representation-aware hardware-constrained strategy that addresses session-level TC through joint model and input optimization. The method co-optimizes network architecture, session length, and header preprocessing; this enables joint control of model complexity, input scale, and data representation within a unified resource-constrained design space. The proposed approach enforces microcontroller-class constraints on memory, model size, and computation, and yields compact models deployable on low-power edge devices. A gateway monitors traffic by aggregating sessions and either performs inference locally or offloads to a low-power edge node. Both scenarios are validated on heterogeneous embedded hardware, including a Raspberry Pi 3B+, STM32 Nucleo-F401RE, and XIAO ESP32-C3, spanning Cortex-A, Cortex-M, and RISC-V processor architectures; measured inference latency ranges from 0.63 to 18.59 ms, with MCU-side inference energy between 0.50 and 1.41 mJ per session. RepTC yields configurations that achieve high accuracy on a variety of established benchmarks: 96.21% on ISCX VPN-nonVPN, 99.62% on USTC-TFC2016, and 99.97% on Edge-IIoTset; at the same time, model size and computational requirements were reduced by up to three orders of magnitude compared with state-of-the-art methods. The results show that representation-aware optimization can improve efficiency while preserving competitive classification performance.
|
| 1244 |
AID: A Framework for AI Infrastructure Dynamics
2610.04801
|
cs.LG
|
Abi Aryan |
A useful model of AI inference infrastructure must specify the system state, the information available to an observer, and the decisions the model is intended to support. We introduce AID (AI Infrastructure Dynamics), a framework for describing this learning p...A useful model of AI inference infrastructure must specify the system state, the information available to an observer, and the decisions the model is intended to support. We introduce AID (AI Infrastructure Dynamics), a framework for describing this learning problem across coupled physical, computational, networking, and serving processes. The formulation allows structured and variable-size state, asynchronous observations, multiple physical timescales, and demand that responds to service. We distinguish representations that support prediction under an existing policy from those that preserve service outcomes under changed actions, and separate both from identifying intervention responses. Two analytical results describe a lower bound on prediction error when available observations cannot distinguish models and a sufficient condition for exact controlled state reduction. These results apply established information and state-abstraction principles to AI infrastructure. We then describe a validation protocol for cache representations, workload histories, measurement availability, and imposed actions.
|
| 1245 |
GitSwarm: Decentralized Compounding Inference
2610.04862
|
cs.LGcs.AI
|
Vedant Shah, Ankur Samanta, Paras Dahal, Mikhail Plekhanov, Carole-Jean Wu |
Long-horizon problem solving and scientific research require computation to accumulate across successive attempts. Partial solutions, experimental findings, and unsuccessful approaches can inform later work, yet most inference-time computation is organized aro...Long-horizon problem solving and scientific research require computation to accumulate across successive attempts. Partial solutions, experimental findings, and unsuccessful approaches can inform later work, yet most inference-time computation is organized around individual trajectories or candidates rather than a persistent body of reusable work. We call this paradigm compounding inference: organizing inference-time computation so that intermediate work persists and can be inspected, extended, combined, or challenged by subsequent computation. We instantiate compounding inference in GitSwarm, an asynchronous system where homogeneous agents independently decide how to advance a task while collaborating through structured persistent memory. Agents explore, experiment, verify, refine, and synthesize previous work in a shared, branch-able Git repository. Atomic commits preserve intermediate artifacts, while explicit semantic dependencies record how later contributions build on work across branches. We evaluate GitSwarm on long-horizon problem solving and sustained GPU-backed experimental research. On IMOProofBench-Advanced, GitSwarm solves all 30 problems in one run using GPT-5.5. On ProgramBench, it achieves a $79.4\%$ mean score, versus $65.1\%$ for the strongest reported baseline under the stated budget. On three neural architecture research tasks (Residual Matrix Transformer, Looped Transformer, NanoChat), GitSwarm improves upon the starting architectures through successive experimentation. Beyond final performance, we measure whether computation accumulates: on ProgramBench, $94.7\%$ of contributions are subsequently built upon, while the selected solution's ancestry covers $82-93\%$ of the contribution graph. These results show that inference-time computation can accumulate across otherwise independent episodes, forming an evolving body of work that subsequent inference can reuse.
|
| 1246 |
When the noncommutative AM-GM inequality holds
2610.04874
|
cs.LG
|
Yimin Zhong |
In this note, we prove that the noncommutative AM-GM inequality holds if $n\ge 2\lceil m/2 \rceil^2$. The motivation comes from counterexamples constructed in [De Sa, Random reshuffling is not always better, NeurIPS2020]. The proof constructs a vertex measure ...In this note, we prove that the noncommutative AM-GM inequality holds if $n\ge 2\lceil m/2 \rceil^2$. The motivation comes from counterexamples constructed in [De Sa, Random reshuffling is not always better, NeurIPS2020]. The proof constructs a vertex measure based on the Chebyshev nodes on the Boolean cube to extract the distinct indices. The main difficulty is that the measure is not positive on non-integer nodes. The key technique comes from [Grigoriev, Complexity of Positivstellensatz proofs for the knapsack, Computational Complexity (2001)] and eventually transforms the problem into a quadrature estimate.
|
| 1247 |
Bounded Reasoning: Cognitive Hierarchy in Human-versus-AI Cyber Defense
2610.04878
|
cs.LG
|
Zahra Aref, Sheng Wei, Narayan B. Mandayam |
Human-agent evaluations often compress interaction into a single performance score, even when human and automated policies adapt differently over time. We study this in a sequential cyber-defense game on an attack graph, where a human or reinforcement-learning...Human-agent evaluations often compress interaction into a single performance score, even when human and automated policies adapt differently over time. We study this in a sequential cyber-defense game on an attack graph, where a human or reinforcement-learning defender protects cloud assets against a Deep Q-Network (DQN) attacker. We compare four defender settings: a human reward-only game operationalizing the DQN information structure, a human reward-plus-transition game operationalizing the Cognitive Hierarchy Theory-driven DQN (CHT-DQN) information structure, an automated DQN defender, and an automated CHT-DQN defender. In the reward-only game, participants receive payoff and reward feedback. In the reward-plus-transition game, they also see attacker-aware transition probabilities from the CHT-DQN model. Across 80 Mechanical Turk participants and matched automated simulations, human defenders adapted to outcomes in a way the automated defenders did not: they were more likely to reselect a node after a successful defense than after a failure, in both games, an asymmetry we interpret as consistent with Prospect Theory and Cumulative Prospect Theory. The 40-round average also mixes early rounds, where the DQN attacker acts mostly at random, with late rounds, where it mostly exploits; in the final stage, mean protection ranks the reward-plus-transition human game above the reward-only human game, then the automated CHT-DQN defender, then the automated DQN defender, an ordering the overall mean hides. By contrast, the reward-plus-transition game does not produce a statistically reliable overall gain in weighted data protection over the reward-only game. These results suggest that evaluating human-agent cyber-defense systems only by an averaged task score can miss behaviorally meaningful differences in adaptation, action allocation, and bounded human reasoning.
|
| 1248 |
Static Bootstrap Placement for Encrypted Language Model Decoding
2610.04912
|
cs.LG
|
Halil Ibrahim Kanpak, Didem Unat |
Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a...Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a token under encryption, and generating text with them requires a client round trip at every generated token. Keeping the loop on the server instead requires selecting and consuming a token under encryption, and placing bootstraps for a loop body that grows with the context. We build AR-HE, which runs the whole loop on the server, selects each token under encryption, retrieves its embedding, and writes it back into the encrypted state. The client sends one prompt and remains offline until the output. One rule places every bootstrap in the run, without search, so the bootstrap cost of a token is a formula in the context length that is known before the run starts. The schedule skips work whose result cannot reach the output, packs bootstraps that share an operand, and keeps the keys and values of past positions in an encrypted cache. With every optimization applied, generating a GPT-2 small token costs 544 seconds on one NVIDIA H100, down from 4715 seconds without optimization. The prompt step before it costs 4630 seconds. The cache alone takes a generated step from 11751 bootstraps to 1072. The formula predicts every step we measured, including steps of a model it was not derived from.
|
| 1249 |
Assembling Insights for Agentic Machine Learning Engineering Systems
2610.04927
|
cs.LGcs.AI
|
Bihui Jin, Yinxi Li, Kaiyuan Wang, Pengyu Nie |
Agentic machine learning engineering (MLE) is an emerging AI4SE application for complex ML tasks and a step toward recursive self-improvement of AI systems. Recent agentic MLE systems show the value of leveraging insights from related MLE tasks: some systems c...Agentic machine learning engineering (MLE) is an emerging AI4SE application for complex ML tasks and a step toward recursive self-improvement of AI systems. Recent agentic MLE systems show the value of leveraging insights from related MLE tasks: some systems condition code generation on expert domain knowledge, which is implicitly curated from peer MLE tasks; some systems have a loop of solving an MLE task, gathering memory to benefit future tasks. However, insight collection and injection remain ad hoc, and existing evaluations often do not control insight sources, making data leakage a threat to validity. To close these gaps, we systematically study how insights from related MLE tasks affect agentic MLE systems, and how their impact depends on insight type and injection method. We introduce MLE-InsightBench, a benchmark of 160 Kaggle competitions across 12 domains. To prevent leakage, one competition per domain is held out as the target, while the remaining 148 serve as sources only when they cannot have been influenced by the target task, such as through later editions or reused solutions. We also develop MLE-InsightForge, an insight assembly agent that explores source competitions, top human solutions, and writeups to construct three insight types: 12 domain insights summarizing best practices, 148 competition insights describing winning solutions, and 16 improvement insights capturing generic debugging and optimization strategies. These insights are injected either as on-demand skills or as an upfront playbook. In our evaluation, the extracted insights improve the normalized private-test score by 123.9% over a no-insight baseline. Controlled experiments further show that the three insight types are complementary; the better injection mode (skill vs. playbook) varies by tasks, with skills more cost-effective overall; and the gains generalize across backend LLMs.
|
| 1250 |
DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling
2610.04933
|
cs.LGcs.AI
|
Seongheon Park, Heecheol Kim, Shulin Tian, Lilika Makabe, Namiko Saito |
Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candid...Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.
|
| 1251 |
Near-Optimal Complexity of Finite-Sum Nonconvex-Strongly-Concave Minimax Optimization
2610.04944
|
cs.LG
|
Qihao Zhou |
We characterize, up to logarithmic factors, the minimax expected query complexity of finite-sum nonconvex-strongly-concave optimization under mean-squared averaged smoothness. For randomized zero-respecting incremental first-order algorithms, the complexity is...We characterize, up to logarithmic factors, the minimax expected query complexity of finite-sum nonconvex-strongly-concave optimization under mean-squared averaged smoothness. For randomized zero-respecting incremental first-order algorithms, the complexity is $\tilde{\Theta}(n+\min\{\sqrt{n}\,\kappa,n^{3/4}\sqrt{\kappa}\}\,L\Delta\varepsilon^{-2})$ throughout $\kappa=L/\mu\ge1$, under exact dual initialization and for $0<\varepsilon^2\le c_0L\Delta$, where $c_0>0$ is universal. This characterization is our main result. A dual-chain construction provides the lower bound and matches the existing Catalyst upper bound for $\kappa\ge\sqrt{n}$. For $1\le\kappa\le\sqrt{n}$, we introduce recursive proximal descent-ascent (RPDA), which retains a recursive gradient estimator across proximal subproblems and attains $\tilde{O}(n+\sqrt{n}\,\kappa L\Delta\varepsilon^{-2})$ with a fixed query budget. Both upper bounds guarantee an expected squared primal gradient at most $\varepsilon^2$. Under individual smoothness, we also prove $\Omega(n+\min\{\kappa,\sqrt{n\kappa}\}\,L\Delta\varepsilon^{-2})$ for every $\kappa\ge1$. It matches the polynomial upper rate for a bilinear subclass when $\kappa\ge n$; the intermediate-condition-number gap remains open.
|
| 1252 |
Optimal Oracle Complexity for Finite-Sum Monotone Inclusions
2610.05038
|
cs.LG
|
Qihao Zhou |
We present an oracle-optimal method for finite-sum monotone inclusions under mean-square Lipschitz continuity. Our switching regularization method finds a point $y$ and a certificate $g\in G(y)$ with $(\mathbb{E}\|F(y)+g\|^2)^{1/2}\le\varepsilon$ using $\mathc...We present an oracle-optimal method for finite-sum monotone inclusions under mean-square Lipschitz continuity. Our switching regularization method finds a point $y$ and a certificate $g\in G(y)$ with $(\mathbb{E}\|F(y)+g\|^2)^{1/2}\le\varepsilon$ using $\mathcal{O}(n+\sqrt{n}LR/\varepsilon)$ expected component evaluations and resolvent evaluations. It removes the additive $n\log n$ cost of restarting a variance-reduced solver at every regularization stage by switching to a centered stochastic proximal iteration at regularization strength $L/\sqrt{n}$. Carrying an operator estimate between the remaining stages limits their total cost to $\mathcal{O}(n)$. A matching $\Omega(n+\sqrt{n}LR/\varepsilon)$ lower bound holds for randomized linear-span component-oracle algorithms with adaptive stopping and expected query budgets. Thus, for $0<\varepsilon\le LR/2$, our method attains the optimal worst-case expected component complexity in this oracle model, up to universal constants.
|
| 1253 |
A Statistical Inference Framework for PMI Estimation and SGNS Word Embeddings
2610.05058
|
cs.LG
|
Zhongqi Fan |
Pointwise Mutual Information (PMI) is a core measure of testing word association, and Skip-gram with Negative Sampling (SGNS) is essentially a method that implicitly factorizes a shifted PMI matrix. However, a systematic and well-rounded characterization of fi...Pointwise Mutual Information (PMI) is a core measure of testing word association, and Skip-gram with Negative Sampling (SGNS) is essentially a method that implicitly factorizes a shifted PMI matrix. However, a systematic and well-rounded characterization of finite-sample uncertainty in PMI estimation remains absent and imperative to venture into. We provide a statistical framework for PMI estimation and its connection to SGNS. We prove consistency, asymptotic unbiasedness, and asymptotic normality of the empirical PMI estimator, derive its variance via the Delta method, and, applying stochastic approximation theory, obtain a variance decomposition for SGNS-based PMI estimation that separates data variance from optimization variance. Simulation experiments validate the Delta method approximation. Real-data experiments on the Brown Corpus (d = 100) reveal that SGNS systematically deviates from the theoretical relationship PMI + log K. The empirical relationship shows an attenuated PMI coefficient, an amplified log K effect, and a positive intercept, indicating systematic bias. Word analogy validation confirms the models are effective. The failure to validate the variance decomposition under low-dimensional conditions does not diminish its theoretical value; rather, it identifies the unbiasedness assumption as the key bottleneck and clarifies the gap between asymptotic theory and practice, providing implications for both practice and theory.
|
| 1254 |
A Contrast-Source Inversion Scheme Based on Stochastic Optimization and Plug-and-Play Regularization
2610.05130
|
cs.LG
|
Lingqi Gao, Hakan Bagci |
An electromagnetic inversion scheme that integrates stochastic optimization (STO) and plug-and-play (PNP) regularization into contrast-source inversion (CSI), termed STO-PNP-CSI, is developed. Standard CSI solves for the contrast source vector of every transmi...An electromagnetic inversion scheme that integrates stochastic optimization (STO) and plug-and-play (PNP) regularization into contrast-source inversion (CSI), termed STO-PNP-CSI, is developed. Standard CSI solves for the contrast source vector of every transmitter at each iteration, which is expensive in a multi-transmitter configuration. STO instead solves for only one randomly selected contrast source vector per iteration, which reduces the per-iteration cost and can help the inversion escape poor local minima and saddle points. The resulting loss of information, however, increases the ill-posedness of the inversion. To counter this, the Swin-Conv-UNet (SCUNet) denoiser is plugged into the CSI scheme as an implicit regularizer, supplying a learned prior that is stronger than conventional hand-crafted ones and stabilizes the reconstruction. The proposed STO-PNP-CSI is applied to both synthetic and experimental data. The results show that it yields accurate reconstructions at substantially lower computational cost than CSI, including under strong nonlinearity and measurement noise.
|
| 1255 |
Locality Sensitive Hashing for p-Exponential Kernels with Applications to Density Estimation
2610.05174
|
cs.LG
|
Barak Gorodissky, Tal Wagner |
A kernel $k(x,y)$ is LSHable if there exists a locality sensitive hashing scheme $H$ such that $k(x,y)=\Pr_{h\sim H}[h(x)=h(y)]$ for all $x,y$. This notion plays a key role in efficient kernel methods in high dimensions. In this work, we show that the $p$-expo...A kernel $k(x,y)$ is LSHable if there exists a locality sensitive hashing scheme $H$ such that $k(x,y)=\Pr_{h\sim H}[h(x)=h(y)]$ for all $x,y$. This notion plays a key role in efficient kernel methods in high dimensions. In this work, we show that the $p$-exponential kernel $k(x,y)=\exp(-\lVert x-y \rVert_p)$ is LSHable in bounded regions for all $1<p\leq2$. Previously, this was known only for $p=1$. Our new "mosaic LSH" scheme is based on a Poisson hyperplane process with hyperplanes sampled as $\ell_1$-biased $p$-stable vectors, for which we develop efficient sampling procedures. As applications, our results yield new and efficient density estimation methods based on LSHability for those $p$-exponential kernels.
|
| 1256 |
On prediction from expert advice with more than five experts
2610.05186
|
cs.LG
|
Jeff Calder, Nadejda Drenska |
We prove that no single rank ordered adversary strategy is globally optimal for the prediction with expert advice problem with six or more experts, in both the geometric stopping and finite time horizon settings. The proof is based on establishing a leading or...We prove that no single rank ordered adversary strategy is globally optimal for the prediction with expert advice problem with six or more experts, in both the geometric stopping and finite time horizon settings. The proof is based on establishing a leading order correction when one expert moves far ahead of the others. This allows us to connect optimal strategies between $n$ and $j<n$ experts and utilize recent results on the exact optimality set for the five expert problem.
|
| 1257 |
Synergizing Drone Delivery Order Pooling and Road Network Monitoring through Monitoring-Task Orderization
2610.05270
|
cs.LG
|
Yulong Hu, Meng Xu, Sen Li, Nikolas Geroliminis |
This paper investigates the real-time dispatch of a shared drone fleet for on-demand food delivery and urban road network monitoring. We consider a courier-drone collaborative setting in which couriers transport orders to launchpads and drones complete the fin...This paper investigates the real-time dispatch of a shared drone fleet for on-demand food delivery and urban road network monitoring. We consider a courier-drone collaborative setting in which couriers transport orders to launchpads and drones complete the final delivery leg to kiosks. Drones may consolidate multiple origin-destination orders within one flight and make monitoring-aware route adjustments to collect real-time traffic information subject to delivery-time constraints. This yields a joint decision problem coupling dynamic order-to-drone matching, multi-order pooling, routing, and time-varying monitoring under fleet-level competition and uncertainty. We propose monitoring-task orderization, which periodically converts road-network nodes with high congestion and stale information into virtual monitoring orders. Pooling these virtual tasks with food-delivery orders creates a unified heterogeneous task set and transforms the coupled matching-and-routing problem into an order-level decision process. Building on this abstraction, we formulate a decentralized graph-interdependent Multi-Agent Markov Decision Process and develop Graph Multi-Agent Q-Learning (Graph-MAQL), which captures localized inter-agent dependencies through bipartite match coordination graphs. Agent-task value estimates are then used as edge weights in a dynamic heterogeneous bipartite matching program for globally feasible execution. Experiments using real-world data reveal strong operational synergy between delivery and monitoring. Monitoring-task orderization improves monitoring performance by 25.1% with less than a 1% reduction in delivery performance, while Graph-MAQL improves the aggregate objective by up to 20.8%, reduces deadline violations by over 40%, and transfers zero-shot to higher demand intensity without retraining.
|
| 1258 |
Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs
2610.05305
|
cs.LGcs.AI
|
Javad Mirzaei, Jeebak Mitra |
Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute an...Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly distributed across multiple GPUs using tensor parallelism (TP), pipeline parallelism (PP), or hybrid parallelism (HB). However, selecting the most effective parallelism strategy remains challenging due to complex interactions among computation, communication, pipeline utilization, sequence length, batch size, and model architecture. Existing approaches largely rely on empirical evaluation and provide limited analytical insight into the trade-offs among these strategies, particularly across the distinct prefill and decoding phases of inference. In this paper, we present a unified analytical framework for modeling distributed LLM inference under TP, PP, and HB. The framework decomposes end-to-end latency into computation, inter-GPU communication, and pipeline bubble overhead, and derives analytical models that capture TP collective communication, PP point-to-point communication, and pipeline utilization as functions of hardware, model, and workload characteristics. The model further characterizes the differing execution behavior of prefill and decoding, explaining why PP-oriented configurations favor compute-intensive prefill while TP-oriented configurations reduce decoding latency by eliminating pipeline bubbles. Experiments with modern LLMs on multi-GPU platforms validate the model and confirm the fundamental compute-communication trade-off across parallelism strategies. The framework provides practical guidance for parallelism selection, capacity planning, and optimization of future LLM serving systems.
|
| 1259 |
Transferable Adversarial Robustness for Speech Foundation Models via Hierarchical Stabilization
2610.05310
|
cs.LGeess.AS
|
Aref Mousavi, Shahab Sherafat, Kiarash Kiani Feriz, Amirparsa Safari, Raoof Zare Moayedi |
Frozen speech foundation models (SFMs) make downstream adaptation efficient: the backbone can stay fixed while a task learns layer fusion and a lightweight classifier. Full adversarial fine-tuning is a standard route to robustness, but generating adversarial e...Frozen speech foundation models (SFMs) make downstream adaptation efficient: the backbone can stay fixed while a task learns layer fusion and a lightweight classifier. Full adversarial fine-tuning is a standard route to robustness, but generating adversarial examples and updating the backbone for every task sacrifices that efficiency. We ask whether robustness can instead be learned before future tasks are known. For a frozen backbone and linear classifier, robustness can be understood through the interaction between representation stability and decision-boundary margin. This leads directly to our design: we stabilize representations across the hidden layers, rather than only the final layer, while preserving clean representations; after clean adaptation selects the layer mixture, we keep it fixed and enlarge only the classifier margin, without downstream adversarial examples. We evaluate Wav2Vec2, HuBERT, and WavLM Large on four tasks under adaptive 30 dB attacks. Across 12 backbone-task pairs, hierarchical robustification improves robust accuracy by 46.4 pp, while margin refinement adds 4.0 pp for 1.1 pp of clean accuracy. Code and configurations are available at https://github.com/arefmousavi/hierarchical-robust-sfm.
|
| 1260 |
CoDance: Learning Reactive and Compliant Human-Humanoid Interaction from Video
2610.05324
|
cs.LGcs.AI
|
Zhuoqun Chen, Shucheng Jia, Boyuan Chen |
Partnered human-humanoid interaction couples locomotion with continuous physical contact. A humanoid needs to coordinate with a person's motion while responding to interaction forces and maintaining stable and natural movement. We present CoDance, a framework ...Partnered human-humanoid interaction couples locomotion with continuous physical contact. A humanoid needs to coordinate with a person's motion while responding to interaction forces and maintaining stable and natural movement. We present CoDance, a framework for learning reactive and compliant human-humanoid interaction from video. We study partnered dancing as a challenging instantiation, where a humanoid coordinates its footsteps with a moving partner and maintains continuous two-hand contact. Given a single video of two human dancers, CoDance retargets their motions into a robot reference and a moving partner. We introduce a multi-link compliance augmentation that transforms the kinematic demonstration into force-aware training data by adapting the robot reference under structured forces at both hands. Policies trained on this data follow the observed partner while preserving the demonstrated locomotion style and responding compliantly to physical interaction. In simulation, the policies adapt their footsteps to changes in the partner and reproduce approximately 80% of the wrist displacement encoded by the augmented demonstrations. On a physical humanoid, CoDance enables sustained two-hand dancing with a human partner including repeated transitions between forward and backward motions.
|
| 1261 |
CT-Miner: Fast and Coarse-Grained Time-Series Pattern Mining via Cartesian Trees
2610.05330
|
cs.LG
|
Hyundong Jin, Hyunki Hong, Yo-Sub Han |
Time series often contain recurring structural patterns, and efficiently mining such patterns into compact representations is essential for scalable analysis of long sequences. Cartesian tree (CT) equivalence provides a well-established structural abstraction ...Time series often contain recurring structural patterns, and efficiently mining such patterns into compact representations is essential for scalable analysis of long sequences. Cartesian tree (CT) equivalence provides a well-established structural abstraction that preserves hierarchical order structure while discarding exact values and fine-grained ordinal variations. By grouping multiple ordinal patterns into a shared structural form, CT equivalence offers a principled way to compress recurring temporal structure. However, mining frequent CT-equivalent patterns at scale remains computationally expensive. A naive pairwise approach repeatedly constructs and counts CT representations over subsequences, requiring $O(n^4)$ time for a sequence of length $n$, which severely limits its applicability to long sequences. We propose a new Cartesian pattern mining algorithm based on a Cartesian suffix tree that compactly organizes CT-equivalent subsequences and reuses shared structural information. Our method reduces exhaustive CT-pattern occurrence collection from $O(n^4)$ to $O(n^2)$ time, and we formally prove the correctness and complexity bounds. We further show that this computational gain translates into effective compact representations. Across diverse time-series datasets, a small set of mined CT patterns preserves meaningful clustering structure, and comparisons with finer-grained order-preserving representations show that CT equivalence reduces redundant ordinal distinctions under limited feature budgets. Our implementation is available at https://github.com/hyundong98/CT-Miner .
|
| 1262 |
VAMPS: Visual and Motor Policies from Sampling-Based Planning
2610.05331
|
cs.LGcs.AI
|
Mohamed Yassine Kabouri, Pietro Noah Crestaz, Quang-Nam Nguyen, Qilong Cheng, Ludovic Righetti |
Learning robot policies directly on physical systems remains difficult because data collection is costly and policy exploration can be unsafe. We introduce Visual and Motor Policies from Sampling-Based Planning (VAMPS), a framework that uses Model Predictive P...Learning robot policies directly on physical systems remains difficult because data collection is costly and policy exploration can be unsafe. We introduce Visual and Motor Policies from Sampling-Based Planning (VAMPS), a framework that uses Model Predictive Path Integral (MPPI) control to train reusable policies without human demonstrations. VAMPS supports two training modes. For one-step proprioceptive policies, it operates iteratively in simulation: the policy warm-starts MPPI, and the refined trajectories provide new supervision as the policy changes. A learned terminal value improves short-horizon planning, while an Implicit Q-Learning (IQL) critic guides the policy update. Iterative refinement outperforms training once on frozen MPPI data, and we transfer the learned locomotion policy to a Unitree Go2. For visuomotor policies, VAMPS operates directly from real-robot data. MPPI uses task-specific state estimates to plan and execute trajectories while recording RGB and sensor observations on a Flexiv Rizon 10S. Action Chunking with Transformers predicts action chunks, reducing the effective prediction horizon, and is trained offline on this fixed dataset. We demonstrate visuomotor pick-and-place and force-aware whiteboard erasing. In the latter task, the policy additionally observes the measured $6$-D wrench and desired normal force. These results show that VAMPS can learn policies either in simulation followed by hardware transfer or directly from autonomously collected real-robot data.
|
| 1263 |
AgentDiscover: Autonomous Discovery with Minimal Search Scaffolding
2610.05334
|
cs.LGcs.AI
|
Mahdi Farahbakhsh, Ilan Sela, Fatemeh Doudi, Vishnu Teja Kunde, Krishna Narayanan |
Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what...Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what it is shown. As models grow more capable, a question arises: does a search strategy chosen by a human before the run scale better than promoting the model from proposer to planner and letting it own the search? The Bitter Lesson suggests that choosing the strategy in advance is the kind of hand-designed structure that general methods eventually outscale. We introduce AgentDiscover, in which a coding agent plans the search using its context as working memory, runs experiments, and records every attempt in a database of ideas, candidates, and their relations. This database serves as the agent's long-term memory and is structured so that the selection rules of classical algorithms such as MAP-Elites and Monte Carlo tree search each reduce to a single query, which the agent is free to use, combine, or replace. A server maintains the database and steers the agent after every submission, keeping it on course over long runs. In our experiments, AgentDiscover is more cost-efficient than existing frameworks, reaching better scores at lower cost. On tasks in kernel engineering, biology, algorithm design, and mathematics, AgentDiscover outperforms prior discovery frameworks. Its programs would have placed first among human competitors in seven past AtCoder heuristic contests, and on eleven mathematical and systems optimization tasks it matches or exceeds every baseline that uses the same model. Our code is available at https://github.com/mhdfb/AgentDiscover.
|
| 1264 |
AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness
2610.05367
|
cs.LGcs.AI
|
Prithwish Jana, Viet Bach Hoang, Logan Luna, Viresh Pati, Akash Singirikonda |
Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistan...Proof auto-formalization translates natural-language (NL) theorems and proofs into a formal language (FL) such as Lean, enabling mechanical verification. Despite rapid progress, research-level proofs often depend on concepts missing from leading proof assistant libraries (e.g., Lean's Mathlib), and successful compilation does not guarantee that a translation preserves the theorem's meaning or the proof's reasoning. Furthermore, aligned NL-FL training data are scarce, and leading agents often rely on costly frontier models and manually engineered harnesses. To address the above issues, we present AIProver, an agentic framework for autonomous proof auto-formalization and proof synthesis (AFPS) that jointly post-trains a 119B open-weight language model and evolves its agentic, tool-calling harness with HarnessEvolve. Verifiers assess type correctness, proof completeness, and semantic correctness, returning rewards and diagnostic certificates that drive model fine-tuning and alternating reinforcement learning via symbolic feedback and HarnessEvolve, a certificate-driven evolutionary search over the whole harness control flow that re-tailors the harness to the updated model. For research-level training and evaluation, we introduce LoCoBench, 58.9k instances from Mathlib, CSLib, Mizar Math Library, and a bounded-arithmetic textbook, with a 771-instance validation split whose theorem-proof pairs have no public Lean formalization. Against 39 frameworks spanning AFPS agents, frontier LLMs, and coding agents, AIProver lifts pass@4 semantic correctness over its Leanstral-1.5 base from 15.7% to 36.7% and outperforms every other open-weight system and Aristotle. As a Claude Code and Codex skill, it lifts their semantic correctness from 41.9% and 34.1% to 79.8% and 62.4%, respectively. Further, it is also 24% cheaper than Numina-Lean-Agent, pushing the accuracy-cost frontier of research-level AFPS.
|
| 1265 |
Learning in Continuous Games from Pairwise Preference Feedback
2610.05428
|
cs.LG
|
Anas Barakat |
We study learning in continuous games when players receive only pairwise preference feedback, revealing which of two actions is preferred but neither payoff values nor preference magnitudes. We first show that standard external regret and coarse correlated equ...We study learning in continuous games when players receive only pairwise preference feedback, revealing which of two actions is preferred but neither payoff values nor preference magnitudes. We first show that standard external regret and coarse correlated equilibria (CCE) are not identifiable from this ordinal information: the same sequence of play can incur zero and linear regret in two ordinally equivalent games, while the distributions that remain CCE across all cardinal representations consistent with the same preferences are exactly those supported on pure Nash equilibria. Motivated by this gap, we develop a first-order ordinal theory based on normalized unilateral preference directions, introducing an ordinal directional regret benchmark and corresponding equilibrium notions. We show that block-normalized pseudogradient dynamics achieve sublinear ordinal regret and, under additional structure, Nash-convergence guarantees. We then use a single-comparison estimator to implement these dynamics from finite pairwise comparisons. With one comparison per player and round, the resulting algorithm achieves sublinear finite-resolution ordinal regret against arbitrary opponent behavior and, in ordinal potential games, almost-sure last-iterate convergence to the Nash set. Our results provide regret, dynamics, and equilibrium guarantees directly from preference feedback without reconstructing cardinal utilities.
|
| 1266 |
The sublevel Flood bifiltration: towards scalable 2-parameter persistent homology
2610.05441
|
cs.LG
|
Matt\'eo Cl\'emot, Julie Digne, Julien Tierny |
Multiparameter persistent homology is a rapidly developing branch of topological data analysis that improves the robustness of single-parameter persistent homology to outliers, while still capturing the metric characteristics of the data. However, a notable li...Multiparameter persistent homology is a rapidly developing branch of topological data analysis that improves the robustness of single-parameter persistent homology to outliers, while still capturing the metric characteristics of the data. However, a notable limitation is its lack of scalability. In this paper, we introduce a novel approach for efficiently computing 2-parameter persistent homology on large point sets. Our work extends the Flood filtration, originally developed for single-parameter persistence. Our construction, called the sublevel Flood bifiltration, offers a scalable approximation of the sublevel offset bifiltration. We show that it benefits from theoretical stability properties and describe how to compute it efficiently. We demonstrate the performance of our approach in classification tasks on low-dimensional synthetic datasets, where density awareness is critical, as well as on real-world time series datasets.
|
| 1267 |
Taylor Representations for Model-Free RL in Networked MDPs
2610.05456
|
cs.LG
|
Salah Chikhi, Abdelhaq Chaoui, Asuman Ozdaglar, Saurabh Amin |
In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over...In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over $N$ agents requires $\Theta(N^n)$ coefficients. We justify these expansions under smooth expected future local rewards with controlled derivatives. Under this condition, finite-speed information propagation and discounting imply that local-critic Taylor coefficients decay exponentially with the graph distance to the farthest agent involved. Discarding distant-agent coefficients and marginalizing then yield scalable local Taylor representations with a bound controlled by graph locality. Building on these representations, we propose a scalable model-free actor--critic algorithm, establishing finite-sample critic and near-stationarity guarantees for a linear LSTD critic. We then introduce a more expressive neural TD parameterization. Unlike prior constructive spectral methods, our approach covers settings without access to a known local dynamics map, such as hidden switched linear--quadratic regulation. Across three control benchmarks, our method matches or outperforms spectral baselines while scaling efficiently to large graphs.
|
| 1268 |
CodeForge-MA: Execution-Verified Multi-Agent Learning with Language-Conditioned LoRA for Multilingual Code Generation
2610.05481
|
cs.LGcs.AI
|
Zhizhou Gu, Xianting Wu, Siyu Gu, Tian Zhang, Kejian Tong |
Large language models for code generation often fail on execution, multilingual coverage, and contamination control, especially under frozen backbone constraints. We present CodeForge-MA, a unified framework that improves code synthesis through a multi-agent d...Large language models for code generation often fail on execution, multilingual coverage, and contamination control, especially under frozen backbone constraints. We present CodeForge-MA, a unified framework that improves code synthesis through a multi-agent data forge, execution verified reinforced instruction tuning, and a language conditioned mixture of LoRA adapters. Four specialized agents, Composer, Reviewer, Executor, and Curator, iteratively refine instruction code pairs, validate them with tests, and filter duplicates and benchmark leakage. During training, we combine masked supervised fine tuning with a test driven reinforcement objective to align generations with executable correctness. For the larger model, we use sparse expert routing over low rank adapters to improve cross language transfer while keeping the base model unchanged at inference. Experiments show that joint data, objective, and adapter design yields robust gains across programming languages.
|
| 1269 |
G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm
2610.05563
|
cs.LGcs.AI
|
Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian |
Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal A...Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.
|
| 1270 |
Moment-Accurate Gaussian Mixtures for Constant-Step Stochastic Approximation
2610.05595
|
cs.LG
|
Xiaoli Li, Wei Biao Wu |
Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornste...Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornstein--Uhlenbeck limits, ruling out quadratic tail mass invisible to weak convergence. For step size $a$, the second-order Wasserstein error is $o(\sqrt a)$, uniformly over invariant laws, using each law's actual root weights. The assumptions combine confinement, descent, finitely many hyperbolic equilibria and root continuity with finite-variance innovations. The result yields observable covariances, expected objective gaps and first-order mean shifts, while allowing singular covariances, compatible saddles and weights without a limit. For additive noise given by a fixed invertible transform of independent standardized Student $t_3$ coordinates, symmetry gives an order-sharp $\sqrt a$ smooth-test bound. Numerical transport calculations demonstrate the value of root-specific covariances; controlled SGD studies assess observable predictions across step sizes, batch sizes and model geometries.
|
| 1271 |
Gaussian Limits for SGD Without Stationary Moments
2610.05599
|
cs.LG
|
Xiaoli Li, Wei Biao Wu |
Temporal dependence can separate the Gaussian approximation of stochastic gradient descent from its stationary moments. For unmodified least-squares SGD, we construct a design with standard Gaussian marginals whose stationary error has every positive moment in...Temporal dependence can separate the Gaussian approximation of stochastic gradient descent from its stationary moments. For unmodified least-squares SGD, we construct a design with standard Gaussian marginals whose stationary error has every positive moment infinite. Independent observations with the same marginals instead give finite stationary variance. Both regimes retain a Gaussian small-step limit. Our general theory establishes pathwise contraction from a finite second design moment, then uses score cancellation and localization to obtain stationary Gaussian and Ornstein--Uhlenbeck limits. Independent Gaussian regression errors yield an exact conditional Gaussian law and total-variation convergence under the same design integrability. Stronger design conditions identify a positive first-order total-variation constant and a deterministic covariance correction with $o(a)$ error. A scalar coverage expansion translates this correction into its inference consequence. Experiments examine distributional error, coverage, and calibration with dependent scores. Together, these results establish precise probability-law approximation beyond moment-based stationary analysis.
|
| 1272 |
Causal Lag Structure Discovery in Confounded Time Series via Orthogonalized Adaptive Estimation
2610.05618
|
cs.LG
|
Hong Kiat Tan, Isaac-Neil Zanoria, James Chen, Haoyang Lyu, Mihai Cucuringu |
Finding which variables cause which others in multivariate time series, and at what lags, is central to science and policy, yet existing methods force a choice between flexible confounder adjustment, data-driven lag selection, and inference that controls the f...Finding which variables cause which others in multivariate time series, and at what lags, is central to science and policy, yet existing methods force a choice between flexible confounder adjustment, data-driven lag selection, and inference that controls the false discovery rate (FDR). ORACLE-VARX does all three in one pipeline. First, double/debiased machine learning (DML) removes nonlinear confounder effects from the outcomes and the lagged series. Second, adaptive causal lag estimation (ACLE) picks the lag order at each time step by sequential significance tests, tracking regime changes. Third, entry-wise $z$-tests with Benjamini--Hochberg correction select directed edges at a target FDR. We prove that in each rolling window, the debiased coefficients are asymptotically normal around a window-averaged target, so their $z$-tests are asymptotically valid. On a synthetic benchmark with time-varying structure and nonlinear confounding, ORACLE-VARX (LightGBM) tracks the true lag order best (RMSE $0.96$ vs $1.1$--$1.5$), has edge FDR $0.047$, close to PCMCI ($0.045$) and below VAR ($0.129$) and VAR-LiNGAM ($0.187$), and forecasts better than all three. On nine U.S. sector ETFs with macroeconomic confounders, it yields interpretable causal graphs whose lag order rises in high-volatility regimes.
|
| 1273 |
Spacecraft Rendezvous Trajectory Generation with Modular Constraints via Diffusion Model Composition
2610.05642
|
cs.LG
|
Mariko A. Storey-Matsutani, Richard Linares |
Emerging mission classes such as on-orbit servicing, satellite inspection, and active debris removal require trajectory design methods that are adaptable to a variety of mission scenarios. We present a diffusion-based trajectory generation approach for rendezv...Emerging mission classes such as on-orbit servicing, satellite inspection, and active debris removal require trajectory design methods that are adaptable to a variety of mission scenarios. We present a diffusion-based trajectory generation approach for rendezvous and proximity operations (RPO) that enables flexible configuration of mission constraints. First, individual energy-based diffusion models are trained to satisfy distinct constraints such as approach cone and sensor line-of-sight from a set of optimized trajectories. Then, at inference time, the learned energy models can be composed with one another, or with an analytically defined energy field, to enforce specific constraint combinations. We validate this framework with the composition of a learned approach cone model and a learned sensor line-of-sight model, as well as a learned approach cone model and synthetic obstacle avoidance model, both of which yield constraint satisfaction rates that are within 1 percentage point of the single-constraint models or higher. These results indicate that our compositional diffusion framework can provide a modular approach to RPO trajectory design and enable reconfiguration for new constraint combinations without requiring model retraining.
|
| 1274 |
An evolutionary origin of collective decision making in humans and machines
2610.05676
|
cs.LG
|
Guocheng Wang, Qi Su, Joshua B. Plotkin |
Groups of individuals can solve collective problems more accurately than any single member, by aggregating their opinions. Recent theoretical work has identified individual-level reward schemes that allow uninformed individuals to evolve collective intelligenc...Groups of individuals can solve collective problems more accurately than any single member, by aggregating their opinions. Recent theoretical work has identified individual-level reward schemes that allow uninformed individuals to evolve collective intelligence from the bottom up, through social learning. Yet these results are restricted to linear prediction problems and simple averaging, while the decision tasks that real groups confront are often non-linear, and the institutions that aggregate opinions are seldom single-layer averages: districts elect representatives who in turn vote on policy, referees advise editors who decide on publication. Here we develop a framework for the evolution of collective intelligence in multi-layer voting populations, where individuals observe limited information and groups recursively aggregate their opinions by majority rule. We prove that single-layer voting cannot solve non-linear classification problems under any individual reward scheme. We then identify a "marginal feedback" payoff structure, which rewards individuals only when their opinion is pivotal in their group, and at every layer above them. This reward scheme induces a layered population to evolve accurate collective solutions to complex, non-linear decision tasks through individual-level peer imitation alone. The collective behavior that emerges is equivalent to a multi-layer perceptron in machine learning. Our results provide a naturalistic account of hierarchical institutions, in which the outsize importance of swing voters is the incentive that sustains collective accuracy; and they identify the credit-assignment rule in machine learning as not just an engineered solution but a natural evolutionary outcome.
|
| 1275 |
RESOLVE: Language-Agnostic Validation of GPU Kernels Through Testing, Reduction, and Proof
2610.05683
|
cs.LG
|
Ashkan Vedadi Gargary, Guido Mart\'inez, Sebastian Burckhardt, Gabriel Ebner, Abhinav Jangda |
AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not suffi...AI systems can now write and optimize production GPU kernels, but validating them remains an important challenge. Evaluating the kernel on a few random inputs and checking that its outputs match a trusted reference kernel within numeric tolerances is not sufficient: races can cause nondeterministic behavior that fails to manifest in tests, and numeric tolerances can hide bugs and cause false positives even after extensive calibration. To address this challenge, we present RESOLVE, which combines testing and formal verification to build a comprehensive kernel validation pipeline. It operates in three steps: First, it tests for nondeterminism using binary instrumentation that perturbs execution timing to expose races. Second, an agent rewrites the candidate and reference kernels to obtain "reduced-concurrency" versions that are simpler to analyze but still produce bitwise-identical outputs in all tests. Third, the reduced kernels are formally analyzed in the F*/Pulse framework and prove that they perform the same computation on real numbers. This sidesteps the need for numeric tolerances. We show that RESOLVE can validate a broad selection of kernels using KernelBench, and prove equivalence across fused GEMMs in three state-of-the-art frameworks and languages: CUTLASS, Triton, and Gluon. It also analyzes mega-kernels, notoriously difficult to validate, and finds four previously unreported issues, including two clear bugs. We show that agents can use RESOLVE to repair the issues, with minimal performance impact, highlighting that agents can optimize aggressively when they can rigorously check their results.
|
| 1276 |
Two-Sample Testing via Path-based Inference
2610.05684
|
cs.LGcs.AI
|
Eshant English, Wei-Cheng Lai, Yanfeng Yang, Kenji Fukumizu, Taiji Suzuki |
Modern deep generative models are primarily studied for their ability to generate realistic samples, yet the generative dynamics they learn can also serve as objects of statistical inference. We develop this idea for two-sample testing, the problem of deciding...Modern deep generative models are primarily studied for their ability to generate realistic samples, yet the generative dynamics they learn can also serve as objects of statistical inference. We develop this idea for two-sample testing, the problem of deciding whether the same distribution generated two finite datasets. Using stochastic interpolants, we connect both distributions to a shared Gaussian bottleneck, so that each half of the resulting path is a Gaussian channel acting on a single population. We prove that the null hypothesis holds if and only if the population denoiser, or equivalently, the velocity fields of the two halves, coincide at any single noise level, which amounts to a reflection symmetry of the path about the bottleneck. Deviations from this symmetry yield a continuum of two-sample witnesses, which we estimate via held-out regression risks on learned denoisers and velocities and aggregate along the path; under an information-theoretic weighting, the aggregated discrepancy equals the Jeffreys divergence between the noise-smoothed distributions. Calibrating the resulting statistics by permutation yields tests that are valid in finite samples for any trained networks and consistent when the fields are learned accurately. On a synthetic benchmark and three image benchmarks, the proposed tests improve power over the strongest baseline by up to 33 percentage points at an equal total sample budget, with the best choice of regression representation and path weighting depending on the data modality. These results show that generative paths provide a principled representation for statistical testing, extending stochastic-interpolant models beyond generation.
|
| 1277 |
Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts
2610.05690
|
cs.LG
|
Yanjun Chen, Yongfeng Zhang, Lanjing Zhang |
Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LL...Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.
|
| 1278 |
Retrieval-Based In-Context Learning: A Domain Adaptation Framework
2610.05717
|
cs.LG
|
Yilun Zhu, Naihao Deng, Yingcong Li, Naichen Shi, Clayton Scott |
In-context retrieval (ICR) is a retrieval-based form of in-context learning (ICL) in which demonstrations are retrieved from a source database based on similarity to the query, rather than sampled independently. In this work, we formulate ICR as a type of doma...In-context retrieval (ICR) is a retrieval-based form of in-context learning (ICL) in which demonstrations are retrieved from a source database based on similarity to the query, rather than sampled independently. In this work, we formulate ICR as a type of domain adaptation problem, where the source distribution $P$ of the database may differ from the target distribution $Q$ of the test query-label pair. We investigate the performance of ICR under a flexible class of distributional shifts that substantially extends prior work \citep{li2024fine,guo2025retrieval}, and establish theoretical guarantees that quantify the benefits and pitfalls of this learning paradigm. Our theory is verified by experiments on synthetic and language tasks.
|
| 1279 |
Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation
2610.05768
|
cs.LGcs.AIcs.SDeess.AS
|
Keren Shao, Ayaka Kawano, Shlomo Dubnov |
We ask: given a retrieved source audio $S$ and a separate reference audio $R$, can we synthesize novel audio $Y$ out of this pair $(S,R)$ such that $Y$ remains acoustically consistent with $S$, while not persistently copying segments of $S$ or $R$? The first c...We ask: given a retrieved source audio $S$ and a separate reference audio $R$, can we synthesize novel audio $Y$ out of this pair $(S,R)$ such that $Y$ remains acoustically consistent with $S$, while not persistently copying segments of $S$ or $R$? The first clause is a well-known goal in Foley audio production, and the second is a well-known issue in neural RAG when $S$ and $R$ are naively injected into neural generators. We show that both clauses can be addressed simultaneously using a method we coin relational synthesis, a variation of concatenative synthesis where target cost is replaced by a relational Gromov-like structural cost. Rather than imitating the content of $R$, relational synthesis exploits it from the "other side of the hill": it transfers the temporal structure and directed amplitude motion of $R$ to reorganize and concatenate the grains of $S$ in a novel manner that protects $S$'s acoustic information. Our experiments show that relational synthesis integrates naturally with neural RAG and produces Foley audio that performs well on metrics measuring temporal agreement, acoustic fidelity, and leakage persistence, while maintaining distribution-level quality and text alignment.
|
| 1280 |
Isotropic Gaussian Processes Improve Vanilla Bayesian Optimization in High Dimensions
2610.05780
|
cs.LG
|
Wei-Ting Tang, Madhav Muthyala, Joel A. Paulson |
High-dimensional Bayesian optimization (BO) often fits Gaussian process (GP) surrogates from far fewer observations than input dimensions. Modern Vanilla BO can perform well in this regime with dimension-aware priors, initialization, and acquisition optimizati...High-dimensional Bayesian optimization (BO) often fits Gaussian process (GP) surrogates from far fewer observations than input dimensions. Modern Vanilla BO can perform well in this regime with dimension-aware priors, initialization, and acquisition optimization, but it typically retains automatic relevance determination (ARD), fitting one lengthscale per input coordinate. We study this modeling choice and propose Iso-BO, a controlled modification that replaces the ARD GP with an isotropic GP using one shared lengthscale while keeping the surrounding BO pipeline matched. For radial kernels, we show that the marginal log likelihood (MLL) depends on the inverse-squared ARD lengthscales only through weighted pairwise distances among the observed inputs. The current design can therefore leave some ARD directions exactly invisible or only weakly constrained by the MLL. Iso-BO removes coordinatewise reweighting and fits a single shared scale instead. Lengthscale-fitting and predictive-density diagnostics show that this finite-data effect appears in practice, including when the data-generating process is anisotropic. Across GP-prior, synthetic, and real-world benchmarks, Iso-BO often improves over matched modern Vanilla BO and remains competitive with the included high-dimensional BO baselines under the tested budgets. Stress tests also show the expected boundary wherein sufficiently strong, learnable anisotropy can favor the more flexible ARD model.
|
| 1281 |
Agentic-ZTA: A Multi-Agent Architecture for Autonomous Zero Trust Enforcement
2610.05782
|
cs.LGcs.AI
|
Shovan Roy, Lopamudra Praharaj, Maanak Gupta, Bhavani Thuraisingham |
Agentic AI is emerging as a promising paradigm for automating complex cybersecurity decisions, yet its use in enforcing zero trust introduces significant challenges in safety, reliability, and policy compliance. This paper presents Agentic AI based zero trust ...Agentic AI is emerging as a promising paradigm for automating complex cybersecurity decisions, yet its use in enforcing zero trust introduces significant challenges in safety, reliability, and policy compliance. This paper presents Agentic AI based zero trust architecture (Agentic-ZTA) that operationalizes the NIST SP 800-207 ZTA architecture control loop through coordinated multi- agent decision pipeline. In the proposed framework, policy knowledge is embedded into a retrieval-augmented generation pipeline and retrieved at inference time as top-k relevant policies. Access requests are intercepted by the Policy Enforcement Point (PEP), enriched with contextual metadata. The request context is routed to a policy engine agent which invokes domain-specialized core agents first followed by supporting agents, if further evaluation needed. AI agents reason over access context, policy constraints and determine trust. The retrieved policies are embedded into agent prompt during inference time and agentic trust scores are aggregated and evaluated by a trust-algorithm, producing the final access decision for enforcement under continuous verification. We implement Agentic-ZTA in a testbed and evaluate it on representative access-control use cases scenarios. Our Agentic-ZTA framework achieves 95.0% accuracy, 93.9% precision, and 96.3% recall, and demonstrate the feasibility of enforcing zero trust using AI agents.
|
| 1282 |
Dimension-Free Decentralized Nonsmooth Nonconvex Stochastic Optimization
2610.05789
|
cs.LG
|
Yuanyu Wan, Lan Xue, Haomin Bai, Tong Wei, Mingli Song |
We investigate decentralized nonsmooth nonconvex stochastic optimization over a network of $n$ nodes, with the goal of finding an $(\delta,\epsilon)$-Goldstein stationary point. The best existing algorithm achieves $O(\delta^{-1}(\epsilon^{-3}+d\epsilon^{-1}))...We investigate decentralized nonsmooth nonconvex stochastic optimization over a network of $n$ nodes, with the goal of finding an $(\delta,\epsilon)$-Goldstein stationary point. The best existing algorithm achieves $O(\delta^{-1}(\epsilon^{-3}+d\epsilon^{-1}))$ sample complexity and $\widetilde{O}(\gamma^{-1/2}\delta^{-1}(\epsilon^{-3}+d\epsilon^{-1}))$ communication complexity, where $d$ is the problem dimension and $\gamma$ is the spectral gap of the communication matrix. However, the polynomial dependence on $d$ can be a major bottleneck in high-dimensional regimes. In this paper, we propose a novel algorithm that achieves $O(\delta^{-1}\epsilon^{-3})$ sample complexity and $\widetilde{O}(\gamma^{-1/2}\delta^{-1}\epsilon^{-3})$ communication complexity. The primary technique is an elegant decentralized online-to-nonconvex conversion that reduces the original problem to a decentralized online convex optimization (D-OCO) problem. A key property of our conversion is that its consensus requirements can be inherited directly from the consensus of the underlying D-OCO decisions. In particular, this property enables us to establish an explicit connection between the dimension dependence and the consensus error, which in turn shows that the polynomial dependence on $d$ can be removed with only logarithmic additional communication.
|
| 1283 |
A Testable Theory of Atomic Features
2610.05794
|
cs.LGcs.AI
|
Kenny Peng, Jon Kleinberg, Nikhil Garg |
We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent at...We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features prevalent in both, and sufficiently large SAEs recover both parent and child features. In contrast to conventional wisdom that SAE features are unstable and "split" as size increases, we find that these predictions hold on SAEs of sizes ranging from 512 to 131,072 trained on two large embedding models. From a theoretical perspective, our results suggest the promise of a scientific theory of representations based on atomic features. Practically, our results suggest the promise of scaling SAEs.
|
| 1284 |
Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL
2610.05807
|
cs.LGcs.AI
|
Hang Xiao, Janet Sung, Zhaoyi Li, Gangzhen Qian, Chuhong Xu |
Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, seque...Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, sequence, and semantic configurations on Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) logs. Random splits place several configurations near the average-precision ceiling, whereas group-disjoint HDFS and chronological BGL evaluation produce lower scores and different observed orderings. At a fixed BGL cutoff, parser choice spans 0.124 in semantic XGBoost mean average precision while preserving its lead over count XGBoost; the earliest rolling period reverses that ordering. A two-factor cross-system ablation contrasts source-only representations with offline transductive access to unlabeled target templates through the representation corpus and inverse document frequency: HDFS-to-BGL mean average precision moves from 0.191 with source-only access to 0.325 with union-corpus, target-IDF access, and the intermediate conditions reveal direction-dependent interactions in average precision and retrieval at fixed review budgets. Component-level profiling separates parsing and representation costs from classifier training, prediction, and storage. Together, these findings connect detector comparisons to the test population, preprocessing state, visible information, and measured pipeline stages, and identify the protocol fields needed alongside a score to support interpretable comparisons of log anomaly detection accuracy and resource use.
|
| 1285 |
Online AutoML: Evaluating Poisoning Attacks on Adversarial Training Defense Strategy in IoT Networks
2610.05810
|
cs.LG
|
Chukwunonso Henry Nwokoye, Khalil El-Khatib, Li Yang |
Machine learning (ML)-powered poisoning attack vectors are adversarial maneuvers whereby an attacker intentionally inserts, corrupts, or alters training data to distort an ML model's learning process. The objective is to diminish model efficacy, instill biases...Machine learning (ML)-powered poisoning attack vectors are adversarial maneuvers whereby an attacker intentionally inserts, corrupts, or alters training data to distort an ML model's learning process. The objective is to diminish model efficacy, instill biases, induce misclassifications, or include concealed backdoors that may be attacked during implementation. In streaming contexts, poisoning attacks pose significant risks since models perpetually update based on incoming streams of data. An assailant may incrementally introduce harmful samples into this data stream, leading the model to assimilate erroneous features over time without timely identification. Therefore, this study is aimed at evaluating the efficacy of the adversarial training (AT) defense approach against poisoning attacks (label flip and noise injection) using an online AutoML pipeline for Internet of Things (IoT) networks. Specifically, poisoning attacks (label flip and noise injection) were applied to streaming-capable AutoML learners (Hoeffding Tree (HT), Leveraging Bagging (LB), Adaptive Random Forest (ARF), Hoeffding Adaptive Tree (HAT), and Streaming Random Patches (SRP)). Under the strongest poisoning rate (PR = 1.0), AT-SRP achieved the highest F1-score against label flip poisoning (0.904), while AT-LB achieved the highest F1-score against noise-injection poisoning (0.933). Finally, several drift detection methods were used for rolling accuracy and prequential evaluation.
|
| 1286 |
Transporting Unsecured Stacked Payloads with a Quadrupedal Robot via Multi-Objective Reinforcement Learning
2610.05819
|
cs.LG
|
Nobuo Namura, Masayuki Hiromoto, Kento Uemura, Hironobu Sasaki, Kanata Suzuki |
Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation ...Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted board without dedicated payload sensors or active carrier mechanisms. To address this trade-off, we propose Payload-Adaptive Multi-Objective Reinforcement learning for Transportation (PAMORT). PAMORT trains a multi-objective base policy conditioned on a preference vector that weights locomotion and payload-stability reward groups, then trains a weight adjuster on the frozen policy to adapt this preference online from proprioception. In simulation, PAMORT achieves comparable or better overall transportation success than a corresponding single-objective baseline across different payload configurations, including an unseen three-box stack, despite training only with two boxes. Real-world experiments on a Unitree Go2 demonstrate zero-shot transfer to slopes and steps at or beyond the training difficulty, with mean success rates of 0.850 for PAMORT and 0.675 for the baseline across eight tasks. These results demonstrate robust unsecured-payload transportation with online adaptation of the locomotion--payload trade-off from proprioceptive information.
|
| 1287 |
Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents
2610.05833
|
cs.LGcs.AI
|
Tiffany Gu, Annie Guan, Manshu Huang, Nitin Rao, Siddhant Shah |
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We s...Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 69.0% with CacheBlend's token top-$k$ policy to 26.1%. When each prompt is evaluated after a different sequence of preceding requests, document-aligned recomputation improves fidelity to full prefill by 34.5-52.5 percentage points over token top-$k$, while both policies achieve approximately 5.7$\times$ median TTFT speedup. Our ablation study shows that, in our rolling-agent workload, contiguity is the main factor associated with robust selective recomputation.
|
| 1288 |
Adaptive-Shot Hybrid Quantum Anomaly Detection for Tactile Internet Security: Reliability-Aware Measurement Allocation Under Resource Constraints
2610.05835
|
cs.LG
|
Mubassir Serneabat Sudipto, Shakil Ahmed, Ashfaq Khokhar, Samir M. Iqbal |
Tactile Internet (TI) security analytics must balance reliable thresholded decisions with constrained computational and measurement resources. We study this tension for finite-shot hybrid quantum anomaly inference and introduce the Adaptive-Shot Variational Qu...Tactile Internet (TI) security analytics must balance reliable thresholded decisions with constrained computational and measurement resources. We study this tension for finite-shot hybrid quantum anomaly inference and introduce the Adaptive-Shot Variational Quantum Circuit (AS-VQC) policy. This validation-calibrated policy begins each record at 128 shots and cumulatively escalates through 256, 512, and 1024 shots only when the finite-shot anomaly score remains close to a validation-selected security threshold. The quantum scorer is evaluated as an off-path security analytics component rather than part of the haptic critical path. Using a 4,875-record CESNET-TimeSeries24-derived aggregate-flow benchmark, leakage-safe random, entity-group-disjoint, and temporal holdouts, and five trained quantum neural network (QNN) checkpoints per holdout, the primary AS-VQC-95 (beta = 0.95) policy averages 129.2, 276.9, and 131.2 shots per record, saving 87.4%, 73.0%, and 87.2% of the uniform 1024-shot baseline (Fixed-1024), respectively. The decision disagreement with analytic (exact-expectation) inference is 0.771%, 0.409%, and 0.635%, lower than both the uniform 128-shot baseline (Fixed-128) and a matched-budget shuffled-allocation control. Fixed-1024 remains more decision-stable, establishing a measurable reliability-resource trade-off rather than cost-free equivalence. A more conservative AS-VQC-99 (beta = 0.99) further reduces disagreement while using fewer than 512 average shots across all holdouts. These results show that finite quantum measurements can be treated as an inference resource and concentrated on boundary-sensitive TI-security decisions while exposing checkpoint-dependent escalation under unseen-entity conditions.
|
| 1289 |
AnchorPose for Geometry-Aware MOF Assembly through Meso-Grained Pose Generation
2610.05843
|
cs.LG
|
Zhonglong Peng, Rui Jiao, Chang Chen, Geng Zhong, Qiuliang Liu |
Predicting metal-organic framework (MOF) structures from given building blocks requires recovering their positions and orientations in a periodic crystal. The spatial effects of rotation errors are geometry-dependent and anisotropic. The same angular error can...Predicting metal-organic framework (MOF) structures from given building blocks requires recovering their positions and orientations in a periodic crystal. The spatial effects of rotation errors are geometry-dependent and anisotropic. The same angular error can produce different atomic displacements depending on block size, shape, and rotation axis. Angular error alone, without reference to the specific block geometry, therefore cannot fully describe the spatial consequences of a pose error. We introduce AnchorPose, a meso-grained pose generation framework that incorporates this geometric dependence into its generative representation. It represents each block through a small set of representative atoms, combines their local geometry with the current spatial state, and generates their coordinates with Bayesian Flow Networks. Known atom correspondences enable rigid alignment to recover complete building-block poses and return geometrically consistent points to the generation process. This design connects point-level spatial prediction with block-level structural constraints. Geometry participates in the pose state and its prediction, while rigid reconstruction preserves intra-block structure without treating all atomic coordinates as assembly variables. On the MOF benchmark, AnchorPose improves single-candidate match rates over the compared block-level and all-atom baselines.
|
| 1290 |
PhaseMatcher: Autoregressive Phase-Set Identification with Spectral Decomposition
2610.05844
|
cs.LG
|
Zhonglong Peng, Qiuliang Liu, Chang Chen, Geng Zhong, Qi Li |
Recovering complete phase sets from powder X-ray diffraction (PXRD) is challenging when weak-phase peaks overlap stronger signals. A natural strategy is to identify phases iteratively, removing the contribution of each identified phase from the observed patter...Recovering complete phase sets from powder X-ray diffraction (PXRD) is challenging when weak-phase peaks overlap stronger signals. A natural strategy is to identify phases iteratively, removing the contribution of each identified phase from the observed pattern before predicting the next. However, even after a phase is correctly identified, misestimating its contribution can distort the residual and cause subsequent errors. We introduce PhaseMatcher, an autoregressive framework for complete phase-set identification with physics-guided spectral decomposition. After each phase prediction, PhaseMatcher re-estimates the contributions of all selected phases and the residual from the original observation and all selected reference patterns, accounting for physically plausible variation between reference patterns and the corresponding phase contributions in the observation. The resulting residual guides subsequent phase identification, while a separate stopping module determines when the phase set is complete. On synthetic mixtures and controlled mixtures constructed from measured single-phase patterns, PhaseMatcher improves complete-set identification over the evaluated baselines. On PhaseMix-135K, it also estimates contributions and residuals more accurately than scalar subtraction.
|
| 1291 |
Finite-Sample Distribution Theory and Efficient Large-Scale Inference for Online Quantile Regression
2610.05869
|
cs.LG
|
Ziyang Wei, Jiaqi Li, Lan Wang, Wei Biao Wu |
This paper studies online quantile regression for large-scale and streaming data using Stochastic SubGradient Descent (SSGD) with constant learning rates. Classical offline inference for quantile regression is computationally and memory intensive. Existing wor...This paper studies online quantile regression for large-scale and streaming data using Stochastic SubGradient Descent (SSGD) with constant learning rates. Classical offline inference for quantile regression is computationally and memory intensive. Existing works of online inference for quantile regression provide only asymptotic guarantees and typically require sub-exponential tail conditions for distribution theory. To bridge these gaps, we introduce new techniques to prove a quenched central limit theorem (CLT) and finite-sample Gaussian approximation for SSGD under a finite-moment assumption. We further show that Ruppert-Polyak averaging with a constant learning rate has a non-vanishing bias and fails to satisfy CLT centering at the population target. Hence we propose suffix averaging to address this issue and establish its finite-sample Gaussian approximation. Based on these results, we provide an efficient online inference method for quantile regression that avoids covariance estimation. Numerical experiments show that our method achieves desirable empirical coverage rates and competitive performance compared to other inference methods. We also apply our approach to U.S. wage data to demonstrate its practical effectiveness.
|
| 1292 |
Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning
2610.05882
|
cs.LG
|
Lars Ankile, Perry Dong, Rohan Bhowmik, Aneesh Muppidi, David D. Yuan |
Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the obje...Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the objects and intervenes when the policy fails. We ask how to maximize improvement from a fixed budget of supervised episodes on high-precision manipulation tasks with wide ranges of object placements. We observe that failures can concentrate in a small subset of initial states, so uniform collection spends much of the operator's time on states the policy already handles. Mulligan makes the initial-state distribution a decision, starting each round's episodes at observed failures and untried states. To further improve data efficiency, we augment interactive imitation learning with a value function trained on all data, including failures that imitation discards. Across three real-world tasks evaluated on 2,550 held-out, blinded episodes and two simulated tasks, Mulligan outperforms uniform initial-state sampling at matched collection budgets, and combined with value-based action selection, HiL-IDQL+Mulligan, improves final real-task success by 10-34 percentage points. With operator interventions, the human-robot team completes 98% of collection episodes, remaining productive while the policy learns. Videos, code, and data are available at https://mulligan.page/.
|
| 1293 |
Incentive Alignment in Online Experimentation
2610.05922
|
cs.LGcs.AI
|
{Ermis Soumalias, Richard Mudd, Abbas Zaidi |
Evaluating the causal effect of new features is a central goal for online platforms. While recent literature addresses limited testing traffic via centralized portfolio optimization, this perspective abstracts away a critical institutional reality: experimenta...Evaluating the causal effect of new features is a central goal for online platforms. While recent literature addresses limited testing traffic via centralized portfolio optimization, this perspective abstracts away a critical institutional reality: experimentation is operationally decentralized. The experimenters who develop new features also dictate which hypotheses to test, and they are typically rewarded based on empirical average treatment effects that are prone to upward bias. Left unchecked, this principal-agent conflict can severely erode platform value, a structural failure that conventional centralized levers, such as significance thresholds and traffic budgets, cannot resolve. By reframing experimentation as an incentive design problem, we demonstrate that two practical mechanisms, sample splitting and shrinkage, can effectively bridge this gap. Sample splitting aligns incentives perfectly at a bounded traffic cost, while shrinkage consumes no additional traffic and guarantees that interventions with negative expected effects are strictly unprofitable to field.
|
| 1294 |
Combining Improvements in Uplink AI-RAN
2610.05936
|
cs.LG
|
Petteri Kela, Dani Korpi, Mikko Honkala |
One of the major transformative factors in 6G will be the integration of Artificial Intelligence (AI) to become a native part of Radio Access Network (RAN). While most physical-layer AI features have so far been evaluated in isolation using link-level simulati...One of the major transformative factors in 6G will be the integration of Artificial Intelligence (AI) to become a native part of Radio Access Network (RAN). While most physical-layer AI features have so far been evaluated in isolation using link-level simulations, their combined behavior in a realistic multi-cell, multi-UE deployment has remained largely unexplored. In this paper, we present system-level performance results when multiple uplink AI features are enabled together, achieved by integrating accurate link-level and system-level simulators. To infer state-of-the-art deep-learning-aided Multiple Input Multiple Output (MIMO) receivers under the dynamic allocations produced by a realistic uplink scheduler, we propose a mirrored data augmentation method that decouples receiver performance from scheduled allocation size. In addition to these Physical Layer (PHY) receiver features, we combine several recent advances in deep reinforcement learning to train uplink power control and link adaptation that outperform a heuristic baseline and further boost the gains obtainable from the AI receiver alone. The system-level results show that the combined AI features improve the mean uplink user throughput by roughly 27% compared to a non-AI baseline, confirming that the individual PHY and Medium Access Control (MAC) AI features provide complementary gains when deployed jointly.
|
| 1295 |
How (and How Not) to Use Data Augmentation in VLA Post-Training
2610.05994
|
cs.LG
|
Bram Grooten, Joaquin Vanschoren |
Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with r...Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with reinforcement learning (RL) has been shown to benefit robustness, but significant room for improvement remains. In this work, we systematically study the effect of image augmentation on VLA post-training. We find that it is crucial to augment only the critic module during RL updates, while leaving the actor's input clean during both rollouts and updates. For $\pi_{0.5}$ and GR00T N1.5 this raises out-of-distribution success on LIBERO-Plus by $7.8$ and $10.0$ points respectively, while augmenting the actor collapses training entirely. We investigate a range of augmentation types and strengths, and provide practical recommendations for improving generalization in VLA post-training.
|
| 1296 |
P3: Persistent Particle Planning for Constrained Diffusion Control
2610.06002
|
cs.LG
|
Hikmet Simsir, Mahyar Fardinfar, Ozgur S. Oguz |
Diffusion models provide expressive priors over trajectories, but adapting these priors to test-time constraints requires maintaining feasibility and consistency across successive control decisions. We introduce Persistent Particle Planning (P3), a sequential ...Diffusion models provide expressive priors over trajectories, but adapting these priors to test-time constraints requires maintaining feasibility and consistency across successive control decisions. We introduce Persistent Particle Planning (P3), a sequential Monte Carlo framework for diffusion control that maintains a weighted population of candidate plans across replanning steps. At each control step, P3 shifts and partially re-noises the candidate trajectories, refines them under the latest observation, and uses constraint-aware weighting and resampling to select among alternative continuations without retraining the diffusion model. We consider denoising and replanning as one Feynman--Kac particle system and analyze it under an idealized repair. We prove that the re-noising depth controls how reliably a kept plan stays on its route, and that keeping a rare, well-separated route takes far fewer plans than rediscovering it by sampling from scratch. Experiments under multiple test-time constraint configurations show that population reuse reduces route switching and improves success without constraint violations. Because P3 refines earlier plans instead of redrawing them, it also needs fewer denoising iterations per replan. On maze-navigation tasks, it plans faster than both regenerated populations and methods that correct a single sampled plan by constrained optimization. Code and pretrained models are available at https://github.com/p3-username/p3-anon.
|
| 1297 |
Polynomial neural surrogates for designing photonic quantum experiments
2610.06032
|
cs.LG
|
Rohit Chaurasiya, Xuemei Gu |
Physics simulators can support the discovery of quantum experiments by predicting the states generated by experimental configurations. When these simulators are computationally expensive, repeated simulator calls can limit the search for experiments that gener...Physics simulators can support the discovery of quantum experiments by predicting the states generated by experimental configurations. When these simulators are computationally expensive, repeated simulator calls can limit the search for experiments that generate a desired quantum state. Here, we develop a physics-inspired polynomial neural surrogate for PyTheus, a graph-based quantum-optics simulator, to predict quantum states and use it to design quantum experiments. Its polynomial activations are motivated by the relation between graph perfect matchings and the resulting state amplitudes. We train separate surrogate models for four-, six-, and eight-photon systems and show that they achieve higher prediction accuracy with fewer trainable parameters than standard multilayer perceptrons. We then use the trained surrogates for inverse design of GHZ, W, and linear-cluster states. For the larger systems, the surrogates also enable faster inverse design than direct optimization with PyTheus. These results suggest that incorporating the underlying physics into neural surrogates can provide an efficient approach to quantum experiment design.
|
| 1298 |
Last-Iterate Convergence Rate of Normalized Gradient Descent under H\"older Smoothness
2610.06070
|
cs.LG
|
Yuki Takezawa, Eduard Gorbunov |
Normalized gradient descent is a widely studied adaptive optimization method. Most existing analyses focus on the best iterate or a weighted average of the iterates, whereas practical implementations typically return the last iterate. In this paper, we study t...Normalized gradient descent is a widely studied adaptive optimization method. Most existing analyses focus on the best iterate or a weighted average of the iterates, whereas practical implementations typically return the last iterate. In this paper, we study the last-iterate convergence of normalized gradient descent for convex, $(\nu,M_\nu)$-H\"older-smooth objectives. For a constant stepsize, we establish an upper bound of $\mathcal{O}\bigl((\log^2(T)/T)^{(1+\nu)/2}\bigr)$, which contains a logarithmic overhead relative to the known $\mathcal{O}\bigl(T^{-(1+\nu)/2}\bigr)$ guarantees for the best and weighted-average iterates. For $\nu = 0$, this overhead is known to be unavoidable. We complement this analysis with numerical results based on the performance estimation problem (PEP), investigating the finite-horizon worst-case behavior in the smooth setting and whether the logarithmic overhead reflects an intrinsic limitation of constant-step normalized gradient descent. We then show that a linearly decreasing stepsize yields a last-iterate guarantee of $\mathcal{O}\bigl(T^{-(1+\nu)/2}\bigr)$, matching the order of the best-iterate/weighted-average guarantees without requiring knowledge of $\nu$ and $M_\nu$.
|
| 1299 |
Quantum data loading from the learned shared structure of real signals
2610.06076
|
cs.LG
|
Pablo Herrero G\'omez, Antonio Jimeno Morenilla, David Mu\~noz-Hern\'andez, Higinio Mora Mora |
Preparing quantum states from classical data can cost more than the computation they serve; most loaders tailor a circuit to each input. Here we show that the signals of a real dataset share structure that can be learned once and reused. Our quantum-native loa...Preparing quantum states from classical data can cost more than the computation they serve; most loaders tailor a circuit to each input. Here we show that the signals of a real dataset share structure that can be learned once and reused. Our quantum-native loader learns a low-dimensional description of a dataset and prepares every signal with one fixed circuit set by a few numbers. Across seven views of five public datasets it meets the targets of the strongest structured loader at equal gate cost with several times fewer numbers per signal. These numbers can be inferred from a random subset: in a preregistered blind replication the subset needed to come within ten per cent of full-signal accuracy stayed constant within a prespecified margin as signals grew sixteenfold, whereas the structured loader needed ever more. It declines what it cannot represent, covering fewer cases than that baseline and no electrocardiogram.
|
| 1300 |
Gaussian Universality and Its Breakdown in Tensor-Network Machine Learning
2610.06080
|
cs.LG
|
Shi-Tuan Wang, Zidu Liu, Li-Wei Yu |
Gaussian-process limits are powerful in describing overparameterized machine learning models, yet their validity in structured tensor-network architectures remains unclear. Here we analytically present a moment-based approach that identifies precise conditions...Gaussian-process limits are powerful in describing overparameterized machine learning models, yet their validity in structured tensor-network architectures remains unclear. Here we analytically present a moment-based approach that identifies precise conditions for the emergence and breakdown of Gaussian universality in tensor-network learning models, with a focus on matrix product states. We prove that in the large bond dimension limit, the learning models with both local and global observables converge to Gaussian processes, with explicit finite-size bounds on higher-order moment deviations. Whereas in the large physical dimension limit, the Gaussian universality no longer persists: while the models with local observables retain Gaussian-process behavior, those global cases exhibit persistent non-Gaussian corrections. Our results reveal that Gaussian-process behavior in tensor-network learning is controlled not only by parameter number, but also by architectural scaling, observable locality, and the spectral properties.
|
| 1301 |
Reinforcement Learning-Based Optimization of Workload-Aware Power Delivery Networks
2610.06148
|
cs.LG
|
Oran Hayes, Maria Pantazi-Kypraiou, Athanasios Tziouvaras, George Stamoulis, Anuj Pathania |
Power Delivery Networks (PDNs) are critical components of modern VLSI chips, providing stable voltage levels while satisfying electromigration (EM) and IR-drop constraints. Conventional PDN design methodologies typically rely on worst-case assumptions, often r...Power Delivery Networks (PDNs) are critical components of modern VLSI chips, providing stable voltage levels while satisfying electromigration (EM) and IR-drop constraints. Conventional PDN design methodologies typically rely on worst-case assumptions, often resulting in over-provisioned networks and inefficient use of resources. This paper presents a reinforcement learning-based framework for the optimization of workload-aware PDNs. The proposed methodology first generates workload-aware PDNs using architectural power traces obtained from system-level simulations. These power traces are mapped to spatial power density distributions, enabling adaptive allocation of PDN resources according to local current demand. A reinforcement learning agent then performs wire-width optimization to minimize PDN area while maintaining EM and voltage integrity constraints. Electrical and reliability metrics are obtained using SPICE-based circuit analysis and EM lifetime estimation. Experimental evaluation is performed on a dataset of workload-aware PDNs generated from 4-, 8-, and 16-core multiprocessor floorplans using PARSEC and SPLASH-2 benchmark workloads. Furthermore, the proposed Deep Q-Network (DQN)-based optimizer reduces the average normalized PDN area by 47\% while satisfying all EM and IR-drop constraints. Compared to simulated annealing, the proposed approach achieves comparable optimization quality while providing approximately 26$\times$ faster optimization.
|
| 1302 |
APOD: reasoning-guided agentic population ordinary differential equation discovery for pharmacological digital twins
2610.06227
|
cs.LGcs.AI
|
Romain Ferrara, Martin Soucail, Victor Gertner, Adil Moussali, Joris Cocquebert |
Establishing ordinary differential equations (ODEs) describing population data is a fundamental part of mathematical modeling in pharmacology, crucial to developing digital twins. However, doing so from sparse, noisy data is a slow, expert-driven task. Existin...Establishing ordinary differential equations (ODEs) describing population data is a fundamental part of mathematical modeling in pharmacology, crucial to developing digital twins. However, doing so from sparse, noisy data is a slow, expert-driven task. Existing automated methods either search a restricted model space or ignore population inter-individual variability. Here we introduce APOD (Agentic Population ODE Discovery), a language-model agent that iteratively reasons over biological knowledge and fit diagnostics in an open-ended search space to discover a population digital twin (PDT), i.e., a shared ODE system with between-subject variability. On synthetic pharmacokinetic and tumor-dynamics benchmarks, APOD recovered ground-truth structures in 94-100\% of runs, 12-fold faster in median than an established library-based search. On real cohorts it converged to valid structures, and proposed a PDT of radioligand-therapy-induced platelet dynamics that predicts thrombocytopenia from first-cycle data and simulates alternative dosing schedules that lower the predicted risk of toxicity.
|
| 1303 |
AUTOPILOT An Advanced Perception, Localization and Path Planning Techniques for Autonomous Vehicles Using YOLOv7 and MiDaS
2610.06232
|
cs.LG
|
Harshkumar Devmurari, Gautham Kuckian, Prajjwal Vishwakarma |
Self driving vehicles have emerged as a reliable technology that has the capability to transform transportation and mobility. The development of self driving cars requires significant advances in a number of areas, including perception, localization, decision ...Self driving vehicles have emerged as a reliable technology that has the capability to transform transportation and mobility. The development of self driving cars requires significant advances in a number of areas, including perception, localization, decision making, and control. This research paper is based on the project implementation of the combination of object detection using YOLO (You Only Look Once), depth sensing using MiDaS for the localization and perception of obstacles, perspective transform, and decision making for path planning in self driving cars. The contemporary state of the technology for object detection, depth sensing, localization, and path planning evaluates the performance of the combined system through simulations and experiments. The results show that the combination of YOLO and MiDaS provides a new robust system for object detection and depth sensing. This research paper contributes to the advancement of self driving car technology and provides new and innovative approaches to the perception and localization of obstacles in the environment. Keywords: YOLO, MiDaS, perception, localization, decision making
|
| 1304 |
Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies
2610.06235
|
cs.LG
|
Shaohan Jiang, Jiahang Cao, Qiduo He, Fengting Deng, Kun Wu |
Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers...Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through scene-preserving instruction interventions, using valid target substitutions, arbitrary nouns, and unrelated sentences while holding the scene fixed. We evaluate vision-language-action (VLA) policies and world-action models (WAMs) in simulation and in real-world experiments. Our analysis addresses three questions: (a) Does task success imply instruction following? When instructions request a different visible object, all evaluated policies predominantly approach and pick up the incorrect original target associated with the scene. (b) Is this failure caused by language insensitivity? Instruction perturbations affect task performance. A layerwise action lens shows intermediate action predictions respond to these perturbations. Linear probes accurately recover instructed targets, indicating modified instructions are encoded despite rarely determining target selection. (c) Why does encoded language fail to control action? Attention analysis indicates weak instruction-token contributions to action generation. Target-token attention can remain focused on the original object, revealing a mismatch between target encoding and visual grounding. UMAP and shared non-negative matrix factorization show target information remains accessible within representations increasingly organized by scene identity. Our findings expose a grounding gap concealed by nominal success and provide a diagnostic framework. They further establish a concrete criterion for progress: policies should reliably follow valid changes in user intent, even when they conflict with scene-favored behavior.
|
| 1305 |
Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries
2610.06269
|
cs.LGcs.AI
|
Chonghe Jiang, Ao Qu, Siyuan Liu, Ruoyun Ma, Zijian Zhou |
Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-tim...Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes expensive when reliable execution requires a large model, since training must maintain gradients, optimizer states, and policy statistics while repeatedly generating long, structured outputs. It also complicates credit assignment: outcome-level verifier feedback must jointly evaluate the high-level strategy and its low-level implementation. In this work, we introduce Guidance-TTT, which separates these roles. A compact guidance model is trained at test time to propose high-level strategic changes, while a frozen execution model implements them as complete executable solutions. At each step, the system selects a promising previously discovered solution, proposes a change, executes and verifies it, and updates only the guidance model using an adaptive group-relative RL objective. This concentrates test-time learning on short strategic decisions while retaining the implementation capability of a substantially stronger model without adapting it. Without web access, Guidance-TTT produces strong solutions across four distinct domains: combinatorial optimization (Polyomino Packing), heuristic programming (AHC058), machine learning (Lasso), and GPU kernel optimization (TriMul). Across these tasks, it outperforms the best solutions reported in prior work while remaining competitive with state-of-the-art results on public online leaderboards. Code is available at https://github.com/Human-Agent-Society/reef/tree/guidance-ttt-support.
|
| 1306 |
GAMBIT: Learning to Plan Continuous Multi-Robot Trajectories
2610.06290
|
cs.LGcs.AI
|
Rishabh Jain, Akmaral Moldagalieva, Lorenzo Magnino, Michael Amir, Keisuke Okumura |
GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to impr...GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to improve overall team performance. Such self-sacrificial behaviours are difficult to capture with manually designed heuristics, particularly in dense, interaction-rich environments. Focusing on double-integrator continuous dynamics, this work studies how to learn such coordinated heuristics over motion primitives for multi-robot trajectory execution. Our framework, GAMBIT, first learns coordinated motion-primitive selection through imitation learning and subsequently fine-tunes the policy through reinforcement learning. We further introduce a safeguarded rollout mechanism with backup trajectories that guarantees collision-free execution at all times. Experiments demonstrate that GAMBIT substantially outperforms a range of baselines, including centralised motion planners and decentralised reactive planners, while exhibiting strong scalability. In particular, it coordinates over a thousand robots with planning latency below a few hundred milliseconds in continuous domains.
|
| 1307 |
Sharp dimensional analysis of midpoint methods for Langevin sampling
2610.06308
|
cs.LG
|
Fan Chen, Sinho Chewi, Jianfeng Lu, Matthew S. Zhang |
We study deterministic and randomized midpoint discretizations of Langevin dynamics for a target $\pi \propto e^{-V}$, where $0 \prec \alpha I\preceq\nabla^2V\preceq\beta I$ and $\kappa=\beta/\alpha$. To achieve $\sqrt\alpha\,W_2\leqslant\varepsilon$, we show ...We study deterministic and randomized midpoint discretizations of Langevin dynamics for a target $\pi \propto e^{-V}$, where $0 \prec \alpha I\preceq\nabla^2V\preceq\beta I$ and $\kappa=\beta/\alpha$. To achieve $\sqrt\alpha\,W_2\leqslant\varepsilon$, we show that deterministic Heun uses at most $\widetilde O(\kappa^{4/3}d^{1/3}\varepsilon^{-2/3})$ gradient queries, and underdamped exponential midpoint uses $\widetilde O(\kappa^{5/4}d^{1/4}\varepsilon^{-1/2})$. The proofs exploit cancellation at stationarity and smoothing using techniques from Malliavin calculus, outperforming previous upper bounds based on standard couplings. At bounded condition number, a lower bound matches the $d$ and $\varepsilon$ powers of both deterministic methods. To contrast, for the randomized midpoint methods and Poisson midpoint with at least two grid points (both overdamped and underdamped variants), a simple Gaussian calculation yields a lower bound $d^{1/3}\varepsilon^{-1/3}$ to get an $\varepsilon$-close sample despite starting at a benign initialization. This shows surprisingly that in high dimensions, deterministic discretizations can outperform their random counterparts.
|
| 1308 |
Watermarking: from Impossibility to Auditable Compliance
2610.06317
|
cs.LG
|
Fernando Delbianco, Fernando Tohm\'e, Hugo Acciarri |
Article 50 (2) of the EU Artificial Intelligence Act requires providers of generative systems to make synthetic outputs machine-readable and detectable, while qualifying the effectiveness, interoperability, robustness, and reliability by technical feasibility,...Article 50 (2) of the EU Artificial Intelligence Act requires providers of generative systems to make synthetic outputs machine-readable and detectable, while qualifying the effectiveness, interoperability, robustness, and reliability by technical feasibility, cost, content-specific limits, and the state of the art. For free-form text, one important implementation route is the implementation of a generative watermarking procedure, which poses a compliance problem that is hard to address. Strong watermarking is impossible against adaptive removal, while ordinary edits attenuate statistical evidence, and unmarked human text may overlap distributionally with machine output. This article develops an auditable alternative. First, it defines a description-length robustness profile. A finite-sample bound shows that detectable bias decays and that the required sample size grows with the inverse square of the decay rate. This replaces an unidentified Shannon-entropy constant with collision entropy. Second, it constructs label-conditional conformal prediction sets with separate false-attribution and false-exclusion levels, reporting ``watermark supported,'' ``not supported,'' or ``inconclusive''. Coverage is obtained as a finite-sample result and is class-conditional under exchangeability. A small reproducible simulation of a tournament watermark confirms both claims and shows that the surviving-token rule overstates the tolerable edit rate roughly twofold. The resulting premarket certificate, signed detector report, and postmarket recalibration protocol operationalize the Commission's 2026 Code of Practice without claiming universal robustness.
|
| 1309 |
IGA-KAN: Isogeometric Analysis with Physics-Informed Closed-Form Kolmogorov-Arnold Networks for Forward and Inverse PDEs
2610.06348
|
cs.LG
|
Sima Naraghi, Kourosh Parand, Amirhossein Sadr, Dara Rahmati |
Isogeometric analysis (IGA) solves partial differential equations accurately on exact NURBS geometry, whereas neural solvers are mesh-free but often orders of magnitude less accurate and typically trained by non-convex optimization without error control. We pr...Isogeometric analysis (IGA) solves partial differential equations accurately on exact NURBS geometry, whereas neural solvers are mesh-free but often orders of magnitude less accurate and typically trained by non-convex optimization without error control. We propose IGA-KAN, which uses local Kolmogorov-Arnold networks, fitted in closed form, to improve the IGA solution instead of replacing it. An IGA Galerkin solve produces u_h; on every knot-vertex patch a Kolmogorov-Arnold ridge model is fitted to the strong form of the equation, the exact boundary data and u_h, and the models are blended by IGA hat functions. With fixed inner functions the fit is one batched linear least-squares problem, without optimizer, learning rate or initialization. An a posteriori safeguard, motivated by a maximum-principle bound, decides where local models are used, keeping the IGA solution elsewhere. On eight benchmarks with exact solutions, five from the literature and one also posed on a domain fitted to a brain slice from MRI, the method reduces the error of IGA, at an unchanged number of Galerkin unknowns, by factors of 4.2 to 90 in L^2 and 4.1 to 220 in H^1 on the reference meshes, and its L^2 error is 6 to 6x10^4 times smaller than that of the best Kolmogorov-Arnold network trained from scratch on the same equations with a fixed budget. In an inverse problem it recovers an unknown constant source from one noise-free observation 167 times more accurately than IGA. The gain is attributed to the superconvergence of local averages of the Galerkin solution.
|
| 1310 |
Latent Similarity Gaussian Processes: A Theory-Grounded Approach to Personalized Suicide-Risk Forecasting for Clinical Decision-Support
2610.06355
|
cs.LG
|
Yaniv Yacoby, Weiwei Pan, Hope Neveux, Taylor C. McGuire, Franchesca Castro-Ramirez |
Forecasting suicide risk is difficult due to the high heterogeneity of patients and the low base rate of suicide-related events (SREs). We present Latent Similarity Gaussian Processes (LSGPs), which embed patients in a continuous latent space to jointly model ...Forecasting suicide risk is difficult due to the high heterogeneity of patients and the low base rate of suicide-related events (SREs). We present Latent Similarity Gaussian Processes (LSGPs), which embed patients in a continuous latent space to jointly model similarity and forecast risk. By selectively drawing information from latent peers, LSGPs better capture individualized risk trajectories, generalizing nomothetic (pooled), idiographic (per-patient), and hierarchical frameworks. Our contributions are: (1) an identifiable two-channel Similarity Kernel; (2) proof that the standard model-fitting algorithm, mean-field variational inference, collapses LSGPs to nomothetic models, along with a fix; and (3) empirical results on intensive longitudinal suicide data showing LSGPs outperform nomothetic, idiographic, and hierarchical models for next-week risk forecasting, with the largest gains in forecasting first-occurrence SREs.
|
| 1311 |
Correct Verdicts, Flawed Reasoning: Structured Auditing of LLM-based Vulnerability Reasoning
2610.06366
|
cs.LG
|
Boyue Caroline Hu, Kaivalya Ahir, Ronghao Ni, Limin Jia |
Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought...Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to obscure logical leaps, hallucinated execution steps, and internal inconsistencies behind plausible prose. Our manual audit reveals that approximately 60% of correct vulnerability verdicts are accompanied by fabricated or unverifiable claims, and free-form explanations allow reasoning errors to evade LLM-as-a-judge evaluation. We present Vulnerability Explanation Reasoning Auditor (VERA), an automated framework for auditing LLM vulnerability reasoning. Rather than accepting free-form text, VERA asks models to output a Structured Reasoning Record (SRR) encoding tracked pointers, memory operations, and state transitions in machine-readable fields. A multi-stage judge audits each SRR against eight reasoning failure modes using deterministic checks, with LLM calls reserved for semantic interpretation. The standardized SRR schema also enables automated mutation testing to benchmark judges at scale without human annotation. Our evaluation shows reasoning flaws occur in correct verdicts just as frequently as incorrect ones, and VERA exposes 87% of reasoning errors that free-form LLM-as-judge systematically miss.
|
| 1312 |
Valid Stopping in Adaptive Generator-Verifier Loops
2610.06432
|
cs.LGcs.AI
|
Mahmoud Hegazy, Michael I. Jordan, Aymeric Dieuleveut |
Numerous agentic workflows are based on a generator-verifier loop: a generator proposes candidates, a cheap verifier scores them, and the workflow terminates when a proposal is verified as good enough. The verifier typically proxies a more costly ground-truth ...Numerous agentic workflows are based on a generator-verifier loop: a generator proposes candidates, a cheap verifier scores them, and the workflow terminates when a proposal is verified as good enough. The verifier typically proxies a more costly ground-truth oracle, and as the generator searches adaptively against it, false acceptances may accumulate. Proposals can pass the proxy but fail under the costlier ground-truth check. We study when to stop these loops while controlling the false discovery rate of the accepted proposals. Our construction introduces tools of independent interest in distribution-free statistical testing and conformal risk control, including analysis of $e$-values constructed through index betting and a novel conformal risk control procedure for non-monotone losses. We validate the approach in synthetic settings and on a protein-design benchmark.
|
| 1313 |
KESurv: A Kernel Ensemble Method for Patient-Specific Survival Prediction
2610.06434
|
cs.LG
|
Rahul Goswami |
Predicting patient-specific survival functions is crucial for clinicians in making informed decisions about patient care and treatment strategies. Among the various models available, the Survival Forest has demonstrated significant effectiveness in numerous sc...Predicting patient-specific survival functions is crucial for clinicians in making informed decisions about patient care and treatment strategies. Among the various models available, the Survival Forest has demonstrated significant effectiveness in numerous scenarios. In this work, we propose an ensemble method that leverages the strengths of the Survival Forest as the master model, complemented by several base models. This ensemble incorporates the Beran estimator, a type of kernel estimator, to enhance predictions of patient-specific survival curves. We evaluated the performance of our proposed model using four distinct healthcare datasets. The results highlight the superiority of our ensemble method over baseline models in both calibration and ranking across most datasets. The findings suggest that our approach offers a more accurate and reliable estimation of patient-specific survival functions, providing a valuable tool for clinical decision-making.
|
| 1314 |
ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing
2610.06514
|
cs.LGcs.AI
|
Fan Li, Xiangyu Gao, Zixuan Liu, Tong Li, Chuanpu Fu |
The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source o...The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this question. We introduce ANT (Agent Network Traffic), a dataset providing agent behavior information at risk, scenario, and behavior primitive granularities alongside network traffic. ANT contains 3,114 execution episodes across 20 tasks and five scenarios, comprising 276,417 bidirectional flows and 40,049 behavior primitive segments organized into 47 macro groups. We establish a benchmark for agent risk identification, scenario recognition, and behavior primitive classification using 13 representative traffic analysis baselines. The results show that existing methods recover useful but uneven behavioral signals. They struggle to identify risk when malicious workflows resemble benign tasks and to distinguish scenarios with similar traffic patterns. Primitive classification is more reliable for frequent macro groups and those with distinctive traffic patterns than for rare or semantically similar groups. ANT provides a common basis for developing more precise auditing and forensic analysis of agent behavior from network traffic. Our data and code are available at https://anonymous.4open.science/r/ant-main-suite-7BC0/.
|
| 1315 |
MIRT: Transformers for Truthful Generative Auctions with Whole-feed Permutation Externalities
2610.06559
|
cs.LG
|
Ali Elahi, Ermis Soumalias, Jason Cheuk Nam Liang, Daniel Yao, Michael J. Curry |
Modern online platforms commonly rank ads and organic content separately before blending them into a feed displayed to the user, overlooking externalities: an item's click-through rate depends on its surrounding content, not only on its own position. Recent le...Modern online platforms commonly rank ads and organic content separately before blending them into a feed displayed to the user, overlooking externalities: an item's click-through rate depends on its surrounding content, not only on its own position. Recent learning-based feed generation mechanisms model some of these cross-type interactions to globally optimize for the whole feed's welfare. However, these approaches either fix the ordering of organic content, or lack exact strategyproofness guarantees for bidders. To combat these shortfalls, we introduce the Maximal-in-Range Transformer (MIRT) mechanism class, which uses a transformer to generate a range of candidate feeds that jointly order ads and organic content, and selects the welfare-maximizing feed in the range. However, there is a tension: strategyproofness requires the generated range to be bid-independent, even though a candidate feed's welfare depends linearly on the bids. Our key technical contribution is a reinforcement learning approach that incorporates both candidate generation and bid-aware selection into training, enabling a bid-independent transformer to learn to generate high-welfare ranges by accounting for both individual feed quality and the collective quality of the range. Additionally, we bound the pseudo-dimension of the MIRT class under hard attention, showing that near-optimal expected welfare is learnable with sample complexity polynomial in the transformer size and only logarithmic in the range size. Empirically, MIRT outperforms the previous non-strategyproof state-of-the-art feed models while remaining exactly strategyproof. Our results show that transformer-based auctions can deliver externality-aware whole-feed optimization without sacrificing exact incentive compatibility, removing a major obstacle to their practical deployment.
|
| 1316 |
The Surrogate Is Not the Reward: Post-Surrogate Primary-Outcome Acquisition in Contextual Bandits
2610.06610
|
cs.LG
|
Kyungbok Lee, Michael R. Kosorok |
We study contextual bandits in which a surrogate is observed after the action but before the learner decides whether to acquire the primary outcome that defines action value and regret. The value of acquiring the primary outcome depends on both decision releva...We study contextual bandits in which a surrogate is observed after the action but before the learner decides whether to acquire the primary outcome that defines action value and regret. The value of acquiring the primary outcome depends on both decision relevance (how much the current outcome matters for comparing policies) and the residual uncertainty after observing the surrogate. The Audited Surrogate Bandit (ASB) learns a contextual policy while allocating a budget of $B$ primary-outcome acquisitions over $T$ rounds. ASB sets a pre-surrogate acquisition level from current decision relevance and, after observing the surrogate, redistributes that level using an estimate of that residual uncertainty. For a finite class of $N$ policies over $K$ actions, ASB incurs $\widetilde O[\sqrt{KT\log N}\{1+\sqrt{T/B}\}]$ regret relative to the best policy in the class. In a two-action family where the surrogate does not reveal the better action, a learner that observes the surrogate before deciding whether to acquire can achieve bounded regret, whereas any learner that must decide before seeing the surrogate incurs $\Omega(T/B)$ worst-case regret under the same budget. Synthetic experiments show that both acquisition factors matter: ASB has lower regret than variants using only decision relevance or only residual uncertainty. On a KuaiRec benchmark of user-video interactions, the regret gap relative to relevance-only acquisition widens and then narrows as the budget grows.
|
| 1317 |
On the Cardinality of Optimal Representations in the Binary-Source Information Bottleneck
2610.06627
|
cs.LG
|
Dier Tang, Jun Chen |
The information bottleneck (IB) seeks a representation $U$ of a source $X$ that retains as much information as possible about a target $Y$, subject to a constraint on $I(U;X)$. A classical argument shows that it suffices to consider representations with at mos...The information bottleneck (IB) seeks a representation $U$ of a source $X$ that retains as much information as possible about a target $Y$, subject to a constraint on $I(U;X)$. A classical argument shows that it suffices to consider representations with at most $|\mathcal{X}|+1$ symbols, and this bound is known to be tight whenever $|\mathcal{X}| \geq 3$. We show that the binary case behaves differently: if $X$ is binary and $Y$ is finite, then for every joint distribution of $(X,Y)$ and every rate constraint, the IB optimum is attained by a binary $U$. Hence the bound $|\mathcal{U}| \leq |\mathcal{X}|+1$ sharpens to $|\mathcal{U}| \leq |\mathcal{X}|$ for binary sources. The proof combines a separating hyperplane argument with the observation that, for a binary source, the ratio of the second derivatives of the two entropy functions involved is concave.
|
| 1318 |
Inverse Cross-spectral Neural Networks for Multivariate Time Series
2610.06630
|
cs.LG
|
Lorenzo Marinucci, Leonardo Di Nino, Gabriele D'Acunto, Paolo Di Lorenzo, Sergio Barbarossa |
CoVariance Neural Networks and their extensions have emerged as effective tools for processing multivariate data, deriving graph shift operators directly from second-order statistics. These architectures, however, are designed for independent and identically d...CoVariance Neural Networks and their extensions have emerged as effective tools for processing multivariate data, deriving graph shift operators directly from second-order statistics. These architectures, however, are designed for independent and identically distributed observations and do not fully capture the joint structure of temporal and cross-variable dependencies in multivariate time series. In this work, we introduce Inverse Cross-Spectral Neural Networks (iCSNNs), a class of graph neural networks for stationary multivariate time series whose shift operators are the inverse cross-spectral density (iCSD) matrices. These operators encode frequency-specific conditional relationships among variables, exploiting the decomposition provided by the spectral representation theorem. Leveraging spectral smoothness, frequencies are grouped into bands sharing a single iCSD operator, yielding a compact parametrisation that retains the frequency-dependent structure of the process. We further propose a joint learning procedure to estimate both the Fourier-domain dependence structure and the iCSNN parameters, adapting the iCSD operators to the downstream task. When tested on synthetic data, iCSNN outperforms baselines from different methodological families.
|
| 1319 |
Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood
2610.06691
|
cs.LGcs.SD
|
Seunghwan Kim, Jinyong Kim, Sooyoung Yang, Youngjin Ko, Myungjoo Kang |
Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean ta...Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises--Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.
|
| 1320 |
A Solvable Model of Adaptive Learning Rate Rescaling: Acceleration, Stability & Scaling
2610.06701
|
cs.LG
|
Itay Lavie, Clarissa Lauditi, Cengiz Pehlevan |
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear. We isolate this mechanism by studying normalized SGD in a random-feature ...A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear. We isolate this mechanism by studying normalized SGD in a random-feature model with power-law teacher and data covariance. Fixed-norm updates induce an effective learning rate that grows as gradients shrink. We derive a dynamical mean-field theory (DMFT) describing the joint dependence of the loss on training time, model width and batch size. Normalization initially accelerates SGD, mapping the power-law exponent $r_{\rm SGD}<1$ to $2r_{\rm SGD}/(1-r_{\rm SGD})$, with exponential convergence at $r_{\rm SGD}=1$ and formal finite-time convergence for $r_{\rm SGD}>1$. At finite step size, however, the same feedback ultimately breaks the acceleration and leads to marginal stability. The late-time theory yields width-limited, edge-of-stochastic-stability (EoSS), and deterministic edge-of-stability (EoS) regimes. These phases determine when larger batches or wider models reduce serial training time at comparable compute. We quantify in which of these phases increased batch size or width can compensate the excess compute use per step by fewer optimization steps to target loss. Linearized ResNet experiments on CIFAR-5M support the predicted acceleration, breakdown, and resource-scaling trends. Together, these results connect normalization-induced acceleration, EoS effects, and width--batch allocation within a solvable theory.
|
| 1321 |
Out-of-control Hamiltonian Learning
2610.06709
|
cs.LG
|
Weiyuan Gong, Muzhou Ma, Sitan Chen, Jordan Cotler, Hsin-Yuan Huang |
Learning the Hamiltonian of a many-body system from its dynamics is a central task in quantum science, yet the algorithms with the strongest provable guarantees assume some level of quantum control--fast, arbitrary single-qubit gates interleaved with time evol...Learning the Hamiltonian of a many-body system from its dynamics is a central task in quantum science, yet the algorithms with the strongest provable guarantees assume some level of quantum control--fast, arbitrary single-qubit gates interleaved with time evolution, and measurements in arbitrary bases--that is beyond the capabilities of near-term analog quantum simulators. Motivated by analog atom- and ion-based platforms, we study Hamiltonian learning under minimal access models. Uniform state preparation and measurements: We first consider the setting where in every experiment, one can rotate each qubit to the same state, perform short-time evolution, and measure every qubit in the same basis. Surprisingly, we show that for generic 2-local Hamiltonians on any interaction graph, all of the parameters can be reconstructed from such experiments. Computational basis state preparation and measurements: We then consider a similarly constrained setting, but where state preparation and measurement are restricted to the computational basis. For nearest-neighbor Hamiltonians with only Pauli $X/Z$ interactions, a class which captures contemporary Rydberg atom platforms, we show that over 1D and 2D rectangular lattices, all of the parameters can be reconstructed from such experiments up to unavoidable gauges. Our protocols introduce new techniques for solving structured polynomial systems over an extensive number of parameters. Taken together, our results suggest that one can learn a great deal from the dynamics of quantum many-body systems even under the most stringent experimental constraints.
|
| 1322 |
BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
2610.06748
|
cs.LGcs.AI
|
Ziyan Wang, Shuqing Shi, James Oldfield, Samuele Marro, Jialin Yu |
In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model (LLM) agents now act for users, raising risks to their money, privacy, and reputatio...In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model (LLM) agents now act for users, raising risks to their money, privacy, and reputation. We introduce BazaarBench, a simulated C2C marketplace and benchmark for evaluating the safety of these agents. It tracks ownership, item condition, and commitments across transactions, combining record checks with rubric-based LLM judgments to identify six failure types across five stages. We run three base markets for 30 simulated days, each with 100 agents using one model and inventories drawn from a public eBay sample. Across 45 continuations, we evaluate five models under ordinary instructions, deadline pressure, or adversarial instructions to exploit other traders. Each continuation runs for seven simulated days from a copy of a market's day-30 state. The tested model controls the same 20 selected agents, retaining their personas, inventories, and histories, while the other 80 keep the base model. All five models attempt to promise the same item to multiple buyers under ordinary instructions. Adding targets and deadlines increases these attempts for every model. Under adversarial instructions, the share of tested sellers' committed transactions completed despite unavailable items or overstated conditions rises from 15.4% to 33.4%, reaching 55.5% for GPT-5.4. Averaged across models and markets, simulated weekly earnings per tested agent rise from USD 20 under ordinary instructions to USD 33 under adversarial instructions. Most of the increase comes from items the sellers never held. We release the simulator, saved market states, evaluation code, and records covering 357,608 agent model calls for evaluating new models and developing safer marketplace agents.
|
| 1323 |
Singular parameters and missing limits in neural PDE solvers
2610.06770
|
cs.LG
|
Daniel Fern\'andez |
Neural solvers for partial differential equations (PDEs) can approach an accurate solution while their parameters grow without bound. In such cases, the limiting solution may have no finite representation in the chosen model, leaving the best loss unattained. ...Neural solvers for partial differential equations (PDEs) can approach an accurate solution while their parameters grow without bound. In such cases, the limiting solution may have no finite representation in the chosen model, leaving the best loss unattained. Our analysis connects missing limits in deep neural tanh- networks to unbounded hidden parameters or increasingly redundant neurons. For a class of models built from translated kernels, we describe the missing functions and recover them by adding kernel derivatives to the model. This completion makes the best approximation attainable under standard assumptions. Numerical studies follow the associated parameter growth and explore how completion affects PDE optimization.
|
| 1324 |
On Learning Optimal Corners in Orthogonal Partially Observable Cooperative Guard Art Galleries
2610.06777
|
cs.LG
|
Yassin Ben Mansour, Edwin Meriaux |
The CADENCE algorithm solves the Partially Observable Cooperative Guard Art Gallery Problem (POCGAGP) with formal coverage and connectivity guarantees, but leaves unspecified which valid corner each agent should be deployed to, a choice that strongly affects e...The CADENCE algorithm solves the Partially Observable Cooperative Guard Art Gallery Problem (POCGAGP) with formal coverage and connectivity guarantees, but leaves unspecified which valid corner each agent should be deployed to, a choice that strongly affects efficiency. We introduce two learned corner-selection heuristics that preserve these guarantees: a CNN scoring candidates on a grid encoding, and a GATv2 network trained with Deep Q-Learning (DQN) on a visibility graph. Across 7,500 runs on random orthogonal environments (50x50 to 250x250), our heuristics outperform baseline CADENCE in both steps to full coverage and peak agent count, with gains growing with scale, and improve on Incremental Self-Deployment (ISDA) baselines in agent utilization while providing guarantees ISDA lacks. Learned corner selection thus improves CADENCE in speed and agent utilization at no cost to its formal properties.
|
| 1325 |
How to scale your HEP ML models: A recipe for robust architecture comparisons at scale
2610.06784
|
cs.LG
|
Matthias Vigl, Nikita Pond, Jackson Barr, Alexander Froch, Dan Guest |
Much of the recent progress in machine learning domains such as language models has come from scaling laws that predict performance as a function of training effort. In high-energy physics (HEP) similar behavior has now been observed. To aid further study, we ...Much of the recent progress in machine learning domains such as language models has come from scaling laws that predict performance as a function of training effort. In high-energy physics (HEP) similar behavior has now been observed. To aid further study, we present a systematic procedure to derive robust scaling laws and compare design choices on the relevant budget axes for HEP tasks. We first validate the full scaling trajectory on toy problems and then apply the procedure to multi-task transformers on the ~11 billion-jet ATLAS JetSet2 dataset, in both the compute- and data-constrained regimes. For the latter, we predict, to the best of our knowledge for the first time, the jointly optimal model size, training horizon, learning rate and batch size under early stopping. At compute-optimal scaling, we recover a near-equal $\sqrt{C}$ dependence of model and dataset size, and find that auxiliary objectives lower the primary jet-classification loss at equal compute budget. Expanding the inputs toward lower-level data systematically lowers the loss while leaving the scaling exponent nearly unchanged. The onset of the power-law regime is itself set by scale: below a threshold in dataset size the loss carries little information about high-compute scaling, underscoring the value of large, high-quality full-simulation datasets as a foundation for scaling studies and the development of foundation models in HEP.
|
| 1326 |
A Response Theory Probe for Learned Stochastic AI Simulators, Tested on Lorenz-63
2610.06798
|
cs.LG
|
Jo\~ao B\"oger, Simon Driscoll, Niccol\`o Zagli, Valerio Lucarini, Francisco Camara Pereira |
Machine-learning emulators of chaotic and stochastic systems are usually validated on forecast skill and long-run statistics. Neither certifies that an emulator responds correctly to forcing, the property that projection and attribution studies rely on. Linear...Machine-learning emulators of chaotic and stochastic systems are usually validated on forecast skill and long-run statistics. Neither certifies that an emulator responds correctly to forcing, the property that projection and attribution studies rely on. Linear response theory makes this testable: the forced response follows from unperturbed correlations through a generalized fluctuation-dissipation relation, and decomposes over the stochastic Ruelle-Pollicott resonances of the Koopman generator. Building on the Koopmanism Response framework, we turn this into a calibrated, mode-resolved test for learned surrogates: each surrogate rollout passes or fails each check, and failure rates are compared with those of independent realizations of the true system. On stochastic Lorenz-63, a three-variable toy model, we evaluate SINDy, an MLP, a reservoir computer, a neural ODE and a neural SDE with learned diffusion, over up to 80 rollouts each. A sparse-regression model with the correct library passes every check at rates consistent with the true system. Invariant-statistics fidelity and response fidelity dissociate in both directions: a quarter of reservoir-computer rollouts pass every invariant-statistics check and match the static susceptibility $\chi(0)$, yet misrepresent the slow relaxation modes, while the neural ODE and SDE rarely meet the invariant-statistics floor but recover those modes in three quarters of rollouts. As expected of a time-integrated quantity dominated here by fast relaxation, $\chi(0)$ does not separate these cases. For a fixed network, the training formulation (one-step drift, flow map, or multi-step through the integrator) decides which of these properties it gets right.
|
| 1327 |
Finding Gaussian Structure in Bosonic States
2610.06810
|
cs.LG
|
Alvan Arulandu, Sitan Chen, Ziyun Chen, Jerry Li, Eric Ma |
We study agnostic tomography of pure bosonic Gaussian states: given copies of an arbitrary $n$-mode bosonic state $\rho$, the goal is to output a pure Gaussian state whose infidelity with $\rho$ is at most $\mathrm{opt} + \epsilon$, where $\mathrm{opt}$ is the...We study agnostic tomography of pure bosonic Gaussian states: given copies of an arbitrary $n$-mode bosonic state $\rho$, the goal is to output a pure Gaussian state whose infidelity with $\rho$ is at most $\mathrm{opt} + \epsilon$, where $\mathrm{opt}$ is the minimum infidelity achievable by any pure Gaussian state. We give efficient protocols achieving this in both the high and low fidelity regimes. When $\mathrm{opt}$ is below some universal constant, our protocol has runtime and copy complexity which is strongly polynomial in $n, 1/\epsilon$ and $\log \log E$, where $E$ is the energy of the closest pure Gaussian state. For arbitrary $\mathrm{opt}$, our protocol uses $(n+1)^{\mathrm{poly}(1/\epsilon)} \mathrm{poly}\left(1+\log\log(E)\right)$ copies and runtime. As a corollary, we obtain the first truly tolerant Gaussianity testing protocol for distinguishing whether $\mathrm{opt} > c + \epsilon$ or $\mathrm{opt} < c - \epsilon$, for any threshold $c\in(0,1)$. We also prove $\mathrm{poly}(n,1/\epsilon)$ runtime is impossible, unless $\mathrm{NP}\subseteq\mathrm{BQP}$. Our protocols follow a shared paradigm: first, we iteratively use general Gaussian measurements combined with techniques from classical robust statistics to obtain a good warm start estimate, then we leverage non-Gaussian measurements to refine this warm start using convex and non-convex optimization methods. Interestingly, we prove that non-Gaussian measurements are necessary to match the strong agnostic guarantees we obtain, and in fact these guarantees are provably superior to what is possible for robustly estimating classical Gaussians.
|
| 1328 |
Direct Intermediate Initialization for Tilted Diffusion Samplers
2610.06834
|
cs.LG
|
Gregory D. Bellchambers |
Some diffusion posterior samplers construct Gaussian-tilted intermediate distributions along the reverse process. We observe that these targets can be pulled back to clean-space posteriors with weaker conditioning, with samples transported analytically to the ...Some diffusion posterior samplers construct Gaussian-tilted intermediate distributions along the reverse process. We observe that these targets can be pulled back to clean-space posteriors with weaker conditioning, with samples transported analytically to the corresponding noisy-space target through a Gaussian bridge. For the sequential Monte Carlo (SMC) sampler MCGDiff, the effective observation variance of this pulled-back problem is up to twice the diffusion-noise variance. We exploit this structure to initialize MCGDiff directly at an intermediate time: an approximate solver samples the softened clean-space posterior, the Gaussian bridge maps these samples to the tilted target, and only the remaining SMC suffix is run. This trades asymptotic consistency for finite-particle performance. With moment-matching posterior sampling (MMPS) as the solver, the hybrid improves sliced Wasserstein distance by roughly $2\times$ at matched particle count on a structured Gaussian-mixture inverse problem, and by more than an order of magnitude when the posterior-relevant mode is rare under the prior. A prior-initialization control, which retains the bridge but drops the clean-space conditioning, shows that on MCGDiff's standard Gaussian-mixture benchmark most of the improvement is insensitive to the conditioning. Conditioning the initialization gives a further consistent gain on the structured problem, and becomes decisive on a rare-mode problem, where resampling cannot repopulate a mode absent from the initial population.
|
| 1329 |
Understanding Certified Training with Interval Bound Propagation
2306.10426
|
cs.LGcs.AI
|
Yuhao Mao, Mark Niklas M\"uller, Marc Fischer, Martin Vechev |
As robustness verification methods are becoming more precise, training certifiably robust neural networks is becoming ever more relevant. To this end, certified training methods compute and then optimize an upper bound on the worst-case loss over a robustness ...As robustness verification methods are becoming more precise, training certifiably robust neural networks is becoming ever more relevant. To this end, certified training methods compute and then optimize an upper bound on the worst-case loss over a robustness specification. Curiously, training methods based on the imprecise interval bound propagation (IBP) consistently outperform those leveraging more precise bounding methods. Still, we lack an understanding of the mechanisms making IBP so successful. In this work, we thoroughly investigate these mechanisms by leveraging a novel metric measuring the tightness of IBP bounds. We first show theoretically that, for deep linear models, tightness decreases with width and depth at initialization, but improves with IBP training, given sufficient network width. We, then, derive sufficient and necessary conditions on weight matrices for IBP bounds to become exact and demonstrate that these impose strong regularization, explaining the empirically observed trade-off between robustness and accuracy in certified training. Our extensive experimental evaluation validates our theoretical predictions for ReLU networks, including that wider networks improve performance, yielding state-of-the-art results. Interestingly, we observe that while all IBP-based training methods lead to high tightness, this is neither sufficient nor necessary to achieve high certifiable robustness. This hints at the existence of new training methods that do not induce the strong regularization required for tight IBP bounds, leading to improved robustness and standard accuracy.
|
| 1330 |
CHAOSMINING: Benchmarking Post-Hoc Attribution with Sparse Informative Features in High Dimensions
2406.12150
|
cs.LGcs.AI
|
Ge Shi, Fangyi Liu, Ziwen Kan, Menglin Liu |
Post-hoc attribution is widely used to identify important model inputs, but evaluating whether these attributions identify truly informative features is difficult because real datasets rarely provide reliable ground truth. We introduce a multimodal benchmark c...Post-hoc attribution is widely used to identify important model inputs, but evaluating whether these attributions identify truly informative features is difficult because real datasets rarely provide reliable ground truth. We introduce a multimodal benchmark containing symbolic tabular, vision, and audio tasks with known informative feature sets. In the main benchmark conditions, informative variables, spatial regions, or channels occupy fixed input coordinates while the remaining inputs provide irrelevant or distracting information. We use the benchmark to study how attribution quality depends on predictive performance, irrelevant-feature burden and structure, model configuration, and attribution mechanism, while separately measuring identification, stability, and computational cost. In most symbolic-data sweeps, informative-set identification co-varies with predictive performance, while the relative ordering of attribution methods remains largely stable. Across modalities, no method dominates all architectures and conditions, and greater attribution complexity does not consistently improve identification. Simple gradient attribution is often competitive at lower computational cost, while the vision and audio results show that architecture and the form of irrelevant content materially affect attribution quality.
|
| 1331 |
Online Linear Programming with Batching
2408.00310
|
cs.LG
|
Haoran Xu, Peter W. Glynn, Yinyu Ye |
We study Online Linear Programming (OLP) with batching. The planning horizon is cut into $K$ batches, and decisions on orders can be delayed to the end of their associated batch. The ability to delay decisions improves operational performance, as measured by r...We study Online Linear Programming (OLP) with batching. The planning horizon is cut into $K$ batches, and decisions on orders can be delayed to the end of their associated batch. The ability to delay decisions improves operational performance, as measured by regret. We study two questions: (1) What is a lower bound on the regret as a function of $K$ and the length of the planning horizon? (2) Which algorithms can achieve this regret lower bound? This paper analyzes these questions when the distribution of the reward has a continuous support. We provide an $\Omega(\log K)$ regret lower bound in the single-resource case, and we provide pricing algorithms having an $O(\log K)$ regret in the setting with Poisson arrivals and multiple types of resources. All the algorithms update the prices at most $K$ times and only delay orders of the first and the last batches. All the regret bounds are independent of the length of the planning horizon. Finally, we study a more realistic large support setting where the number of distinct order types is finite but scales with the total number of orders. We prove that our $\Theta(\log K)$ bounds still hold for the batching operation of a multisecretary problem with discretized uniform rewards in this large support setting, provided the support grows fast enough. This suggests that the continuous support setting that is the paper's focus serves as a useful theoretical surrogate for the more realistic large finite support setting.
|
| 1332 |
Efficient Hierarchical Transformers for Representing Log Data
2408.16803
|
cs.LGcs.AI
|
Zhichao Hou, Mina Ghashami, Mikhail Kuznetsov, Zifan Zhang, Ali Torkamani |
Transformers have gained widespread acclaim for their versatility in handling diverse data structures, yet their application to log data remains underexplored. Log data, characterized by its hierarchical, dictionary-like structure, poses unique challenges when...Transformers have gained widespread acclaim for their versatility in handling diverse data structures, yet their application to log data remains underexplored. Log data, characterized by its hierarchical, dictionary-like structure, poses unique challenges when processed using conventional transformer models. Traditional methods often rely on manually crafted templates for parsing logs, a process that is labor-intensive and lacks generalizability. Additionally, the linear treatment of log sequences by standard transformers neglects the rich, nested relationships within log entries, leading to suboptimal representations and excessive memory usage. To address these issues, we introduce HLogformer, a novel hierarchical transformer framework specifically designed for log data. HLogformer leverages the hierarchical structure of log entries to significantly reduce memory costs and enhance representation learning. Unlike traditional models that treat log data as flat sequences, our framework processes log entries in a manner that respects their inherent hierarchical organization. This approach ensures comprehensive encoding of both fine-grained details and broader contextual relationships. Our contributions are threefold: First, HLogformer is the first framework to design a dynamic hierarchical transformer tailored for dictionary-like log data. Second, it dramatically reduces memory costs associated with processing extensive log sequences. Third, comprehensive experiments demonstrate that HLogformer more effectively encodes hierarchical contextual information, proving to be highly effective for downstream tasks such as synthetic anomaly detection and product recommendation.
|
| 1333 |
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
2409.18778
|
cs.LGcs.AI
|
Joseph Cotnareanu, Zhanguang Zhang, Hui-Ling Zhen, Yingxue Zhang, Mark Coates |
Efficiently determining the satisfiability of a boolean equation -- known as the SAT problem for brevity -- is crucial in various industrial problems. Recently, the advent of deep learning methods has introduced significant potential for enhancing SAT solving....Efficiently determining the satisfiability of a boolean equation -- known as the SAT problem for brevity -- is crucial in various industrial problems. Recently, the advent of deep learning methods has introduced significant potential for enhancing SAT solving. However, a major barrier to the advancement of this field has been the scarcity of large, realistic datasets. The majority of current public datasets are either randomly generated or extremely limited, containing only a few examples from unrelated problem families. These datasets are inadequate for meaningful training of deep learning methods. In light of this, researchers have started exploring generative techniques to create data that more accurately reflect SAT problems encountered in practical situations. These methods have so far suffered from either the inability to produce challenging SAT problems or time-scalability obstacles. In this paper we address both by identifying and manipulating the key contributors to a problem's ``hardness'', known as cores. Although some previous work has addressed cores, the time costs are unacceptably high due to the expense of traditional heuristic core detection techniques. We introduce a fast core detection procedure that uses a graph neural network. Our empirical results demonstrate that we can efficiently generate problems that remain hard to solve and retain key attributes of the original example problems. We show via experiment that the generated synthetic SAT problems can be used in a data augmentation setting to provide improved prediction of solver runtimes.
|
| 1334 |
Nonlinear Equilibrium Transitions in a Potential Game Model for Federated Learning
2411.11793
|
cs.LG
|
Kang Liu, Ziqi Wang, Enrique Zuazua |
In federated learning (FL), a central server typically allocates training efforts to clients. However, from a market-oriented perspective, clients may independently choose their training efforts based on rational self-interest. To study this setting, we propos...In federated learning (FL), a central server typically allocates training efforts to clients. However, from a market-oriented perspective, clients may independently choose their training efforts based on rational self-interest. To study this setting, we propose a potential game framework in which each client's payoff is determined by its individual effort and the rewards provided by the server. The rewards are influenced by the collective efforts of all clients and can be modulated by a reward factor. We first establish the existence of Nash equilibria (NEs) and then investigate their uniqueness in a stationary setting. We show that the NEs depend nonlinearly on the reward factor and exhibit a nonsmooth transition at a critical value, where the stationary potential loses strict curvature, leading to nonunique NEs and a jump between low-effort and high-effort branches. Furthermore, we prove the convergence of the best-response algorithm for computing NEs in our FL game. Finally, we apply the clients' rational efforts derived from the NEs to FL training with various datasets and models, thereby validating the effectiveness of the identified critical reward factor. The source code is available at https://github.com/DCN-FAU-AvH/FL-Potential-Game
|
| 1335 |
How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies
2502.03698
|
cs.LG
|
Akansha Kalra, Basavasagar Patil, Guanhong Tao, Daniel S. Brown |
Imitation learning, also known as learning from demonstrations, is a popular approach to train AI models; however, the vulnerability of these models to adversarial attacks remains underexplored. We present the first systematic study of adversarial attacks, acr...Imitation learning, also known as learning from demonstrations, is a popular approach to train AI models; however, the vulnerability of these models to adversarial attacks remains underexplored. We present the first systematic study of adversarial attacks, across a range of both classic and recently proposed imitation learning algorithms, including Vanilla Behavior Cloning (Vanilla BC), LSTM-GMM, Implicit Behavior Cloning (IBC), Diffusion Policy (DP), and Vector-Quantized Behavior Transformer (VQ-BET). We study the vulnerability of these methods to white-box, grey-box and black-box adversarial perturbations. Our experiments reveal that most existing methods are highly vulnerable to these attacks, including black-box transfer attacks that transfer across algorithms. White-box attacks cause at least a 65% reduction in average task success across all evaluated tasks and algorithms, while the black-box transfer attacks reduce task success by up to 88% on Lift, 99% on Can, and 100% on Square. To the best of our knowledge, we are the first to study and compare the vulnerabilities of different popular imitation learning algorithms to both white-box and black-box attacks. Our findings highlight the vulnerabilities of modern imitation learning algorithms, paving the way for future work in addressing such limitations. Videos and code are available at https://sites.google.com/view/uap-attacks-on-bc.
|
| 1336 |
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
2503.14858
|
cs.LGcs.AI
|
Kevin Wang, Ishaan Javali, Micha{\l} Bortkiewicz, Tomasz Trzci\'nski, Benjamin Eysenbach |
Scaling up self-supervised learning has driven breakthroughs in language and vision, yet comparable progress has remained elusive in reinforcement learning (RL). In this paper, we study building blocks for self-supervised RL that unlock substantial improvement...Scaling up self-supervised learning has driven breakthroughs in language and vision, yet comparable progress has remained elusive in reinforcement learning (RL). In this paper, we study building blocks for self-supervised RL that unlock substantial improvements in scalability, with network depth serving as a critical factor. Whereas most RL papers in recent years have relied on shallow architectures (around 2 - 5 layers), we demonstrate that increasing the depth up to 1024 layers can significantly boost performance. Our experiments are conducted in an unsupervised goal-conditioned setting, where no demonstrations or rewards are provided, so an agent must explore (from scratch) and learn how to maximize the likelihood of reaching commanded goals. Evaluated on simulated locomotion and manipulation tasks, our approach increases performance on the self-supervised contrastive RL algorithm by $2\times$ - $50\times$, outperforming other goal-conditioned baselines. Increasing the model depth not only increases success rates but also qualitatively changes the behaviors learned. The project webpage and code can be found here: https://wang-kevin3290.github.io/scaling-crl/.
|
| 1337 |
Ordinary Least Squares as an Attention Mechanism
2504.09663
|
cs.LG
|
Philippe Goulet Coulombe |
I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. The connection comes from viewing OLS as a similarity-based prediction rule in a ...I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. The connection comes from viewing OLS as a similarity-based prediction rule in a learned embedding space. In this representation, least squares does not estimate coefficients per se. Instead, it selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products. This maps directly onto the query-key-value structure of attention mechanisms. I then discuss extensions to dimensionality reduction, nonlinearity, and time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that nonlinear Attention Regression performs competitively against standard machine learning baselines. In the reverse direction, I replace the attention sublayer of a transformer for tabular data with an explicit regression on polynomial features. The resulting model performs comparably to the standard transformer at a fraction of its parameter count.
|
| 1338 |
KITINet: KInetic Theory Inspired Inter-Channel Information Exchange for Residual Networks
2505.17919
|
cs.LG
|
Mingquan Feng, Yifan Fu, Tongcheng Zhang, Yu Jiang, Yixin Huang |
Residual connections are foundational to modern neural networks, yet the standard additive update provides no explicit mechanism for structured inter-channel interaction. We introduce KITINet (KInetics Theory Inspired Network), a training-time module with no a...Residual connections are foundational to modern neural networks, yet the standard additive update provides no explicit mechanism for structured inter-channel interaction. We introduce KITINet (KInetics Theory Inspired Network), a training-time module with no additional trainable parameters that augments residual updates with stochastic pairwise feature interactions inspired by collision sampling in kinetic theory. KITINet reshapes channel groups as particles and mixes their features according to relative distance and velocity, while leaving the underlying residual layer trainable in the usual way. The module is disabled at inference, so the deployed model is identical to the original backbone and incurs no additional inference cost. Across language pre-training and fine-tuning, image classification, and PDE operator learning, KITINet demonstrates broad and robust performance gains across diverse tasks and architectures. We further observe enhanced parameter condensation across several settings and develop a simplified stochastic analysis that provides a mechanistic account of how collision-inspired interactions can promote condensation dynamics.
|
| 1339 |
Distributionally Robust Deep Q-Learning
2505.19058
|
cs.LG
|
Chung I Lu, Julian Sester, Aijia Zhang |
We propose a novel distributionally robust $Q$-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model uncertainty. The uncertainty is taken int...We propose a novel distributionally robust $Q$-learning algorithm for the non-tabular case accounting for continuous state spaces where the state transition of the underlying Markov decision process is subject to model uncertainty. The uncertainty is taken into account by considering the worst-case transition from a ball around a reference probability measure. To determine the optimal policy under the worst-case state transition, we solve the associated non-linear Bellman equation by dualising and regularising the Bellman operator with the Sinkhorn distance, which is then parameterised with deep neural networks. This approach allows us to modify the Deep Q-Network algorithm to optimise for the worst case state transition. We illustrate the tractability and effectiveness of our approach through several applications, including a portfolio optimisation task based on S\&{P}~500 data. We also establish convergence guarantees for exact and approximate robust fitted $Q$-iteration, decompose the numerical RDQN error into interpretable components, and discuss extensions to compact continuous action sets.
|
| 1340 |
Learning Interpretable Differentiable Logic Networks for Tabular Regression
2505.23615
|
cs.LG
|
Chang Yue, Niraj K. Jha |
Neural networks (NNs) achieve outstanding performance in many domains; however, their decision processes are often opaque and their inference can be computationally expensive in resource-constrained environments. We recently proposed Differentiable Logic Netwo...Neural networks (NNs) achieve outstanding performance in many domains; however, their decision processes are often opaque and their inference can be computationally expensive in resource-constrained environments. We recently proposed Differentiable Logic Networks (DLNs) to address these issues for tabular classification based on relaxing discrete logic into a differentiable form, thereby enabling gradient-based learning of networks built from binary logic operations. DLNs offer interpretable reasoning and substantially lower inference cost. We extend the DLN framework to supervised tabular regression. We first redesign the final output layer (the SumLayer) to support continuous targets. More critically, we find the original two-phase training procedure used for classification is suboptimal for regression, and thus develop a unified, single-stage optimization procedure. We also demonstrate that temperature annealing of the network's differentiable relaxations is decisive for achieving stable convergence and high accuracy. We evaluate the resulting model on 15 public regression benchmarks, comparing it with modern neural networks and classical regression baselines. Regression DLNs match or exceed baseline accuracy while preserving interpretability and fast inference. Our results show that DLNs are a viable, cost-effective alternative for regression tasks, especially where model transparency and computational efficiency are important.
|
| 1341 |
Anomaly Detection and Generation with Diffusion Models: A Survey
2506.09368
|
cs.LGcs.AI
|
Yang Liu, Jing Liu, Chengfang Li, Rui Xi, Wenchao Li |
Anomaly detection (AD) plays a pivotal role across diverse domains, including cybersecurity, finance, healthcare, and industrial manufacturing, by identifying unexpected patterns that deviate from established norms in real-world data. Recent advancements in de...Anomaly detection (AD) plays a pivotal role across diverse domains, including cybersecurity, finance, healthcare, and industrial manufacturing, by identifying unexpected patterns that deviate from established norms in real-world data. Recent advancements in deep learning, specifically diffusion models (DMs), have sparked significant interest due to their ability to learn complex data distributions and generate high-fidelity samples, offering a robust framework for unsupervised AD. In this survey, we comprehensively review anomaly detection and generation with diffusion models (ADGDM), presenting a tutorial-style analysis of the theoretical foundations and practical implementations and spanning images, videos, time series, tabular, and multimodal data. Crucially, unlike existing surveys that often treat anomaly detection and generation as separate problems, we highlight their inherent synergistic relationship. We reveal how DMs enable a reinforcing cycle where generation techniques directly address the fundamental challenge of anomaly data scarcity, while detection methods provide critical feedback to improve generation fidelity and relevance, advancing both capabilities beyond their individual potential. A detailed taxonomy categorizes ADGDM methods based on anomaly scoring mechanisms, conditioning strategies, and architectural designs, analyzing their strengths and limitations. We final discuss key challenges including scalability and computational efficiency, and outline promising future directions such as efficient architectures, conditioning strategies, and integration with foundation models (e.g., visual-language models and large language models). By synthesizing recent advances and outlining open research questions, this survey aims to guide researchers and practitioners in leveraging DMs for innovative AD solutions across diverse applications.
|
| 1342 |
Online Conformal Abstention Under Adversarial Bandit Feedback
2506.14067
|
cs.LG
|
Minjae Lee, Yoonjae Jung, Sangdon Park |
As interactive generative systems are increasingly deployed in real-world applications, their tendency to generate unreliable or false responses raises serious concerns. Conformal abstention mitigates this risk by ensuring that the system answers only when con...As interactive generative systems are increasingly deployed in real-world applications, their tendency to generate unreliable or false responses raises serious concerns. Conformal abstention mitigates this risk by ensuring that the system answers only when confident. However, real-world deployments typically provide only partial user feedback (e.g., thumbs up/down) on the selected response and often operate in non-stationary or adversarial environments, for which effective learning methods are largely missing. To bridge this gap, we propose ExAUL, a novel online learning framework for conformal abstention with adversarial and partial feedback. Technically, we introduce (i) a novel conversion lemma that translates the regret of any bandit algorithm into a selective risk (SR) bound, and (ii) feedback unlocking, a strategy that exploits the structure of conformal abstention to extract additional learning signals from partial feedback. We prove that ExAUL achieves a regret bound of $O(\sqrt{T \ln |\Hs|})$, which translates into an $O(\sqrt{T})$ bound on SR violation control, matching the controllability of full-information settings despite receiving only partial feedback. While applicable to general generative tasks, we demonstrate the efficacy of ExAUL for ensuring the reliability of Large Language Models (LLMs) through empirical validation on question-answering tasks across diverse non-stationary and adversarial settings. Our results demonstrate that ExAUL robustly controls the SR while maintaining competitive answering coverage.
|
| 1343 |
Cooperative Sheaf Neural Networks
2507.00647
|
cs.LG
|
Andr\'e Ribeiro, Ana Luiza Ten\'orio, Juan Belieni, Amauri H. Souza, Diego Mesquita |
Sheaf diffusion has recently emerged as a promising design pattern for graph representation learning due to its inherent ability to handle heterophilic data and avoid oversmoothing. Meanwhile, cooperative message passing has also been proposed as a way to enha...Sheaf diffusion has recently emerged as a promising design pattern for graph representation learning due to its inherent ability to handle heterophilic data and avoid oversmoothing. Meanwhile, cooperative message passing has also been proposed as a way to enhance the flexibility of information diffusion by allowing nodes to independently choose whether to propagate/gather information from/to neighbors. A natural question ensues: is sheaf diffusion capable of exhibiting this cooperative behavior? Here, we provide a negative answer to this question. In particular, we show that existing sheaf diffusion methods fail to achieve cooperative behavior due to the lack of message directionality. To circumvent this limitation, we introduce the notion of cellular sheaves over directed graphs and characterize their in- and out-degree Laplacians. We leverage our construction to propose Cooperative Sheaf Neural Networks (CSNNs). Theoretically, we characterize the receptive field of CSNN and show it allows nodes to selectively attend (listen) to arbitrarily far nodes while ignoring all others in their path, potentially mitigating oversquashing. Our experiments show that CSNN presents overall better performance compared to prior art on sheaf diffusion as well as cooperative graph neural networks.
|
| 1344 |
EduAlign: Aligning Educational LLMs for Helpfulness, Personalization, and Creativity
2507.20335
|
cs.LGcs.AI
|
Siyu Song, Ye Lu, Wentao Liu, Ruohua Zhang, Tao Liu |
Educational language models should pursue three complementary objectives: helpfulness through responsible guidance, personalization to learner needs, and creativity that supports exploration. General-purpose models do not consistently satisfy these educational...Educational language models should pursue three complementary objectives: helpfulness through responsible guidance, personalization to learner needs, and creativity that supports exploration. General-purpose models do not consistently satisfy these educational requirements, motivating explicit alignment with all three objectives. We present EduAlign, a framework that combines educational reward modeling with multi-objective reinforcement learning. To obtain training signals for these goals across educational scenarios, we develop a rubric-guided data synthesis pipeline that iteratively refines scoring criteria and generation prompts for each subject, grade band, and task type. The resulting supervision trains a multidimensional reward model, Edu-RM, to predict helpfulness, personalization, and creativity scores. To jointly optimize these educational objectives, we introduce Ideal-point Policy Optimization (IdealPO), a multi-objective reinforcement learning algorithm. IdealPO combines ideal-point marginal credit with continuous conflict weighting to guide policy updates that account for uneven progress and conflicting contributions across objectives. Experiments with 7B and 32B models show that EduAlign achieves the highest mean ratings on all three dimensions from both education experts and a model judge, alongside the highest aggregate scores on EduBench, ELMES, and EduValues among the compared methods. Reward-model validation and component ablations further support the effectiveness of the framework.
|
| 1345 |
OCSVM-Guided Representation Learning for Unsupervised Anomaly Detection
2507.21164
|
cs.LGcs.AI
|
Nicolas Pinon (MYRIAD), Robin Trombetta (MYRIAD), Carole Lartizien (MYRIAD) |
Unsupervised anomaly detection (UAD) aims to detect anomalies without labeled data, a necessity in many machine learning applications where anomalous samples are rare or not available. Most state-of-the-art methods fall into two categories: reconstruction-base...Unsupervised anomaly detection (UAD) aims to detect anomalies without labeled data, a necessity in many machine learning applications where anomalous samples are rare or not available. Most state-of-the-art methods fall into two categories: reconstruction-based approaches, which often reconstruct anomalies too well, and decoupled representation learning with density estimators, which can suffer from suboptimal feature spaces. While some recent methods attempt to couple feature learning and anomaly detection, they often rely on surrogate objectives, restrict kernel choices, or introduce approximations that limit their expressiveness and robustness. To address this challenge, we propose a novel method that couples representation learning the exact One-Class SVM (OCSVM) optimization problem, through a custom loss formulation that directly aligns latent features with the OCSVM decision boundary. The model is evaluated on two tasks: a benchmark based on MNIST-C, and a challenging brain MRI lesion detection task. Unlike most methods that focus on large, hyperintense lesions at the image level, our approach succeeds to target small, non-hyperintense lesions, while we evaluate voxel-wise metrics, addressing a more clinically relevant scenario. Both experiments evaluate a form of robustness to domain shifts, including corruption types in MNIST-C and texture or population age variations in MRI. Results demonstrate performance and robustness of our proposed model, highlighting its potential for general UAD and real-world medical imaging applications. The source code is available at https://github.com/Nicolas-Pinon/uad\_ocsvm\_guided\_repr\_learning.
|
| 1346 |
Disentangling Bias by Modeling Intra- and Inter-modal Causal Attention for Multimodal Sentiment Analysis
2508.04999
|
cs.LG
|
Menghua Jiang, Yuxia Lin, Baoliang Chen, Haifeng Hu, Yuncheng Jiang |
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and visual data. However, existing methods often suffer from spurious correlations both within and across modalities...Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and visual data. However, existing methods often suffer from spurious correlations both within and across modalities, leading models to rely on statistical shortcuts rather than true causal relationships, thereby undermining generalization. To mitigate this issue, we propose a Multi-relational Multimodal Causal Intervention (MMCI) framework, which leverages the backdoor adjustment from causal theory to address the confounding effects of such shortcuts. Specifically, we first model the multimodal inputs as a multi-relational graph to explicitly capture intra- and inter-modal dependencies. Then, we apply an attention mechanism to separately estimate and disentangle the causal features and shortcut features corresponding to these intra- and inter-modal relations. Finally, by approximating backdoor adjustment, we stratify the shortcut features and dynamically combine them with the causal features to encourage MMCI to produce stable predictions under distribution shifts. Extensive experiments on several standard MSA datasets and out-of-distribution (OOD) settings demonstrate that our method effectively suppresses biases and improves performance.
|
| 1347 |
Grad-CAM for Visualizing Attention Regions of PCA and SVM Layers in Convolutional Neural Networks
2508.11880
|
cs.LG
|
Yuto Omae, Hirotaka Takahashi |
Convolutional Neural Networks (CNNs) are an effective approach for classification tasks, particularly when the training dataset is large. Although CNNs have long been considered a black-box classification method, they can be used as a white-box method through ...Convolutional Neural Networks (CNNs) are an effective approach for classification tasks, particularly when the training dataset is large. Although CNNs have long been considered a black-box classification method, they can be used as a white-box method through visualization techniques such as Grad-CAM. When the training samples are limited, incorporating a Principal Component Analysis (PCA) layer and/or a Support Vector Machine (SVM) classifier into a CNN can effectively improve the classification performance. However, a conventional Grad-CAM cannot be directly applied to PCA and/or SVM layers. Generating attention regions for PCA and/or SVM layers in CNNs is important to facilitate the development of white-box methods. Therefore, we propose ``PCA-Grad-CAM'', a method for visualizing attention regions in PCA feature vectors, and ``SVM-Grad-CAM'', a method for visualizing attention regions in an SVM classifier layer. Solving a closed-form Jacobian problem comprising partial derivatives from the last convolutional layer to the PCA and/or SVM layers is necessary to complete the proposed methods analytically. In this paper, we present the exact closed-form Jacobian and visualization results of the proposed methods applied to several major datasets. In addition, the insertion and deletion metrics, which are major evaluation metrics in explainable AI, were applied to PCA- and SVM-Grad-CAM. The results suggest that the proposed method can successfully visualize the attention regions of PCA and SVM.
|
| 1348 |
FedUHD: Unsupervised Federated Learning using In-Memory Hyperdimensional Computing
2508.12021
|
cs.LG
|
You Hak Lee, Keming Fan, Xiaofan Yu, Quanling Zhao, Tianqi Zhang |
Unsupervised federated learning (UFL) enables privacy-preserving distributed training without data labeling, yet practical deployment remains challenging due to non-IID data, high computational and communication costs at edge devices, and sensitivity to commun...Unsupervised federated learning (UFL) enables privacy-preserving distributed training without data labeling, yet practical deployment remains challenging due to non-IID data, high computational and communication costs at edge devices, and sensitivity to communication noise. We propose FedUHD, the first UFL framework based on Hyperdimensional Computing (HDC). On the client side, FedUHD employs kNN-based cluster hypervector removal to mitigate non-IID effects by filtering detrimental local outliers. On the server side, cluster-aware HDC aggregation leverages cluster-level statistics to stabilize learning across heterogeneous clients. To further improve efficiency, we design a compute-in-memory (CIM) accelerator based on a novel phase-change memory (PCM) device, integrated with lightweight ASIC digital modules to execute the client-side HDC pipeline within the accelerator. The intrinsic robustness of HDC to low precision and device variations enables efficient mapping onto analog PCM crossbars, exploiting massive parallelism while minimizing data movement. Experimental results show that FedUHD achieves comparable accuracy to state-of-the-art neural network-based UFL methods across all datasets. On HAR and CIFAR10/100, FedUHD delivers an average 2,239x speedup and 1,542x higher energy efficiency on GPU. In addition, FedUHD reduces communication cost by up to 176x on HAR and CIFAR10/100 and demonstrates greater robustness than Orchestra under communication noise. Compared to GPU implementation of FedUHD, the proposed PCM-based accelerator provides an additional 4.07x speedup and three orders of magnitude higher energy efficiency on average. Furthermore, the results demonstrate the benefit of PCM over RRAM as a CIM substrate.
|
| 1349 |
EEGDM: Label-Efficient EEG Representation Learning with Generative Diffusion Model
2508.14086
|
cs.LG
|
Jia Hong Puah, Sim Kuan Goh, Ziwei Zhang, Zixuan Ye, Chow Khuen Chan |
Electroencephalography (EEG) is a critical tool for monitoring brain activity and diagnosing neurological disorders such as epilepsy. However, learning meaningful representations from raw EEG signals remains challenging due to limited annotations, substantial ...Electroencephalography (EEG) is a critical tool for monitoring brain activity and diagnosing neurological disorders such as epilepsy. However, learning meaningful representations from raw EEG signals remains challenging due to limited annotations, substantial inter-subject variability, and complex temporal dynamics. Recent EEG foundation models (FMs) have demonstrated promising performance through transformer-based architectures and large-scale self-supervised pretraining, yet they often incur substantial computational costs and exhibit diminishing returns with increasing model and dataset scale, limiting their practicality in clinical settings. To address these challenges, we propose EEGDM, a diffusion-based EEG representation learning framework. Specifically, EEGDM introduces a Structured State-Space Model for Diffusion Pretraining (SSMDP) that effectively captures long-range temporal dependencies through generative diffusion training on unlabeled EEG data. The representations are subsequently leveraged for downstream tasks via our Latent Fusion Module (LFM), which integrates multi-layer latent features of SSMDP. We evaluate EEGDM on three EEG benchmarks spanning EEG event classification (TUEV), seizure detection (CHB-MIT), and seizure classification (IIIC). Compared with existing state-of-the-art methods, including EEG FMs, EEGDM achieves competitive performance across diverse tasks and datasets, exceeding existing methods in most cases while requiring substantially fewer samples for both pretraining and downstream adaptation. These results demonstrate that diffusion-based learning can unlock more discriminative EEG representations, with direct implications for epilepsy diagnosis and management. Our source code and pretrained checkpoints are publicly available at: https://github.com/jhpuah/EEGDM.
|
| 1350 |
Minority Collective Action for User-Side Fairness
2508.15374
|
cs.LG
|
Omri Ben-Dov, Samira Samadi, Amartya Sanyal, Alexandru \c{T}ifrea |
Machine learning models often preserve biases present in training data, leading to unfair treatment of certain minority groups. Despite an array of existing firm-side bias mitigation techniques, they typically incur utility costs and require organizational buy...Machine learning models often preserve biases present in training data, leading to unfair treatment of certain minority groups. Despite an array of existing firm-side bias mitigation techniques, they typically incur utility costs and require organizational buy-in. Recognizing that many models rely on user-contributed data, end-users can induce fairness through the framework of Algorithmic Collective Action, where a coordinated minority group strategically relabels its own data to enhance fairness, without altering the firm's training process. We propose three practical, model-agnostic methods to approximate ideal relabeling and validate them on real-world datasets. Our findings show that a subgroup of the minority can substantially reduce unfairness with a small impact on the overall prediction error.
|
| 1351 |
Learning Interpretable Differentiable Logic Networks for Time-Series Classification
2508.17512
|
cs.LG
|
Chang Yue, Niraj K. Jha |
Differentiable logic networks (DLNs) have shown promising results in tabular domains by combining accuracy, interpretability, and computational efficiency. In this work, we apply DLNs to the domain of TSC for the first time, focusing on univariate datasets. To...Differentiable logic networks (DLNs) have shown promising results in tabular domains by combining accuracy, interpretability, and computational efficiency. In this work, we apply DLNs to the domain of TSC for the first time, focusing on univariate datasets. To enable DLN application in this context, we adopt feature-based representations relying on Catch22 and TSFresh, converting sequential time series into vectorized forms suitable for DLN classification. Unlike prior DLN studies that fix the training configuration and vary various settings in isolation via ablation, we integrate all such configurations into the hyperparameter search space, enabling the search process to select jointly optimal settings. We then analyze the distribution of selected configurations to better understand DLN training dynamics. We evaluate our approach on 51 publicly available univariate TSC benchmarks. The results confirm that classification DLNs maintain their core strengths in this new domain: they deliver competitive accuracy, retain low inference cost, and provide transparent, interpretable decision logic, thus aligning well with previous DLN findings in the realm of tabular classification and regression tasks.
|
| 1352 |
Physics-informed GNN for medium-high voltage AC power flow with edge-aware attention and line search correction operator
2509.22458
|
cs.LGcs.AI
|
Changhun Kim, Timon Conrad, Redwanul Karim, Julian Oelhaf, David Riebesel |
Physics-informed graph neural networks (PIGNNs) have emerged as fast AC power-flow solvers that can replace the classic NewtonRaphson (NR) solvers, especially when thousands of scenarios must be evaluated. However, current PIGNNs still need accuracy improvemen...Physics-informed graph neural networks (PIGNNs) have emerged as fast AC power-flow solvers that can replace the classic NewtonRaphson (NR) solvers, especially when thousands of scenarios must be evaluated. However, current PIGNNs still need accuracy improvements at parity speed; in particular, the soft constraint on the physics loss is inoperative at inference, which can deter operational adoption. We address this with PIGNN-Attn-LS, combining an edge-aware attention mechanism that explicitly encodes line physics via per-edge biases to form a fully differentiable knownoperator layer inside the computation graph, with a backtracking line-search-based globalized correction operator that restores an operative decrease criterion at inference. Training and testing use a realistic High-/Medium-Voltage scenario generator, with NR used only to construct reference states. On held-out HV cases consisting of 4-32-bus grids, PIGNN-Attn-LS achieves a test RMSE of 0.00033 p.u. in voltage and 0.08 deg in angle, outperforming the PIGNN-MLP baseline by 99.5% and 87.1%, respectively. With streaming micro-batches, it delivers 2-5x faster batched inference than NR on 4-1024-bus grids.
|
| 1353 |
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
2509.23410
|
cs.LGcs.AI
|
Younes Hourri, Mohammad Mozaffari, Maryam Mehri Dehnavi |
Large language models (LLMs) deliver impressive performance but incur prohibitive memory and compute costs at deployment. Model pruning is an effective way to reduce these overheads, yet existing approaches face challenges: unstructured sparsity, where nonzero...Large language models (LLMs) deliver impressive performance but incur prohibitive memory and compute costs at deployment. Model pruning is an effective way to reduce these overheads, yet existing approaches face challenges: unstructured sparsity, where nonzeros can appear anywhere, preserves accuracy but yields irregular access patterns that prevent GPU acceleration, while semi-structured 2:4 sparsity is hardware-friendly but enforces a rigid 50% pattern that degrades model quality. To bridge this gap, we introduce PATCH, a hybrid sparsity framework that enables a continuous sparsity ratio between 0% and 50%. PATCH partitions weight matrices into tiles, assigning each tile to be either dense or 2:4 sparse via a learnable mask selection mechanism. This design provides fine-grained control over accuracy-acceleration tradeoffs and supports non-uniform sparsity across layers, leading to superior overall quality. Across models from 0.5B to 13B parameters, PATCH consistently narrows the gap to dense accuracy while delivering practical speedups. For instance, on LLaMA-2 7B with an A6000 GPU, PATCH achieves 1.18x-1.38x end-to-end speedup over dense baselines while improving accuracy by 0.37%-2.96% compared to the state-of-the-art 2:4 pruning method, MaskLLM.
|
| 1354 |
GeoFunFlow: Geometric function flow matching for joint probabilistic inference of physical fields and complex geometries
2509.24117
|
cs.LG
|
Yajie Ji, Sifan Wang, Zhikai Wu, David van Dijk, Lu Lu |
Inverse problems governed by partial differential equations (PDEs) arise widely in science and engineering, but are often ill-posed and limited by sparse, noisy observations. In many applications, measurements reveal only part of the physical state, while the ...Inverse problems governed by partial differential equations (PDEs) arise widely in science and engineering, but are often ill-posed and limited by sparse, noisy observations. In many applications, measurements reveal only part of the physical state, while the domain geometry may also be unknown even though it shapes the observed response. Joint field and geometry inference across varying computational domains and discretizations remains challenging, whereas many existing machine learning approaches are designed for known geometries and deterministic field reconstruction. Here, we introduce GeoFunFlow, a probabilistic framework that unifies field reconstruction on known domains and joint field and geometry inference on unknown domains. GeoFunFlow combines a geometric function autoencoder (GeoFAE) with flow matching in the latent space to model a joint distribution over physical fields and geometries. GeoFAE establishes a common representation across spatial discretizations that captures the relationship between physical fields and domain geometries, with unknown geometry represented by a signed distance function. The resulting representation allows observations to guide both field reconstruction and geometry recovery, while latent rectified flow enables efficient conditional sampling and spatially resolved uncertainty quantification. A calibration procedure further provides geometry uncertainty estimates with interpretable empirical coverage. Across seven benchmarks spanning porous media flow, fluid mechanics, and optical tomography, GeoFunFlow accurately recovers fields and geometries across complex, variable, and unknown domains while quantifying spatially resolved conditional uncertainty.
|
| 1355 |
Bayesian Distributional Models of Executive Functioning
2510.00387
|
cs.LG
|
Robert Kasumba, Zeyu Lu, Dom CP Marticorena, Mingyang Zhong, Paul Beggs |
This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian Distributional Active LEarning (DALE) perform in comparison to conventional Independent Maximum Likelihood Estim...This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian Distributional Active LEarning (DALE) perform in comparison to conventional Independent Maximum Likelihood Estimation (IMLE). DLVM integrates observations across multiple executive function tasks and individuals, allowing parameter estimation even under sparse or incomplete data conditions. To establish known-ground truth, we uniformly sample individual sessions from a neural network learned latent space and map them to distributional cognitive performance across different tasks. The individual test-items are then sampled from these distributions using either DALE, random procedure or a standard fixed battery approach. When given the same set of observations, DLVM consistently outperformed IMLE, especially under smaller amounts of data, and converges faster to highly accurate estimates of the true distributions. In a second set of analyses, DALE adaptively guided sampling to maximize information gain, outperforming random sampling and fixed test batteries, particularly within the first 80 trials. These findings establish the advantages of combining DLVM's cross-task inference with DALE's optimal adaptive sampling, providing a principled basis for more efficient cognitive assessments.
|
| 1356 |
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
2510.00915
|
cs.LGcs.AI
|
Xin-Qiang Cai, Wei Wang, Feng Liu, Tongliang Liu, Gang Niu |
Reinforcement Learning with Verifiable Rewards (RLVR) replaces costly human labeling with automated verifiers. To reduce verifier hacking, many RLVR systems binarize rewards to $\{0,1\}$, but imperfect verifiers inevitably introduce \emph{false negatives} (rej...Reinforcement Learning with Verifiable Rewards (RLVR) replaces costly human labeling with automated verifiers. To reduce verifier hacking, many RLVR systems binarize rewards to $\{0,1\}$, but imperfect verifiers inevitably introduce \emph{false negatives} (rejecting correct answers) and \emph{false positives} (accepting incorrect ones). We formalize verifier unreliability as a stochastic reward channel with asymmetric noise rates $\rho_0$ and $\rho_1$ -- the FP rate and the FN rate, respectively. From this abstraction we derive two lightweight corrections: (i) a \emph{backward} correction that yields an unbiased surrogate reward and thus an unbiased policy-gradient estimator in expectation, and (ii) a \emph{forward} correction that reweights score-function terms so the expected update aligns with the clean gradient direction and requires only the FN rate. We implement both as lightweight hooks in a group relative policy optimization pipeline, both corrections improve RLVR for math reasoning under synthetic and real verifier noise, with the forward variant being more stable under heavier noise. Finally, an appeals mechanism with a lightweight LLM verifier estimates the FN rate online and further improves performance.
|
| 1357 |
Neural Bayesian Filtering
2510.03614
|
cs.LGcs.AI
|
Christopher Solinas, Radovan Haluska, David Sychrovsky, Finbarr Timbers, Nolan Bard |
Sequential estimation under partial observability requires tracking beliefs that may be high-dimensional, multimodal, and non-Gaussian. Classical Bayesian filters generalize zero-shot to any system whose dynamics can be evaluated, but their representations sca...Sequential estimation under partial observability requires tracking beliefs that may be high-dimensional, multimodal, and non-Gaussian. Classical Bayesian filters generalize zero-shot to any system whose dynamics can be evaluated, but their representations scale poorly: parametric filters struggle to capture multimodality, and particle filters require exponentially many particles in the state dimension. Generative Distribution Embeddings (GDEs) learn compact representations of complex distributions but have no sequential update. We present Neural Bayesian Filtering (NBF), which represents beliefs as embeddings and approximates their Bayesian update. At each step, NBF samples particles from a learned conditional generator, propagates them through the dynamics, and embeds the propagated set, weighted by the likelihood of the new observation. Regenerating the set from the embedding at each step, rather than resampling a surviving pool, mitigates the risk of impoverishment that degrades particle filters. The system dynamics and likelihood enter only through propagation and weighting, so NBF retains zero-shot adaptability to new dynamics as long as the beliefs they produce fall within the family the embedding was trained on. Training the GDE to represent that family requires only the states realized at training time. We validate NBF on Lorenz-96 and a pursuit--evasion domain, showing graceful scaling with state dimension, accurate tracking of multimodal posteriors, and zero-shot generalization to dynamics and likelihoods held out from training.
|
| 1358 |
Climate Surrogates for Scalable Multi-Agent Reinforcement Learning: A Case Study with CICERO-SCM
2510.07971
|
cs.LG
|
Oskar Bohn Lassen, Serio Angelo Maria Agriesti, Filipe Rodrigues, Blaz Kurnik, Francisco Camara Pereira |
Climate policy analysis requires models that capture multi-gas climate effects, but such models are too slow to embed in reinforcement learning loops at scale. In collaboration with the European Environment Agency, we develop a multi-agent reinforcement learni...Climate policy analysis requires models that capture multi-gas climate effects, but such models are too slow to embed in reinforcement learning loops at scale. In collaboration with the European Environment Agency, we develop a multi-agent reinforcement learning (MARL) framework that integrates a higher-fidelity climate surrogate as the environment transition, enabling regional agents to learn policies under multi-gas dynamics. We train a recurrent surrogate on $20{,}000$ multi-gas emission pathways to emulate CICERO-SCM. The surrogate achieves near-simulator accuracy (global-mean temperature RMSE $\approx\!4\!\times \!10^{-4}\,\mathrm{K}$) with $\sim\!1000\times$ faster one-step inference and yields $>\!100\times$ end-to-end MARL training speed-up. We show policy agreement with the simulator in tractable settings and propose a replay- and rank-consistency test (Kendall's $\tau$) for assessing policy fidelity when simulator-in-the-loop training is infeasible. This enables large-scale multi-agent policy experiments while retaining high-fidelity multi-gas climate response.
|
| 1359 |
Spectral Analysis of Molecular Features: When Richer Features Do Not Guarantee Better Generalization
2510.14217
|
cs.LG
|
Asma Jamali, Tin Sum Cheng, Rodrigo A. Vargas-Hern\'andez |
The spectral properties of feature embeddings offer critical insights into model generalization and representation quality. While deep learning models are widely used for molecular property prediction, kernel methods remain competitive in low-data regimes, yet...The spectral properties of feature embeddings offer critical insights into model generalization and representation quality. While deep learning models are widely used for molecular property prediction, kernel methods remain competitive in low-data regimes, yet their spectral behavior is largely unexplored. We present the first comprehensive spectral analysis of kernel ridge regression across diverse representations, including molecular fingerprints (ECFP), pretrained transformers, graph neural networks, and 3D descriptors, evaluated on QM9, MoleculeNet, and Therapeutics Data Commons benchmarks. Surprisingly, richer spectral features do not consistently yield better generalization performance, contradicting common representation heuristics used in self-supervised learning (SSL). Across four spectral metrics, only ECFP-based kernels show a strictly positive correlation with performance, particularly SSE, ID, and SR on MoleculeNet. Global 3D representations exhibit mixed behavior; transformer-based representations are consistently positive across all four metrics, whereas local 3D representations show a negative trend and no significant correlations. Truncation analysis further emphasizes this disparity: for local 3D representations on thermodynamic targets, fewer than 2% of eigenvalues, and occasionally as few as 0.02%, are needed to recover 95% of performance, whereas ECFP and transformer kernels require significantly more. By demonstrating a strong dependence on both task and representation, our results challenge the heuristic that richer spectra inherently improve generalization, providing new guidance for evaluating representations in SSL and in label-limited scientific tasks.
|
| 1360 |
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
2510.18814
|
cs.LGcs.AI
|
Mengqi Li, Lei Zhao, Anthony Man-Cho So, Ruoyu Sun, Xiao Li |
Can a language model improve reasoning by learning from its own imperfect responses, without rewards or teacher-provided solutions? We present Self-evolving Post-Training (SePT), a simple method that alternates temperature-controlled self-generation with next-...Can a language model improve reasoning by learning from its own imperfect responses, without rewards or teacher-provided solutions? We present Self-evolving Post-Training (SePT), a simple method that alternates temperature-controlled self-generation with next-token likelihood training. Each round uses the updated model to generate new training responses, with one response per prompt by default and no correctness filtering. Across six mathematical benchmarks, SePT improves a temperature-selected no-training baseline by 11.4 and 6.7 AVG points on Qwen2.5-Math-7B and Qwen2.5-7B, respectively, where AVG averages Pass@1, Pass@8 and Pass@32 across benchmarks. We analyze how sampling temperature shapes the learning signal and investigate the value of the resulting responses. Responses from a SePT-trained model improve a student initialized from the original weights, while their reasoning prefixes help an unchanged model complete solutions, even when matched in length to prefixes from a colder initial model. Comparing next-token predictions at identical contexts also reveals changes in token rankings that decoding-temperature adjustment cannot reproduce. Further evaluations across nine starting models, general reasoning and code generation examine broader applicability. Together, these results show that reward-free self-training can improve both a model's predictions and the supervision it provides. Our code is available at https://github.com/ElementQi/SePT.
|
| 1361 |
Learning Granger Causality under Latent Confounding via Intervention-Induced Heterogeneity
2510.19138
|
cs.LGcs.AI
|
Ziyi Zhang, Shaogang Ren, Xiaoning Qian, Nick Duffield |
Granger causality characterizes directed predictive dependencies in multivariate time series, but recovering such dependencies becomes challenging in the presence of latent confounding. Cross-environment invariance provides a natural source of information in h...Granger causality characterizes directed predictive dependencies in multivariate time series, but recovering such dependencies becomes challenging in the presence of latent confounding. Cross-environment invariance provides a natural source of information in heterogeneous settings, yet invariance alone can be insufficient: when latent-to-observed mechanisms remain stable, hidden confounders can induce predictive dependencies that are just as invariant as genuine Granger-causal relations. We show that interventions provide an additional source of identifying information by inducing structured variation in observed mechanisms, while stable latent pathways need not exhibit the same cross-environment changes. In practice, however, neither the intervened environments nor the affected mechanisms are known. We propose GRACE, a framework for learning Granger causality under latent confounding from intervention-induced heterogeneity. GRACE decomposes multivariate dynamics into a shared Granger mechanism, sparse environment-specific deviations that capture edge-level interventions, and a latent component that accounts for confounding. Under a linear generative model, we show that GRACE can recover which environments intervene on a given edge when the edge is perturbed in at least one but fewer than half of the environments and the intervention effect is sufficiently large to survive sparsity shrinkage; the recovered intervention pattern then provides a certificate for the corresponding Granger causal edge. Experiments on synthetic and real-world time series demonstrate improved Granger causal structure recovery under latent confounding and unknown interventions.
|
| 1362 |
DeNoise: Learning Robust Graph Representations for Unsupervised Graph-Level Anomaly Detection
2511.04086
|
cs.LGcs.AI
|
Qingfeng Chen, Haojin Zeng, Jingyi Jie, Shichao Zhang, Debo Cheng |
With the rapid growth of graph-structured data in critical domains, unsupervised graph-level anomaly detection (UGAD) has become a pivotal task. UGAD seeks to identify entire graphs that deviate from normal behavioral patterns. However, most Graph Neural Netwo...With the rapid growth of graph-structured data in critical domains, unsupervised graph-level anomaly detection (UGAD) has become a pivotal task. UGAD seeks to identify entire graphs that deviate from normal behavioral patterns. However, most Graph Neural Network (GNN) approaches implicitly assume that the training set is clean, containing only normal graphs, which is rarely true in practice. Even modest contamination by anomalous graphs can distort learned representations and sharply degrade performance. To address this challenge, we propose DeNoise, a robust UGAD framework explicitly designed for contaminated training data. It jointly optimizes a graph-level encoder, an attribute decoder, and a structure decoder via an adversarial objective to learn noise-resistant embeddings. Further, DeNoise introduces an encoder anchor-alignment denoising mechanism that fuses high-information node embeddings from normal graphs into all graph embeddings, improving representation quality while suppressing anomaly interference. A contrastive learning component then compacts normal graph embeddings and repels anomalous ones in the latent space. Extensive experiments on eight real-world datasets demonstrate that DeNoise consistently learns reliable graph-level representations under varying noise intensities and significantly outperforms state-of-the-art UGAD baselines.
|
| 1363 |
ProDER: A Continual Learning Approach for Fault Classification and Localization in Evolving Smart Grids
2511.05420
|
cs.LGcs.AI
|
Emad Efatinasab, Nahal Azadi, Davide Dalle Pezze, Gian Antonio Susto, Chuadhry Mujeeb Ahmed |
Data-driven fault diagnosis models for smart grids are usually trained once on a fixed dataset, whereas in operation new fault types appear and monitoring is extended to new grid zones. Retraining from scratch on all accumulated data is costly, while naively u...Data-driven fault diagnosis models for smart grids are usually trained once on a fixed dataset, whereas in operation new fault types appear and monitoring is extended to new grid zones. Retraining from scratch on all accumulated data is costly, while naively updating the model on new data causes catastrophic forgetting. To address this problem, we formulate fault type classification and fault zone localization as continual learning (CL) problems and design four evaluation scenarios on the IEEE 13-node test feeder, three class-incremental and one domain-incremental. We then propose Prototype-based Dark Experience Replay (ProDER), which extends DER++ with prototype attraction and prototype-level repulsion losses that stabilize the feature space, temperature-scaled logit distillation, and a prototype-aware replay memory that retains both core and boundary samples of each class. ProDER achieves the highest accuracy among the tested CL methods in all scenarios, with an average accuracy of 58.2\%, 6.6 points above the strongest competing method (DPDMR, 51.6\%) and only 3.2 points below joint training (61.4\%). Per scenario, it improves over the strongest competitor by 4.2 to 7.4 points and closes the gap to joint training to as little as 1.0 point in fault type classification, while matching it in fault zone localization. Moreover, it remains the best method when the replay buffer is substantially reduced. These results show that prototype-guided replay is an effective, memory-bounded way to keep fault diagnosis models up to date as the grid evolves, while validation on field measurements remains a necessary next step.
|
| 1364 |
Scalable Decision Making for Games of Imperfect Information
2511.07312
|
cs.LGcs.AI
|
Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter |
Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and sear...Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts, top-human-level play at Stratego---a board wargame with hidden information on a massive scale---has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin---achieving, to our knowledge, the first superhuman result in the game's history---while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.
|
| 1365 |
Global Optimization on Graph-Structured Data via Gaussian Processes with Spectral Representations
2511.07734
|
cs.LGcs.AI
|
Shu Hong, Yongsheng Mei, Mahdi Imani, Tian Lan |
Bayesian optimization (BO) is a powerful framework for optimizing expensive black-box objectives, yet extending it to graph-structured domains remains challenging due to the discrete and combinatorial nature of graphs. Existing approaches often rely either on ...Bayesian optimization (BO) is a powerful framework for optimizing expensive black-box objectives, yet extending it to graph-structured domains remains challenging due to the discrete and combinatorial nature of graphs. Existing approaches often rely either on full graph topology, which is impractical for large or partially observed graphs, or on incremental local exploration, which can lead to slow convergence. We introduce a scalable framework for global optimization over graphs that builds Gaussian process (GP) surrogates from low-rank spectral representations inferred from sparse structural observations. The method fits a low-rank graph surrogate to the observed connections and uses its spectrum to define a kernel over all nodes, so that function evaluations and structural observations jointly refine a global model, enabling efficient global search and principled uncertainty estimation even with limited data. We further show how a limited budget of structural observations can be allocated according to its expected value for the optimization. We also provide theoretical analysis establishing conditions under which the surrogate geometry is identified under different sampling regimes, and bounding the optimization error incurred under an estimated geometry. Experiments on synthetic and real-world datasets demonstrate that our approach achieves lower regret and improved optimization performance compared to prior methods.
|
| 1366 |
Gradient descent dynamics for deep equilibrium models
2511.16976
|
cs.LG
|
Sanjit Dandapanthula, Aaditya Ramdas |
Deep equilibrium models (DEQs) have recently emerged as a powerful paradigm for training infinitely deep weight-tied neural networks that achieve state of the art performance across many modern machine learning tasks. Despite their practical success, theoretic...Deep equilibrium models (DEQs) have recently emerged as a powerful paradigm for training infinitely deep weight-tied neural networks that achieve state of the art performance across many modern machine learning tasks. Despite their practical success, theoretically understanding the gradient descent dynamics for training DEQs remains an area of active research. In this work, we rigorously study the gradient descent dynamics for DEQs in the simple setting of linear models and single-index models, filling several gaps in the literature. We prove a matrix conservation law for linear DEQs which implies that the parameters remain trapped on spheres along gradient flow and use this property to show that gradient flow remains well-conditioned for all time. We then prove linear convergence of gradient descent to a global minimizer for linear DEQs and deep equilibrium single-index models under appropriate initialization and with a sufficiently small step size. In these simple settings, we also study the effect of common Neumann approximations used to reduce the cost of backpropagation; for linear DEQs, we show that increasing the number of terms in the Neumann approximation can prevent convergence to zero risk, even with exact forward solves. For nonlinear single-index DEQs, we prove linear convergence for fixed Neumann approximations, including Jacobian-free backpropagation under suitable assumptions. Finally, we validate our theoretical findings through experiments.
|
| 1367 |
Mixed Data Clustering Survey and Challenges
2512.03070
|
cs.LGcs.AI
|
Guillaume Guerard, Sonia Djebali |
The advent of the big data paradigm has transformed how industries manage and analyze information, ushering in an era of unprecedented data volume, velocity, and variety. Within this landscape, mixed-data clustering has become a critical challenge, requiring i...The advent of the big data paradigm has transformed how industries manage and analyze information, ushering in an era of unprecedented data volume, velocity, and variety. Within this landscape, mixed-data clustering has become a critical challenge, requiring innovative methods that can effectively exploit heterogeneous data types, including numerical and categorical variables. Traditional clustering techniques, typically designed for homogeneous datasets, often struggle to capture the additional complexity introduced by mixed data, underscoring the need for approaches specifically tailored to this setting. Hierarchical and explainable algorithms are particularly valuable in this context, as they provide structured, interpretable clustering results that support informed decision-making. This paper introduces a clustering method grounded in pretopological spaces. In addition, benchmarking against classical numerical clustering algorithms and existing pretopological approaches yields insights into the performance and effectiveness of the proposed method within the big data paradigm.
|
| 1368 |
Spectral Embedding via Chebyshev Bases for Robust DeepONet Approximation
2512.09165
|
cs.LG
|
Muhammad Abid, Omer San |
Deep Operator Networks (DeepONets) have emerged as a powerful framework for data-driven operator learning, providing flexible surrogates for nonlinear mappings arising in partial differential equations (PDEs). However, the standard trunk network, which operate...Deep Operator Networks (DeepONets) have emerged as a powerful framework for data-driven operator learning, providing flexible surrogates for nonlinear mappings arising in partial differential equations (PDEs). However, the standard trunk network, which operates directly on raw spatial or spatiotemporal coordinates through fully connected layers, often struggles to represent sharp gradients, boundary layers, and other non-periodic solution structures on bounded domains. To address these limitations, we introduce the Spectral-Embedded Deep Operator Network (SEDONet), a novel DeepONet architecture in which the trunk is driven by a fixed Chebyshev spectral dictionary instead of coordinate inputs. This non-periodic spectral embedding provides a principled inductive bias for bounded domains, enabling the learned operator to capture fine-scale features that are difficult for Fourier-based or MLP-only trunks to represent. SEDONet is evaluated on the 2-D Poisson equation, 1-D Burgers' equation, 1-D advection-diffusion equation, Allen-Cahn equation, Lorenz-96 chaotic system, and Darcy flow, covering elliptic, hyperbolic, parabolic, chaotic, and multiscale problems. Across all benchmarks, SEDONet consistently achieves the lowest or statistically comparable relative $L^2$ errors among DeepONet, FEDONet, and SEDONet, with improvements of up to 54% over the baseline DeepONet and consistent gains over Fourier-embedded variants on bounded, non-periodic problems. Energy spectrum analyses further demonstrate that SEDONet more accurately preserves intermediate- and high-frequency solution structures. The proposed framework provides a simple, parameter-neutral modification to DeepONets, offering a robust and computationally efficient spectral approach for surrogate modeling of nonlinear operators in scientific computing.
|
| 1369 |
Towards Optimal Valve Prescription for Transcatheter Aortic Valve Replacement (TAVR) Surgery: A Machine Learning Approach
2512.09198
|
cs.LGcs.AI
|
Phevos Paschalidis, Vasiliki Stoumpou, Lisa Everest, Yu Ma, Talhat Azemi |
Transcatheter Aortic Valve Replacement (TAVR) has emerged as a prominent, minimally invasive treatment for patients with severe aortic stenosis, a life-threatening cardiovascular condition. Multiple transcatheter heart valves (THV) have been approved for use i...Transcatheter Aortic Valve Replacement (TAVR) has emerged as a prominent, minimally invasive treatment for patients with severe aortic stenosis, a life-threatening cardiovascular condition. Multiple transcatheter heart valves (THV) have been approved for use in TAVR, but current guidelines regarding valve type prescription remain a topic of ongoing debate within the medical community. We propose a data-driven clinical support tool to identify the optimal valve type with the objective of minimizing the risk of permanent pacemaker implantation (PPI), a predominant postoperative complication. We synthesize a novel dataset, combining U.S. and Greek patient populations, that integrates data from three distinct sources (patient demographics, computed tomography scans, echocardiograms) while harmonizing the different encoding processes specific to each country's record system. We propose leaf-level analysis to leverage the heterogeneity of the patient populations and avoid benchmarking against uncertain counterfactual risk estimates. The final prescriptive model shows a reduction in PPI rates of 26% and 16% compared to the current standard of care in our internal U.S. population and external, Greek validation set, respectively. To the best of our knowledge, this work represents the first unified, personalized prescription strategy for THV selection in TAVR.
|
| 1370 |
Contrastive Time Series Forecasting with Anomalies
2512.11526
|
cs.LGcs.AI
|
Joel Ekstrand, Zahra Taghiyarrenani, Slawomir Nowaczyk |
Time series forecasting predicts future values from past data. In real-world settings, some anomalous events have lasting effects and influence the forecast, while others are short-lived and should be ignored. Standard forecasting models fail to make this dist...Time series forecasting predicts future values from past data. In real-world settings, some anomalous events have lasting effects and influence the forecast, while others are short-lived and should be ignored. Standard forecasting models fail to make this distinction, often either overreacting to noise or missing persistent shifts. We propose Co-TSFA (Contrastive Time Series Forecasting with Anomalies), a regularization framework that learns when to ignore anomalies and when to respond. Co-TSFA generates input-only and input-output augmentations to model forecast-irrelevant and forecast-relevant anomalies, and introduces a latent-output alignment loss that ties representation changes to forecast changes. This encourages invariance to irrelevant perturbations while preserving sensitivity to meaningful distributional shifts. Experiments on the Traffic and Electricity benchmarks, as well as on a real-world cash-demand dataset, demonstrate that Co-TSFA improves performance under anomalous conditions while maintaining accuracy on normal data. An anonymized GitHub repository with the implementation of Co-TSFA is provided and will be made public upon acceptance.
|
| 1371 |
Amortized Acquisition Optimization for Bayesian Optimization with Variational Mutual Information
2601.08172
|
cs.LG
|
Farhad Mirkarimi |
Bayesian optimization (BO) of expensive black-box functions is traditionally addressed with Gaussian processes (GPs), which scale cubically with observations, or Bayesian neural networks (BNNs), which incur costly posterior sampling and inner-loop acquisition ...Bayesian optimization (BO) of expensive black-box functions is traditionally addressed with Gaussian processes (GPs), which scale cubically with observations, or Bayesian neural networks (BNNs), which incur costly posterior sampling and inner-loop acquisition optimization. We propose VBO-MI (Variational Bayesian Optimization with Mutual Information), a fully gradient-based BO framework that requires no explicit GP prior or fixed parametric posterior family over objective function and treats it as a strict black box. An actor-critic architecture pairs an action-net with a variational critic that estimates information gain, eliminating the acquisition optimization bottleneck and achieving up to $10^{2}\times$ fewer FLOPs than BNN-BO baselines. A lightweight surrogate network further reduces real function queries to one batch per iteration. We establish consistency guarantees and evaluate VBO-MI on synthetic benchmarks (Ackley, Levy, Griewank) and real-world tasks (Rover Trajectory, Lunar Lander, Pest Control), demonstrating competitive or superior performance over the baselines.
|
| 1372 |
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
2601.17063
|
cs.LGcs.AI
|
Byeongju Kim, Jungwan Lee, Donghyeon Han, Hoi-Jun Yoo, Sangyeob Kim |
Recently, Mixture-of-Experts (MoE) models have gained attention for efficiently scaling large language models. Although these models are extremely large, their sparse activation enables inference to be performed by accessing only a fraction of the model at a t...Recently, Mixture-of-Experts (MoE) models have gained attention for efficiently scaling large language models. Although these models are extremely large, their sparse activation enables inference to be performed by accessing only a fraction of the model at a time. This property opens the possibility of on-device inference of MoE, which was previously considered infeasible for such large models. Consequently, various systems have been proposed to leverage this sparsity and enable efficient MoE inference for edge devices. However, previous MoE inference systems like Fiddler[8] or DAOP[13] rely on DRAM-based offloading and are not suitable for memory constrained on-device environments. As recent MoE models grow to hundreds of gigabytes, RAM-offloading solutions become impractical. To address this, we propose FlashMoE, a system that offloads inactive experts to SSD, enabling efficient MoE inference under limited RAM. FlashMoE incorporates a lightweight ML-based caching strategy that adaptively combines recency and frequency signals to maximize expert reuse, significantly reducing storage I/O. In addition, we built a user-grade desktop platform to demonstrate the practicality of FlashMoE. On this real hardware setup, FlashMoE improves cache hit rate by up to 51% over well-known offloading policies such as LRU and LFU, and achieves up to 2.6x speedup compared to existing MoE inference systems.
|
| 1373 |
Stability and Generalization of Stochastic Nonconvex Optimization under Heavy-Tailed Noise: From Gradient Norm to Goldstein Stationarity
2601.19730
|
cs.LG
|
Hongxu Chen, Ke Wei, Xiaoming Yuan, Luo Luo |
The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most existing works on this phen...The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most existing works on this phenomenon focus on the convergence of optimization errors, while the analysis for generalization bounds under the heavy-tailed gradient noise remains limited. In this paper, we develop stability-based frameworks to establish dimension-independent generalization bounds of nonconvex optimization under heavy-tailed noise for both smooth and nonsmooth settings. Specifically, we introduce a truncation argument to achieve the generalization error bound based on the algorithmic stability under finite $p$th moment conditions with $p\in(1,2]$. For the smooth case, we apply our framework to derive population gradient norm bounds for several popular stochastic optimization algorithms, including clipped and normalized stochastic gradient descent, as well as their mini-batch and momentum variants. For the nonsmooth case, we establish a population Goldstein stationarity guarantee for online-to-nonconvex conversion methods. It is worth noting that our theory for the nonsmooth setting is established by introducing an on-average gradient stability for the smooth surrogate function, which does not require the additional weak convexity assumption used in existing analysis based on the Moreau envelope.
|
| 1374 |
Missing-Data-Induced Phase Transitions in Spectral Partial Least Squares
2601.21294
|
cs.LG
|
Anders Gj{\o}lbye, Emma Kargaard, Ida Kargaard, Lina Skerath, Hiba Nassar |
Spectral partial least squares (PLS-SVD) estimates the directions shared by two views of the same samples. We analyze it using a rank-one regression model, with entries of both views missing completely at random and filled with zeros. Missing response entries ...Spectral partial least squares (PLS-SVD) estimates the directions shared by two views of the same samples. We analyze it using a rank-one regression model, with entries of both views missing completely at random and filled with zeros. Missing response entries weaken the signal, whereas missing design entries also tilt the recovered direction. The setting is the proportional regime, with Gaussian response noise and a design that is column-orthogonal before masking. For incoherent designs and planted directions, we prove that as the signal diverges, the squared overlap with the truth approaches $1$ on the response side but only a ceiling on the design side. This ceiling is strictly below $1$ whenever design entries are missing, unless samples of vanishing leverage carry the signal. For this estimator, a stronger signal therefore cannot remove the design-side error. Under equal leverage and an explicit spectral condition on the masked design, which we prove for designs taken from nested Hadamard matrices, we establish a phase transition and derive the recovery curves. Under the same conditions, the critical signal strength is exactly proportional to $\rho_y^{-1/2}$, with $\rho_y$ the fraction of response entries retained. A heuristic replica derivation reproduces the equations behind these curves, and simulations follow them, including on random orthogonal designs outside the verified class. In semi-synthetic experiments on real biological designs with unequal leverage, the predicted ceiling varies across directions, and high-signal recovery tracks it. These results bring the theory of PLS-SVD closer to what is known for PCA under homogeneous missingness and show that the loss of signal-to-noise ratio observed there does not, by itself, describe missing design entries.
|
| 1375 |
Learning-to-Defer in Non-Stationary Time Series via Switching State-Space Models
2601.22538
|
cs.LG
|
Yannis Montreuil, Letian Yu, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi |
Learning-to-defer (L2D) lets a predictor decide, at each round, whether to issue its own forecast or pay for an expert's. In non-stationary time series this decision must keep adapting, although deployment reveals only the consulted expert's forecast while a h...Learning-to-defer (L2D) lets a predictor decide, at each round, whether to issue its own forecast or pay for an expert's. In non-stationary time series this decision must keep adapting, although deployment reveals only the consulted expert's forecast while a historical archive records every expert with the target. L2D-SLDS learns from this archive a switching state-space model of the target and all expert forecasts, whose shared and expert-specific states describe how experts move together and apart. Its predictive law supplies the internal forecast and the expected cost of every consultation, which a greedy router minimizes, and one consultation also updates the beliefs about unconsulted and unavailable experts. We prove sublinear regret against a changing conditional-risk oracle without exploration, when the candidate models are accurate and either the archive separates them or live feedback reveals cost differences. On three real datasets, L2D-SLDS has the lowest cost among eight bandit routers and adapts its consultation rate to the fee.
|
| 1376 |
Lethe: Principled Dual-Stream Update for Persistent Knowledge Erasure in Federated Unlearning
2601.22601
|
cs.LG
|
Wentai Wu, Hanwei Tan, Yijun Quan, Haixia Peng, Ligang He |
Federated unlearning (FU) aims to erase knowledge from a global model. Existing studies commonly assume that federated collaboration terminates after unlearning, overlooking a deployment-realistic scenario where training continues on the remaining clients afte...Federated unlearning (FU) aims to erase knowledge from a global model. Existing studies commonly assume that federated collaboration terminates after unlearning, overlooking a deployment-realistic scenario where training continues on the remaining clients after deletion requests are fulfilled. In this work, we identify a critical failure mode, termed knowledge resurfacing, revealing that continued training on retained data alone can reactivate unlearned knowledge in a few rounds. Empirically, we demonstrate that many state-of-the-art FU methods are prone to knowledge resurfacing. We then propose Lethe, a novel unlearning method for persistent knowledge erasure in federated settings. In each iteration, Lethe operates on a forget stream from the unlearning client and a retain stream from the retained clients. It redirects unlearning updates toward a region where the two streams are anti-aligned, discouraging retained-data training from moving back toward the forgotten knowledge. Consequently, Lethe ensures stronger unlearning persistence during subsequent federated training. Extensive experiments across diverse models, datasets, and unlearning levels validate that Lethe supports all levels of unlearning in a unified manner across both CV and NLP tasks, demonstrating consistently low RR, below 1% in most cases, even after an extremely long horizon of follow-up training.
|
| 1377 |
PlatoLTL: Scaling LTL-Guided Multi-Task RL
2601.22891
|
cs.LG
|
Jacques Cloete, Mathias Jackermeier, Ioannis Havoutis, Alessandro Abate |
Linear temporal logic (LTL) has emerged as a powerful formalism for specifying structured, temporally extended tasks in multi-task reinforcement learning (RL). However, while existing approaches in LTL-guided multi-task RL demonstrate success in simple environ...Linear temporal logic (LTL) has emerged as a powerful formalism for specifying structured, temporally extended tasks in multi-task reinforcement learning (RL). However, while existing approaches in LTL-guided multi-task RL demonstrate success in simple environments, they suffer from challenges in representation learning and exploration efficiency when applied to high-dimensional environments with parameterized specifications. We present PlatoLTL, which elevates state-of-the-art methods to address both challenges. We model atomic propositions as instances of atomic predicates and inject task parameters directly into the goal embedding to enable efficient generalization. We also leverage simple priors on closeness to predicate satisfaction to enable and accelerate learning of complex tasks without biasing the optimal policy. We validate our approach on challenging environments including robotic manipulation and multi-drone navigation.
|
| 1378 |
Divide-and-Conquer CoT: RL for Reducing Latency via Parallel Reasoning
2601.23027
|
cs.LG
|
Arvind Mahankali, Kaiyue Wen, Tengyu Ma |
Long chain-of-thought reasoning (Long CoT) is now fundamental to state-of-the-art LLMs, especially in mathematical reasoning. However, LLM generation is highly sequential, and long CoTs lead to a high latency. We propose to train Divide-and-Conquer CoT (DC-CoT...Long chain-of-thought reasoning (Long CoT) is now fundamental to state-of-the-art LLMs, especially in mathematical reasoning. However, LLM generation is highly sequential, and long CoTs lead to a high latency. We propose to train Divide-and-Conquer CoT (DC-CoT) to reduce the latency. With DC-CoT, the model can act as a director that identifies distinct subtasks that can be performed in parallel in its reasoning process, and then spawns workers to execute the subtasks. Our goal is to achieve high accuracy, with a low longest path length, which is a theoretical measure of the latency needed for the response. We start with a long CoT base model (DeepScaleR-1.5B-Preview), and first use SFT with a small curated demonstration set to initialize its ability to spawn workers in a certain format. Because SFT degrades the accuracy significantly, we design a multi-stage RL algorithm, with various data filtering strategies, to recover the accuracy while decreasing the longest path length. Across several benchmarks including AIME 2024 and HMMT 2025, DC-CoT achieves similar accuracy as DeepScaleR-1.5B-Preview while decreasing longest path length by 35-40%. Our code, SFT dataset and models are publicly available at https://github.com/amahankali10/DC_CoT_RL_for_Low_Latency_CoT_with_Parallel_Reasoning.
|
| 1379 |
Quantum Model Parallelism for MRI-Based Classification of Alzheimer's Disease Stages
2602.00128
|
cs.LG
|
Emine Akpinar, Murat Oduncuoglu |
With increasing life expectancy, AD has become a major global health concern. While classical AI-based methods have been developed for early diagnosis and stage classification of AD, growing data volumes and limited computational resources necessitate faster, ...With increasing life expectancy, AD has become a major global health concern. While classical AI-based methods have been developed for early diagnosis and stage classification of AD, growing data volumes and limited computational resources necessitate faster, more efficient approaches. Quantum-based AI methods, which leverage superposition and entanglement principles along with high-dimensional Hilbert space, can surpass classical approaches' limitations and offer higher accuracy for high-dimensional, heterogeneous, and noisy data. In this study, a Quantum-Based Parallel Model (QBPM) architecture is proposed for the efficient classification of AD stages using MRI datasets, inspired by the principles of classical model parallelism. The proposed model leverages quantum advantages by employing two distinct quantum circuits, each incorporating rotational and entanglement blocks, running in parallel on the same quantum simulator. The classification performance of the model was evaluated on two different datasets to assess its overall robustness and generalization capability. The proposed model demonstrated high classification accuracy across both datasets, highlighting its overall robustness and generalization capability. Results obtained under high-level Gaussian noise, simulating real-world conditions, further provided experimental evidence for the model's applicability not only in theoretical but also in practical scenarios. Moreover, compared with five different classical transfer learning methods, the proposed model demonstrated its efficiency as an alternative to classical approaches by achieving higher classification accuracy and comparable execution time while utilizing fewer circuit parameters. The results indicate that the proposed QBPM architecture represents an innovative and powerful approach for the classification of stages in complex diseases such as Alzheimer's.
|
| 1380 |
STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs
2602.02180
|
cs.LG
|
Weikang Meng, Liangyu Huo, Yadan Luo, Jiawen Guan, Jingyi Zhang |
Linearizing pretrained large language models (LLMs) primarily relies on intra-layer hybrid attention mechanisms to alleviate the quadratic complexity of standard softmax attention. Existing methods perform token routing based on sliding-window partitions, resu...Linearizing pretrained large language models (LLMs) primarily relies on intra-layer hybrid attention mechanisms to alleviate the quadratic complexity of standard softmax attention. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local-global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training-inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2% relative improvement over prior linearized attention methods on long-context benchmarks. Code is available at this URL.
|
| 1381 |
Routing-Aware Safety Alignment for Mixture-of-Experts Models
2602.04448
|
cs.LGcs.AI
|
Jiacheng Liang, Yuhui Wang, Tanqiu Jiang, Ting Wang |
Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we o...Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts. To address this challenge, we propose RASA, a routing-aware expert-level alignment framework that explicitly repairs Safety-Critical Experts while preventing routing-based bypasses. RASA identifies experts disproportionately activated by successful jailbreaks, selectively fine-tunes only these experts under fixed routing, and subsequently enforces routing consistency with safety-aligned contexts. Across two representative MoE architectures and a diverse set of jailbreak attacks, RASA achieves near-perfect robustness, strong cross-attack generalization, and substantially reduced over-refusal, while preserving general capabilities on benchmarks such as MMLU, GSM8K, and TruthfulQA. Our results suggest that robust MoE safety alignment benefits from targeted expert repair rather than global parameter updates, offering a practical and architecture-preserving alternative to prior approaches.
|
| 1382 |
Multi-Level Strategic Classification: Incentivizing Improvement through Promotion and Relegation Dynamics
2602.11439
|
cs.LG
|
Ziyuan Huang, Lina Alkarmi, Mingyan Liu |
Strategic classification studies the problem where self-interested individuals or agents manipulate their response to obtain favorable decision outcomes made by classifiers, typically turning to dishonest actions when they are less costly than genuine efforts....Strategic classification studies the problem where self-interested individuals or agents manipulate their response to obtain favorable decision outcomes made by classifiers, typically turning to dishonest actions when they are less costly than genuine efforts. While existing studies on sequential strategic classification primarily focus on optimizing dynamic classifier weights, we depart from these weight-centric approaches by analyzing the design of classifier thresholds and difficulty progression within a multi-level promotion-relegation framework. Our model captures the critical inter-temporal incentives driven by an agent's farsightedness, skill retention, and a leg-up effect where qualification and attainment can be self-reinforcing. We characterize the agent's optimal long-term strategy and demonstrate that a principal can design a sequence of thresholds to effectively incentivize honest effort. Crucially, we prove that under mild conditions, this mechanism enables agents to reach arbitrarily high levels solely through genuine improvement efforts.
|
| 1383 |
Same Pieces, Different Servers: A Tetris Benchmark for AI Agents as Served
2603.02348
|
cs.LGcs.AI
|
Haochuan Wang |
An agent meets a model as served: through an endpoint with a price card, a shared cache and other tenants, or on whatever hardware a self-hosted model runs. Benchmarks rank the weights. We introduce a Tetris benchmark that measures what agents get from models ...An agent meets a model as served: through an endpoint with a price card, a shared cache and other tenants, or on whatever hardware a self-hosted model runs. Benchmarks rank the weights. We introduce a Tetris benchmark that measures what agents get from models as served: every move is scored against an oracle, and every agent receives the same pieces. In five pre-specified experiments with nine open-weight models on one serverless provider, plus an open decision model self-hosted on a CPU, we find that price and size do not predict decision quality; that resending history costs almost nothing when cached input is free, would cost eleven to twelve times more if it were not, and makes play worse; that deployments keep between 4 and more than 64 agent contexts warm, in line with their throughput rather than the model's KV-cache size; that an agent's own long requests slow its slowest short decisions more than tenfold, which a simple admission rule cuts by 59% without hurting play; and that the decision model plays mid-pack but takes 27 s per move on a CPU, about 650 times longer than reported on a GPU. Architecture predicts some of what an agent sees; the deployment sets the rest.
|
| 1384 |
Dissecting Quantization Error: A Concentration-Alignment Perspective
2603.04359
|
cs.LGcs.AI
|
Marco Federici, Boris van Breugel, Paul Whatmough, Markus Nagel |
Quantization can drastically increase the efficiency of large language and vision models, but typically incurs an accuracy drop. Recently, function-preserving transforms (e.g. rotations, Hadamard transform, channel-wise scaling) have been successfully applied ...Quantization can drastically increase the efficiency of large language and vision models, but typically incurs an accuracy drop. Recently, function-preserving transforms (e.g. rotations, Hadamard transform, channel-wise scaling) have been successfully applied to reduce post-training quantization error, yet a principled explanation remains elusive. We analyze linear-layer quantization via the signal-to-quantization-noise ratio (SQNR), showing that for uniform integer quantization at a fixed bit width, SQNR decomposes into (i) the concentration of weights and activations (capturing spread and outliers), and (ii) the alignment of their dominant variation directions. This reveals an actionable insight: beyond concentration - the focus of most prior transforms (e.g. rotations or Hadamard) - improving alignment between weight and activation can further reduce quantization error. Motivated by this, we introduce block Concentration-Alignment Transforms (CAT), a lightweight linear transformation that uses a covariance estimate from a small calibration set to jointly improve concentration and alignment, approximately maximizing SQNR. Experiments across several LLMs show that CAT consistently matches or outperforms prior transform-based quantization methods at 4-bit precision, confirming the insights gained in our framework.
|
| 1385 |
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
2603.07904
|
cs.LG
|
Zihao Zheng, Hangyu Cao, Chibang Tao, Zhihao Mao, Lingyue Zhang |
Vision-language-action (VLA) models achieve favorable task performance, yet runtime errors in closed-loop execution evolve with alternating updates of actions and observations. Existing VLA quantization methods mainly trade off inference speed and model perfor...Vision-language-action (VLA) models achieve favorable task performance, yet runtime errors in closed-loop execution evolve with alternating updates of actions and observations. Existing VLA quantization methods mainly trade off inference speed and model performance while rarely investigating how quantization alters runtime errors. By comparing error distributions between full-precision and quantized models, we identify distinct evolutionary patterns for the two types of errors: quantization enlarges the variance of translation error distributions and increases their dispersion, whereas the distribution center of rotation errors gradually shifts across execution steps, demonstrating cumulative drift. Motivated by this observation, we rethink the optimal quantization strategy for VLA models and propose \textit{DyQ-VLA}, a runtime-error-aware quantization framework, which dynamically selects activation precision according to execution steps and tracks as well as compensates rotation errors via accumulated quantization residuals. Corresponding operators and runtime adaptation strategies are devised within the framework to enable dynamic-precision execution. Experiments show that \textit{DyQ-VLA} achieves a 1.88 to 1.93 inference speedups while maintaining comparable or higher average task success rates. Moreover, it reduces mean execution steps by 17.7% to 19.8%. Our code is here: https://anonymous.4open.science/r/DyQ-VLA-7F51/.
|
| 1386 |
Physics-informed neural operator for parametric phase-field modelling of interfacial degradation and microstructural evolution
2603.09693
|
cs.LG
|
Nanxi Chen, Airong Chen, Rujin Ma |
Predicting interfacial degradation and microstructural evolution through phase-field modelling is computationally intensive, particularly for parametric studies. While neural operators such as the Fourier neural operator (FNO) offer a promising route to accele...Predicting interfacial degradation and microstructural evolution through phase-field modelling is computationally intensive, particularly for parametric studies. While neural operators such as the Fourier neural operator (FNO) offer a promising route to acceleration, the absence of physical constraints compromises their generalisation and long-term stability. Here, we present PF-PINO, a physics-informed neural operator framework that embeds phase-field governing equation residuals directly into the training loss, together with gradient normalisation loss balancing and a staggered training scheme for coupled multi-field systems. PF-PINO is validated on four benchmarks including pencil-electrode corrosion, electro-polishing corrosion, dendritic solidification, and spinodal decomposition, collectively probing parametric generalisation across orders-of-magnitude kinetic regimes, non-periodic boundaries with stochastic morphologies, coupled multi-field dynamics with out-of-distribution extrapolation, and long-term autoregressive stability under complex pattern formation. Across all benchmarks, PF-PINO reduces relative $L^2$ errors by up to an order of magnitude over FNO and achieves sub-mesh-size interface accuracy, establishing a data-efficient surrogate for high-throughput parametric phase-field assessment in computational materials science and engineering applications.
|
| 1387 |
PDE-SSM: A Spectral State Space Approach to Spatial Mixing in Diffusion Transformers
2603.13663
|
cs.LG
|
Eshed Gal, Moshe Eliasof, Eldad Haber |
The success of vision transformers-especially for generative modeling-is limited by the quadratic cost and weak spatial inductive bias of self-attention. We propose PDE-SSM, a spatial state-space block that replaces attention with a learnable convection-diffus...The success of vision transformers-especially for generative modeling-is limited by the quadratic cost and weak spatial inductive bias of self-attention. We propose PDE-SSM, a spatial state-space block that replaces attention with a learnable convection-diffusion-reaction partial differential equation. This operator encodes a strong spatial prior by modeling information flow via physically grounded dynamics rather than all-to-all token interactions. Solving the PDE in the Fourier domain yields global coupling with near-linear complexity of $O(N \log N)$, delivering a principled and scalable alternative to attention. We integrate PDE-SSM into a flow-matching generative model to obtain the PDE-based Diffusion Transformer PDE-SSM-DiT. Empirically, PDE-SSM-DiT matches or exceeds the performance of state-of-the-art Diffusion Transformers while substantially reducing compute. Our results show that, analogous to 1D settings where SSMs supplant attention, multi-dimensional PDE operators provide an efficient, inductive-bias-rich foundation for next-generation vision models.
|
| 1388 |
DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management
2603.19621
|
cs.LGcs.AI
|
Yaqi Xie, Xinru Hao, Jiaxi Liu, Will Ma, Linwei Xin |
Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyp...Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyperparameters used during training. In this paper, we show that by imposing policy regularizations, grounded in classical inventory concepts such as "Base Stock", we can significantly accelerate hyperparameter tuning and improve the final performance of several DRL methods. We report details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall. We also include extensive synthetic experiments, which show that policy regularizations reshape the narrative on what is the best DRL method for inventory management.
|
| 1389 |
NASimJax: A GPU-Accelerated Policy Learning Framework for Penetration Testing
2603.19864
|
cs.LG
|
Raphael Simon, Jos\'e Carrasquel, Elli Makdis Antoun, Wim Mees, Pieter Libin |
Penetration testing - the practice of simulating cyberattacks to identify vulnerabilities - is a complex sequential decision-making task that is inherently partially observable and features large action spaces. Existing RL simulators for this domain are CPU-bo...Penetration testing - the practice of simulating cyberattacks to identify vulnerabilities - is a complex sequential decision-making task that is inherently partially observable and features large action spaces. Existing RL simulators for this domain are CPU-bound and fixed to narrow scenarios, making it infeasible to train policies that generalize across networks. We present NASimJax, a JAX-native framework that formulates penetration testing as a Contextual POMDP and introduces a network generation pipeline producing structurally diverse, guaranteed-solvable scenarios. The framework reaches up to 80$\times$ higher environment throughput than previous simulators, enabling experiments on larger networks and tractable hyperparameter searches. We provide PPO and PQN baselines and conduct the first systematic evaluation of unsupervised environment design for penetration testing. We find that Prioritized Level Replay and ACCEL handle dense training distributions better than Domain Randomization, and that training on sparser topologies yields an implicit curriculum that improves generalization - even to topologies denser than those seen during training. A recurrent PPO variant confirms that these distributional findings are not an artifact of feed-forward policies. The code is available at: https://github.com/raphsimon/NASimJax.
|
| 1390 |
High dimensional theory of two-phase optimizers
2603.26954
|
cs.LG
|
Atish Agarwala |
The trend towards larger training setups has brought a renewed interest in partially asynchronous two-phase optimizers which optimize locally and then synchronize across workers. Additionally, recent work suggests that the one-worker version of one of these al...The trend towards larger training setups has brought a renewed interest in partially asynchronous two-phase optimizers which optimize locally and then synchronize across workers. Additionally, recent work suggests that the one-worker version of one of these algorithms, DiLoCo, shows promising results as a (synchronous) optimizer. Motivated by these studies we present an analysis of LA-DiLoCo, a simple member of the DiLoCo family, on a high-dimensional linear regression problem. We show that the one-worker variant, LA, provides a different tradeoff between signal and noise than SGD, which is beneficial in many scenarios. We also show that the multi-worker version generates more noise than the single worker version, but that this additional noise generation can be ameliorated by appropriate choice of hyperparameters. We conclude with an analysis of SLA -- LA with momentum -- and show that stacking two momentum operators gives an opportunity for acceleration via a non-linear transformation of the "effective'' Hessian spectrum, which is maximized for Nesterov momentum. Altogether our results show that two-phase optimizers represent a fruitful new paradigm for understanding and improving training algorithms.
|
| 1391 |
Pre-Deployment Complexity Estimation for Federated Perception Systems
2603.28282
|
cs.LGcs.AI
|
KMA Solaiman, Shafkat Islam, Ruy de Oliveira, Bharat Bhargava |
Edge AI systems increasingly rely on federated learning to train perception models in distributed, privacy-preserving, and resource-constrained environments. Before training, however, practitioners often lack practical tools for estimating task difficulty in t...Edge AI systems increasingly rely on federated learning to train perception models in distributed, privacy-preserving, and resource-constrained environments. Before training, however, practitioners often lack practical tools for estimating task difficulty in terms of expected accuracy and communication effort. We present a classifier-agnostic, pre-deployment framework that combines intrinsic data properties such as dimensionality, sparsity, and heterogeneity, with client-distribution composition to estimate learning complexity in federated perception systems. Using federated learning as a representative distributed training setting, we examine how learning difficulty varies across different federated configurations. Experiments on three MNIST variants show strong negative correlations between the combined complexity metric and maximum and average federated accuracy, while the intrinsic and distributed components exhibit consistent relationships with communication effort. These findings suggest that complexity estimation can serve as a practical diagnostic tool for resource planning, dataset assessment, and feasibility evaluation in edge-deployed perception systems.
|
| 1392 |
GUIDE: Reinforcement Learning for Behavioral Action Support in Type 1 Diabetes
2604.00385
|
cs.LG
|
Saman Khamesian, Sri Harini Balaji, Di Yang Shi, Stephanie M. Carpenter, Daniel E. Rivera |
Type 1 diabetes (T1D) management requires continuous adjustment of insulin and lifestyle behaviors to maintain blood glucose within a safe target range. Although automated insulin delivery (AID) systems have improved glycemic outcomes, many patients still fail...Type 1 diabetes (T1D) management requires continuous adjustment of insulin and lifestyle behaviors to maintain blood glucose within a safe target range. Although automated insulin delivery (AID) systems have improved glycemic outcomes, many patients still fail to achieve recommended clinical targets. Current reinforcement learning (RL)-based methods focus primarily on insulin-only treatment and do not provide behavioral recommendations for glucose control. To address this gap, we propose GUIDE, an RL-based decision-support framework designed to complement AID technologies by providing structured behavioral recommendations defined by intervention type, magnitude, and timing, including bolus insulin administration and carbohydrate intake events. GUIDE integrates a patient-specific glucose predictor trained on real-world continuous glucose monitoring data and supports offline and online RL algorithms within a unified environment. The algorithms are evaluated across 25 individuals from the AZT1D dataset and 12 individuals from the OhioT1DM dataset. Among the evaluated algorithms, CQL-BC achieved the best performance on both datasets, with mean time-in-range values of 84.18 $\pm$ 19.89% on AZT1D and 77.64 $\pm$ 8.81% on OhioT1DM. It also maintained time-below-range values of 0.43 $\pm$ 1.27% and 2.81 $\pm$ 3.56%, respectively. Behavioral analysis yielded mean cosine similarities of 0.767 $\pm$ 0.128 on AZT1D and 0.774 $\pm$ 0.162 on OhioT1DM, indicating that the learned policy preserves key structural characteristics of patient action patterns. These findings demonstrate the potential of conservative offline RL with a structured behavioral action space to provide personalized and behaviorally plausible decision support for diabetes management.
|
| 1393 |
Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method
2604.01279
|
cs.LGcs.AI
|
Samuel Bright-Thonney, Thomas R. Harvey, Andre Lukas, Jesse Thaler |
We introduce Sven (Singular Value dEsceNt), a new optimization algorithm for neural networks that exploits the natural decomposition of loss functions into a sum over individual data points, rather than reducing the full loss to a single scalar before computin...We introduce Sven (Singular Value dEsceNt), a new optimization algorithm for neural networks that exploits the natural decomposition of loss functions into a sum over individual data points, rather than reducing the full loss to a single scalar before computing a parameter update. Sven treats each data point's residual as a separate condition to be satisfied simultaneously, using the Moore-Penrose pseudoinverse of the loss Jacobian to find the minimum-norm parameter update that best satisfies all conditions at once. In practice, this pseudoinverse is approximated via a truncated singular value decomposition, retaining only the $k$ most significant directions. We show that Sven can be understood as a natural gradient method generalized to the overparametrized regime, recovering natural gradient descent in the underparametrized limit. We test Sven on a variety of regression and classification tasks, including small-scale language modeling with transformers, and find that it is competitive with leading baselines such as Adam, Muon, and K-FAC. We also discuss Sven's memory overhead, which presents a barrier to scaling under a naive implementation, and introduce an optimized implementation that keeps memory usage on par with standard baselines under mild restrictions on model architecture. Beyond standard machine learning benchmarks, we anticipate that Sven will find natural application in scientific computing settings where custom loss functions decompose into several conditions.
|
| 1394 |
Tensor-based computation of the Koopman generator via operator logarithm
2604.07685
|
cs.LG
|
Tatsuya Kishimoto, Jun Ohkubo |
Identifying governing equations of nonlinear dynamical systems from data is challenging. While sparse identification of nonlinear dynamics (SINDy) and its extensions are widely used for system identification, operator-logarithm approaches use the logarithm to ...Identifying governing equations of nonlinear dynamical systems from data is challenging. While sparse identification of nonlinear dynamics (SINDy) and its extensions are widely used for system identification, operator-logarithm approaches use the logarithm to avoid time differentiation, enabling larger sampling intervals. However, they still suffer from the curse of dimensionality. Then, we propose a data-driven method to compute the Koopman generator in a low-rank tensor train (TT) format by taking logarithms of Koopman eigenvalues while preserving the TT format. Experiments on 4-dimensional Lotka-Volterra and 10-dimensional Lorenz-96 systems show accurate recovery of vector field coefficients and scalability to higher-dimensional systems.
|
| 1395 |
TempusBench: An Evaluation Framework for Time-Series Forecasting
2604.11529
|
cs.LG
|
Denizalp Goktas, Gerardo Ria\~no-Brice\~no, Omkar Tekawade, Alif Abdullah, Md Yasif Jamal |
Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent open-source models demonstrate the promis...Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent open-source models demonstrate the promise of TSFMs, the field lacks a comprehensive and community-accepted model evaluation framework. We see at least four major issues impeding progress on the development of such a framework. First, existing evaluation frameworks comprise benchmark forecasting tasks derived from often outdated datasets (e.g., M3), many of which lack clear metadata and overlap with the corpora used to pre-train TSFMs. Second, these frameworks evaluate models along a narrowly defined set of benchmark forecasting tasks, such as forecast horizon length or domain, but overlook core statistical properties such as non-stationarity and seasonality. Third, domain-specific models (e.g., XGBoost) are often compared unfairly, as existing frameworks do not enforce a systematic and consistent hyperparameter tuning convention for all models. Fourth, visualization tools for interpreting comparative performance are lacking. To address these issues, we introduce TempusBench, an open-source evaluation framework for TSFMs. TempusBench consists of 1) new datasets which are not included in existing TSFM pretraining corpora, 2) a set of novel benchmark tasks that go beyond existing ones, 3) a model evaluation pipeline with a standardized hyperparameter tuning protocol, and 4) a tensorboard-based visualization interface. We provide access to our code on GitHub: https://github.com/Smlcrm/TempusBench and maintain a live leaderboard at https://smlcrm.com/tempusbench.
|
| 1396 |
Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization
2604.14765
|
cs.LG
|
Mathias Dus (IRMA) |
We present a geometric framework for Reinforcement Learning (RL) that views policies as maps into the Wasserstein space of action probabilities. First, we define a Riemannian structure induced by stationary distributions, providing rigorous guarantees for thei...We present a geometric framework for Reinforcement Learning (RL) that views policies as maps into the Wasserstein space of action probabilities. First, we define a Riemannian structure induced by stationary distributions, providing rigorous guarantees for their existence in a general context. We then define the tangent space of policies and characterize the geodesics, specifically addressing the measurability of vector fields mapping from the state space to the tangent space of probability measures over the action space. Next, we formulate a general RL optimization problem and construct a gradient flow using Otto's calculus. We compute the gradient and the Hessian of the energy (i.e., the expected cumulative cost), providing a formal second-order analysis. Finally, we illustrate the method with numerical examples: we compute the gradient directly from our theoretical formalism for low-dimensional problems, and we illustrate the numerical applicability of our formal framework to high-dimensional continuous control by parameterizing the policy with a neural network optimized via an ergodic approximation of the cost.
|
| 1397 |
Chronax: A Jax Library for Forecasting and Conformal Inference
2604.16719
|
cs.LG
|
Xan Carey, Arsh Singh, Yash Deshmukh, Sunit Jadhav, Omkar Tekawade |
Time-series forecasting is central to many scientific and industrial domains, such as energy systems, climate modeling, finance, and retail. While forecasting methods have evolved from classical statistical models to automated, and neural approaches, the surro...Time-series forecasting is central to many scientific and industrial domains, such as energy systems, climate modeling, finance, and retail. While forecasting methods have evolved from classical statistical models to automated, and neural approaches, the surrounding software ecosystem remains anchored to the traditional Python numerical stack. Existing libraries rely on interpreter-driven execution and object-oriented abstractions, limiting composability, large-scale parallelism, and integration with modern differentiable and accelerator-oriented workflows. Meanwhile, today's forecasting increasingly involves large collections of heterogeneous time series data, irregular covariates, and frequent retraining, placing new demands on scalability and execution efficiency. JAX offers an alternative paradigm to traditional stateful numerical computation frameworks based on pure functions and program transformations such as just-in-time compilation and automatic vectorization, enabling end-to-end optimization across CPUs, GPUs, and TPUs. However, this modern paradigm has not yet been fully incorporated into the design of forecasting systems. We introduce Chronax, a JAX-native time-series forecasting library that rethinks forecasting abstractions around functional purity, composable transformations, and accelerator-ready execution. By representing preprocessing, modeling, and multi-horizon prediction as pure JAX functions, Chronax enables scalable multi-series forecasting, model-agnostic conformal uncertainty quantification, and seamless integration with modern machine learning and scientific computing pipelines.
|
| 1398 |
Untrained CNNs Exceed Backpropagation in V1 Alignment at High Evaluation Resolution: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI
2604.16875
|
cs.LG
|
Nils Leutenegger |
Revised version (v4): noise ceiling replaced, analyses restricted to stimulus pairs presented in different fMRI runs, repaired evaluation of all five conditions. We compare four learning rules (backpropagation (BP), feedback alignment (FA), predictive coding (...Revised version (v4): noise ceiling replaced, analyses restricted to stimulus pairs presented in different fMRI runs, repaired evaluation of all five conditions. We compare four learning rules (backpropagation (BP), feedback alignment (FA), predictive coding (PC), and spike-timing-dependent plasticity (STDP)) in identical convolutional networks against human fMRI from THINGS-fMRI (720 stimuli, 3 subjects) using representational similarity analysis (RSA), together with an untrained random-weights baseline that isolates the contribution of architecture. Models are trained on 32 px CIFAR-10 and evaluated at 224 px; results are averaged across 5 seeds, with uncertainty from resampling stimuli. At V1, the untrained baseline exceeds BP ($\rho = 0.053$ vs. $0.022$; $\Delta\rho = +0.031$, 95% CI $[0.019, 0.043]$), and the same holds at V2. This advantage depends on the evaluation resolution: in the 5-seed model set of the resolution study (gap $+0.030$ at 224 px), it vanishes when the models are evaluated at their 32 px training resolution ($\Delta\rho = -0.001$, 95% CI $[-0.011, 0.009]$; arXiv:2608.12408). At LOC, BP, PC and STDP exceed the untrained baseline, FA is borderline, and the trained rules do not differ from each other. At IT, no two conditions can be distinguished at the low reliability of the data (leave-one-subject-out lower bound 0.021). FA has the lowest alignment at V1. At V1, an untrained network thus matches BP at the training resolution and exceeds it at 224 px; learning-rule differences are resolvable only at an intermediate area.
|
| 1399 |
Same Methods, Different Rankings: Trainable Depth as an Evaluation Variable in Continual Learning
2604.21927
|
cs.LG
|
Paul-Tiberiu Iordache, Elena Burceanu, Mihai Dascalu |
Continual learning (CL) examines how models learn a sequence of tasks while retaining previously learned knowledge. Despite substantial progress in benchmarking CL methods, comparative evaluations typically keep the fine-tuning regime fixed. In this paper, we ...Continual learning (CL) examines how models learn a sequence of tasks while retaining previously learned knowledge. Despite substantial progress in benchmarking CL methods, comparative evaluations typically keep the fine-tuning regime fixed. In this paper, we argue that the fine-tuning regime, defined by the trainable parameter subspace, is itself a key evaluation variable. We formalize adaptation regimes as projected optimization over fixed trainable subspaces, showing that changing the trainable depth alters the effective update signal through which both current task fitting and knowledge preservation operate. This analysis motivates the hypothesis that method comparisons need not be invariant across regimes. We test this hypothesis in task incremental CL while considering 5 trainable depth regimes and 5 standard methods: online EWC, LwF, SI, GEM, and DER. We find that the relative ranking of methods is not consistently preserved across regimes when evaluating across 5 benchmark datasets, namely MNIST, Fashion MNIST, KMNIST, QMNIST, and CIFAR-100, and across 11 task orders per dataset. We further show that deeper adaptation regimes are associated with larger update magnitudes, higher forgetting, and a stronger relationship between the two. These results show that comparative conclusions in CL can depend strongly on the chosen fine-tuning regime, motivating regime-aware evaluation protocols that treat trainable depth as an explicit experimental factor.
|
| 1400 |
SPLICE: Latent Diffusion over JEPA Embeddings for Conformal Time-Series Inpainting
2605.00126
|
cs.LG
|
Arnaud Zinflou |
Long gaps in time series are difficult to reconstruct reliably, and in power systems imputed load values inform dispatch and planning. Existing generative imputers are accurate but provide no finite-sample reliability guarantees under changing conditions. We i...Long gaps in time series are difficult to reconstruct reliably, and in power systems imputed load values inform dispatch and planning. Existing generative imputers are accurate but provide no finite-sample reliability guarantees under changing conditions. We introduce SPLICE (Self-supervised Predictive Latent Inpainting with Conformal Envelopes), which couples a JEPA encoder, a conditional latent bridge, an hourly-conditioned decoder, and Adaptive Conformal Inference (ACI). Controlled ablations indicate that long-horizon fidelity is driven more by the predictive representation than by generative complexity: under a matched bridge and decoder, JEPA outperforms reconstructive VAE and contrastive TS2Vec encoders, and once the representation is fixed a 5-step flow sampler matches or improves on a 50-step DDIM sampler at a 5-10x speedup. The same ordering is recovered on twelve hydrological catchments. Across thirteen load datasets at 91-day gaps, SPLICE attains the lowest mean Load-only MSE (0.056), winning 9 of 12 non-degenerate datasets against seven established baselines, and the best mean CRPS (0.182, 7.8% below the strongest competitor), while ACI holds 93-95% empirical coverage where static calibration under-covers by up to 7.5 percentage points. A pooled JEPA encoder transfers to four held-out datasets with only bridge and decoder adaptation.
|
| 1401 |
Perturb and Correct: Post-Hoc Ensembles using Affine Redundancy
2605.01632
|
cs.LG
|
Eleanor Quint |
Models that are nearly indistinguishable on in-distribution data can behave very differently under distribution shift. We introduce Perturb-and-Correct (P&C), a post-hoc method for constructing epistemically diverse predictors from a single pretrained netw...Models that are nearly indistinguishable on in-distribution data can behave very differently under distribution shift. We introduce Perturb-and-Correct (P&C), a post-hoc method for constructing epistemically diverse predictors from a single pretrained network. P&C applies random hidden-layer perturbations with a least-squares correction in the following affine layer, producing predictors that agree on calibration data while remaining free to disagree away from it. We analyze this mechanism through the post-correction residual that remains in the activation of the layer corrected using calibration data, yielding a leverage-based preservation bound and a covariance interpretation of ensemble disagreement. Empirically, P&C achieves a strong ID/OOD tradeoff across MuJoCo dynamics prediction and CIFAR-10 OOD detection, and scales to a pretrained ViT-B/16 on ImageNet-1K. In matched comparisons with Deep Ensembles, P&C substantially reduces construction and storage cost. Our findings highlight the potential in further exploiting overparameterization as a strength of deep learning models.
|
| 1402 |
Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum
2605.02043
|
cs.LG
|
Tehila Dahan, Roie Reshef, Sharon Goldstein, Kfir Y. Levy |
Asynchronous SGD enables scalable distributed training but suffers from stale gradients. When delays depend on the data, slower samples may be underrepresented, biasing training toward faster samples. We introduce ordered momentum, a unified framework that att...Asynchronous SGD enables scalable distributed training but suffers from stale gradients. When delays depend on the data, slower samples may be underrepresented, biasing training toward faster samples. We introduce ordered momentum, a unified framework that attains the best-known rates for smooth convex and non-convex objectives under both data-independent and data-dependent delays. Notably, we establish (i) the first convergence guarantee for smooth convex objectives with data-dependent delays and, among analyses of data-dependent delays, the first to (ii) benefit from parallelization and (iii) match the tight data-independent rate, with a leading stochastic term independent of the number of workers. Finally, we derive robust learning rates that simplify hyperparameter tuning across convex and non-convex settings.
|
| 1403 |
Learning to Theorize the World from Observation
2605.03413
|
cs.LGcs.AI
|
Doojin Baek, Gyubin Lee, Junyeob Baek, Hosung Lee, Sungjin Ahn |
What does it mean to understand the world? Contemporary world models often operationalize understanding as accurate future prediction in latent or observation space. Developmental cognitive science, however, suggests a different view: human understanding emerg...What does it mean to understand the world? Contemporary world models often operationalize understanding as accurate future prediction in latent or observation space. Developmental cognitive science, however, suggests a different view: human understanding emerges through the construction of internal theories of how the world works, even before mature language is acquired. Inspired by this theory-building view of cognition, we introduce Learning-to-Theorize, a learning paradigm for inferring explicit explanatory theories of the world from raw, non-textual observations. We instantiate this paradigm with the Neural Theorizer (NEO), a World Theory Model, that induces latent programs as a learned Language of Thought and executes them through a shared transition model. In NEO, a theory is represented as an executable, compositional program whose learned primitives can be systematically recombined to explain novel phenomena. Experiments show that this formulation enables explanation-driven generalization, allowing observations to be understood in terms of the programs that generate them.
|
| 1404 |
Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts
2605.04952
|
cs.LG
|
Klaus-Rudolf Kladny, Maximilian Mordig, Bernhard Sch\"olkopf, Michael Muehlebach |
Sparse Mixture-of-experts (MoE) models enable scalable transformer architectures by activating only a subset of experts per token. Recent evidence suggests that performance improves with increasingly granular experts, i.e., many small experts instead of a few ...Sparse Mixture-of-experts (MoE) models enable scalable transformer architectures by activating only a subset of experts per token. Recent evidence suggests that performance improves with increasingly granular experts, i.e., many small experts instead of a few large ones. However, this regime substantially increases routing cost, which can dominate computation and memory. We introduce adaptive inverted-index routing for MoE (AIR-MoE), an inverted-index-inspired routing architecture based on vector quantization (VQ). In a first stage, AIR-MoE performs coarse shortlisting by assigning tokens to VQ codewords to construct a candidate set of experts. In a second stage, fine scoring computes exact routing scores restricted to this shortlist. This two-stage procedure approximates true top-K routing while avoiding full expert scoring and, in contrast to prior work, imposing no structural constraints on expert parameters. AIR-MoE serves as a drop-in replacement for standard routers and requires no modifications to the model architecture or loss function. We further provide a lower bound on the mass recall achieved by AIR-MoE that yields insights into the inner workings. Empirically, AIR-MoE improves the perplexity-FLOPs trade-off over existing MoE routers in regimes with up to more than a million tiny experts, and achieves consistent perplexity improvements over the best existing baseline.
|
| 1405 |
When Can Voting Help, Hurt, or Change Course? Exact Structure of Binary Test-Time Aggregation
2605.05592
|
cs.LG
|
Yi Liu |
Increasing a voting budget can improve or reduce population accuracy, and an initial gain or loss need not persist. We give a complete structural classification of these curves for binary unweighted majority under infinitely exchangeable correctness indicators...Increasing a voting budget can improve or reduce population accuracy, and an initial gain or loss need not persist. We give a complete structural classification of these curves for binary unweighted majority under infinitely exchangeable correctness indicators, with an arbitrary latent success-probability law $\Pi$. The classification is encoded by the signed voting signature $ \omega_\Pi(B)=\int \mathbf{1}\{q(1-q)\in B\}(2q-1)\,\Pi(\,\mathrm{d} q). $ The complete odd-budget curve and this signature determine each other: the initial accuracy gives its total signed mass, and successive budget increments give its higher moments. Reflection-symmetric differences between latent laws are exactly the changes invisible to voting. A weighted total-variation constraint characterizes every attainable signature, and a minimum-mass allocation to the two reflected branches, completed by symmetric probability mass, gives every population producing its curve. Consequently, the full curve identifies the latent population precisely when no symmetric mass remains to redistribute.
|
| 1406 |
ProtoSSL: Self-Supervised Pretraining and Downstream Transfer for Projection-Based Prototype Models
2605.06943
|
cs.LG
|
Steven Song, Sahil Sethi, Brett Beaulieu-Jones, Robert L. Grossman |
In domains where both predictive performance and interpretability are essential, deep neural networks achieve strong results but provide limited insight into how their predictions are made. Projection-based prototype networks address this limitation by groundi...In domains where both predictive performance and interpretability are essential, deep neural networks achieve strong results but provide limited insight into how their predictions are made. Projection-based prototype networks address this limitation by grounding predictions in similarity to representative training examples, enabling case-based explanations and global prototype inspection. However, existing approaches rely on label supervision, tying prototypes to a specific task and requiring large labeled datasets. We introduce ProtoSSL, a framework for pretraining a foundational latent prototype bank on unlabeled data and transferring it to downstream tasks to create interpretable, projection-based prototype models. Our key idea is to separate motif discovery from label alignment. ProtoSSL first learns a transferrable prototype bank using a self-supervised objective applied directly to prototype activations, and then aligns these prototypes to downstream tasks through a novel efficient assignment procedure. Across six electrocardiography (ECG) datasets, ProtoSSL improves label efficiency, outperforming supervised prototype baselines in low-data regimes with as few as 256 labeled examples; with fine-tuning, ProtoSSL outperforms supervised prototype baselines at full dataset scale. In a human evaluation study, ProtoSSL produces prototypes and prototype-based explanations that are judged more favorably than those learned with direct label supervision. We further show that the framework extends to audio classification. Thus, ProtoSSL enables both learning foundational prototypes from unlabeled data before the downstream label space is known, and subsequent assignment to new tasks to create interpretable, projection-grounded prototypes.
|
| 1407 |
PLOT: Progressive Localization via Optimal Transport in Neural Causal Abstraction
2605.06979
|
cs.LGcs.AI
|
Jonathn Chang, Arya Datla, Ziv Goldfeld |
Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with low-level neural computation through interchange intervention analysis. Finding such an alignment, however, often requires fitting and ev...Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with low-level neural computation through interchange intervention analysis. Finding such an alignment, however, often requires fitting and evaluating separate learned mappings across many candidate neural locations. We introduce PLOT, a gradient-free approach for joint correspondence discovery via a global matching of intervention effects. PLOT represents abstract and neural interventions by geometric signatures of their effects on the shared task output and fits an optimal transport coupling between the two collections. The coupling can be calibrated directly into an executable intervention handle or used progressively to select coarse parent sites for finer localization, reducing the cost of searching broad collections of neural sites. Experiments on hierarchical equality, binary addition, and multiple-choice question answering (MCQA) demonstrate fast and accurate direct handles without learning intervention rotations and show that progressive localization can improve accuracy under intervention-size constraints while reducing computational cost. PLOT can also guide gradient-based subspace learning, with PLOT-guided distributed alignment search (DAS) attaining accuracy comparable to full DAS at approximately $3.6\times$ and $20\times$ lower serial runtime on binary addition and MCQA, respectively. PLOT thus separates the search for candidate correspondences from optional subspace learning and provides a common framework for constructing and testing neural intervention handles for a given causal model.
|
| 1408 |
Mask2Cause: Temporal Causal Discovery Beyond Causality in Mean
2605.07280
|
cs.LGcs.AI
|
Omar Muhammad, Pasupuleti Dhruv Shivkant, Deepak N. Subramani |
One approach to discovering causal relationships in multivariate time series is to ask whether a variable's history improves prediction of another variable. When this improvement is measured solely by reductions in optimal squared error, the criterion misses r...One approach to discovering causal relationships in multivariate time series is to ask whether a variable's history improves prediction of another variable. When this improvement is measured solely by reductions in optimal squared error, the criterion misses relationships that affect the target's conditional variance without changing its conditional mean. We propose Mask2Cause, an end-to-end Transformer framework for causal discovery through conditional mean and variance prediction. Each variable's history is embedded as a separate token, and a single learnable adjacency matrix, shared across layers and attention heads, controls information exchange between tokens. Gaussian negative log-likelihood and a sparsity penalty on this matrix jointly train the forecasting model and graph. Our population analysis connects the optimal log-loss reduction to causally conditioned directed information, providing an information-theoretic motivation for the training objective. We introduce Mixed Physics, a benchmark with separately controlled mean and variance dependencies, to evaluate recovery of mean-only and variance-only edges. In its pure-variance setting, Mask2Cause achieves 1.000 off-diagonal AUROC, compared with 0.521 for a corresponding squared-error model. Experiments on linear, chaotic, biological, and real-data-derived benchmarks demonstrate competitive graph recovery, while the inferred graphs support downstream tasks such as causal pruning for efficient forecasting and root-cause localization.
|
| 1409 |
TIDES: Implicit Time-Awareness in Selective State Space Models
2605.09742
|
cs.LGcs.AI
|
Taylan Soydan, Miguel A. Bessa, Dirk Mohr, Rui Barreira |
Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step $\Tilde{\Delta}$ a learned function of the input. However, in doing so, $\Tilde{\Delta}$ no longer equals the physical time gap $\D...Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step $\Tilde{\Delta}$ a learned function of the input. However, in doing so, $\Tilde{\Delta}$ no longer equals the physical time gap $\Delta$ between consecutive observations, limiting the ability of these models to handle irregular time series. Continuous time SSMs, such as S5, keep $\Tilde{\Delta}\equiv\Delta$ and therefore handle irregular timestamps natively, but their dynamics remain linear time invariant (LTI), limiting per token expressivity. We propose \textbf{TIDES}, a selective SSM variant that reconciles selective and continuous architectures by moving input dependence off the step size and onto the diagonal state matrix. As a result, $\Tilde{\Delta}\equiv\Delta$ as in S5, allowing the model to handle irregular timestamps natively without sacrificing the per-token expressivity that makes selective SSMs effective. We show this on a novel \emph{Fading Flash} experimental benchmark, a compact controlled diagnostic for sequence models that jointly tests input dependence and extrapolation to out-of-distribution $\Delta$ values, and isolates the distinct failure modes of current state-of-the-art architectures that TIDES avoids by construction. On large-scale benchmarks, TIDES sets the new best average rank on UEA time series classification and the Physiome ODE regression benchmark, and matches or exceeds the reference baseline model on 6 of 8 natively irregular datasets from astronomy, agriculture, neuromorphic sensing, and climate events. Code available at: \url{https://github.com/TaylanSoydan/TIDES}.
|
| 1410 |
Twincher: Bijective Representation Learning for Robust Inversion of Continuous Systems
2605.13470
|
cs.LG
|
Arkady Gonoskov |
Perception, state estimation, and action planning in physical systems often reduce to inverting a continuous forward map $p \mapsto y = f(p)$ (a renderer, a physics simulator, or a learned surrogate) that can be queried but not inverted in closed form. Learned...Perception, state estimation, and action planning in physical systems often reduce to inverting a continuous forward map $p \mapsto y = f(p)$ (a renderer, a physics simulator, or a learned surrogate) that can be queried but not inverted in closed form. Learned direct inverses are limited by an approximation error that decreases only gradually with the amount of data, whereas iterative solvers that minimize the misfit in observation space can stall at spurious stationary points. We study an intermediate route: learning a representation $r(y)$ in which the composite map $p \mapsto r(f(p))$ is a well-conditioned bijection and which is insensitive to perturbations of $y$ normal to the data manifold. In such coordinates, the inverse problem becomes a square nonlinear system whose least-squares formulation has no spurious interior stationary points, so that Newton-type iterations using the exact forward model can refine the solution to the accuracy permitted by the observations; the learned model must be qualitatively correct rather than accurate. We introduce Twincher, an architecture composed of localized, area-preserving pairwise rotations ("twinches"), and explore training possibilities with Jacobian-level objectives, a domain-growing curriculum, and memory-efficient routines for propagating derivative tensors. On a family of synthetic forward maps with controllable nonconvexity, once the query budget suffices to form a bijective representation, the worst-case residual over $10^3$ test targets reaches machine precision within five refinement steps, whereas a parameter-matched baseline that combines an MLP inverse with Gauss-Newton refinement improves only gradually with the budget. A pose-estimation example from noisy depth maps illustrates inference errors that scale linearly with, and vanish together with, the observation noise. We release an open-source CPU/GPU Python implementation.
|
| 1411 |
A Novel Schur-Decomposition-Based Weight Projection Method for Stable State-Space Neural-Network Architectures
2605.14489
|
cs.LG
|
Sergio Vanegas, Lasse Lensu, Fredy Ruiz |
Building black-box models for dynamical systems from data is a challenging problem in machine learning, especially when asymptotic stability guarantees are required. In this paper, we introduce a novel stability-ensuring and backpropagation-compatible projecti...Building black-box models for dynamical systems from data is a challenging problem in machine learning, especially when asymptotic stability guarantees are required. In this paper, we introduce a novel stability-ensuring and backpropagation-compatible projection scheme based on the Schur decomposition for the state matrix of linear discrete-time state-space layers, as well as an alternative pre-factorized formulation of the methodology. The proposed methods dynamically project the quasi-triangular factor of the state matrix's real Schur decomposition onto its nearest stable peer, ensuring stable dynamics with minimal overparameterization. Experiments on synthetic linear systems demonstrate that, despite a marginal increase in computational complexity, the method achieves accuracy and convergence rates comparable to those of state-of-the-art stable-system identification techniques. Furthermore, the lower weight count facilitates convergence during training without sacrificing accuracy in stacked neural-network architectures with static nonlinearities targeting real-world datasets. These results suggest that the Schur-based projection provides a numerically robust framework for identifying complex dynamics on par with the state of the art while satisfying strict asymptotic-stability requirements.
|
| 1412 |
On the Stability of Growth in Structural Plasticity: Forward-Active yet Backward-Starved
2605.15435
|
cs.LG
|
Lute Lillo, Nick Cheney |
Standard deep-learning pipelines usually choose the network architecture before training and keep it fixed throughout optimization. In contrast, a model can also be adapted by editing its structure during training, for example by pruning existing hidden units ...Standard deep-learning pipelines usually choose the network architecture before training and keep it fixed throughout optimization. In contrast, a model can also be adapted by editing its structure during training, for example by pruning existing hidden units or growing new ones; however, growth is not simply the inverse of pruning. Pruning selects among units trained from initialization, whereas growth inserts new capacity into an already specialized optimization trajectory. We isolate this insertion problem and show that newborn units can be forward-active yet backward-starved, receiving substantially weaker gradient signal than incumbent units. This asymmetry is weak in small MLPs but emerges in more challenging convolutional settings, where \textsc{Grow} reaches competitive final architectures despite weaker trajectory-level performance. Extending structural edits to residual networks reveals a strong dependence on edit location: head-localized growth produces concentrated but task-aligned representations and faster bounded label remapping, whereas growth throughout the residual hierarchy yields lower effective feature support, and a lower bounded adaptation endpoint. Interventions targeting optimizer state, insertion, selection, and trainability improve newborn integration, but not necessarily final subnetwork quality. Across continual-learning benchmarks, growth is most competitive when newborn units have sufficient time and trainability to integrate. Therefore, \textsc{Grow} should be evaluated not only by the architecture it ultimately discovers, but by whether late-arriving capacity becomes usable before a new structural change occurs during training.
|
| 1413 |
GraViti: Graph-Level Variational Autoencoders with Relaxed Permutation Invariance
2605.16668
|
cs.LGcs.AI
|
Roman Bresson, Konstantinos Divriotis, Johannes F. Lutzeyer, Iakovos Evdaimon, Michalis Vazirgiannis |
We introduce GraViti, a transformer-based graph-level variational autoencoder that encodes entire graphs into single, fixed-dimensional latent vectors rather than per-node embeddings, yielding a graph-level latent space that supports smooth interpolation, prop...We introduce GraViti, a transformer-based graph-level variational autoencoder that encodes entire graphs into single, fixed-dimensional latent vectors rather than per-node embeddings, yielding a graph-level latent space that supports smooth interpolation, property-guided search, and other downstream tasks beyond the reach of node-level approaches. GraViti achieves state-of-the-art reconstruction accuracy on large molecular graph datasets while offering competitive single-step generative performance. Our central finding concerns where permutation invariance is needed in such a model, and where it is not. The encoder remains permutation-invariant throughout, ensuring the model generalizes to graphs regardless of input order. The reconstruction loss, however, does not need to be: when training data is provided in a consistent node order, comparing predictions to targets directly in that order, without a graph-matching step, is sufficient. Relaxing invariance in the loss alone lets GraViti avoid the cubic-cost matching step required by prior graph-level autoencoders, reducing complexity to quadratic while improving reconstruction fidelity. We show the resulting latent space is chemically meaningful, supporting controlled molecular editing and property optimization, recovering established chemical trends, and enabling direct regression of graph-level properties from latent embeddings.
|
| 1414 |
A PAC-Bayesian View of Generalisation for Physics-Informed Machine Learning
2605.26341
|
cs.LG
|
Thien V. Nguyen, Amaury Habrard, Benjamin Guedj |
Physics-informed machine learning (PIML) integrates mechanistic knowledge, typically through partial differential equations (PDEs), into data-driven models. Despite strong empirical performance, its statistical generalisation properties remain poorly understoo...Physics-informed machine learning (PIML) integrates mechanistic knowledge, typically through partial differential equations (PDEs), into data-driven models. Despite strong empirical performance, its statistical generalisation properties remain poorly understood, especially for regression with unbounded losses. We develop a PAC-Bayesian framework for PIML that provides high-probability generalisation guarantees under potentially unbounded losses. Exploiting the structure of physics-informed objectives, we derive component-wise bounds whose complexity scales with the input-gradient energy of each loss, establishing a direct link between physical regularity and generalisation. We further introduce a PAC-Bayesian calibration procedure that yields computable gradient-based complexities while controlling rare large-gradient events through a residual-tail correction. Adopting a multi-task view of data fidelity, PDE residuals, initial conditions, and boundary conditions, we obtain a refined certificate with a single PAC-Bayesian complexity penalty, avoiding the looseness of independently bounding each component. Building on this certificate, we propose a bound-aware learning procedure that promotes empirical accuracy, proximity to a physics-informed prior, and low input-gradient complexity. Experiments on six PDE benchmarks yield substantially tighter certificates than bounded-loss and sub-Gaussian alternatives, while ablations quantify the effects of posterior-training data, calibration data, gradient envelopes, and hypothesis localisation.
|
| 1415 |
Function-Valued Causal Influence in Nonlinear Time Series
2605.26408
|
cs.LG
|
Valentina V. Kuskova, Dmitry Zaytsev, Michael Coppedge |
Causal discovery in time series is increasingly performed using nonlinear machine-learning models, yet the resulting causal relationships are almost always summarized by scalar edge scores. We argue that this practice obscures the true object learned by nonlin...Causal discovery in time series is increasingly performed using nonlinear machine-learning models, yet the resulting causal relationships are almost always summarized by scalar edge scores. We argue that this practice obscures the true object learned by nonlinear autoregressive models: a state-dependent function whose effect varies across regimes, magnitudes, and contexts. We formalize function-valued causal influence for additive, contribution-decomposable architectures and show that scalar causal scores constitute a severe information bottleneck, conflating between-state variation with within-state residual noise. Using Neural Additive Vector Autoregression as a representative architecture, we introduce a practical framework based on Individual Conditional Expectation for estimating causal response functions directly from trained models. Through controlled synthetic experiments, we demonstrate that edges with indistinguishable scalar scores can exhibit qualitatively different functional behaviors, including monotonic, thresholded, saturating, and sign-changing effects. An applied case study on democratic development further shows that function-valued analysis reveals regime-specific and asymmetric causal structure systematically missed by score-centric approaches.
|
| 1416 |
SPHERE-JEPA: Spherical Prediction with Homogeneous Embeddings
2605.26900
|
cs.LG
|
L{\'e}o Nicollier (CB, ATT), Max Dunitz (CB, ATT), Marc Pic (ATT) |
A fundamental open question in self-supervised learning (SSL) is the explicit characterization of the optimal geometry of the learned representations. Recently, LeJEPA identified isotropic Gaussian embeddings as optimal for minimizing downstream prediction ris...A fundamental open question in self-supervised learning (SSL) is the explicit characterization of the optimal geometry of the learned representations. Recently, LeJEPA identified isotropic Gaussian embeddings as optimal for minimizing downstream prediction risk in Euclidean spaces. However, the corresponding problem for distributions supported on lower-dimensional manifolds, such as the hypersphere, remains unexplored. In this work, we demonstrate that extending this minimax analysis to smooth distributions on Riemannian manifolds fundamentally changes the optimal solution. We show that, under a worst-case formulation, both k-nearest neighbors and kernel ridge regression induce hyperspherical uniformity. More precisely, we show that uniform distributions on manifolds are optimal for k-nearest neighbors, and that the uniform distribution on the sphere is optimal for kernel ridge regression with both the exponential dot-product kernel and the linear kernel. This theoretical insight reveals a fundamental limitation of Gaussian embeddings: their non-uniform density induces anisotropic k-NN neighborhoods, severely biasing the estimator. To correct this, we introduce SPHERE-JEPA, a theoretically grounded SSL framework. We adapt LeJEPA's Cram{\'e}r-Wold projection mechanism to enforce hyperspherical uniformity rather than a Gaussian prior. Empirically, SPHERE-JEPA yields significant improvements, boosting texture retrieval mAP by over 6%, while consistently matching or outperforming LeJEPA on standard benchmarks-including a +1.8% linear probing gain on ImageNet-1K (ViT-B/14).
|
| 1417 |
The Model Shape Behind Scaling Laws: A Gradient Superposition Perspective
2605.27989
|
cs.LG
|
Wenjie Sun, Jinning Yang, Shuai Zhang, Mengnan Du |
Neural scaling laws guide resource allocation, but how should a fixed parameter budget be divided between depth and width? Motivated by inverse-width representation-error scaling under strong superposition, we study this allocation through the network's end-to...Neural scaling laws guide resource allocation, but how should a fixed parameter budget be divided between depth and width? Motivated by inverse-width representation-error scaling under strong superposition, we study this allocation through the network's end-to-end response. The Neural Feature Ansatz links weight and gradient geometry, motivating the Frobenius norm $F$ of the input average gradient outer product (AGOP) as an observable of response strength and overlap. By reproducing Anthropic's toy-model double-descent experiment, we find that $F$ and test reconstruction loss follow an approximately affine relationship across the full sample-size sweep and share an interpolation peak. Through Gaussian minimum-norm regression experiments, we further show that this correlation arises from a fitted-noise amplification term shared by prediction error and $F$. By pretraining language models spanning different shapes and parameter budgets, we find a strong positive correlation between loss and $F$ among the best shapes selected by validation loss. This supports interpreting effective parameter allocation as preserving task-relevant computation while reducing the amplification of error responses. Our results identify error-response amplification as a mechanism linking gradient superposition to generalization, providing a functional account of depth-width allocation and efficient parameter utilization. The code is availale at: $\hyperlink{https://github.com/wenjie1835/Shape_Scaling}{https://github.com/wenjie1835/Shape_Scaling}$.
|
| 1418 |
Zero Collapse: A Failure Mode of Policy Gradient Methods in Discontinuous Reward Environments
2605.30896
|
cs.LG
|
Nishant Kumar, Enrique Areyan Viqueira, Amy Greenwald |
Policy-gradient learning can deteriorate even after substantial initial improvement. We study this behavior in a repeated first-price auction setting with a sharp threshold between winning and losing bids. We identify a failure mode, zero collapse, in which a ...Policy-gradient learning can deteriorate even after substantial initial improvement. We study this behavior in a repeated first-price auction setting with a sharp threshold between winning and losing bids. We identify a failure mode, zero collapse, in which a policy shifts into persistently near-zero reward as informative outcomes become rarely sampled. Experiments with REINFORCE and a TD state-value actor-critic illustrate this failure and evaluate practical mitigations. Adaptive update-magnitude control and smooth-output parameterization support sustained REINFORCE learning, while a learned baseline substantially stabilizes the fixed-rate comparison. Some actor-critic runs still collapse under these interventions. A diagnostic near one observed collapse reveals severe continuation-value errors that reverse the relative desirability of actions across the winning threshold. These results illustrate how thresholded outcomes, policy updates, and inaccurate value estimates can contribute to zero collapse, and demonstrate practical mitigations in the studied setting.
|
| 1419 |
Trajectory-Aware Best-Arm Identification for Local Search Allocation in Bayesian Optimization and Beyond
2605.31050
|
cs.LG
|
Nobuo Namura, Sho Takemori |
Bayesian optimization (BO) is effective for expensive black-box optimization, but its performance often degrades on multimodal or high-dimensional problems. A promising remedy is to run multiple local search processes, such as trust region--based BO, and alloc...Bayesian optimization (BO) is effective for expensive black-box optimization, but its performance often degrades on multimodal or high-dimensional problems. A promising remedy is to run multiple local search processes, such as trust region--based BO, and allocate evaluations to those with high potential. However, current best values do not always reflect the final performance of each process. We propose a trajectory-aware best-arm identification (BAI) framework that uses optimization history to guide this allocation. The proposed method extrapolates improvement trajectories, estimates final performance, and progressively eliminates low-potential candidates. For trust region--based BO, we theoretically show that the proposed BAI-guided allocation accelerates convergence to the global optimum under mild assumptions. Experiments on synthetic and real-world benchmarks demonstrate that our method improves BO on multimodal problems. We also show that the same mechanism can enhance other local or population-based optimizers, suggesting its potential as a general extension strategy for multimodal black-box optimization.
|
| 1420 |
Lightweight CNN-Based Anomaly Detection for High Voltage Converter Modulators in the Spallation Neutron Source
2605.31259
|
cs.LG
|
Alberto D. Cencillo, Leonardo Concepci\'on, Juli\'an Luengo, Isaac Triguero |
Time series anomaly detection is central to monitoring industrial and scientific equipment, where anomalous sensor readings can precede failures. As such equipment usually carries several sensors, a detector must also decide how to model the dependencies betwe...Time series anomaly detection is central to monitoring industrial and scientific equipment, where anomalous sensor readings can precede failures. As such equipment usually carries several sensors, a detector must also decide how to model the dependencies between channels. In many physical systems, normal behaviour is defined by how groups of sensors evolve together, and an anomaly can appear as the coupling or decoupling of sensors while each signal remains normal. Anomalies in popular public benchmarks, however, are mostly visible in single channels, so these benchmarks cannot show whether modelling inter-channel dependencies improves detection. We study this on the High Voltage Converter Modulators (HVCMs) of the Spallation Neutron Source (SNS), whose fourteen sensors are physically coupled and whose faults are labelled by type. Keeping a convolutional neural network (CNN) fixed, we vary only whether its blocks combine channels before or after temporal filtering, and measure detection for each SNS subsystem and fault family. The effect of this choice depends on how visible each fault family is on a single channel. Families clearly visible on one channel are detected equally well by all variants, whereas weakly separable families, such as flux faults, gain consistently when channels are combined first. Modelling the dependencies between channels therefore helps only where individual channels do not already reveal the fault. The best detector reaches an average area under the precision--recall curve of 0.816, competitive with published HVCM detectors at roughly half the parameters of a standard CNN.
|
| 1421 |
GEM: Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
2606.00140
|
cs.LGcs.AI
|
Jonas Henry Grebe, Tobias Braun, Anna Rohrbach, Marcus Rohrbach |
While the rapid adoption of multimodal generative models offers immense potential, it has also increased the risks of harmful content synthesis, deepfakes, and copyright infringements. To address these challenges, concept erasure has emerged as a prospective s...While the rapid adoption of multimodal generative models offers immense potential, it has also increased the risks of harmful content synthesis, deepfakes, and copyright infringements. To address these challenges, concept erasure has emerged as a prospective safeguard. However, as the field gradually transitions from U-Net-based diffusion models to Rectified Flow Transformers, erasure research has struggled to keep pace. In this work, we introduce GEM, a simple but highly effective erasure framework for Rectified Flow models. As part of our contribution, we establish a principled bridge between trajectory-based unlearning grounded in Generative Flow Networks and classic teacher-guided erasure: we translate trajectory-based signals into a teacher-guided flow-matching setup that unifies the strengths of both paradigms. Concretely, a teacher provides complementary attraction and repulsion signals that we combine into a single geometric guidance objective, yielding targeted suppression of unwanted concepts while preserving benign generation. Our code is available at https://github.com/multimodal-ai-lab/GEM.
|
| 1422 |
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
2606.00686
|
cs.LG
|
Maryam Hashemzadeh, Jerry Huang, Minseon Kim, Marc-Alexandre C\^ot\'e, Sarath Chandar |
The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmful prompts. While effective at reducing immediate toxicity, this approach fundamentally constricts the model'...The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmful prompts. While effective at reducing immediate toxicity, this approach fundamentally constricts the model's epistemological scope, resulting in over-cautious systems that output uninformative blanket refusals to sensitive yet benign queries. In this work, we challenge the orthodoxy that unsafe data must be discarded. We propose a dialectical approach to alignment, positing that unsafe data encodes rich, domain specific knowledge critical for nuanced, safe, and informative generation. To operationalize this, we introduce SafeMoE, a Mixture-of-Experts (MoE) framework that isolates unsafe knowledge into domain-specific Low-Rank Adapters (LoRA experts) trained exclusively on harmful corpora. To synthesize safety from these unsafe primitives, we train a lightweight gating network using a minimal, highly curated set of safe-informative responses. During inference, this router dynamically orchestrates the unsafe experts, effectively steering the generation trajectory to harness their deep domain knowledge while strictly enforcing safety constraints. Extensive empirical evaluations across stringent safety benchmarks demonstrate that SafeMoE is not only safer, achieving over a 20% relative improvement in safe response rate (more than a 15% absolute gain), but also produces more informative responses when safety and harmfulness are of paramount concern. Furthermore, the routing mechanism exhibits strong zero-shot generalization to unseen domains and broader safety tasks without domain-specific supervision. Our findings suggest a paradigm shift in alignment: true safety requires not the masking of unsafe knowledge, but its controlled integration.
|
| 1423 |
A Goal-Set Characterization of Task Composition in the Boolean Task Algebra
2606.04053
|
cs.LGcs.AI
|
Eduardo Terr\'es-Caballero, Herke van Hoof |
The Boolean Task Algebra (BTA) provides a principled framework for zero-shot task composition in reinforcement learning by equipping goal-reaching tasks with Boolean operations. We revisit its structural assumptions and formalize a collapse in the space of opt...The Boolean Task Algebra (BTA) provides a principled framework for zero-shot task composition in reinforcement learning by equipping goal-reaching tasks with Boolean operations. We revisit its structural assumptions and formalize a collapse in the space of optimal extended Q-value functions: in deterministic MDPs, every such function is fully determined by the universal and empty tasks. This makes the logarithmic set of base tasks proposed in the original BTA formulation redundant. Building on this observation, we introduce a goal-set-based composition method that performs logical operations on goal sets and reconstructs composed value functions by selecting slices from the universal and empty value functions. This reduces learning costs for standard BTA and reduces composition time for both BTA and Skill Machines, while preserving policy performance. Experiments across tabular, visual, function-approximation, and continuous-control domains show that learning additional base tasks does not yield better performance. Finally, we study the stochastic setting and provide a counterexample showing that this collapse need not hold, that is, optimal composition may require accounting for exponentially many policies in the number of goals. Code is available at \url{https://github.com/EduardoTerres/GoalSetBTA}.
|
| 1424 |
A Unifying View of Attention Sinks: From Mechanisms to Architectural Interventions
2606.08105
|
cs.LG
|
Lukas Fesser*, Mozes Jacobs*, Thomas Fel*, Andy Keller, Sham Kakade |
When attention concentrates on a single token, a sink, what is the model actually computing? Attention sinks are ubiquitous in softmax transformers, yet this shared visual signature can hide fundamentally different algorithms. We show that visually similar sin...When attention concentrates on a single token, a sink, what is the model actually computing? Attention sinks are ubiquitous in softmax transformers, yet this shared visual signature can hide fundamentally different algorithms. We show that visually similar sink patterns can reflect two distinct mechanisms: (i) adaptive nop, where a head suppresses its update by routing to a null token, and (ii) broadcast, where a sink aggregates and redistributes global information. Each mechanism leaves distinct traces (nop-sinks exhibit negligible value norms; broadcast sinks induce low-rank outputs), which we formalize on synthetic tasks and use to derive practical diagnostics. Applied to pretrained vision transformers, these diagnostics reveal that both mechanisms exist at scale: sinks transition from CLS in early layers to patches in deeper layers and concentrate in specialized heads. Causal interventions further connect these signatures to near-null suppression and shared residual contributions. We then use architectural interventions to show how these computations can be reorganized: gating eliminates detected nop-like sinks but increases broadcast-like sinks, registers relocate rather than remove sink computation, and our position-free global pathway provides an explicit route for shared communication that reduces the broadcast-like sinks induced by gating. On dense probes, combining gating with the global pathway gives the strongest results among the tested variants, despite retaining some broadcast-like sinks. Overall, we find that the same attention pattern can reflect two very different computations, and that effective intervention depends not only on identifying the computation, but also on providing architectural alternatives through which the model can reorganize it.
|
| 1425 |
Ambiguous Strategic Classification
2606.10137
|
cs.LG
|
Ivri Hikri, Nir Rosenfeld |
A common assumption in strategic classification is that the classifier is public knowledge. However, it remains unclear whether, and why, a system would choose to commit to full disclosure. We study a setting in which regulation requires the system to disclose...A common assumption in strategic classification is that the classifier is public knowledge. However, it remains unclear whether, and why, a system would choose to commit to full disclosure. We study a setting in which regulation requires the system to disclose some, but not all, of the information. This induces a learning task in which the learner must jointly optimize the classifier and the uncertainty surrounding it. To this end, we adopt from robust mechanism design the notion of ambiguity, which in our setting allows the learner to reveal a set or range of possible classifiers, while privately choosing which of them to ultimately realize. We investigate how ambiguity affects the learning task, develop efficient algorithms for computing best-responses and training, and empirically explore strategic learning and its outcomes in this novel setting and using our approach.
|
| 1426 |
Learning the Context of Errors: Black-Box Online Adaptation of Time Series Foundation Models
2606.14222
|
cs.LG
|
Xilin Dai, Yiding Liu, Hongjie Xia, Yifan Hu, Zewei Dong |
The rapid evolution of Time Series Foundation Models (TSFMs) has advanced zero-shot forecasting across diverse domains. Inspired by the current form of Large Language Models, future TSFMs may be offered as commercialized, closed-source API services. However, m...The rapid evolution of Time Series Foundation Models (TSFMs) has advanced zero-shot forecasting across diverse domains. Inspired by the current form of Large Language Models, future TSFMs may be offered as commercialized, closed-source API services. However, many existing online adaptation methods still rely on white-box access for parameter fine-tuning or gradient backpropagation. This paradigm mismatch raises a question: In black-box online adaptation for TSFMs, what should we learn? We answer this with an insight: the predictive errors of the base model are conditioned on both the input and output of the base model (i.e., the context of errors). To validate this insight, we propose ORCA (Online Residual Contextual Adaptation). We conduct extensive experiments across 5 state-of-the-art TSFMs and 8 datasets to demonstrate the effectiveness of our approach. Furthermore, through ablation studies, we quantitatively analyze the impact of different adapter learning hypotheses on the final adaptation performance in black-box online adaptation. Code available at https://github.com/Fifthky/ORCA.
|
| 1427 |
Your Privacy My Cloak: Backdoor Attacks on Differentially Private Federated Learning
2606.17035
|
cs.LG
|
Xiaolin Li, Ning Wang, Ninghui Li, Wenhai Sun |
Prior research suggests that differential privacy (DP) can enhance the robustness of federated learning (FL) against backdoor attacks. In this paper, we challenge this assumption. Through an empirical analysis of two baseline attack strategies, we uncover a fu...Prior research suggests that differential privacy (DP) can enhance the robustness of federated learning (FL) against backdoor attacks. In this paper, we challenge this assumption. Through an empirical analysis of two baseline attack strategies, we uncover a fundamental tension in DP-FL: while bypassing DP allows state-of-the-art defenses to detect and filter malicious updates, complying with DP inadvertently masks their distinguishing statistical characteristics. Consequently, existing defenses become ineffective as DP reduces the raw backdoor signal. Building on this masking effect, we propose RING, a novel attack that explicitly exploits DP to conceal malicious contributions while maximizing attack impact. By collaboratively crafting adversarial perturbations, compromised clients reconstruct a strong backdoor signal during aggregation without triggering anomaly detection. RING operates as a perturbation layer that is agnostic to the underlying backdoor technique, making it broadly applicable and composable with existing attacks -- a property that significantly amplifies the threat it poses to DP-FL. Extensive evaluations across four image and text datasets under non-iid distributions show that RING achieves an average attack success rate of 90.3% against six state-of-the-art defenses under a moderate privacy budget, an improvement of up to 38.4% over baseline strategies. Finally, we evaluate potential countermeasures and find that mitigating this threat incurs significant utility trade-offs, exposing a fundamental security gap in the deployment of differentially private FL.
|
| 1428 |
Machine learning modeling of hit-into-play probability and quantification of pitch sequence effects
2606.17345
|
cs.LGcs.AI
|
Ryota Takamido, Hiroki Nakamoto |
Although pitch sequencing is a central topic in baseball analytics, previous studies have primarily focused on improving predictive performance for the final pitch or its outcome, leaving the role of preceding pitches insufficiently examined. To address this i...Although pitch sequencing is a central topic in baseball analytics, previous studies have primarily focused on improving predictive performance for the final pitch or its outcome, leaving the role of preceding pitches insufficiently examined. To address this issue, this study conducted a machine learning-based analysis to quantify the importance of preceding pitches and deepen the tactical understanding of baseball. A Transformer-based model was designed and trained to predict whether a target pitch would result in an in-play outcome or a swinging strikeout. Based on the model, two analyses were conducted, including an ablation study to evaluate the contribution of preceding pitches to predictive performance and a replacement analysis to quantify changes in predicted in-play probability when individual pitches were replaced by alternatives. The results suggest that hitting outcome predictions mostly depended on the attributes of the final pitch, and including preceding pitches had little effect on the predictive performance of the model. Moreover, the effect of pitch replacement on the probability output of the model for preceding pitches ranged from approximately 0.11 to 0.14, compared with approximately 0.44 to 0.45 for the final pitch. Therefore, preceding pitches played only a supplementary role in slightly modifying the predicted probability of a ball being put into play, although even such small effects may be practically meaningful.
|
| 1429 |
Score Approximation for Diffusion Models on Arbitrary Low-Dimensional Structures
2606.19894
|
cs.LG
|
Xinhe Mu, Zaijiu Shang, Zhaoqi Zhou, Chuan Zhou, Qi Meng |
Score-based diffusion models have achieved remarkable empirical success, motivating extensive theoretical work to establish their foundations. However, existing complexity bounds for score approximation, a vital step in diffusion modeling, rely on rigid constr...Score-based diffusion models have achieved remarkable empirical success, motivating extensive theoretical work to establish their foundations. However, existing complexity bounds for score approximation, a vital step in diffusion modeling, rely on rigid constraints such as Lipschitz continuous scores or lower bounded densities. This severely limits their applicability to real-world perceptual data, where singularities, sharp boundaries, and disjoint clusters routinely violate such restrictive assumptions. We bridge this gap between theory and practice, presenting the first universal score approximation theorem applicable to any compactly supported distribution in $\mathbb{R}^n$. Using a novel discretization technique that directly models the underlying distribution, we prove that the neural network complexity is governed by the support's upper Minkowski dimension $d$ rather than the ambient dimension $n$. Furthermore, by leveraging the inherent smoothing of Gaussian kernels, we show that even for irregular, fractal distributions, an $\epsilon$ approximation error can be achieved with $\mathcal{O}(\epsilon^{-{d/(M+1)}})$ local Taylor modules each at size of $\mathcal{O}(n^M)$, an $\epsilon$-scaling rate previously achieved only for distributions with $(M+1)$-H\"older smoothness. Thus, we reveal a possible mechanism by which score-based diffusion models represent non-smooth data distributions.
|
| 1430 |
GLACIER: Rethinking Mass Spectrum Prediction as an Object Detection Problem
2606.29161
|
cs.LG
|
Rui-Xi Wang, Runzhong Wang, Connor W. Coley |
Predicting tandem mass spectra (MS/MS) from molecular structures represents a central task in analytical chemistry with direct relevance to clinical metabolomics, systems biology, and adjacent disciplines. In this work, we revisit the problem through the lens ...Predicting tandem mass spectra (MS/MS) from molecular structures represents a central task in analytical chemistry with direct relevance to clinical metabolomics, systems biology, and adjacent disciplines. In this work, we revisit the problem through the lens of object detection on molecular graphs. Molecular fragmentation, a central step in MS/MS prediction, can be approximated as detecting a set of subgraphs (i.e., fragments) and their associated spectral contributions. Existing fragment-based models follow a two-stage paradigm -- first generating candidate fragments and then scoring them -- analogous to two-stage R-CNNs in computer vision. Towards higher accuracy and faster inference, we introduce GLACIER, a single-stage transformer-based fragment detection neural network for molecular graphs. This unified formulation eliminates the need for candidate enumeration, enabling scalable and globally consistent modeling of molecular fragmentation. GLACIER is faster and more accurate than existing state-of-the-art by a significant margin, achieving 70.0% and 69.7% Top-1 retrieval accuracy with and without contrastive finetuning on the MassSpecGym dataset (from the previous SOTA of 64.0%) and 52.5% and 38.5% respectively on the NIST'20 dataset (from 33.2%). Furthermore, GLACIER provides nearly 8-fold inference speedup over our prior two-stage model. Code is available at https://github.com/coleygroup/ms-pred
|
| 1431 |
Anchored or Drifting: What Recursive Self-Generation Reveals About Training Data
2606.31991
|
cs.LGcs.AI
|
Wojciech {\L}apacz, Stanis{\l}aw Pawlak |
Large generative models are known to memorize their training data, posing severe privacy risks. Yet, current methods to detect training membership typically rely on the weak signals of a single forward pass. In this work, we find that training samples and unse...Large generative models are known to memorize their training data, posing severe privacy risks. Yet, current methods to detect training membership typically rely on the weak signals of a single forward pass. In this work, we find that training samples and unseen (held-out) data follow visibly different trajectories under recursive self-generation -- repeatedly feeding a model's output back as its next input. Held-out samples \emph{drift}: they lose the specifics of the original within a few steps. Training samples stay \emph{anchored}, degrading far more slowly. We show that the membership signal this produces holds across model modalities, architectures and scales, spanning language, diffusion, and autoregressive vision models. Furthermore, these recursive trajectories provide a signal that raises membership inference TPR at $1\%$ FPR for nearly every attack we evaluate, roughly doubling it on the weakest baselines and still improving the strongest, which shows that a model's behavior under recursion carries membership evidence that a single query does not.
|
| 1432 |
NFTR: From Provable Mode-Averaging to Geodesic Subgoal Selection in Offline Goal-Conditioned RL
2607.07855
|
cs.LG
|
Erdemt Bao, Xing Lei, Jun Chen, Xuetao Zhang, Donglin Wang |
Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone. This rule has two coupled failure modes. Optimistic bias treats lucky stochastic outcomes as skillful choices, and mode collaps...Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone. This rule has two coupled failure modes. Optimistic bias treats lucky stochastic outcomes as skillful choices, and mode collapse reduces a multi-modal subgoal distribution to a single Gaussian mean that often falls in unreachable regions. We propose NFTR (Normalizing Flow subgoal policies with Triangle-slack Reweighting). A conditional Normalizing Flow replaces the Gaussian policy. A closed-form mode-averaging result identifies the Gaussian limitation, while conditional flows support multi-modal weighted maximum likelihood and direct sampling. A triangle slack score, computed from a jointly trained MRN distance head whose triangle inequality is guaranteed architecturally, corrects the AWR weight multiplicatively and measures a detour inside the learned geometry without requiring exact distance recovery. Triangle-slack vanishes on geodesics in deterministic MDPs and remains a conservative upper bound on composability violation under stochastic dynamics. The RWDR objective preserves AWR's population-level monotonic improvement and admits a three-term suboptimality decomposition. On OGBench the flow carries the larger share of the gain, while the full co-trained configuration adds task-dependent gains. The combined method avoids Gaussian mode averaging and performs well on stochastic tasks. GitHub page: https://github.com/erdemtbao/NFTR
|
| 1433 |
Exact and Certified Data Shapley for Weighted Nearest Neighbor Regression
2607.11956
|
cs.LGcs.AI
|
Zongye Lyu, Lihua Peng |
Data Shapley can be computed exactly and efficiently for unweighted K-nearest-neighbor (KNN) models, the basis of the popular KNN-Shapley. For weighted KNN, however, efficient exact algorithms are known only for hard-label classification. In regression, the pr...Data Shapley can be computed exactly and efficiently for unweighted K-nearest-neighbor (KNN) models, the basis of the popular KNN-Shapley. For weighted KNN, however, efficient exact algorithms are known only for hard-label classification. In regression, the prediction is a weighted average whose normalization term changes with every subset. The best exact algorithm then takes O(N^K) time for N training points. We show that this normalization term is not an obstacle. A subset affects the prediction only through two numbers: the total weight of its K nearest neighbors and their total weighted target. Once the weights are discretized, subsets with equal total weight share their normalization term and can be summed as a group. This gives an exact algorithm that is near-linear in N and polynomial in K and the number of weight levels, for real targets under the squared loss. The same idea gives exact values for soft-label classifiers in the same time, and speeds up the earlier exact algorithm for hard labels by a factor of N. For continuous weights, we give a deterministic approximation with a certified error bound, without discretizing the weights. On the downside, we prove that some exact values take exponentially many bits to write down, so the dependence on weight precision cannot be removed in general. In experiments on real datasets with up to 100,000 training points and with 10 or 100 validation points, we compare the exact values with Monte Carlo estimates computed in the same time. Monte Carlo ranks the points less and less like the exact values as the training set grows. Two of its runs also disagree on a large share of the top 10% of points.
|
| 1434 |
PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs
2607.12111
|
cs.LGcs.AI
|
Jing Liu, Kun Yang, Yan Wang, Dingkang Yang, Xiaoshuai Hao |
Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. Within distributed network environments, Multimodal Large Language Mode...Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. Within distributed network environments, Multimodal Large Language Models (MLLMs) serve as cognitive engines for edge devices, yet federated fine-tuning faces substantial challenges in balancing global knowledge aggregation with local adaptation under heterogeneous network conditions. Conventional federated protocols typically rely on uniform parameter aggregation, which conflates domain-invariant features with client-specific nuances, thereby resulting in suboptimal personalization and excessive communication overhead. To address these challenges, we propose PFAdapter, a communication-efficient framework introducing hierarchical LoRA decomposition to explicitly separate adapter parameters into global-shared and local-private components. Query and key projections are assigned to global synchronization for capturing universal multimodal semantics across the network, while value and output projections remain localized for edge-specific adaptation. Additionally, orthogonality regularization based on the Frobenius norm enforces strict separation between these components, preventing redundant feature learning. Selective aggregation protocols synchronize only global-shared components across the federated network, preserving local expertise and reducing communication costs by nearly 50%. Extensive experiments on VQA-RAD, SLAKE, Hateful Memes, and CrisisMMD datasets demonstrate that PFAdapter consistently outperforms state-of-the-art baselines, achieving accuracy improvements ranging from 2.4% to 4.8% across diverse edge intelligence tasks. Consequently, our framework establishes an efficient solution for agentic AI deployment in resource-constrained communication networks.
|
| 1435 |
The Dynamics of Discoverability: How Trajectories and Priors Shape Equation Recovery
2607.18490
|
cs.LG
|
Matteo Gallo, Fabio Anselmi, Paolo Lazzari |
How the dynamical regime of the observed system affects equation discovery has mainly been investigated through comparisons across systems. However, such comparisons vary both the equations and the dynamics, confounding the effect of the regime with the diffic...How the dynamical regime of the observed system affects equation discovery has mainly been investigated through comparisons across systems. However, such comparisons vary both the equations and the dynamics, confounding the effect of the regime with the difficulty of recovering the equations symbolically. We separate the two by varying the forcing of Lorenz-84, moving its fully observed post-transient trajectories through fixed-point, periodic, and chaotic regimes while preserving the equations' functional form. Within each regime, we separately vary the amount of data, the noise, and prior knowledge of which terms the equations contain. We then measure how well two complementary approaches, sparse regression over a fixed library of candidate functions (SINDy) and an evolutionary search over symbolic expression trees (PySR), recover the true equations' terms and coefficients. We find that recovery depends on whether the sampled states distinguish combinations of candidate functions: equations remain poorly recovered from fixed-point data even when the candidate set contains only the true terms. We link the effect of the dynamics on both algorithms to one object: the moment matrix of the candidate functions under the invariant measure of the regime. Small eigenvalues mark weakly distinguishable combinations of candidate functions: we show that more data, less noise, and more prior knowledge can mitigate the resulting recovery difficulties, while a zero eigenvalue makes distinct equations indistinguishable on the visited states. Hence, a more precise prior needs less informative data, with consequences for data collection, method design, and evaluation.
|
| 1436 |
In-Context Time Series Classification with Random Convolutional Features
2607.19234
|
cs.LG
|
Joscha C\"uppers, Jilles Vreeken |
Time series classification is central to domains such as medical signal analysis, industrial monitoring, and sensor-based activity recognition, where class information manifests as localized shapes, specific frequencies, temporal shifts, or complex cross-chann...Time series classification is central to domains such as medical signal analysis, industrial monitoring, and sensor-based activity recognition, where class information manifests as localized shapes, specific frequencies, temporal shifts, or complex cross-channel interactions. Random convolutional transforms capture these diverse patterns by converting time series into rich, fixed-dimensional feature representations that can be processed by standard tabular classifiers. While these representations are traditionally paired with simple linear models, we investigate whether a pretrained tabular foundation model can exploit them more effectively and how its performance depends on the available data and inference budget. We propose MASHT, a pipeline that combines MultiRocket and Hydra features with an in-context tabular foundation model. Our approach uses a pretrained tabular foundation model to bypass task-specific model training, requiring only feature extraction and direct inference. Extensive experiments demonstrate that MASHT matches state-of-the-art time series classification baselines on univariate tasks, achieving a lower average rank than HIVE-COTE 2.0. On multivariate datasets, MASHT remains highly competitive with the strongest reference methods. Controlled resource experiments show that compact feature tables retain most of the accuracy at substantially lower runtime, while TabPFN outperforms a matched linear baseline across the evaluated label budgets on univariate tasks. These results highlight practical trade-offs between predictive performance, labeled data, and inference cost.
|
| 1437 |
How Much Can a LoRA Adapter Memorize? Measuring Adapter Capacity in Bits
2607.21351
|
cs.LG
|
Kaizhen Tan, Heqing Du, Yang Feng |
LoRA adapters are often shared on their own, and the amount of information they can hold about their training data bounds what sharing them can reveal. Following Morris et al. (2026), we measure this amount in bits by training adapters on frozen language model...LoRA adapters are often shared on their own, and the amount of information they can hold about their training data bounds what sharing them can reveal. Following Morris et al. (2026), we measure this amount in bits by training adapters on frozen language models to memorize random token sequences. LoRA adapters store 2 to 3.4 bits per trainable parameter, less than full fine-tuning of the same model. The number of parameters alone does not set this amount: at equal size, adapters on MLP layers store more than adapters on attention layers, and a randomly initialized frozen model supports as much storage as a pretrained one once the scale of its output logits can be trained. We then train the same adapter on a task with supervised fine-tuning (SFT) and with GRPO. At the same accuracy, SFT stores about three times as much information specific to its training examples, and it stores most of the bits of secrets planted in the training data, which GRPO does not store because it rarely samples them.
|
| 1438 |
Experimentation and Commitment under Reward Shifts
2607.23432
|
cs.LG
|
Puping Jiang (Philip), Wei Tang (Philip), Renyu (Philip), Zhang |
Decision-makers in learning environments face a dilemma when their short-term optimal actions may not favor their long-term benefits the most. To understand the fundamental tradeoff behind the dilemma, we study adaptive experimentation with post-commitment rew...Decision-makers in learning environments face a dilemma when their short-term optimal actions may not favor their long-term benefits the most. To understand the fundamental tradeoff behind the dilemma, we study adaptive experimentation with post-commitment reward shifts. During an experiment phase, the decision-maker may adaptively test multiple options; during a subsequent commitment phase, the decision-maker must commit to a single option, whose reward may differ from its pre-commitment reward. We propose the Reserved Arm Eliminations for Commitment (RAEC) algorithm, which reserves a predetermined portion of the experiment phase to identify the best post-shift option while using the remaining rounds to minimize short-run regret. We establish regret upper bounds for RAEC across all parameter regimes and matching minimax lower bounds, providing a tight characterization of the cost of balancing short-term performance and long-term commitment. A key implication is that deciding in advance how much of the experiment phase to reserve for the commitment decision is sufficient to achieve the best possible worst-case regret rate; adapting this amount as more data are observed does not improve the rate. We further study extensions with structural knowledge of reward shifts and with concave commitment rewards and portfolio choice. Numerical experiments confirm that our proposed algorithms achieve the regret predicted by our theory and outperform other baselines.
|
| 1439 |
Rethinking the Effectiveness of Contrastive Decoding in Mitigating Hallucinations in MLLMs
2607.25196
|
cs.LG
|
Arnav Bendre, Guneesh Gupta, Shreyansh Modi, Kavish Grover, Chayan Aggarwal |
Multimodal large language models frequently describe objects that are not present in an image. Contrastive decoding has become a widely used inference-time remedy, and it is assumed to work by contrasting a normal forward pass against a hallucination-prone one...Multimodal large language models frequently describe objects that are not present in an image. Contrastive decoding has become a widely used inference-time remedy, and it is assumed to work by contrasting a normal forward pass against a hallucination-prone one so that hallucinated content is suppressed. However, whether the reported benchmark improvements actually arise from this mechanism has not been examined. Our key observation is that a correction which suppresses hallucination must depend on whether the model is hallucinating, and that this dependence can be measured directly inside the model. Following this idea, we analyse the internal computation of three contrastive decoding methods on both discriminative and generative tasks, and compare each against controls that remove the contrastive term while preserving its effect on the output. The correction turns out to be unrelated to hallucination at every layer, the gains in captioning come instead from a constraint that narrows sampling towards greedy decoding, and an unstructured perturbation of the same magnitude performs at least as well. Experiments across three models and standard hallucination benchmarks show that these improvements reflect a change in decoding behaviour rather than genuine hallucination mitigation.
|
| 1440 |
Training nGPT
2608.01284
|
cs.LGcs.AI
|
Ilya Loshchilov, Boris Ginsburg |
The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. In this paper, we describe a practical training recipe for nGPT and evaluate it on modern ...The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. In this paper, we describe a practical training recipe for nGPT and evaluate it on modern hybrid Mamba-2--Transformer Mixture-of-Experts (MoE) models. The recipe introduces Logit Gradient Preconditioning, Logarithmic Learning Rate Decay, GatedAdamW, angular update control, and optional exploration mechanisms. Compared with an unnormalized model of the same hybrid MoE architecture trained with AdamW, the 30B-total-parameter nGPT model reaches the same validation loss using approximately half as many training tokens. The recipe scales across the models considered, which contain up to 30B total parameters.
|
| 1441 |
An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules
2608.01383
|
cs.LG
|
Yichao Cai, Javen Qinfeng Shi |
Masked prediction learns to infer missing variables from visible context selected by a mask schedule. When does small excess risk guarantee recovery of the true joint distribution? We study this question on finite product spaces under masked-block log loss, wi...Masked prediction learns to infer missing variables from visible context selected by a mask schedule. When does small excess risk guarantee recovery of the true joint distribution? We study this question on finite product spaces under masked-block log loss, with conditionals induced by a single joint distribution. To quantify recovery, we introduce an $\varepsilon$-identifiability modulus measuring the worst-case error among joint distributions with population excess risk at most $\varepsilon$. For data with separated modes pinned down by sufficiently large visible contexts, schedules that always retain such contexts can permit nonvanishing mode-weight errors while incurring exponentially small excess risk. An information decomposition explains why: the loss captures mode-weight mismatch only where the visible context leaves the mode uncertain. We prove two-sided bounds showing that, over a fixed range of mode weights, sensitivity to mode reweighting is governed by residual mode uncertainty averaged over the mask schedule. Assigning schedule mass to low-visibility masks that retain this uncertainty yields recovery bounds within the reweighting family. Beyond this family, positive full-mask probability characterizes uniform control of joint KL divergence by excess risk. Exact calculations and controlled gradient optimization validate these predictions.
|
| 1442 |
ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces
2608.01968
|
cs.LG
|
Kunal Kumar Pant, Nithin Nagaraj |
Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks. Yet frozen transformer input-embedding spaces may also be examined through their responses to a controlled dete...Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks. Yet frozen transformer input-embedding spaces may also be examined through their responses to a controlled deterministic probe before contextual computation or task-specific adaptation. Guided by this response-based view, we introduce ChaosProbe, a deterministic neurochaos-inspired method for constructing response-based fingerprints of frozen transformer input-embedding spaces. For each prompt-level embedding matrix, ChaosProbe applies a chaotic trajectory-based transformation and summarizes its Firing Rate and Entropy channel responses with complementary representation-level measures, producing a fixed-length signature for each model. In a bounded proof-of-concept study of $80$ neutral prompts and four pretrained models---GPT-2, DistilGPT2, BERT-base-uncased, and RoBERTa-base---Pearson correlation, Spearman correlation, and cosine similarity each recover all four same-family nearest-neighbor assignments and both expected mutual family pairs. Euclidean distance recovers three of the four assignments and one of the two mutual family pairs. Paired bootstrap resampling supports the stability of the Pearson and Spearman pairings over the observed prompt set, and signature-validity checks show that constant or collapsed responses do not dominate the reported fingerprints. These results provide a cohort-dependent proof of concept that deterministic neurochaotic response signatures can expose broad structure among frozen transformer input-embedding spaces.
|
| 1443 |
CLEARMIND: Closed-Form Laplacian Embeddings of Arakelov-Green Resistance for Morphological Intrinsic Neuronal Distances
2608.04460
|
cs.LGcs.AI
|
Yuyang Zhang, Weihan Xu, Xuehai Zhou, Shucheng Cao, Qihuang Zhang |
Tropical geometry turns a metric graph into a flat torus, its tropical Jacobian, and measures distances between the images of its points by a closest vector problem on the period lattice. We prove that for points of the graph this problem is solved by geodesic...Tropical geometry turns a metric graph into a flat torus, its tropical Jacobian, and measures distances between the images of its points by a closest vector problem on the period lattice. We prove that for points of the graph this problem is solved by geodesics: the squared tropical polarization distance equals the path metric minus the Arakelov-Green (AG) distance, an effective-resistance kernel given in closed form by the Laplacian pseudoinverse. Equivalently, the AG distance and the path metric are the minimal energies of real and of integral unit flows, and the squared tropical distance is exactly their integrality gap. The full tropical distance matrix is thus computable in cubic time. On this identity we build CLEARMIND, a training-free descriptor of 3D neuronal morphology. A reconstruction is reduced to its branching skeleton, neurite tips near the soma are joined to it, short bridges are contracted, and the AG matrix of the result is summarized by its spectrum. Every step has an exact algebraic description; for instance, the period matrix records the shared soma-to-tip path lengths, and the spectrum has exactly one positive eigenvalue, so its absolute values determine it. The signature needs no lattice search, is invariant to vertex order, rigid motions and subdivision, and is Lipschitz in the edge lengths. On ACT-4, JML-4 and BIL-6, appending the 64-dimensional spectrum improves a point-cloud GNN and MorphVAE on every dataset, by up to 20.3 points; a Tree-LSTM with the spectrum outperforms every reimplemented baseline; and an MLP on the spectrum alone is a competitive classifier that surpasses every reimplemented deep baseline on ACT-4. On BREC, a GIN with AG eigenvector encodings distinguishes 70.0% of the pairs beyond the 1-WL limit, including 33% of the CFI pairs, more than PPGN (23%) and I$^2$-GNN (21%), and the spectrum alone separates 217 of 400 pairs without training.
|
| 1444 |
Constrained Graph Diffusion for Mixed Integer Optimization
2608.13079
|
cs.LG
|
Vincenzo Di Vito, Yusuf Guven, Babak Badnava, Deepjyoti Deka, Kaarthik Sundar |
This paper proposes a novel learning-based approach to approximately solve instances of mixed-integer optimization problems. These problems are computationally challenging, as they require jointly determining discrete and continuous decisions while satisfying ...This paper proposes a novel learning-based approach to approximately solve instances of mixed-integer optimization problems. These problems are computationally challenging, as they require jointly determining discrete and continuous decisions while satisfying complex combinatorial constraints. problem-agnostic and can accommodate a broad class of mixed-integer optimization problems through suitable projection operators. We introduce Constrained Graph Diffusion (CGD), a learning-based framework that approximately solves recurring instances of such problems by learning a conditional distribution over their discrete decisions. CGD uses a graph-based diffusion model and incorporates constraint information directly into the reverse diffusion process, steering intermediate predictions toward the feasible region throughout generation. By operating on continuous relaxations of the discrete variables, CGD defines a differentiable constrained generation pathway up to terminal discrete recovery. Once the discrete decision is recovered and fixed, a numerical optimizer solves the remaining continuous problem, avoiding online combinatorial search over the binary variables while retaining numerical optimization for continuous completion. We evaluate CGD on AC-OPF with branch switching and discrete portfolio optimization, demonstrating substantial improvements in feasibility and solution quality over learning-based baselines while achieving speedups of up to $543\times$ over state-of-the-art MIP solvers on large instances.
|
| 1445 |
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
2608.16072
|
cs.LGcs.AI
|
Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu |
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed wei...Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. However, when optimizing multiple reward objectives, existing methods typically scalarize the reward vector with a fixed weighted sum before group-wise standardization. We show that this design leads to two fundamental problems: rollouts with distinct reward profiles can receive identical advantages, and all objectives are optimized with fixed relative weights regardless of their current level of saturation. Consequently, an objective whose rewards are already near their upper bound can retain substantial influence when its rewards still vary within rollout groups. We introduce Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization (SA-MRPO), which standardizes each reward objective independently and adaptively discounts its contribution according to a batch-level estimate of objective saturation. This changes the relative contribution of each objective according to its observed reward headroom. We further derive an exact condition under which saturation aware reweighting reverses the sign of a rollout's aggregate advantage. Across mathematical reasoning with two- and three-objective reward combinations, SA-MRPO improves the harder correctness objective over GDPO in 12 of 15 benchmark comparisons, with gains of up to $5\%$ on AIME24. On adaptive reasoning it improves accuracy on all five benchmarks, by $3.8\%$ on average and up to $9.2 \%$ on AMC23, and on coding benchmarks it improves pass rate by up to $2.3\%$, while in all settings maintaining the easier objectives near their already satisfied levels. Additional experiments characterize the accompanying reward tradeoffs and sensitivity to corrupted rewards.
|
| 1446 |
Cone Rayleigh Levels: Finite Perturbations and Certified Control
2608.27122
|
cs.LG
|
Yavdat Sh. Il'yasov, Nur F. Valeev |
We study the reuse of positive trial profiles under finite perturbations of nonsymmetric matrix pencils B-\lambda G. The lower and upper cone Rayleigh levels need not coincide and are defined without requiring positive eigenvectors. In the positive orthant wit...We study the reuse of positive trial profiles under finite perturbations of nonsymmetric matrix pencils B-\lambda G. The lower and upper cone Rayleigh levels need not coincide and are defined without requiring positive eigenvectors. In the positive orthant with positive diagonal $G$, we derive computable perturbation bounds and determine the exact worst-case trial gap over prescribed independent entrywise perturbation classes. This yields the largest uniform radius meeting a given trial-gap tolerance for fixed profiles, certifying trial-value accuracy without recomputation. Experiments on a nonnegative operator and five signed matrices compare sufficient and optimal radii, directional thresholds, and cold- and warm-start recomputation. Several signed cases exhibit severely limited uniform profile reuse despite the optimality of the radius. Exact rational checks verify the reported bounds and worst-case constructions for stored numerical inputs. A finite-budget linear program and differentiable constraints illustrate applications to verified control and learning.
|
| 1447 |
DCCQ: From Ordered Bernoulli Levels to Critical-Line Geometry: Integer Quantization, Bernoulli Residual Phase, and Prime-Power Spectra
2609.03801
|
cs.LG
|
Y. Kenan Y{\i}lmaz |
We present the discrete--complex complement quotient (DCCQ) framework through the ordered Bernoulli-word kernel $f(p,n,k)=p^k(1-p)^{n-k}$ and its inverse-integer level sets. The binary level $2^{-n}$ selects $p=1/2$ as the unique real split-independent anchor....We present the discrete--complex complement quotient (DCCQ) framework through the ordered Bernoulli-word kernel $f(p,n,k)=p^k(1-p)^{n-k}$ and its inverse-integer level sets. The binary level $2^{-n}$ selects $p=1/2$ as the unique real split-independent anchor. Imposing the additional continuation rule $1-z=\overline z$ gives the conjugation-symmetric line $z=1/2+iu$. On the central normalized branch, the inverse-level split $\kappa=k/n$ also has real part $1/2$. The quadratic coordinate $Q(z)=1/4+u^2$ identifies complementary conjugate points, has minimum $1/4$, and has an explicit two-sign inverse; restricting its values to integers defines an arithmetic quantization. For supplied positive critical-line zero ordinates $\gamma_k$, the induced level has the exact decomposition \[ L_k=\frac14+\gamma_k^2=N_k+\delta_k,\qquad \delta_k=\widetilde B_1\!\left(\gamma_k^2+\frac34\right), \] where $N_k=\lfloor L_k+1/2\rfloor$ and $\widetilde B_1(x)=\{x\}-1/2$. The integer--phase pair $(N_k,e^{2\pi i\delta_k})$ retains the level and the positive ordinate. Prime-exponent coordinates describe the factorization of the integer component, while conditional Fourier moments formulate an open question about the residual phases. The finite numerical diagnostics are descriptive. A distinct complex exponent $s$ gives an inverse-level representation of the supplied term $m^{-s}$; direct substitution cancels the auxiliary coordinates. These terms assemble into the classical Dirichlet series and Euler product for $\Re s>1$. The framework organizes exact coordinate identities, classical zeta connections, and numerical diagnostics; it establishes neither conditional equidistribution nor a zero-location theorem.
|
| 1448 |
Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICU
2609.08992
|
cs.LG
|
Athanasios Papastathopoulos-Katsaros, Alexandra Stavrianidi, Zhandong Liu |
False ventricular tachycardia (VT) alarms are a leading contributor to alarm fatigue in intensive care units. We propose a deep learning framework combining a 1D SE-ResNet with ICU-realistic data augmentations and a physics-informed auxiliary reconstruction ta...False ventricular tachycardia (VT) alarms are a leading contributor to alarm fatigue in intensive care units. We propose a deep learning framework combining a 1D SE-ResNet with ICU-realistic data augmentations and a physics-informed auxiliary reconstruction task based on the three-element Windkessel hemodynamic model, implemented as a differentiable forward simulation. By requiring the network's latent representation to produce physiologically plausible arterial pressure waveforms, artifact-driven ECG patterns are penalized while true VT remains coherent across modalities. Evaluated on the VTaC benchmark under a strict real-time protocol (10-second pre-alarm window), our method achieves a 5-point Challenge Score improvement over prior state-of-the-art. Ablation studies confirm that the physics-informed objective is the primary performance driver, providing gains in accuracy, 2x label efficiency, and more localized and clinically meaningful ECG segments.
|
| 1449 |
Online Inverse Integer Linear Optimization via Small-Gradient Skipping: Constant Regret and Finite Mistakes
2609.09809
|
cs.LG
|
Akira Kitaoka |
In online inverse linear optimization, the learner predicts a weight at each round, observes the optimal action of the agent, and updates its prediction. In the general setting, the gap of $\log T$ between the regret upper bound $O(d \log T)$ and the lower bou...In online inverse linear optimization, the learner predicts a weight at each round, observes the optimal action of the agent, and updates its prediction. In the general setting, the gap of $\log T$ between the regret upper bound $O(d \log T)$ and the lower bound $\Omega(d)$ is unresolved (here $T$ is the total number of rounds and $d$ is the dimension). When the action set is M-convex, the regret is known to be bounded by $O(d \log d)$, but the method attaining it computes a center of gravity at every round. This paper therefore proposes Small-Gradient Skipping (SGS), a mechanism that skips the update at rounds without a mistake, and applies it to online gradient descent, the online Newton step, and MetaGrad. When the forward problem is an integer linear program with a unique optimal solution, the number of mistakes is bounded, for all three, by a quantity independent of $T$; and for the online Newton step and for MetaGrad with SGS, the dimension dependence of the regret becomes $O(d^2)$, that is, the factor $\log T$ is removed. Moreover, when the action set is M-convex, the regret is bounded efficiently without computing a center of gravity.
|
| 1450 |
DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy Prediction
2609.10796
|
cs.LG
|
Yingfan Xu, Tieming Liu, Ye Liang |
Pretrained diabetic retinopathy (DR) prediction models differ in their input fields, serialization formats, preprocessing requirements, and output semantics. Making these models accessible through a common clinical interface therefore requires explicit coordin...Pretrained diabetic retinopathy (DR) prediction models differ in their input fields, serialization formats, preprocessing requirements, and output semantics. Making these models accessible through a common clinical interface therefore requires explicit coordination between the user interface and the inference service. We designed and implemented DR-LabStack, a React-Flask web system integrating four externally developed pretrained models: RuleFit, Pruned RuleFit, Elaborative XGBoost, and Two-level Ensemble. A shared form retrieves ordered model features, renders model-specific numerical and categorical controls, and constructs a positional input vector. Backend adapters load heterogeneous artifacts and apply the ensemble's accompanying scaler, while a common JSON response supports binary classification display alongside method and source information. Functional evaluation on September 8, 2026 used copied application files and real model artifacts in a documented isolated environment. All four models loaded and exposed their 14-, 6-, 8-, and 25-field contracts. Sixty-two Flask test-client requests characterized service behavior; 12 limited-vector checks confirmed invocation-path and threshold consistency. Twenty-four browser-component scenarios with mocked transport verified input ordering and result rendering and characterized input-validation behavior. The resulting system demonstrates a reusable interaction and serving workflow for heterogeneous DR models. The contribution is web-system design, integration, and software functionality; clinical effectiveness and clinician usability require separate evaluation.
|
| 1451 |
Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning
2609.12424
|
cs.LG
|
Taoran Liang, Yang Liu, Shang Luo, Yingguang Yang, Rongrong Zhang |
Long-horizon language-model agents trained with reinforcement learning oftenreceive sparse outcome rewards that do not reveal which decisions along a tra-jectory deserve credit. Episode-level advantages provide coarse trajectory-widecredit, while step-level co...Long-horizon language-model agents trained with reinforcement learning oftenreceive sparse outcome rewards that do not reveal which decisions along a tra-jectory deserve credit. Episode-level advantages provide coarse trajectory-widecredit, while step-level comparisons offer finer resolution with context-dependentestimation noise. We propose Granularity-Adaptive Credit Assignment (GACA),a critic-free method that adaptively mixes episode- and step-level credit for eachdecision during policy optimization. GACA normalizes the sampled response'smean per-token negative log-likelihood (NLL) within each trajectory and uses theresulting criticality score to determine the step-specific mixture. The computationreuses rollout log-probabilities without additional training rollouts or model eval-uations. Our analysis characterizes optimal score-dependent mixing and derivesconditions linking expected NLL to a lower bound on the preferred step-levelweight. Across ALFWorld and WebShop with 1.5B and 7B backbones, GACAachieves the highest reported mean success rates among the compared methods,while introducing negligible additional computation.
|
| 1452 |
A Measured Communication-Quality Frontier for Federated LoRA Fine-Tuning with Adaptive Phase-Switching
2609.13512
|
cs.LGcs.AI
|
Jerry Adams Franklin |
Federated fine-tuning of large language models with low-rank adaptation (LoRA) reduces the number of trainable parameters, but communication remains the dominant cost, and protocols are usually compared by parameter-count ratios rather than by measured bytes. ...Federated fine-tuning of large language models with low-rank adaptation (LoRA) reduces the number of trainable parameters, but communication remains the dominant cost, and protocols are usually compared by parameter-count ratios rather than by measured bytes. This paper measures per-round upload and download bytes for five federated LoRA protocols, three of them from prior work, and places them on a single communication-quality frontier scored by held-out instruction-following loss. The frontier has a knee. ReverseAdaptive, which learns both LoRA factors before freezing one once the relative improvement in training loss falls below a dimensionless threshold, sits at that knee: it cuts measured round-trip communication by 40.5% relative to FLoRA at a held-out loss cost of 0.0063, and beats FFA-LoRA, which freezes that factor at initialization, by 0.0182 in held-out loss, more than twenty times the largest per-method seed standard deviation. The same threshold carries to LLaMA-3.2-3B without retuning, where it saves 30.0%.
|
| 1453 |
Data-Free On-Policy Distillation: How Far Can We Go Without External Data?
2609.14193
|
cs.LGcs.AI
|
Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu |
On-policy distillation (OPD) is increasingly applied to frontier foundation model post-training. Prior work in this area has largely focused on algorithmic advances, yet it remains unclear how much OPD depends on its training questions and, in particular, how ...On-policy distillation (OPD) is increasingly applied to frontier foundation model post-training. Prior work in this area has largely focused on algorithmic advances, yet it remains unclear how much OPD depends on its training questions and, in particular, how far this dependence can be reduced. Across two representative single-teacher OPD settings, we find that training on 8 real prompts yields performance comparable to training on 17k problems, while datasets differing substantially in measured difficulty and initial distillation gap yield similar outcomes. Our analyses suggest two complementary explanations: repeated sampling could allow even a few prompts to expose substantial teacher supervision, while OPD transfers generalizable reasoning capabilities beyond dataset-specific knowledge. Building on these observations, we next propose a data-free on-policy distillation (DF-OPD) setting to investigate whether the system can supply the training questions itself, eliminating the need for external data. With 64 self-generated questions obtained without seed examples, DF-OPD yields performance comparable to full-data OPD in both single-teacher settings. This finding also holds in multi-teacher OPD: across mathematics, code, and instruction following, 1k generated questions achieve performance comparable to training on approximately 7k real post-training examples. We further explore whether OPD can operate even without explicit training questions. The experiments show that this is effective only in limited cases, where the student unexpectedly generates and answers its own questions, thereby reducing the process to an implicit form of DF-OPD. Together, these findings invite a reassessment of the role of training data in on-policy distillation. Code is available at \url{https://github.com/Ryuki661/DF-OPD}
|
| 1454 |
Hybrid coupling with numerics-informed neural networks and the overlapping Schwarz alternating method
2609.17841
|
cs.LG
|
George Chumbipuma, Irina Tezaur, Alejandro Diaz, Beatrice Riviere |
We develop a hybrid modeling framework for coupling pre-trained numerics-informed neural networks (NINNs) with classical full order models (FOMs) using the overlapping Schwarz alternating method. We consider the two-dimensional advection-diffusion equation in ...We develop a hybrid modeling framework for coupling pre-trained numerics-informed neural networks (NINNs) with classical full order models (FOMs) using the overlapping Schwarz alternating method. We consider the two-dimensional advection-diffusion equation in the advection-dominated, Peclet-number 10^6 regime. We first demonstrate that, unlike the corresponding physics-informed neural network (PINN), a monolithic NINN can be accurately trained on our model problem without domain decomposition. We then employ overlapping multiplicative Schwarz as a deployment mechanism for coupling a pre-trained, subdomain-local NINN with a neighboring FOM, with the NINN weights held fixed throughout the Schwarz iteration. We consider two training approaches for the subdomain-local NINNs: a top-down approach, in which boundary data are obtained from a coupled Schwarz solve on the full domain with a FOM on each subdomain (FOM-FOM Schwarz), and a bottom-up approach, in which boundary traces are generated synthetically on the NINN subdomain without requiring any full-domain solves. The resulting hybrid NINN-FOM solutions agree closely with the corresponding FOM-FOM Schwarz solutions, with the top-down and bottom-up training approaches yielding comparable accuracy.
|
| 1455 |
Storing Is Not Remembering: LSTM-UT and Bounded Gated Memory for Looped Transformers
2609.19521
|
cs.LG
|
Aras Kavuncu, Muhammad Burhan Hafez |
Recurrent-depth Transformers reuse one block across many steps, so information needed later must survive repeated rewriting of the hidden state. A natural remedy is to keep more history. We show that, in controlled cellular-automaton tasks, making history avai...Recurrent-depth Transformers reuse one block across many steps, so information needed later must survive repeated rewriting of the hidden state. A natural remedy is to keep more history. We show that, in controlled cellular-automaton tasks, making history available is not the same as making it usable. Using Rule 30, where the correct state is known at every recurrent step, we test depth extrapolation and de- layed recall, the recovery of an earlier state after further computation. CoTFormer, which caches keys and values from every earlier step, extrapolates less far and recalls less accurately than a Block Universal Transformer (BUT) that keeps only its current state. Interventions show that its retained history can pull a corrected trajectory back toward failure, and that the cache block written at the requested step is neither necessary nor sufficient for recall. We introduce LSTM-UT, which adds a small, bounded, gated cell state to the shared block. Trained to depth 12, LSTM-UT keeps 99.7% exact-row accuracy at depth 60, where BUT gets no row fully correct, and one checkpoint stays above 99.95% at depth 1,000. It also improves delayed recall over both baselines, and the advantage largely persists at near-matched parameter counts. On these tasks, a small state under learned control proved more useful than a complete but unaddressed history. In OpenWebText2 language modelling, LSTM-UT outperforms BUT and, at equal width, reaches slightly lower perplexity than CoTFormer while CoTFormer needs up to 91% more training time per step; against a parameter-matched CoTFormer, LSTM-UT comes within 0.6 perplexity.
|
| 1456 |
LumoTree: Path-Parallel Speculative Verification for Hybrid Language Models
2609.23900
|
cs.LG
|
Zhiyuan Ma |
Tree speculative decoding for recurrent-hybrid language models requires each accepted path to maintain consistent recurrent state, convolution history, and attention caches. We present LumoTree, a GPU serving design that coordinates verification and commitment...Tree speculative decoding for recurrent-hybrid language models requires each accepted path to maintain consistent recurrent state, convolution history, and attention caches. We present LumoTree, a GPU serving design that coordinates verification and commitment through a shared tree descriptor. The verifier processes independent paths in parallel and reuses recurrent state tiles across local updates. Boundary states connect dependent paths. Accepted-path replay publishes the selected continuation into native running state, alongside coordinated convolution gathering and attention-cache remapping. Fused selection, device-resident acceptance, graph replay, and tree-aware attention integrate this organization into the serving cycle. The design separates temporary branch computation from persistent request state while retaining native prefix-cache interfaces. We evaluate numerical behavior, serving performance, and coding-agent outcomes on NVIDIA DGX Spark.
|
| 1457 |
On the SoS Certifiability of Log-Concave Distributions
2609.30105
|
cs.LG
|
Aleksandr Storozhenko |
We prove that for every isotropic log-concave distribution $P$ on $\mathbb{R}^d$ and every even $m\ge2$, the polynomial $(Cm)^m\|v\|_2^m - \mathbb{E}_{X\sim P}\langle X,v\rangle^m$ is a sum of squares, where $C>0$ is a universal constant. This improves on t...We prove that for every isotropic log-concave distribution $P$ on $\mathbb{R}^d$ and every even $m\ge2$, the polynomial $(Cm)^m\|v\|_2^m - \mathbb{E}_{X\sim P}\langle X,v\rangle^m$ is a sum of squares, where $C>0$ is a universal constant. This improves on the Poincar\'e-dependent bounds (Kothari and Steinhardt, 2017), recovering the optimal moment bounds for log-concave distributions. As an immediate corollary, we obtain computationally efficient algorithms with dimension-free error guarantees for a wide range of statistical estimation problems. Our proof proceeds by using stochastic localization to decompose $P$ as an average of strongly log-concave measures, whose centered moments admit the subgaussian certificates (Diakonikolas et al., 2025). With a covariance-adapted choice of localization, we show that a fourth-moment certificate derived from the variance inequality for quadratic forms (Letwin, 2026) suffices to control this averaging at every even degree.
|
| 1458 |
GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for De Novo Peptide Sequencing
2609.30542
|
cs.LG
|
Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed |
De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, a...De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed reconstruction task, without considering which missing fragments are most relevant to decoder errors. Moreover, existing peak representations do not explicitly model mass differences between peaks, despite their fundamental importance. We introduce GyroNovo, a framework with two main contributions. First, we use decoder errors observed during training to adapt the imputation objective, prioritizing fragments associated with frequent decoding errors. We further use the decoder error distribution to construct easy and hard augmented views of each spectrum, enabling the decoder to learn under varying degrees of spectral corruption and missing-fragment severity. Second, we introduce a mass-aware inductive bias into self-attention by using rotary embeddings to encode pairwise mass differences between spectral peaks. Together, these components align missing-fragment recovery with decoder behavior while explicitly incorporating the mass relationships that underlie peptide fragmentation. At inference time, GyroNovo retains a standard encoder-imputer-decoder architecture and requires neither additional inputs nor auxiliary search procedures. Experiments on NovoBench show gains of about 9 percentage points in peptide-level precision and 7 percentage points in amino-acid-level precision over the state-of-the-art baseline. Code: https://github.com/UBC-NLP/gyronovo.
|
| 1459 |
EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
2609.30929
|
cs.LG
|
Takumi Fujimoto, Hiroaki Nishi |
Completed forecasts provide residual feedback, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficie...Completed forecasts provide residual feedback, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Each channel shares the endpoint across component-wise online ridge regressions using current-forecast coefficients. We evaluate eight multivariate series, DLinear and PatchTST, three seeds, and two training variants: 96 matched fixed-base conditions at a 24-step horizon. EPOC achieves mean condition-wise reductions in mean squared error (MSE) and mean absolute error (MAE) of 15.40% and 9.35% from the uncorrected base, respectively, with a median of 6,352 B in retained auxiliary arrays. EPOC also outperforms the $\delta$-Adapter, COSA, FAC, and OMPB in paired MSE on most conditions while retaining less state. Full ELF achieves the largest mean MSE reduction, 19.29%, but its median retained state is 474,048 B ($\times$75 relative to EPOC). Equal-size summary controls favor the endpoint by 1.65--2.20% in paired MSE; a coefficient-reconstructed endpoint yields similar accuracy to the observed endpoint, highlighting its role as a shared input. Increasing the retained DCT component count from 4 to 8 adds 1.00 percentage point of MSE reduction for 5,728 B. On jointly trained bases, EPOC lowers MSE by 16.69--20.15% relative to globally blended TEFL-style adapters applied to the same base. Code and numerical records are available at [https://github.com/keiotakmin/endpoint-preserving-residual-correction](https://github.com/keiotakmin/endpoint-preserving-residual-correction).
|
| 1460 |
The Linear Representation Hypothesis for Vision-Language-Action Models
2609.30996
|
cs.LGcs.AI
|
Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han |
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vis...The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in a VLA evolves jointly with the system dynamics: the representation influences the actions selected by the policy, which alter the physical state and, in turn, the next representation. In this paper, we develop a theoretical, signature-based formulation of the LRH for VLA that unifies representations and policies. On the representation side, we establish the existence of representations from which the future evolution of a QoI under a candidate action trajectory can be recovered via linear probing. On the policy side, we introduce a signature generalized linear model for stochastic action chunks. This structure yields a monotonic change in the expected future QoI along linear paths in natural parameter space, enabling linear steering. We first validate this structure in a controlled oracle setting, then examine whether the same probing and steering mechanisms emerge in a pretrained VLA.
|
| 1461 |
Trust Guided Decision Transformer
2609.31586
|
cs.LG
|
Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee |
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays el...Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.
|
| 1462 |
Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase
2609.31630
|
cs.LG
|
Yanhai Zhang, Jie Zhang, Chi Xu |
Replay-based continual learning rehearses past data either in an offline phase, during which the agent stops acting, or interleaved with the live stream, where it perturbs the computation serving the current input. An agent that learns in deployment can afford...Replay-based continual learning rehearses past data either in an offline phase, during which the agent stops acting, or interleaved with the live stream, where it perturbs the computation serving the current input. An agent that learns in deployment can afford neither. Motivated by local sleep, the use-dependent off periods of individual cortical circuits in awake animals, we show that replay can instead be written into the degrees of freedom the current input leaves unused. In a network with k-winner-take-all hidden layers, confining replay updates to synapses whose presynaptic unit is silent or whose postsynaptic unit is inactive leaves the hidden computation on the current batch invariant: exactly so for silent and suppressed units, and for all but 0.3% of samples in practice. A refractory rule under which units that have just fired sit out the next competition doubles the width of this channel and carries most of the accuracy. On class-incremental split-MNIST the resulting learner, with no offline phase, matches or exceeds the best offline rehearsal schedule and outperforms experience replay, ER-ACE and unmasked interleaved replay, each re-tuned under the same micro-batch schedule. Against DER++ the comparison splits by protocol: with five epochs per task DER++ leads by 1.5 points once it runs under that schedule, and in a single pass over the stream, the regime closest to the agent deployment setting, the system leads it by 1.6 while offline rehearsal falls 15 points behind. On split CIFAR-10 it again leads offline rehearsal and experience replay, but trails ER-ACE and DER++ by two to three points. The construction is not tied to the local learner: on a backprop network with k-winner-take-all hidden layers under the same schedule, refractory rotation adds half a point in five epochs and three in a single pass, and isolation again costs nothing on top of it.
|
| 1463 |
Model-Agnostic Online Certificate-Driven Calibration for Time Series Forecasting Under Distribution Shift
2609.31960
|
cs.LG
|
Chenfeng Huang, Zixuan Ma, George Michailidis |
Time series out-of-distribution generalization requires forecasters to remain reliable when deployment dynamics differ from training conditions due to covariate shift, concept shift, and temporal dependence. Probably Approximately Correct Bayesian domain adapt...Time series out-of-distribution generalization requires forecasters to remain reliable when deployment dynamics differ from training conditions due to covariate shift, concept shift, and temporal dependence. Probably Approximately Correct Bayesian domain adaptation provides computable certificates by decomposing target risk into a source risk term, a source-to-target mismatch term, and a complexity term, but standard analyses rely on independent sampling and distributional stability, assumptions that are violated in time series by serial dependence and nonstationary shift. We propose a model-agnostic online martingale Probably Approximately Correct Bayesian framework that yields finite-sample certificates under temporal dependence and distribution shift. The certificate replaces independent-sample concentration with martingale concentration that adapts to loss scale and predictable variation. We use the certificate as a surrogate regularizer for online calibration by training a gated residual Bayesian head on top of a fixed forecasting backbone, producing a corrective update that reverts to the backbone prediction when the gate is closed. Online calibration combines a source risk anchor, a posterior-shift penalty, and a time-adaptive mismatch term computed from target windows observed before forecasting. It follows a predict-then-update protocol in which outcomes become available only after forecasting and are used to update subsequent predictions. Experiments across convolutional, attention-based, and large language model-based forecasters show improved stability and accuracy under covariate and concept shift.
|
| 1464 |
Traceprop: Training Data Attribution in a Single Pass
2609.32380
|
cs.LG
|
Amit Nautiyal |
Data attribution tools used in practice, TRAK, LoGRA/LogIX, and EK-FAC, run after training: they recompute per-sample gradients in a separate pass over the training set, and curvature-aware variants need a further pass to estimate covariance. Traceprop records...Data attribution tools used in practice, TRAK, LoGRA/LogIX, and EK-FAC, run after training: they recompute per-sample gradients in a separate pass over the training set, and curvature-aware variants need a further pass to estimate covariance. Traceprop records projected per-sample gradients and K-FAC covariance statistics inside the training backward pass itself. On Pythia-1B LoRA fine-tunes on an A100, this adds 2.6% to training time, against 10.7% for LogIX's best inline configuration, random-init (4.1x lower, one-sided Mann-Whitney p = 9e-5, n = 10 per arm); at Pythia-6.9B the numbers are 11.3% and 27.0% (p = 0.004). LogIX's default configuration, PCA-init, is both slower and lower quality than random-init, so we compare against random-init throughout. On a small transformer where LDS is measurable, Traceprop matches LogIX's best configuration: pooled LDS difference +0.0022, 95% CI [-0.0038, +0.0075]. Two things do not work: at Pythia-160M/SST-2, every method is indistinguishable from a low noise ceiling, and on planted-backdoor and mislabel detection, gradient attribution does not beat gradient norm or representation similarity. The contribution here is systems, not a new estimator: LogIX's attribution quality, in one pass instead of two.
|
| 1465 |
When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability
2609.33169
|
cs.LG
|
Xingjian Li, Jianhua Z. Huang, Junli Duan |
Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping...Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matters. Holding the forward computation fixed and varying only derivative edges, we measure parameter gradients, the updates the optimizer applies, and continued training in a Transformer vessel-trajectory model and a quadrotor tracking policy. In the vessel model, detaching the key-value cache shrank the gradient to about a tenth of its norm, with little rotation, when gradients flowed through all earlier physical states, but barely changed it under one-step physical credit. In this strongly clipped regime the optimizer, not the gradient, set how far updates differed: global-norm clipping removed most of the gradient difference between memory-cut graphs, whereas AdamW turned a 2% gradient difference between two placements of the cut into update differences of up to 31% at the step where the placement was switched. In a quadrotor trained from initialization with 0.20 m/s velocity noise, removing memory raised tracking error by 43% and cutting memory gradients raised it by 32%; at low noise the cut's mean cost exceeded the value of memory. Two-step truncation segments gave no measurable gain, although with hidden velocity a two-step window captured most of the value of memory; eight-step segments removed half to three quarters of the cost. Switching the cut on only for the last fifth of training understated its cost about threefold at 0.20-0.30 m/s, but not at low noise or with hidden velocity. These results suggest measuring the cost of a memory cut by training with it from initialization, and comparing backward graphs by the updates the optimizer applies rather than by raw gradients.
|
| 1466 |
Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
2609.33754
|
cs.LGcs.AI
|
Simeon Allmendinger, Domenique Zipperling, Burhanettin Bahadir Kibar, Niklas K\"uhl |
Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analyt...Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analytics. Federated learning enables collaboration without direct data sharing but does not resolve minority-class scarcity. Synthetic data generation can help, yet lightweight methods are interpolation-bound, while generative models require substantial data and computation. Existing collaborative generative approaches often rely on federated learning, imposing considerable organization-side training burdens. In this paper, we examine CollaFuse as a collaborative diffusion-based alternative for fraud detection and evaluate it across five fraud datasets. Compared with classical oversampling, local generative baselines, and centralized diffusion benchmarks, CollaFuse does not achieve the highest local fidelity but improves downstream fraud detection more consistently across most datasets. These findings suggest that synthetic data create analytical value less through local realism than through transferable cross-organizational structure.
|
| 1467 |
Oracle-Efficient Online Classification with Stochastic Inputs and Adversarial Outputs
2609.33760
|
cs.LG
|
Gon Buzaglo, Elad Hazan |
We consider binary prediction with i.i.d. contexts from an unknown distribution and adaptively chosen losses. We show that a simple Follow-the-Perturbed-Leader algorithm using a Gaussian perturbation for each observed context achieves $\widetilde O(\sqrt{T\log...We consider binary prediction with i.i.d. contexts from an unknown distribution and adaptively chosen losses. We show that a simple Follow-the-Perturbed-Leader algorithm using a Gaussian perturbation for each observed context achieves $\widetilde O(\sqrt{T\log N})$ regret for a class of $N$ experts, while requiring one optimization-oracle call per round and no explicit enumeration of the class. For an infinite hypothesis class $\mathcal H$, the same algorithm achieves $\widetilde O(\sqrt{T\operatorname{VC}(\mathcal H)})$ regret. This resolves an open problem posed by Lazaric and Munos (2012), showing that hybrid classification is computationally as easy as statistical learning. As an application, we reduce the problem of contextual bandits with $K$ actions to classification through uniform exploration, achieving $\widetilde O(K^{2/3}T^{2/3}(\log N)^{1/3})$ regret. This matches the best known dependence on the horizon while removing the context-distribution access required by prior oracle-efficient methods.
|
| 1468 |
Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
2609.33901
|
cs.LGcs.AI
|
Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron, Nadav Dym |
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empiric...Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical foundations remain limited. In this work, we study when finite probe-based representations are sufficient for learning neural functionals. We establish general identification and universality results for probing, and show that using intermediate hidden representations can provide significantly more informative representations than relying only on final outputs. Motivated by these results, we introduce HIDDENPROBE, a simple architecture for learning from hidden probe responses. Across a range of neural functional benchmarks, including both MLPs and Transformers, HIDDENPROBE consistently improves over existing probing methods and achieves state-of-the-art performance. Our code is publicly available on GitHub.
|
| 1469 |
Predicting Delayed Train Trajectories on the Dutch Railway Network: Explainable AI Evaluation of Topological, Operational and Weather Features with Tree Based Ensemble Methods
2609.34692
|
cs.LG
|
Jia Long Bao, Ali Mohammed Mansoor Alsahag, Seyed Sahand Mohammadi Ziabari |
The reliable prediction of passenger train delays is a critical component of railway management. While contemporary research frequently attempts to maximize absolute accuracy by deploying opaque deep learning architectures, the underlying data mechanics drivin...The reliable prediction of passenger train delays is a critical component of railway management. While contemporary research frequently attempts to maximize absolute accuracy by deploying opaque deep learning architectures, the underlying data mechanics driving longitudinal predictive decay remain underexplored. Consequently, this study provides an explainable temporal robustness analysis of network-wide railway delay prediction. Focusing on the Dutch railway network, this research utilizes interpretable tree-based ensembles to integrate granular topological, environmental, and operational features. The overarching finding establishes that while feature-rich tree-based models improve simultaneous (within-month) prediction, predictive performance systematically degrades when evaluated across non-simultaneous (future) months. Furthermore, multi-horizon SHAP and dispersion analyses explicitly link this degradation to environmental feature volatility and instability within the statistical target definition. Ultimately, this study demonstrates that richer feature sets alone are insufficient to resolve long-term forecasting constraints, underscoring the necessity to transition toward dynamic, season-aware architectures anchored by absolute operational boundaries.
|
| 1470 |
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
2609.35297
|
cs.LG
|
Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horv\'ath |
Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in ...Pretraining a language model takes enormous compute, and the right optimizer can save a good part of it. Muon's spectral step gives a stronger direction than a sign step, but it is expensive. Every step runs Newton-Schulz iterations on the full matrix and, in distributed training, an extra all-reduce. Sign steps, as in Lion and Signum, are cheap and stay local to each device. We propose LionMuon, which takes one Muon step every $P$ iterations and Lion steps in between, with a single dual-EMA momentum buffer shared by both. Muon's compute and communication are paid once per $P$ steps, and the optimizer state is half of AdamW's. A single-EMA variant, SignMuon, already improves on Muon. We prove complexity bounds under heavy-tailed noise in which the period sets an interpolation between Muon's and Lion's smoothness and noise constants, and which say when LionMuon is faster than both. On 124M and 355M models trained on FineWeb, LionMuon with $P=2$ and $P=5$ reaches a lower loss than Muon, AdamW, Lion and Signum at the same number of tokens. Under 4-GPU data-parallel training it reaches Muon's final loss with a third less wall-clock on PCIe, and it beats the communication-efficient Muon variants Dion and MuonBP on loss at no more exposed communication, while keeping the exact gradient. Code: https://github.com/brain-lab-research/lion-muon
|
| 1471 |
Know Thyself, Teach Thyself: Internal Information Flow for Selective Self-Distillation
2609.36695
|
cs.LG
|
Rui Wang, Ruijie Wang, Bo Chen, Jiangxuan Long, Yingyu Liang |
Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced c...Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced changes should be learned. Existing methods typically improve teacher-generated data or select training examples in isolation, leaving the information transferred between these stages unmeasured. We introduce InFlow, a retrieval-guided on-policy self-distillation framework that models this process as potential-to-realized information flow. InFlow first retrieves potentially informative sources using certainty-calibrated hidden-state trajectories, then measures their realized effect through the Jensen--Shannon divergence between the teacher's initial and retrieval-conditioned answer beliefs. Examples with larger belief shifts are selected for on-policy distillation. Our analysis formalizes the information optimized by retrieval and selection and relates the answer-level shift to the teacher--student distillation gap. Across four open-weight language models and three knowledge domains, InFlow achieves the strongest cross-model average among the compared selection methods, with ablations supporting both stages of the framework. Our code is available at https://github.com/1240148048/INFLOW.
|
| 1472 |
Learning Macroscopic Dynamics without Reconstructing Microscopic States
2609.37392
|
cs.LG
|
Zhichao Han, Yue Zhao, Qianxiao Li |
Modeling the temporal evolution of macroscopic properties of complex systems is an important scientific task. To predict this evolution without full microscopic simulation, a common approach encodes microstates into compact latent states, learns their evolutio...Modeling the temporal evolution of macroscopic properties of complex systems is an important scientific task. To predict this evolution without full microscopic simulation, a common approach encodes microstates into compact latent states, learns their evolution, and reads out macroscopic predictions from the latent trajectory. These latent states are often learned through microstate reconstruction. However, with limited latent capacity, reconstruction can favor high-variance microscopic details over information needed for macroscopic prediction. Yet jointly learning latent states and their transition without reconstruction often fails to obtain latent dynamics that support accurate macroscopic prediction. We show that this failure can arise from latent scale collapse: shrinking the latent state scale reduces training loss while macroscopic evolution error remains large. Here, we propose a reconstruction-free framework to learn latent states with their dynamics for prescribed macroscopic prediction. Training alternates between updating the latent representation with the transition and next-state latent targets fixed, and updating the transition with the latent representation fixed. At inference, the trained model predicts macroscopic states recursively from an initial microstate. Our theoretical analysis characterizes reconstruction misalignment and scale collapse under joint training, and gives a sufficient condition for local convergence to correct latent dynamics for our method. Experiments on epidemic spreading on a lattice, mixing of two particle species, and polymer stretching demonstrate that the proposed method achieves substantially better macroscopic prediction over baselines.
|
| 1473 |
Probability Contracts: Accuracy, Coherence, and Decisions Across LLM Interfaces
2609.37470
|
cs.LG
|
Han Chen, Yingrui Li |
Equivalent probability requests can lead to different decisions even when both reports are valid. We introduce probability contracts, a benchmark that connects exact finite-world posteriors, validated event alignment, interface coherence, and failure-aware dec...Equivalent probability requests can lead to different decisions even when both reports are valid. We introduce probability contracts, a benchmark that connects exact finite-world posteriors, validated event alignment, interface coherence, and failure-aware decision evaluation. Across four model-interface configurations on 1,000 worlds, Kev has lower aggregate canonical posterior error than Jev but larger complement and coarsening residuals; accuracy ordering varies by stratum. Jev's Event and Choice interfaces change the binary action on 32.8% of valid pairs at defer cost 0.10. Post-hoc analyses show that disagreement certifies only 11-52% of mean binary pair error and does not consistently outperform confidence for selection. An action-region characterization and a standard scoring-rule identity explain averaging's expected Brier guarantee relative to random interface selection, but not a decision-loss guarantee at each cost. The loss contrast takes both signs on a 99-cost grid for every configuration; small penalties where both policies beat deferral have pointwise intervals containing zero. Secondary checks specified before collection include a separate 400-root cohort, where Event/Choice effects remain configuration-dependent. A joint surface-order and answer-ID intervention shifts posttrained probabilities. Both bounded reasoning arms yield no valid probability reports, leaving their probability accuracy undefined. Probability contracts make these distinctions measurable by evaluating event semantics, posterior error, coverage, and decision cost together.
|
| 1474 |
Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure
2609.37841
|
cs.LG
|
Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki |
Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including...Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length $N$; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed approximate conditional oracle. Under explicit regularity conditions and uniform Hellinger error bounds, for any fixed target accuracy $\varepsilon\in(0,1/8]$ and sufficiently large $N$, our sampler achieves seed-averaged total-variation error at most $\varepsilon$, with total masked-state submissions and sequential depth both bounded by $\widetilde{O}(N^C \varepsilon^{-a})$ for constants $0<C<1$ and $a>0$. These guarantees use polynomial vocabulary size and an edge-response lower bound set by $N$ and $\varepsilon$. The sampler shares evaluations of hypothetical reveals across dependence tests to identify safe parallel batches without requiring full recovery of the hidden forest. A tunable parameter trades probing cost against irreversible commit rounds. In the same class, any admissible irreversible product-commit sampler attaining the same seed-averaged accuracy requires $\Omega(N^c \varepsilon^b)$ counterfactual submissions or commit rounds in the worst case, for constants $c,b>0$.
|
| 1475 |
Distilling Diffusion Score Discrepancy for Efficient Training Data Attribution
2609.38776
|
cs.LGcs.AI
|
Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi |
Training data attribution for diffusion models aims to identify the training samples that influence a generated instance, but existing methods either require costly per-sample gradient computation or query-specific model optimization. Moreover, most methods at...Training data attribution for diffusion models aims to identify the training samples that influence a generated instance, but existing methods either require costly per-sample gradient computation or query-specific model optimization. Moreover, most methods attribute changes in a proxy loss rather than changes in the actual model's generative behavior. We address these limitations by formulating attribution directly with a local score discrepancy measure, which applies to any diffusion variant (including DDPM, EDM, and flow matching), and by showing that such measure can be estimated without retraining, as a preconditioned gradient similarity. We instantiate this estimator as Training-data Influence via score Discrepancy (TID), which uses Kronecker-factored curvature to avoid random projections and per-sample gradient storage. We then distill TID into TIDE, a forward-only student trained online to reproduce the teacher's rankings from the diffusion model's internal activations. Under counterfactual evaluation on CIFAR-10, ArtBench-10, and MS-COCO, TID matches or outperforms state-of-the-art approaches, while TIDE retains most of TID's accuracy at four to five orders of magnitude lower per-query cost, attributing generated samples in milliseconds and faster than the generation itself.
|
| 1476 |
Characterizing High Bandwidth Flash for LLM Serving
2609.39131
|
cs.LG
|
Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang |
Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads comp...Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand accelerator memory capacity for large language model (LLM) serving, but its access costs and limited write endurance complicate its use. We evaluate HBF for high-throughput agentic serving across system design and scheduling choices to understand when additional capacity improves serving performance and energy efficiency. We introduce an HBM-HBF-host hierarchical storage system and buffered cache-aware scheduling, and use trace-driven simulations to analyze their effects on performance, energy consumption, and HBF write lifetime. Across the evaluated workloads, the fastest HBF-augmented systems reduce completion time by 36.1-87.7% relative to HBM-only systems. Modeled energy savings reach 59.1%, with benefits depending on the workload and weight placement. Buffered cache-aware scheduling extends estimated HBF write lifetime from 1.21 to 14.82 years in the evaluated configuration. These results demonstrate the importance of coordinating data placement and scheduling to improve serving efficiency while sustaining a practical HBF write lifetime.
|
| 1477 |
Right Answer, Wrong Mechanism: Detecting Pernicious Divergence in Causal Interventions
2609.39243
|
cs.LG
|
Beiming Liu, Minjie Chen |
Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural di...Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural distribution, and that such divergence is sometimes harmless and sometimes pernicious: it can recruit pathways the model never uses on natural inputs, so that an intervention produces the expected answer through the wrong mechanism. No method currently tells the two cases apart. We make this question testable by planting hidden pathways inside pretrained language models; the pathways are silent on every benchmark prompt by construction, so which interventions depend on them is known exactly. Across 72 configurations and 100,800 interventions on GPT-2 small, we find three things. (i) Nearest-neighbour and local-PCA distances at the intervention site, as used in prior work, score below chance (AUROC 0.35-0.47) at picking out interventions that give the right answer through a planted pathway. (ii) Hidden-Pathway Contribution (HPC), a label-free test that clamps downstream units to the regime of natural runs with the same output and measures how much of the decision disappears, flags pathway-dominated interventions with AUROC >= 0.99 when the pathway shows up as unit-level out-of-regime activity, but fails when every unit stays within its natural range, which we identify as the open problem. (iii) Optimised interventions actively seek hidden pathways: on a gender task, DAS routes 90-95% of its successes through planted pathways for three of four families, and a downstream on-manifold penalty cuts this share to under 5% at a cost of 6-11 points of success rate. In unmodified GPT-2, successful interventions show almost no unit-level out-of-regime reliance.
|
| 1478 |
Wavelet Flow Matching for Time Series
2609.39374
|
cs.LGcs.AI
|
Lucas Poinsignon, Jorge da Silva Goncalves, Samuel Ruiperez-Campillo, Julia E. Vogt |
Synthetic time series are increasingly used for data augmentation, privacy-preserving data sharing, and downstream model development, yet faithfully reproducing both multi-scale temporal structure and cross-channel dependencies remains challenging. We study mu...Synthetic time series are increasingly used for data augmentation, privacy-preserving data sharing, and downstream model development, yet faithfully reproducing both multi-scale temporal structure and cross-channel dependencies remains challenging. We study multivariate time-series generation through flow matching in the wavelet domain. By operating on multilevel discrete wavelet coefficients rather than directly in the time domain, the model represents coarse structure and progressively finer details at separate scales. Their naturally different variances further induce an implicit coarse-to-fine generative process without requiring an explicit multi-scale schedule. Since the transform acts independently on each channel, we pair it with a channel-token transformer whose attention directly models cross-channel dependencies. Across seven benchmark datasets and four sequence lengths, our method is best or tied on a majority of dataset-metric combinations, with the largest and most consistent improvements in Context-FID and discriminative score.
|
| 1479 |
T-ARC: Topology-Aware Randomized Clustering via Distributionally Robust Stochastic Block Models
2609.39466
|
cs.LG
|
Serena Grazia De Benedictis, Andersen Ang, Nicoletta Del Buono, Flavia Esposito, Laura Selicato |
In this work, we introduce a new clustering method, namely T-ARC (Topology-Aware Randomized Clustering), that corrects the geometric bias of K-means by embedding topological information directly into the optimization objective. Building on the assumption that ...In this work, we introduce a new clustering method, namely T-ARC (Topology-Aware Randomized Clustering), that corrects the geometric bias of K-means by embedding topological information directly into the optimization objective. Building on the assumption that the data admits an underlying hidden structure modeled via a latent graph, the idea is to uncover this information through the interplay between the standard K-means data-fidelity term and a graph-cut penalty, which discourages cluster assignments inconsistent with the connectivity structure of the data. To render this coupling tractable, the latent graph is modeled as a random realization from a Stochastic Block Model (SBM), whose scalar parameter is optimized within a Distributionally Robust Optimization (DRO) framework, yielding a closed-form proximal update. Both SBM and DRO are informed by a persistence-based similarity matrix derived from zero-dimensional persistent homology ($H_0$), which translates the multiscale connectivity structure of the data into a pairwise topological prior. The overall optimization proceeds via Block Coordinate Descent; convergence is established through a global Lyapunov functional: the deterministic blocks satisfy monotonic descent, while the stochastic graph update satisfies descent in expectation, so that the expected energy converges. Experiments on synthetic datasets with non-convex geometries and on random subsets of Fashion-MNIST show that T-ARC recovers latent topological structures where K-means fails, achieving the highest accuracy on curved and interleaved clusters while remaining competitive, and markedly more stable than K-means, on real data.
|
| 1480 |
Can Domain Generalization be Guaranteed in Small-Sample Learning?
2609.39512
|
cs.LG
|
Hong Zheng |
The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation and generalization. Structural Risk Minimization (SRM) has long been regarded as a principled solution under t...The small-sample learning problem remains a fundamental challenge in machine learning because limited training data lead to unstable model estimation and generalization. Structural Risk Minimization (SRM) has long been regarded as a principled solution under the classical i.i.d. assumption. However, domain generalization (DG) violates this assumption, leaving the theoretical role of SRM in DG largely unexplored. To bridge this gap, we establish the first theoretical guarantees for SRM in DG under mild assumptions. Specifically, based on the concept of stability, we derive learning consistency and generalization error bounds and prove that these bounds become tight when the hypotheses satisfy the stability condition. Building upon this, under a specific hypothesis space assumption, we establish stability, learning, and generalization bounds for SRM. We further discuss the applicability of these bounds to deep learning. This work establishes theoretical foundations for SRM under distribution shifts and sheds light on the design of robust DG algorithms in small-sample scenarios.
|
| 1481 |
Towards Better Exploration in Sequential Test-Time Scaling
2609.39632
|
cs.LG
|
Joseph Rance, Fabio Pizzati, Juil Sock, Woody Bayliss, Marc G\'orriz Blanch |
Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the mo...Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, models, and benchmarks, we find that 53.8% of sequential scaling trajectories enter an attractor within four iterations. Second, we show that a simple model-mixing intervention helps escape attractors. This reduces the attractor hit rate by 21.2 percentage points on average, expands solution coverage beyond a compute-matched parallel baseline, and improves accuracy of recursive self-aggregation by at least 2.2 percentage points. Our results motivate refocusing long-horizon test-time scaling from parallel methods to sequential methods that improve previous answers.
|
| 1482 |
Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse
2609.39634
|
cs.LGcs.AI
|
Nima H. Siboni |
Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral ...Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation. The ratio is therefore restored exactly, from log-probabilities PPO already computes. Autoregressive generation and canonical-order constructive optimization are both history-injective. The exact correction pays importance-sampling variance that grows with the horizon, so we generalize it to a one-parameter family with PPO ($\alpha{=}0$) and the full correction ($\alpha{=}1$) as endpoints: a single bias--variance knob. A gradient-level analysis of the unclipped surrogate identifies two channels the correction acts through and three conditions under which it carries signal; an enumerable testbed confirms the conditions' predictions. On hard credit-assignment scheduling tasks, a short corrected warmup with aggressive early sample reuse learns faster than PPO and than the same reuse uncorrected; the marginal gain grows with task difficulty ($+0.02$ to $+0.09$ learning-curve AUC), and the early win over PPO tracks the prefix bias that reuse incurs. A correction held throughout, or applied where clipping already contains the reuse bias, is null to harmful.
|
| 1483 |
Nous: Learning and Certifying Memory Decisions Before Source Calibration
2610.00094
|
cs.LGcs.AI
|
Pranav Singh |
Agent memory systems update state decisions from reports whose reliability may be unknown. Existing analyses of source estimation do not determine when a policy can be learned or its improvement certified without identifying the reporting channel. We study the...Agent memory systems update state decisions from reports whose reliability may be unknown. Existing analyses of source estimation do not determine when a policy can be learned or its improvement certified without identifying the reporting channel. We study these three tasks using the same observed records. For a specified hidden Markov family with continuous source uncertainty, decision learning and powered certification have quadratic sample complexity, whereas fixed-precision source estimation has quartic complexity. We characterize a sharp identified interval for policy gain under an unknown shared-background channel and derive finite-sample certificates under bounded history dependence and conditional copying. Independently trained witness regions support general history spaces, and disagreement-conditioned auditing improves power for sparse revisions. For dependent histories, prediction-count-preserving batches cancel the unknown reporting background and admit conditional certificates. A MultiWOZ 2.4 evaluation uses text-processing policies on 1,000 human-written test dialogues with simulated audits. Balanced batches retain 2.26 percentage points of the full candidate's 7.34 percentage-point mean gain and obtain more positive certificates under weak audits. These results establish task-specific information requirements and provide an auditable policy-revision framework for Nous.
|
| 1484 |
Four Ways to Grow a Classifier and Why One of Them Cannot Learn
2610.00180
|
cs.LGcs.AI
|
Cagri Temel |
Constructive classifiers add structure while they train: a level to a tree, a unit to a hidden layer, a split at a leaf. This paper asks what each of four such growth decisions buys, measured under one protocol on 24 datasets, and gives an exact diagnosis and ...Constructive classifiers add structure while they train: a level to a tree, a unit to a hidden layer, a split at a leaf. This paper asks what each of four such growth decisions buys, measured under one protocol on 24 datasets, and gives an exact diagnosis and fix for the one that buys nothing. The diagnosis concerns the natural way to deepen a soft decision tree: turn every leaf into a gate whose two children inherit the parent's class distribution, so the function is unchanged. I prove that this leaves the gradient of every new gate identically zero and with the gate at 1/2, gives the two children identical gradients, so under this construction the added level can never learn. That predicts where it costs: nothing on two-class problems, where two leaves already suffice and a great deal where more classes need more leaves. Measured, the cost is -0.2 points over 13 binary datasets and 40.2 points over 7 multi-class ones, and it tracks the number of classes (Spearman 0.64), not the number of features (0.03). The fix is any perturbation of the children. Neither its size nor its direction matters: a residual-directed initialisation changes accuracy by +0.20 points against noise. The other three decisions each buy one thing. Fitting a new hidden unit to the residual before installing it buys a smaller network but not a more accurate one. Splitting one leaf at a time buys sparsity; it lost accuracy until I found that the split started its two children identical and untrained, the same defect in another place; with the children inheriting the parent and the symmetry broken, per-leaf growth comes within 2.3 points of a complete tree using 23% of its splits. Requiring statistical significance before a node gets a more expressive split buys nothing. Every number comes from the measurement scripts.
|
| 1485 |
The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation
2610.00423
|
cs.LGcs.AI
|
S. Aaron McClendon, Jorge Gallego-Feliciano, Antonios Saravanos |
Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-traj...Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled interventions, we trace their emergence, growth, and consolidation. Sink-carrying channels vary across random seeds but stabilize early within each run. Over longer training, surrounding channels erode and the sink concentrates onto a few redundant carriers. Across ablations, gradient attenuation follows the sink token's collective root-mean-square magnitude rather than any single channel, making collective scale central to understanding their effects. Our central result is that weight decay causally controls the turnover of global activation scale. In controlled continuations, removing decay near the peak allows this scale to keep rising, whereas retaining it produces decline even at constant learning rate. We develop a balance model for the rise and peak of massive-activation magnitude, in which AdamW-preconditioned growth opposes weight decay. Sweeping the decay coefficient $\lambda$ shifts peak timing approximately log-linearly and yields peak magnitudes scaling approximately as $\lambda^{-1/2}$, consistent with this balance. Optimizer measurements further show that preconditioning sustains the large-channel cohort against decay even when raw maintaining forces are too small to do so. Together, these findings connect the observed life cycle to scale-regulating training dynamics and establish weight decay as a training-time lever on activation magnitude.
|
| 1486 |
SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning
2610.00838
|
cs.LGcs.AI
|
Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na |
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individ...Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
|
| 1487 |
Variational Streaming Flow: Probabilistic Forecasting in Physical Time
2610.00976
|
cs.LG
|
Hans Hao-Hsun Hsu, Minseon Gwak, Soon Hoe Lim, Pan Li, N. Benjamin Erichson |
Probabilistic forecasting is important for predicting complex dynamical systems because intrinsic randomness and incomplete observations can cause the same observed state to evolve into multiple plausible futures. While flow matching is a flexible approach for...Probabilistic forecasting is important for predicting complex dynamical systems because intrinsic randomness and incomplete observations can cause the same observed state to evolve into multiple plausible futures. While flow matching is a flexible approach for probabilistic forecasting, it is computationally expensive. Streaming flow (SF) reformulates this approach to model temporal evolution efficiently by learning a continuous velocity field directly in physical time. However, SF learns a deterministic velocity field. Thus, it provides only a single future trajectory for a given fixed initial state and observation history. To overcome this limitation, we introduce Variational Streaming Flow (VSF). Our approach learns a latent distribution that is conditioned on the dynamics of interest. In turn, this enables probabilistic forecasting. Importantly, we retain the computational efficiency of SF by generating in physical time. Across deterministic and stochastic dynamical systems, VSF demonstrates superior predictive accuracy and distributional fidelity. We demonstrate the advantage for both long-horizon rollouts exceeding 1,000 steps, and settings with bifurcating dynamics. Moreover, VSF can be integrated into existing Joint-Embedding Predictive Architecture (JEPA)-based world models as a plug-and-play predictor to improve temporal dynamics and goal-directed success rate in navigation, motion planning, and manipulation.
|
| 1488 |
Kernelized Activation Steering
2610.01062
|
cs.LG
|
Laziz U. Abdullaev, Minh-Hieu Pham, Bach Do, Khoat Than, Tan M. Nguyen |
Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector ...Activation steering provides a simple, training-free mechanism for controlling attributes of generative models such as sentiment, style, and helpfulness. However, standard approaches such as Difference-in-Means apply a single input-independent steering vector across all activations, limiting expressivity and ignoring the local geometry of the activation space. We propose Kernelized Activation Steering (KAS), a unifying framework that lifts activation steering into a reproducing kernel Hilbert space. KAS formulates steering as an optimization problem expressed purely via kernel evaluations, yielding an implicit, activation-dependent steering score without constructing explicit feature maps. Unlike DiM, KAS induces locally adaptive steering: each activation is modified according to its relative position with respect to source and target reference sets, producing a nonlinear steering field over the representation space. Importantly, DiM is recovered as a special case under a linear kernel, while richer kernels enable geometry-aware interventions. Across standard activation steering tasks, including jailbreaking LLMs and image style control, KAS outperforms or is on par with the existing methods.
|
| 1489 |
Port-Hamiltonian Neural Networks for Systems with Multiple Asymptotically Stable Equilibria
2610.01356
|
cs.LG
|
Simon Heilig, Jens P\"uttschneider, Mohammad Itani, Asja Fischer, Timm Faulwasser |
Stable port-Hamiltonian neural networks certify asymptotic stability by construction. Yet, their Hamiltonian is a global Lyapunov function with a single global minimum, so they can represent only dynamic systems with one attractor. We demonstrate that this exc...Stable port-Hamiltonian neural networks certify asymptotic stability by construction. Yet, their Hamiltonian is a global Lyapunov function with a single global minimum, so they can represent only dynamic systems with one attractor. We demonstrate that this excludes even simple systems with energy landscapes forming a double well, and we overcome the restriction by parametrising the Hamiltonian as a product of Bregman divergences generated by one input-convex network. We prove that the resulting model is locally Lyapunov stable, that the coexistence of stable equilibria forces additional non-asymptotically-stable equilibria to exist, that all equilibria lie in a bounded region, and under a hyperbolicity assumption that almost-everywhere stability holds. On three systems our approach is able to recover the energy surface characteristics and improve the convergence speed by 1.8$\times$-8.5$\times$.
|
| 1490 |
Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction
2610.01645
|
cs.LG
|
Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem, Syed Muhammad Ibne Zulfiker, Rishov Paul |
Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-ut...Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity--gender--insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.
|
| 1491 |
pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows
2610.01663
|
cs.LG
|
Tong Chen, Maximilian Holsman, Lin Zhao, Pranam Chatterjee |
Biomolecular therapeutics often start from known sequences and require targeted editing to improve multiple properties while satisfying hard biochemical and manufacturability constraints. However, existing generative methods do not jointly support multi-object...Biomolecular therapeutics often start from known sequences and require targeted editing to improve multiple properties while satisfying hard biochemical and manufacturability constraints. However, existing generative methods do not jointly support multi-objective optimization, hard feasibility, and sequence editing in discrete, variable-length biological spaces. In this work, we introduce Pareto-Constrained Molecule Editing (pCoMole), a framework built on discrete flow matching that steers a pre-trained Edit Flow toward user-specified preferences while enforcing terminal feasibility. pCoMole defines a feasibility-gated terminal distribution using an augmented Tchebycheff utility and realizes the resulting preference tilt through a Doob-h transform of the underlying edit process. To make this construction practical, we approximate the required harmonic function using short Monte Carlo rollouts over candidate edits, yielding an efficient guided editor with provable preference consistency. We validate pCoMole by shrinking GFP while retaining fluorescence-related properties, shortening diverse Cas9 orthologs while preserving PAM specificity, and compressing peptide binders into short peptidomimetics that optimize seven drug-related properties under hard constraints. In wet lab testing, two 229-residue pCoMole-designed eGFP variants retained clear green fluorescence in BL21 cells after 10 deletions, with either one or two substitutions. Together, pCoMole enables constraint-aware, Pareto-aligned editing of biomolecular sequences in discrete, variable-length spaces.
|
| 1492 |
Learning PDE Dynamics between Submanifolds Using Green's Observation Operators
2610.01697
|
cs.LG
|
Jan Tauberschmidt, Jephte Abijuru, Naukshatro Bose, Samuel Okon, Sophie Fellenz |
Many physical systems are driven and observed only on lower-dimensional submanifolds of a larger spatial domain, while their dynamics are governed by the ambient medium occupying that domain. Examples include laser-heated parts imaged by an infrared camera, an...Many physical systems are driven and observed only on lower-dimensional submanifolds of a larger spatial domain, while their dynamics are governed by the ambient medium occupying that domain. Examples include laser-heated parts imaged by an infrared camera, and ground-level emissions measured on a sensor plane. Full-domain solvers, however, compute the entire volume for every new source although only the observation submanifold is needed, and black-box surrogates do not exploit that the ambient medium remains fixed. We introduce the \emph{Green's Observation Operator (GObO)}, which maps the ambient medium once to the Green's kernel of a linear PDE restricted to the source and observation submanifolds. New sources then cost one lower-dimensional integral and no network evaluation. Exponential rates in the kernel yield an exact finite streaming state with horizon-independent memory; we prove its stability and an approximation rate for the restricted heat kernel. On three-dimensional heat conduction and advection--diffusion with collocated and distinct source and observation geometries, GObO trained on static sources predicts responses to moving sources zero-shot with 4--8$\times$ lower error than black-box surrogates, at 1.4\,ms per query after a single conditioning pass. The same kernel transfers across resolutions and admits corrections for mild nonlinearities, including radiative losses and temperature-dependent conductivity, without retraining, at the cost of lower in-distribution accuracy.
|
| 1493 |
Removing spurious minima for planar features by skip connections
2610.01728
|
cs.LGcs.AI
|
Jakob Paul Zimmermann, Andrei Balakin, Moritz Grillo, Georg Loho |
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setti...Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.
|
| 1494 |
Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs
2610.02657
|
cs.LG
|
Wentao Lu, Tianyu Zhu, Jesse Clark |
Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been comp...Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent's weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the frozen-tower model scores 71.60 on HumanEval pass@10 against 6.19 for the in-place model, an 11.6x improvement. At the same budget it also keeps 95% of the parent's GSM8K score and 99% of its MMLU-Pro score. A dense-parent experiment reproduces the HumanEval separation. Within the two-tower design at about 500M tokens, freezing the context tower retains substantially more MMLU-Pro performance than training it, while both give similar observed HumanEval scores. Our theoretical analysis establishes that both conversion classes contain an exact sampler for the AR parent under a hard attention mask and left-to-right commitment of one position per round. Under a shared loss, freezing removes the gradient contribution through the context states. Furthermore, evaluation protocol substantially affects a published 500B-token conversion's scores in both directions across tasks, while its AR parent's scores vary by less than three points, so comparing dLLMs needs a common protocol. These results show that, in the tested low-budget regime, the frozen-tower configuration retains substantially more of the parent's generation performance than in-place conversion.
|
| 1495 |
SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents
2610.03372
|
cs.LG
|
Shangyang Wu, Shuai Zhao, Ziyue Zhu, Jinyang Wu, Anh Tuan Luu |
Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as s...Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
|
| 1496 |
Personal VAD: Speaker-Conditioned Voice Activity Detection
1908.04284
|
cs.LGeess.AS
|
Shaojin Ding, Quan Wang, Shuo-yiin Chang, Li Wan, Ignacio Lopez Moreno |
In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only triggers for the target us...In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only triggers for the target user, which helps reduce the computational cost and battery consumption, especially in scenarios where a keyword detector is unpreferable. We achieve this by training a VAD-alike neural network that is conditioned on the target speaker embedding or the speaker verification score. For each frame, personal VAD outputs the probabilities for three classes: non-speech, target speaker speech, and non-target speaker speech. Under our optimal setup, we are able to train a model with only 130K parameters that outperforms a baseline system where individually trained standard VAD and speaker recognition networks are combined to perform the same task.
|
| 1497 |
A Unifying Perspective on Descent Directions for Constrained Minimization and Switching Criteria
2006.08426
|
cs.LG
|
Swati Gupta, Hassan Mortagy, Sebastian Pokutta |
Descent directions are fundamental to first-order methods for constrained convex optimization. In this work, we provide a unified perspective of these directions as they arise within the projected and conditional gradient methods. We show that the directional ...Descent directions are fundamental to first-order methods for constrained convex optimization. In this work, we provide a unified perspective of these directions as they arise within the projected and conditional gradient methods. We show that the directional derivative of the Euclidean projection of the negative gradient is the steepest feasible descent direction, and induces a natural optimality measure independent of polytope-specific geometric constants. We characterize the properties of the projection of the gradient over the polytope, and show that Frank-Wolfe vertices are essentially the limit of this projection curve. Motivated by these insights, we develop hybrid algorithms that interpolate between projected gradient descent and conditional gradients, using computable switching criteria. The proposed methods achieve linear convergence without relying on the pyramidal width or similar constants. We further show that the number of breakpoints in the projection curve over the simplex and the hypercube is linear in the dimension.
|
| 1498 |
Textual Echo Cancellation
2008.06006
|
cs.LGeess.AS
|
Shaojin Ding, Ye Jia, Ke Hu, Quan Wang |
In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intellige...In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intelligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS playback as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to enhancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC).
|
| 1499 |
Understanding over-squashing and bottlenecks on graphs via curvature
2111.14522
|
cs.LG
|
Jake Topping, Francesco Di Giovanni, Benjamin Paul Chamberlain, Xiaowen Dong, Michael M. Bronstein |
Most graph neural networks (GNNs) use the message passing paradigm, in which node features are propagated on the input graph. Recent works pointed to the distortion of information flowing from distant nodes as a factor limiting the efficiency of message passin...Most graph neural networks (GNNs) use the message passing paradigm, in which node features are propagated on the input graph. Recent works pointed to the distortion of information flowing from distant nodes as a factor limiting the efficiency of message passing for tasks relying on long-distance interactions. This phenomenon, referred to as 'over-squashing', has been heuristically attributed to graph bottlenecks where the number of $k$-hop neighbors grows rapidly with $k$. We provide a precise description of the over-squashing phenomenon in GNNs and analyze how it arises from bottlenecks in the graph. For this purpose, we introduce a new edge-based combinatorial curvature and prove that negatively curved edges are responsible for the over-squashing issue. We also propose and experimentally test a curvature-based graph rewiring method to alleviate the over-squashing.
|
| 1500 |
GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular Packing
2406.07579
|
cs.LGcs.AI
|
Tianyang Xue, Lin Lu, Yang Liu, Mingdong Wu, Hao Dong |
2D irregular packing is a classic combinatorial optimization problem with various applications, such as material utilization and texture atlas generation. Due to its NP-hard nature, conventional numerical approaches typically encounter slow convergence and hig...2D irregular packing is a classic combinatorial optimization problem with various applications, such as material utilization and texture atlas generation. Due to its NP-hard nature, conventional numerical approaches typically encounter slow convergence and high computational costs. Previous research (GFPack) introduced a generative method for gradient-based packing, providing early evidence of its feasibility but faced limitations such as insufficient rotation support, poor boundary adaptability, and high overlap ratios. In this paper, we propose GFPack++, a deeply investigated framework that adopts attention-based geometry and relation encoding, enabling more comprehensive modeling of complex packing relationships. We further design a constrained gradient and a weighting function to enhance both the feasibility of the produced solutions and the learning effectiveness. Experimental results on multiple datasets demonstrate that GFPack++ achieves higher space utilization, supports continuous rotation, generalizes well to arbitrary boundaries, and infers orders of magnitude faster than previous approaches. Codes for this paper are at https://github.com/TimHsue/GFPack-pp.
|
| 1501 |
Answer Set Networks: Casting Answer Set Programming into Deep Learning
2412.14814
|
cs.LGcs.AI
|
Arseny Skryagin, Daniel Ochs, Philipp Deibert, Simon Kohaut, Devendra Singh Dhami |
Although Answer Set Programming (ASP) allows constraining neural-symbolic (NeSy) systems, its employment is hindered by the prohibitive costs of computing stable models and the CPU-bound nature of state-of-the-art solvers. To this end, we propose Answer Set Ne...Although Answer Set Programming (ASP) allows constraining neural-symbolic (NeSy) systems, its employment is hindered by the prohibitive costs of computing stable models and the CPU-bound nature of state-of-the-art solvers. To this end, we propose Answer Set Networks (ASN), a NeSy solver. Based on Graph Neural Networks (GNN), ASNs are a scalable approach to ASP-based Deep Probabilistic Logic Programming (DPPL). Specifically, we show how to translate ASPs into ASNs and demonstrate how ASNs can efficiently solve the encoded problem by leveraging GPU's batching and parallelization capabilities. Our experimental evaluations demonstrate that ASNs outperform state-of-the-art CPU-bound NeSy systems on multiple tasks. Simultaneously, we make the following two contributions based on the strengths of ASNs. Namely, we are the first to show the finetuning of Large Language Models (LLM) with DPPLs, employing ASNs to guide the training with logic. Further, we show the "constitutional navigation" of drones, i.e., encoding public aviation laws in an ASN for routing Unmanned Aerial Vehicles in uncertain environments.
|
| 1502 |
Convergence of Statistical Estimators via Mutual Information Bounds
2412.18539
|
cs.LG
|
El Mahdi Khribch, Pierre Alquier |
Recent advances in statistical learning theory have revealed profound connections between mutual information (MI) bounds, PAC-Bayesian theory, and Bayesian nonparametrics. This work introduces a mutual information bound for statistical models, and derives from...Recent advances in statistical learning theory have revealed profound connections between mutual information (MI) bounds, PAC-Bayesian theory, and Bayesian nonparametrics. This work introduces a mutual information bound for statistical models, and derives from it convergence rates for fractional posteriors, for their variational approximations, and for the maximum likelihood estimator. The observations are assumed independent but not identically distributed, and the model is not assumed well-specified, so that the bounds are oracle inequalities and cover regression with a fixed or a conditioned design; the independent and identically distributed, well-specified case is recovered by dropping an index. We illustrate the method on two applications. In the Gaussian sequence model the rate is minimax in both the radius of the Sobolev ball and the sample size, which only appears through its product with the temperature. In logistic regression, where the model is not conjugate and the variational approximation is computed by a stochastic gradient method, the bound applies to the output of the algorithm rather than to an idealized minimizer, and attains the parametric order in the Renyi risk with no logarithmic factor, the statistical and the optimization error being separated.
|
| 1503 |
Linear Bandits beyond Inner Product Spaces, the case of Bandit Optimal Transport
2502.07397
|
cs.LG
|
Lorenzo Croissant (CREST, FAIRPLAY, ENSAE Paris) |
Linear bandits have long been a central topic in online learning, with applications ranging from recommendation systems to adaptive clinical trials. Their general learnability has been established when the objective is to minimise the inner product between a c...Linear bandits have long been a central topic in online learning, with applications ranging from recommendation systems to adaptive clinical trials. Their general learnability has been established when the objective is to minimise the inner product between a cost parameter and the decision variable. While this is highly general, this reliance on an inner product structure belies the name of \emph{linear} bandits, and fails to account for problems such as Optimal Transport. Using the Kantorovich formulation of Optimal Transport as an example, we show that an inner product structure is \emph{not} necessary to achieve efficient learning in linear bandits. We propose a refinement of the classical OFUL algorithm that operates by embedding the action set into a Hilbertian subspace, where confidence sets can be built via least-squares estimation. Actions are then constrained to this subspace by penalising optimism. The analysis is completed by leveraging convergence results from penalised (entropic) transport to the Kantorovich problem. Up to this approximation term, the resulting algorithm achieves the same trajectorial regret upper bounds as the OFUL algorithm, which we turn into worst-case regret using functional regression techniques. Its regret interpolates between $\tilde{\mathcal O}(\sqrt{T})$ and ${\mathcal O}(T)$, depending on the regularity of the cost function, and recovers the parametric rate $\tilde{\mathcal O}(\sqrt{dT})$ in finite-dimensional settings.
|
| 1504 |
Transformer Based Time-Series Forecasting for Stock
2502.09625
|
cs.LG
|
Shuozhe Li, Zachery B Schulwolf, Risto Miikkulainen |
To the naked eye, stock prices are considered chaotic, dynamic, and unpredictable. Indeed, it is one of the most difficult forecasting tasks that hundreds of millions of retail traders and professional traders around the world try to do every second even befor...To the naked eye, stock prices are considered chaotic, dynamic, and unpredictable. Indeed, it is one of the most difficult forecasting tasks that hundreds of millions of retail traders and professional traders around the world try to do every second even before the market opens. With recent advances in the development of machine learning and the amount of data the market generated over years, applying machine learning techniques such as deep learning neural networks is unavoidable. In this work, we modeled the task as a multivariate forecasting problem, instead of a naive autoregression problem. The multivariate analysis is done using the attention mechanism via applying a mutated version of the Transformer, "Stockformer", which we created.
|
| 1505 |
Backdoor Attacks on Discrete Graph Diffusion Models
2503.06340
|
cs.LG
|
Jiawen Wang, Samin Bin Karim, Yuan Hong, Binghui Wang |
Diffusion models have demonstrated remarkable generative capabilities in continuous data domains such as images and videos. Recently, discrete graph diffusion models (DGDMs) have extended this success to graph generation, achieving state-of-the-art performance...Diffusion models have demonstrated remarkable generative capabilities in continuous data domains such as images and videos. Recently, discrete graph diffusion models (DGDMs) have extended this success to graph generation, achieving state-of-the-art performance. However, deploying DGDMs in safety-critical applications, such as drug discovery, poses significant risks without a thorough understanding of their security vulnerabilities. In this work, we conduct the first study of backdoor attacks on DGDMs, a potent threat that manipulates both the training and generation phases of graph diffusion. We begin by formalizing the threat model and then design a backdoor attack that enables the compromised model to: 1) generate high-quality, benign graphs when the backdoor is not activated, 2) produce effective, stealthy, and persistent backdoored graphs when triggered, and 3) preserve fundamental graph properties (permutation equivariance and exchangeability) even under attack. We validate 1) and 2) empirically, both with and without backdoor defenses, and support 3) through theoretical analysis inspired by prior work.
|
| 1506 |
Deep Fair Learning: Task-Aware Fair Representations via Joint Distance-Covariance Regularization
2504.06470
|
cs.LG
|
Enze Shi, Yiqun Xiao, Linglong Kong, Bei Jiang |
Ensuring fairness is essential as machine learning increasingly informs consequential decisions. However, many fairness-aware methods focus on the outputs of individual predictors, without directly controlling sensitive information retained in the underlying r...Ensuring fairness is essential as machine learning increasingly informs consequential decisions. However, many fairness-aware methods focus on the outputs of individual predictors, without directly controlling sensitive information retained in the underlying representations. We propose Deep Fair Learning (DFL), which combines distance covariance regularization with predictive loss to jointly learn representations and downstream predictors, promoting fairness at both levels while preserving task-relevant information. Its marginal and class-conditional formulations target independence and separation, respectively. Under suitable regularity conditions, we establish non-asymptotic joint excess-risk rates and convergence of the learned representation up to natural invariances. We further derive fairness-inheritance bounds linking representation-level dependence to downstream disparities over suitable predictor classes, extending fairness guarantees beyond the jointly trained predictor. Experiments on tabular, text, and image benchmarks show that DFL achieves lower fairness gaps than competing methods in many evaluated settings while maintaining competitive predictive accuracy, with fairness gains largely preserved after downstream retraining.
|
| 1507 |
Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
2504.21072
|
cs.LGcs.AI
|
Tobias Braun, Jonas Henry Grebe, Patrick Mohr, Marcus Rohrbach, Anna Rohrbach |
The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwa...The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to sexually explicit imagery. To mitigate such risks, prior work has proposed concept erasure methods that aim to sever unwanted concepts from the model via fine-tuning, yet it remains unclear whether these approaches truly remove all links to the harmful concept or merely conceal superficial connections. In this work, we reveal a critical vulnerability, the Erasure Evasion Backdoor (EEB): an adversary binds a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. We show that both black-box and white-box adversaries can instantiate this threat. Across six state-of-the-art erasure methods, including robust ones that explicitly search for alternative representations of the target concept, EEB consistently exposes harmful content: up to 82% success against celebrity-identity unlearning, up to 94% for object erasure, and up to 16 times amplification of explicit-content exposure. While EEB uncovers a blind spot in current erasure methods, it also provides a diagnostic tool for stress-testing future concept erasure techniques. Our code is available at https://github.com/multimodal-ai-lab/EEB.
|
| 1508 |
Graph-Based Floor Separation Using Node Embeddings and Clustering of WiFi Trajectories
2505.08088
|
cs.LGcs.AI
|
Rabia Yasa Kostas, Kahraman Kostas |
Vertical localization, particularly floor separation, remains a major challenge in indoor positioning systems operating in GPS-denied multistory environments. This paper proposes a fully data-driven, graph-based framework for blind floor separation using only ...Vertical localization, particularly floor separation, remains a major challenge in indoor positioning systems operating in GPS-denied multistory environments. This paper proposes a fully data-driven, graph-based framework for blind floor separation using only Wi-Fi fingerprint trajectories, without requiring prior building information or knowledge of the number of floors. In the proposed method, Wi-Fi fingerprints are represented as nodes in a trajectory graph, where edges capture both signal similarity and sequential movement context. Structural node embeddings are learned via Node2Vec, and floor-level partitions are obtained using K-Means clustering with automatic cluster number estimation. The framework is evaluated on multiple publicly available datasets, including a newly released Huawei University Challenge 2021 dataset and a restructured version of the UJIIndoorLoc benchmark. Experimental results demonstrate that the proposed approach effectively captures the intrinsic vertical structure of multistory buildings using only received signal strength data. By eliminating dependence on building-specific metadata, the proposed method provides a scalable and practical solution for vertical localization in indoor environments.
|
| 1509 |
Using Echo-State Networks to Reproduce Rare Events in Chaotic Systems
2505.16208
|
cs.LGcs.AI
|
Anton Erofeev, Balasubramanya T. Nadiga, Ilya Timofeyev |
Echo-State Networks (ESNs) are known to reproduce trajectories and main properties of chaotic attractors. Here, we investigate the question of whether ESNs can reproduce rare events in a system with a stationary distribution with slowly decaying bounded tails....Echo-State Networks (ESNs) are known to reproduce trajectories and main properties of chaotic attractors. Here, we investigate the question of whether ESNs can reproduce rare events in a system with a stationary distribution with slowly decaying bounded tails. In particular, we consider the chaotic four-dimensional competitive Lotka--Volterra system, and quantify extremes using the Generalized Extreme Value distribution. In stationary simulations, the ESN reproduces the stationary distributions and rare event statistics with high accuracy. In addition, the ESN also reproduces rare event statistics for non-equilibrium ensemble simulations for initial conditions relatively close to the underlying attractor. However, the performance of the ESN deteriorates as the ensemble spread increases and trajectories explore regions of phase space not represented in the training data.
|
| 1510 |
Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings
2508.17092
|
cs.LGcs.AI
|
Yahya Badran, Christine Preisach |
Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many KT models rely on knowledge concepts (KCs), which represent the skills required for each item. However, some of these mode...Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many KT models rely on knowledge concepts (KCs), which represent the skills required for each item. However, some of these models are vulnerable to label leakage, a phenomenon in which the input data inadvertently reveal the correct answer, particularly in datasets with multiple KCs per question. We propose a straightforward yet effective solution to prevent label leakage by masking ground-truth labels during input embedding construction whenever such leakage could occur. To accomplish this, we introduce a dedicated MASK label, inspired by masked language modeling (e.g., BERT), to replace ground-truth labels. In addition, we introduce Recency Encoding, which encodes the step-wise distance between the current item and its most recent previous occurrence. This distance is important for modeling learning dynamics such as forgetting, which is a fundamental aspect of human learning, yet it is often overlooked in existing models. Recency Encoding demonstrates improved performance over traditional positional encodings on multiple KT benchmarks. We show that incorporating our embeddings into KT models such as DKT, DKT+, AKT, and SAKT consistently improves prediction accuracy across multiple benchmarks. The approach is both efficient and widely applicable.
|
| 1511 |
Uniform-in-time convergence bounds for Persistent Contrastive Divergence algorithms
2510.01944
|
cs.LG
|
Paul Felix Valsecchi Oliva, O. Deniz Akyildiz, Andrew Duncan |
We propose a continuous-time formulation of a noisy persistent contrastive divergence (PCD)-like method for maximum likelihood estimation (MLE) of unnormalised densities. Our approach couples parameter updates and sampling of the parametrised density in a mult...We propose a continuous-time formulation of a noisy persistent contrastive divergence (PCD)-like method for maximum likelihood estimation (MLE) of unnormalised densities. Our approach couples parameter updates and sampling of the parametrised density in a multiscale system of stochastic differential equations (SDEs). From this formulation, we derive non-asymptotic bounds for weak test-function errors between the resulting numerical schemes and the MLE point target. The error is decomposed into numerical discretisation, slow-fast averaging, and finite-temperature concentration terms. We also introduce an efficient implementation based on explicit stabilized integrators and establish corresponding long-time error estimates. This leads to a novel method for training energy-based models (EBMs) with quantitative error guarantees.
|
| 1512 |
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
2510.04455
|
cs.LGcs.AI
|
Akira Kitaoka |
Data-driven inverse optimization for mixed-integer linear programs (MILPs), which seeks to learn an objective function and constraints consistent with observed decisions, is important for building accurate mathematical models in a variety of domains, including...Data-driven inverse optimization for mixed-integer linear programs (MILPs), which seeks to learn an objective function and constraints consistent with observed decisions, is important for building accurate mathematical models in a variety of domains, including power systems and scheduling. However, to the best of our knowledge, existing data-driven inverse optimization methods primarily focus on learning objective functions under known constraints, and learning both objective functions and constraints from data for MILPs remains largely unexplored. In this paper, we propose a two-stage approach for a class of inverse optimization problems in which the objective is a linear combination of given feature functions and the constraints are parameterized by unknown functions and thresholds. Our method first learns the constraints and then, conditioned on the learned constraints, estimates the objective-function weights. On the theoretical side, we provide finite-sample guarantees for solving the proposed inverse optimization problem. To this end, we develop statistical learning tools for pseudo-metric spaces under sub-Gaussian assumptions and use them to derive a learning-theoretic framework for inverse optimization with both unknown objectives and constraints. On the experimental side, we demonstrate that our method successfully solves inverse optimization problems on scheduling instances formulated as ILPs with up to 100 decision variables.
|
| 1513 |
Multi-Marginal Schr\"odinger Bridge Matching
2510.16587
|
cs.LG
|
Byoungwoo Park, Juho Lee |
Understanding the continuous evolution of populations from discrete temporal snapshots is a critical research challenge, particularly in fields like developmental biology and systems medicine where longitudinal tracking of individual entities is often impossib...Understanding the continuous evolution of populations from discrete temporal snapshots is a critical research challenge, particularly in fields like developmental biology and systems medicine where longitudinal tracking of individual entities is often impossible. Such trajectory inference is vital for unraveling the mechanisms of dynamic processes. While Schr\"odinger Bridge (SB) offer a potent framework, their traditional application to pairwise time points can be insufficient for systems defined by multiple intermediate snapshots. This paper introduces Multi-Marginal Schr\"odinger Bridge Matching (MSBM), a novel algorithm specifically designed for the multi-marginal SB problem. MSBM extends iterative Markovian fitting (IMF) to effectively handle multiple marginal constraints. This technique ensures robust enforcement of all intermediate marginals while preserving the continuity of the learned global dynamics across the entire trajectory. Empirical validations on synthetic data and real-world single-cell RNA sequencing datasets demonstrate the competitive or superior performance of MSBM in capturing complex trajectories and respecting intermediate distributions, all with notable computational efficiency.
|
| 1514 |
Quantum Information Ordering and Differential Privacy
2511.01467
|
cs.LG
|
Naqueeb Ahmad Warsi, Ayanava Dasgupta, Masahito Hayashi |
We study quantum differential privacy (QDP) by defining a notion of the order of informativeness between two pairs of quantum states. In particular, we show that if the hypothesis testing divergence of the one pair dominates over that of the other pair, then t...We study quantum differential privacy (QDP) by defining a notion of the order of informativeness between two pairs of quantum states. In particular, we show that if the hypothesis testing divergence of the one pair dominates over that of the other pair, then this dominance holds for every $f$-divergence. This approach completely characterizes $(\varepsilon,\delta)$-QDP mechanisms by identifying the most informative $(\varepsilon,\delta)$-DP quantum state pairs. We apply this to study precise limits for privatized hypothesis testing and privatized quantum parameter estimation, including tight upper-bounds on the quantum Fisher information under QDP. Finally, we establish near-optimal contraction bounds for differentially private quantum channels with respect to the Hockey-Stick divergence.
|
| 1515 |
Two Americas of Well-Being: Divergent Rural-Urban Patterns of Life Satisfaction and Happiness from 2.6 B Social Media Posts
2511.10542
|
cs.LG
|
Stefano Maria Iacus, Giuseppe Porro |
Using 2.6 billion geolocated social-media posts (2014-2022) and a fine-tuned generative language model, we construct county-level indicators of life satisfaction and happiness for the United States. We document an apparent rural-urban paradox: rural counties e...Using 2.6 billion geolocated social-media posts (2014-2022) and a fine-tuned generative language model, we construct county-level indicators of life satisfaction and happiness for the United States. We document an apparent rural-urban paradox: rural counties express higher life satisfaction while urban counties exhibit greater happiness. We reconcile this by treating the two as distinct layers of subjective well-being, evaluative vs. hedonic, showing that each maps differently onto place, politics, and time. Republican-leaning areas appear more satisfied in evaluative terms, but partisan gaps in happiness largely flatten outside major metros, indicating context-dependent political effects. Temporal shocks dominate the hedonic layer: happiness falls sharply during 2020-2022, whereas life satisfaction moves more modestly. These patterns are robust across logistic and OLS specifications and align with well-being theory. Interpreted as associations for the population of social-media posts, the results show that large-scale, language-based indicators can resolve conflicting findings about the rural-urban divide by distinguishing the type of well-being expressed, offering a transparent, reproducible complement to traditional surveys.
|
| 1516 |
Improving Forecasts of Suicide Attempts for Patients with Little Data
2511.18199
|
cs.LG
|
Genesis Hang, Annie Chen, Hope Neveux, Matthew K. Nock, Yaniv Yacoby |
Ecological Momentary Assessment (EMA) studies provide real-time data on suicidal thoughts and behaviors, but forecasting suicide attempts remains challenging: attempts are rare, and the pathways patients take to them are heterogeneous. Here, we investigate a c...Ecological Momentary Assessment (EMA) studies provide real-time data on suicidal thoughts and behaviors, but forecasting suicide attempts remains challenging: attempts are rare, and the pathways patients take to them are heterogeneous. Here, we investigate a cohort of patients from an EMA study with recorded suicide-related events. We show that a single model fit to all patients forecasts poorly, while idiographic (per-patient) models show improvement but overfit for those with little data. Based on this result, one may hypothesize that patients should be partitioned into subgroups---this way, similar patients' data can be pooled together to improve forecasts. However, we show that grouping patients at random already improves forecasts, with performance increasing monotonically with the number of groups. Moreover, we show that grouping patients by demographics yields worse forecasts than random groupings. From these results, we hypothesize that patient similarity is continuous, rather than discrete, and must be inferred from the data. This motivated us to use Latent Variable Multiple Output Gaussian Processes (LVMOGPs), adapted to our data. Preliminary results show that, even without careful kernel design, LVMOGPs already match the strongest baseline models on most metrics, and their latent spaces yield a similarity between patients that we can inspect directly. Because the cohort is conditioned on the outcome and the splits are not temporal, we read these results as evidence that idiographic structure exists and can be recovered, not as deployable forecasting performance---an area for future work.
|
| 1517 |
PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic
2511.20586
|
cs.LGcs.AI
|
Koffi Ismael Ouattara, Ioannis Krontiris, Theo Dimitrakos, Dennis Eisermann, Houda Labiod |
Trustworthiness has become a key requirement for deploying artificial intelligence in safety-critical applications, yet conventional metrics such as accuracy fail to capture uncertainty or the reliability of predictions, particularly under adversarial or degra...Trustworthiness has become a key requirement for deploying artificial intelligence in safety-critical applications, yet conventional metrics such as accuracy fail to capture uncertainty or the reliability of predictions, particularly under adversarial or degraded conditions. This paper introduces the Parallel Trust Assessment System (PaTAS), a framework for modeling and propagating trust in neural networks using Subjective Logic (SL). PaTAS operates in parallel with standard neural computation through Trust Nodes and Trust Functions that propagate input, parameter, and activation trust across the network, refining parameter trust during training (Parameter Trust Update) and returning a per-inference trust opinion, quantifying belief, disbelief, and uncertainty, on each prediction (Inference-Path Trust Assessment). Experiments on real-world and adversarial datasets show that these estimates are interpretable and stable and complement accuracy: PaTAS flags predictions that remain accurate yet rely on parameters learned from corrupted data, detects adversarially patched inputs without knowledge of the trigger, and audits training-label corruption without any test data. Compared with output-only uncertainty baselines and dedicated input- and feature-space detectors, out-of-distribution and corruption detection is carried by the conformity-derived input opinions PaTAS consumes, within a regime that a per-dataset diagnostic identifies in advance; trust propagation adds model-side signals these detectors lack, and the composed trust degrades gracefully where any single signal fails. The mechanism extends to convolutional networks. PaTAS thereby provides a foundation for transparent, quantifiable trust reasoning across the AI lifecycle.
|
| 1518 |
Revisiting the Broken Symmetry Phase of Solid Hydrogen: A Neural Network Variational Monte Carlo Study
2512.17703
|
cs.LG
|
Shengdu Chai, Chen Lin, Xinyang Dong, Yuqiang Li, Wanli Ouyang |
The crystal structure of high-pressure solid hydrogen remains a fundamental open problem. Although the research frontier has mostly shifted toward ultra-high pressure phases above 400 GPa, we show that even the broken symmetry phase observed around 130~GPa req...The crystal structure of high-pressure solid hydrogen remains a fundamental open problem. Although the research frontier has mostly shifted toward ultra-high pressure phases above 400 GPa, we show that even the broken symmetry phase observed around 130~GPa requires revisiting due to its intricate coupling of electronic and nuclear degrees of freedom. Here, we develop a first principle quantum Monte Carlo framework based on a deep neural network wave function that treats both electrons and nuclei quantum mechanically within the constant pressure ensemble. Our calculations reveal an unreported ground-state structure candidate for the broken symmetry phase with $Cmcm$ space group symmetry, and we test its stability up to 96 atoms. The predicted structure quantitatively matches the experimental equation of state and gives the closest x-ray diffraction peak-position match among the tested candidates. Furthermore, our group-theoretical analysis provides a symmetry-counting compatibility check between the $Cmcm$ structure and existing Raman and infrared spectroscopic data. Crucially, static density functional theory calculation reveals the $Cmcm$ structure as a dynamically unstable saddle point on the Born-Oppenheimer potential energy surface, demonstrating that a full quantum many-body treatment of the problem is necessary. These results shed new light on the phase diagram of high-pressure hydrogen and call for further experimental verifications.
|
| 1519 |
Critical Points of Degenerate Metrics on Algebraic Varieties: A Tale of Overparametrization
2512.21029
|
cs.LG
|
Giovanni Luca Marchetti, Erin Connelly, Paul Breiding, Kathl\'en Kohn |
We study the critical points over an algebraic variety of an optimization problem defined by a quadratic objective that is degenerate. This scenario arises in machine learning when the dataset size is small with respect to the model, and is typically referred ...We study the critical points over an algebraic variety of an optimization problem defined by a quadratic objective that is degenerate. This scenario arises in machine learning when the dataset size is small with respect to the model, and is typically referred to as overparametrization. Our main result relates the degenerate optimization problem to a nondegenerate one via a projection. In the highly-degenerate regime, we find that a central role is played by the ramification locus of the projection. Additionally, we provide tools for counting the number of critical points over projective varieties, and discuss specific cases arising from deep learning. Our work bridges tools from algebraic geometry with ideas from machine learning, and it extends the line of literature around the Euclidean distance degree to the degenerate setting.
|
| 1520 |
Towards causal effect estimation with learned instrument representations
2602.10370
|
cs.LG
|
Frances Dean, Jenna Fields, Radhika Bhalerao, Marie Charpignon, Ahmed Alaa |
Instrumental variable (IV) methods mitigate bias from unobserved confounding in observational causal inference but rely on the availability of a valid instrument, which can often be difficult or infeasible to identify in practice. In this paper, we propose a r...Instrumental variable (IV) methods mitigate bias from unobserved confounding in observational causal inference but rely on the availability of a valid instrument, which can often be difficult or infeasible to identify in practice. In this paper, we propose a representation learning approach that constructs instrumental representations from observed covariates, which could enable IV-based estimation even in the absence of an explicit instrument. Our model (ZNet) achieves this through an architecture that mirrors the structural causal model of IVs; it decomposes the ambient feature space into confounding and instrumental components, and is trained by enforcing it empirical conditions corresponding to the defining properties of valid instruments (i.e., relevance and exclusion restriction). ZNet is compatible with a wide range of downstream two-stage IV estimators of causal effects. Our experiments demonstrate that ZNet (i) can recover ground-truth instruments when they already exist in the ambient feature space and (ii) constructs candidate latent instruments in the embedding space when no explicit IVs are available. However, instrument representations have inherent challenges which warrant caution and further research. This work explores when ZNet might be used as a module for causal inference in general observational settings.
|
| 1521 |
Greedy Multi-Path Block Verification for Faster Decoding in Speculative Sampling
2602.16961
|
cs.LG
|
Rahul Thomas, David Hidary, Arka Pal |
The goal of $L$-step speculative decoding is to accelerate autoregressive decoding of a target model by using a cheaper draft model to generate a candidate path of $L$ tokens. Based on a verification algorithm involving target and draft model probabilities, a ...The goal of $L$-step speculative decoding is to accelerate autoregressive decoding of a target model by using a cheaper draft model to generate a candidate path of $L$ tokens. Based on a verification algorithm involving target and draft model probabilities, a prefix of the candidate sequence is accepted, and an additional correction token is sampled from a residual distribution to ensure that the final output adheres to the target distribution. While standard speculative decoding uses a verification algorithm which is independent at each token on the path, a recent extension called block verification uses a joint condition involving all sampled on-path probabilities. Block verification (BV) was shown to be optimal over all verification algorithms which use only on-path probabilities, improving on standard speculative decoding. In this work, we first show that block verification is optimal even over verification algorithms that use off-path probabilities, by constructing an information-agnostic linear program (LP). Further, we can extend our LP to the setting where the draft model samples multiple candidate paths, and use it to construct a natural class of multi-path block verification generalizations. While computing the optimal algorithm in this class is not tractable, by considering a stricter class of greedy algorithms, we can formulate an efficient method called greedy multi-path block verification (GBV). Empirically, GBV can improve block efficiency by over 30% and reduce decoding walltimes by over 15% relative to BV. On Llama-3 70B, GBV can improve the end-to-end decoding throughput over SOTA multi-path verification methods by more than 15%.
|
| 1522 |
Understanding Gap-Dependent Regret for Optimism-Based Reinforcement Learning with Linear Function Approximation
2602.20297
|
cs.LG
|
Haochen Zhang, Zhong Zheng, Lingzhou Xue |
We study gap-dependent regret for reinforcement learning with linear function approximation. While prior works have established gap-dependent guarantees in this setting, existing analyses do not apply to algorithms that achieve the nearly minimax-optimal worst...We study gap-dependent regret for reinforcement learning with linear function approximation. While prior works have established gap-dependent guarantees in this setting, existing analyses do not apply to algorithms that achieve the nearly minimax-optimal worst-case regret bound $\tilde{O}(d\sqrt{H^3K})$, where $d$ is the feature dimension, $H$ is the horizon length, and $K$ is the number of episodes. We bridge this gap by establishing the first gap-dependent regret bound for the nearly minimax-optimal algorithm LSVI-UCB++ (He et al., 2023), with an expected regret bound $\tilde{O}(d^2H^3/\Delta_{\min}+d^6H^5)$, improving the dependence on both $d$ and $H$ in the leading gap-dependent term compared with previous results. To understand the exploration cost induced by optimism, we establish a structural lower bound $\Omega(d^2H^3/\Delta_{\min})$ for a broad class of algorithms based on persistent ellipsoidal optimism. When specialized to LSVI-UCB++, this result shows that the leading dependence of our upper bound on $d$, $H$, and $\Delta_{\min}$ is tight up to logarithmic factors. Beyond this algorithmic class, we establish a general gap-dependent lower bound $\Omega(dH^3\log K/\Delta_{\min})$ for arbitrary learning algorithms, showing that the logarithmic dependence on $K$ and the cubic dependence on $H$ are intrinsic to gap-dependent expected regret in linear MDPs. Together, our results substantially narrow the gap between upper and lower bounds and provide a sharper characterization of gap-dependent learning and optimism-based exploration with linear function approximation.
|
| 1523 |
RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms
2603.05026
|
cs.LG
|
Kenan Li, Rongzhi Li, Linghao Zhang, Qirui Jin, Liao Zhu |
Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck. In this work, we introduce RepoLaunch, a novel...Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck. In this work, we introduce RepoLaunch, a novel agentic framework that automatically resolves dependencies, compiles source code, and extracts test results across diverse programming languages and operating systems. RepoLaunch achieves a 78% build success rate, outperforming the Python/Linux-only prior system by 18%. To demonstrate its application, we further present a fully automated pipeline for SWE dataset creation driven by RepoLaunch, which only requires human input at the task-design stage. RepoLaunch is open-sourced, and its automated task-generation pipeline has been adopted by several recent works on agentic benchmarking and training.
|
| 1524 |
The Value of Information in Resource-Constrained Pricing
2603.24974
|
cs.LG
|
Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi |
Firms that price perishable resources -- airline seats, hotel rooms, seasonal inventory -- now routinely use demand predictions, but these predictions vary widely in quality. Under hard capacity constraints, acting on an inaccurate prediction can irreversibly ...Firms that price perishable resources -- airline seats, hotel rooms, seasonal inventory -- now routinely use demand predictions, but these predictions vary widely in quality. Under hard capacity constraints, acting on an inaccurate prediction can irreversibly deplete inventory needed for future periods. We study how prediction uncertainty propagates into dynamic pricing decisions with linear demand, stochastic noise, and finite capacity. A certified demand forecast with known error bound~$\epsilon^0$ specifies where the system should operate: it shifts regret from $O(\sqrt{T})$ to $O(\log T)$ when $\epsilon^0 \lesssim T^{-1/4}$, and we prove this threshold is tight. A misspecified surrogate model -- biased but correlated with true demand -- cannot set prices directly but reduces learning variance by a factor of $(1-\rho^2)$ through control variates. The two mechanisms compose: the forecast determines the regret regime; the surrogate tightens estimation within it. All algorithms rest on a boundary attraction mechanism that stabilizes pricing near degenerate capacity boundaries without requiring non-degeneracy assumptions. Experiments confirm the phase transition threshold, the variance reduction from surrogates, and robustness across problem instances.
|
| 1525 |
Planning for Change: Reinforcement Learning Combined with Bounded Extremum Seeking for Robotic Control under Distribution Shift
2604.01142
|
cs.LG
|
Shaifalee Saxena, Rafael Fierro, Alexander Scheinker |
Reinforcement learning has shown strong performance in robotic manipulation, but learned policies often degrade in performance when test conditions differ from the training distribution. This limitation is especially important for reliable robot planning and c...Reinforcement learning has shown strong performance in robotic manipulation, but learned policies often degrade in performance when test conditions differ from the training distribution. This limitation is especially important for reliable robot planning and control in contact-rich tasks such as pushing and pick-and-place, where changes in goals, contact conditions, or robot dynamics can drive the system out-of-distribution at inference time. In this paper, we investigate a hybrid controller that combines reinforcement learning with bounded extremum seeking (ES) to improve robustness under such conditions. In the proposed approach, deep deterministic policy gradient (DDPG) policies are trained under standard conditions on the robotic pushing and pick-and-place tasks, and are then combined with bounded ES during deployment. The RL policy provides fast manipulation behavior, while bounded ES ensures robustness of the overall controller to time variations when operating conditions depart from those seen during training. The resulting controller is evaluated under several out-of-distribution settings, including time-varying goals and spatially varying friction patches. Crucially, under reasonable assumptions, we prove finite time RL-based convergence of a robotic arm to an object after which bounded ES achieves finite-time convergence to a bounded velocity moving target. We provide analytic formulas for both convergence times, and an ability to tune the convergence times by our choice of gains.
|
| 1526 |
EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild
2604.01554
|
cs.LG
|
Yiming Fan (The Ohio State University), Jun Yeon Won (The Ohio State University), Ding Zhu (The Ohio State University), Melih Sirlanci (The Ohio State University), Mahdi Khalili (The Ohio State University) |
Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance. In the past few decades, numerous models and tools have been developed for this a...Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance. In the past few decades, numerous models and tools have been developed for this application; however, due to the lack of a comprehensive universal benchmark in this field, researchers have struggled to compare different models effectively. Existing datasets are limited in scope, often focusing on a narrow set of transformations or types of binaries, and fail to reflect the full diversity of real-world applications. We introduce EXHIB, a benchmark comprising five realistic datasets collected from the wild, each highlighting a distinct aspect of the BFSD problem space. We evaluate 9 representative models spanning multiple BFSD paradigms on EXHIB and observe performance degradations of up to 30% on firmware and semantic datasets compared to standard settings, revealing substantial generalization gaps. Our results show that robustness to low- and mid-level binary variations does not generalize to high-level semantic differences, underscoring a critical blind spot in current BFSD evaluation practices.
|
| 1527 |
Learning Over-Relaxation Policies for ADMM with Convergence Guarantees
2604.26932
|
cs.LG
|
Junan Lin, Paul J. Goulart, Luca Furieri |
The Alternating Direction Method of Multipliers (ADMM) is a widely used method for structured convex optimization, and its practical performance depends strongly on the choice of penalty and relaxation parameters. Motivated by settings such as Model Predictive...The Alternating Direction Method of Multipliers (ADMM) is a widely used method for structured convex optimization, and its practical performance depends strongly on the choice of penalty and relaxation parameters. Motivated by settings such as Model Predictive Control (MPC), where one repeatedly solves related optimization problems with fixed structure and changing parameter values, we propose learning online updates of the relaxation parameter to improve average performance on problem classes of interest, while guaranteeing that asymptotic convergence is not compromised for the worst-case realization of such problems. This choice is computationally attractive in the Operator Splitting Quadratic Program (OSQP)-like architectures, since adapting relaxation does not trigger the matrix refactorizations associated with penalty updates. We establish convergence guarantees for ADMM with time-varying penalty and relaxation parameters under mild assumptions, and show on benchmark quadratic programs that the resulting learned policies improve both iteration count and wall-clock time on average over baseline OSQP.
|
| 1528 |
On the Influence of the Feature Computation Budget on Per-Instance Algorithm Selection for Black-Box Optimization
2605.04954
|
cs.LG
|
Koen van der Blom, Diederick Vermetten |
Per-instance algorithm selection (PIAS) takes advantage of complementarity between a set of algorithms by deciding which algorithm to run on a given instance. This decision is based on features of the instances, which, in the context of black-box optimization ...Per-instance algorithm selection (PIAS) takes advantage of complementarity between a set of algorithms by deciding which algorithm to run on a given instance. This decision is based on features of the instances, which, in the context of black-box optimization (BBO), require a part of the optimization budget to be computed. This raises two questions: (a) from which fraction of the budget spent on feature computation does PIAS become worth it for BBO, and (b) which fraction of the budget optimizes the tradeoff between feature accuracy and PIAS performance. To this end, we perform a broad study where PIAS with varying sampling budgets for feature computation is compared to the single best algorithm on a broad range of algorithm selection scenarios. These scenarios consist of two portfolio sizes, three problem sets, 4 dimensionalities, and 10 target budgets. We find that PIAS is viable for the majority of tested scenarios, even when as much as a quarter of the total budget is spent on feature computation. The tradeoff for the fraction of the budget spent on feature computation to maximize the benefit of PIAS is highly dependent on the specific AS scenario. Further, on average 20 percent of PIAS loss to the virtual best solver is explained by the budget spent on feature computation, highlighting the importance of properly accounting for the feature budget.
|
| 1529 |
Active Learning for Communication Structure Optimization in LLM-Based Multi-Agent Systems
2605.05703
|
cs.LGcs.AI
|
Huchen Yang, Xinghao Dong, Dan Negrut, Jin-Long Wu |
Optimizing the communication structure of large language model based multi-agent systems (LLM-MAS) has been shown to improve downstream performance and reduce token usage. Existing methods typically rely on randomly sampled training tasks. However, tasks may d...Optimizing the communication structure of large language model based multi-agent systems (LLM-MAS) has been shown to improve downstream performance and reduce token usage. Existing methods typically rely on randomly sampled training tasks. However, tasks may differ substantially in difficulty and domain, and thus they are not equally informative for updating communication structure, making optimization often unstable and highly sensitive to the particular training set. To actively identify the most valuable tasks for communication-structure optimization, we propose an ensemble-based information-theoretic task selection framework. The proposed method estimates task informativeness by how much a candidate task changes the distribution over graph parameters, using ensemble Kalman inversion as an efficient and derivative-free approximation of the corresponding Bayesian update. The resulting estimator is especially suitable for black-box and noisy multi-agent systems. To enhance scalability, we construct a compact candidate pool through embedding-based representative selection and combine the informative selection with surrogate modeling and batch Thompson sampling. We validate the proposed framework across both benign and adversarial settings and multiple task formats. It consistently outperforms random training, demonstrating more effective task selection and greater overall cost efficiency.
|
| 1530 |
Implicit Target Shift in Online Learning: Characterization and Correction
2605.07886
|
cs.LG
|
Ziyan Li, Naoki Hiratani |
Online learning from a stream of data is a defining feature of intelligence, yet modern machine learning systems often struggle in this setting, especially under distributional shift. To understand its basic properties, we study the relationship between online...Online learning from a stream of data is a defining feature of intelligence, yet modern machine learning systems often struggle in this setting, especially under distributional shift. To understand its basic properties, we study the relationship between online and offline learning in the context of kernel regression by deriving a closed-form expression for the function learned by online kernel regression. We reveal that online kernel regression is equivalent to offline regression with shifted, inaccurate target outputs. Conversely, we show that by compensating for this implicit target shift in the teaching signal through target correction, online kernel-based learning can provably learn the same predictor as its offline counterpart. We derive both a closed-form expression for this target correction and an iterative form that can be applied sequentially. Applying this framework to continual image classification tasks on domain-incremental Split CIFAR-10 and CORe50, we show that online stochastic gradient descent with iteratively corrected targets outperforms learning with the true targets. This work therefore provides a basic framework for analyzing and improving online learning in non-stationary environments from the perspective of implicit target shift.
|
| 1531 |
Mutual Information Optimal Density Control of Linear Systems and Generalized Schr\"{o}dinger Bridges with Reference Refinement
2605.09349
|
cs.LG
|
Shoju Enami, Kenji Kashima |
We consider a mutual information (MI) regularized optimal density control of a discrete-time linear system. MI optimal control has been utilized for exploration in reinforcement learning and privacy protection in control. MI regularization induces stochasticit...We consider a mutual information (MI) regularized optimal density control of a discrete-time linear system. MI optimal control has been utilized for exploration in reinforcement learning and privacy protection in control. MI regularization induces stochasticity in the policy, which poses challenges for applications of MI optimal control in safety-critical scenarios. To remedy this situation, we impose Gaussian density constraints at specified times to directly control state uncertainty. For this MI optimal density control problem, we propose an alternating optimization algorithm and investigate its convergence properties. In addition, we reveal a relationship between the MI optimal density control problem and a so-called generalized Schr\"{o}dinger bridge problem associated with the discrete-time linear system. Based on the results of MI optimal density control and this relationship, we also investigate alternating optimization for the Schr\"{o}dinger bridge problem and its convergence properties.
|
| 1532 |
Debiasing Message Passing to Mitigate Popularity Bias in GNN-based Collaborative Filtering
2605.11145
|
cs.LG
|
Md Aminul Islam, Ahmed Sayeed Faruk, Sourav Medya, Elena Zheleva |
Collaborative filtering (CF) models based on graph neural networks (GNNs) achieve strong performance in recommender systems by propagating user-item signals over interaction graphs. However, they are susceptible to popularity bias, since skewed interactions an...Collaborative filtering (CF) models based on graph neural networks (GNNs) achieve strong performance in recommender systems by propagating user-item signals over interaction graphs. However, they are susceptible to popularity bias, since skewed interactions and repeated message passing across high-order neighborhoods amplify the influence of popular items while suppressing long-tail ones. Existing debiasing approaches, including re-weighting objectives, regularization, causal methods, and post-processing, are less effective in GNN-based settings because they do not directly counteract bias propagated through the aggregation process, and recent in-aggregation weighting methods often rely on static heuristics or unstable embedding estimates. We propose Debiasing Popularity Amplification in Aggregation (DPAA), a popularity debiasing framework for GNN-based CF that integrates adaptive, representation-aware interaction weighting and layer-wise weighting directly into message passing. DPAA assigns interaction-level weights from a representation-based popularity signal, stabilized by a smooth transition from pre-trained to evolving model embeddings during training. It further introduces a layer-wise weighting that amplifies higher-order neighborhoods, surfacing long-range interactions with diverse and underexposed items. Experiments on real-world and semi-synthetic datasets show that DPAA outperforms state-of-the-art popularity bias correction methods for GNN-based CF.
|
| 1533 |
An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing
2605.13221
|
cs.LGcs.AI
|
Hanwen Zhang, Dusit Niyato, Wei Zhang, Xin Lou, Malcolm Yoke Hean Low |
In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint operation forms a hybrid scheduling problem, where physical logistics decisions are coupled with computational task scheduli...In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint operation forms a hybrid scheduling problem, where physical logistics decisions are coupled with computational task scheduling. In this paper, UAVs collect finished products from manufacturing stations and transport them back to a central depot. Meanwhile, computational tasks generated by industrial sensor devices at these stations are processed locally, at UAVs, or offloaded via UAVs to the cloud. This coupling makes the problem challenging. A UAV can provide MEC services only during its service window at a station, so routing decisions directly determine when UAV-assisted offloading is available. Routing decisions also affect the UAV energy budget and the availability of onboard computing and communication resources for computational task execution under task deadline constraints. To address this, we propose an agentic-AI-assisted optimization framework with two components. First, we develop an agentic AI that combines large language models, retrieval-augmented generation, and chain-of-thought reasoning to translate user input into an interpretable mathematical formulation for the hybrid scheduling problem. Second, we design a hierarchical deep reinforcement learning approach based on proximal policy optimization (PPO), where the upper layer learns UAV routing and the lower layer optimizes per-slot task execution and resource allocation. Simulation results show that the proposed framework yields more consistent formulations, while the hierarchical PPO achieves full product collection in 99.6% of the last 500 episodes and maintains a 100% deadline satisfaction rate, with more stable performance than the advantage actor-critic approach.
|
| 1534 |
Analog RF Computing: A New Paradigm for Energy-Efficient Edge AI Over MU-MIMO Systems
2605.14331
|
cs.LGcs.AI
|
Wentao Yu, Vincent W. S. Wong |
Modern edge devices increasingly rely on neural networks for intelligent applications. However, conventional digital computing-based edge inference requires substantial memory and energy consumption. In analog radio frequency (RF) computing, a base station (BS...Modern edge devices increasingly rely on neural networks for intelligent applications. However, conventional digital computing-based edge inference requires substantial memory and energy consumption. In analog radio frequency (RF) computing, a base station (BS) encodes the weights of the neural networks and broadcasts the RF waveforms to the clients. Each client reuses its passive mixer to multiply the received weight-encoded waveform with a locally generated input-encoded waveform. This enables wireless receivers to perform the matrix-vector multiplications (MVMs) that account for most of the computation burden in edge inference with ultra-low energy consumption. Unlike conventional downlink transmissions which are optimized for communications, analog RF computing requires a computing-centric physical layer that controls both the analog MVM accuracy and the energy consumption for inference. Motivated by this, in this paper, we propose a physical layer design framework for analog RF computing in MU-MIMO wireless systems. We derive tractable models for computing accuracy and energy consumption for inference, formulate a joint BS beamforming and client-side scaling problem subject to computing accuracy, transmit power, and hardware constraints, and develop a low-complexity algorithm to solve the non-convex problem. The proposed design provides client- and layer-specific accuracy control for both uniform- and mixed-precision inference. Simulations under 3GPP specifications show that analog RF computing can significantly reduce client-side energy consumption by nearly two orders of magnitude compared to digital computing, while mixed-precision inference requires even lower energy consumption than uniform-precision inference. Overall, these results establish analog RF computing over wireless networks as a promising paradigm for energy-efficient edge inference.
|
| 1535 |
CompoSE: Compositional Synthesis and Editing of 3D Shapes via Part-Aware Control
2605.19350
|
cs.LG
|
Habib Slim, Shariq Farooq Bhat, Mohamed Elhoseiny, Yifan Wang, Mike Roberts |
Creating and editing high-quality 3D content remains a central challenge in computer graphics. We address this challenge by introducing CompoSE, a novel method for Compositional Synthesis and Editing of 3D shapes via part-aware control. Our method takes as inp...Creating and editing high-quality 3D content remains a central challenge in computer graphics. We address this challenge by introducing CompoSE, a novel method for Compositional Synthesis and Editing of 3D shapes via part-aware control. Our method takes as input a set of coarse geometric primitives (e.g., bounding boxes) that represent distinct object parts arranged in a particular spatial configuration, and synthesizes as output part-separated 3D objects that support localized granular (i.e., compositional) editing of individual parts. The key insight that enables our method is our use of a diffusion transformer architecture that alternates between processing each part locally and aggregating contextual information across parts globally, and features a novel conditioning technique that ensures strong adherence to the user's input. Importantly, our method learns to infer part semantics and symmetries directly from the user's coarse layout guidance, and does not require part-level text prompts. We demonstrate that our method enables powerful part-level editing capabilities, including context-aware substitution, addition, deletion, and style-preserving resizing operations. We show through extensive experiments that our method significantly outperforms existing approaches on guided synthesis, as measured by objective metrics and LLM-based evaluations.
|
| 1536 |
PAC Learning with Bandit Feedback: Sharp Sample Complexity in the Realizable Setting
2605.25678
|
cs.LG
|
Steve Hanneke, Qinglin Meng, Shay Moran, Amirreza Shaeiri |
We study the problem of multiclass PAC learning with bandit feedback in the realizable setting. In this framework, there is an unknown data distribution over an instance space $\mathcal{X}$ and a label space $\mathcal{Y}$, as in classical multiclass PAC learni...We study the problem of multiclass PAC learning with bandit feedback in the realizable setting. In this framework, there is an unknown data distribution over an instance space $\mathcal{X}$ and a label space $\mathcal{Y}$, as in classical multiclass PAC learning, but the learner does not observe the labels of the i.i.d. training examples. Instead, in each round, it receives an unlabeled instance, predicts its label, and receives bandit feedback indicating only whether the prediction is correct. Despite this restriction, the goal remains the same as in classical PAC learning. We provide a general characterization of the optimal sample complexity of this problem, sharp for every non-trivial concept class, up to logarithmic factors. Our characterization is based on a new combinatorial dimension, termed the bandit $\mathrm{DS}$ dimension, defined via generalized combinatorial structures we call pseudo-boxes. These extend the pseudo-cubes underlying the $\mathrm{DS}$ dimension by allowing a different number of neighbors in each coordinate. In contrast to the $\mathrm{DS}$ dimension, which governs the full-information setting by counting the number of coordinates in the pseudo-cube, the bandit $\mathrm{DS}$ dimension aggregates the number of neighbors across coordinates, leading to a characterization in which the sample complexity scales with the total number of neighbors. We also propose a general learning algorithm achieving the upper bound, based on an algorithmic principle called ListCascade, which connects bandit learning to list learning and may be of independent interest.
|
| 1537 |
Dr-CiK: A Benchmark testing Deep Research for Context-Aided Forecasting
2605.27904
|
cs.LGcs.AI
|
Yihong Tang, Andrew Robert Williams, Arjun Ashok, Vincent Zhihao Zheng, Shubham Gupta |
Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively discovered from noisy, heterogeneous information sources. Yet existing context-aided forecasting benchmarks typ...Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively discovered from noisy, heterogeneous information sources. Yet existing context-aided forecasting benchmarks typically assume that the supporting context is already provided, leaving open whether agents can identify it on their own. Therefore, we introduce Dr-CiK, a benchmark for evaluating whether agents can retrieve forecasting-relevant supporting context from a document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and generate forecasts supported by that evidence. Through context ablations and evaluations of state-of-the-art deep research and forecasting methods paired together, we show that high-quality context substantially improves forecasting performance in Dr-CiK. However, existing DR methods recover limited ground-truth supporting evidence, are frequently misled by distractors, and can cause forecasters to perform worse with retrieved context than without context. Our results motivate research on agents that search for the right context to predict the future.
|
| 1538 |
SWIM: Compact Environment Representation for Reinforcement Learning Human Swimming
2605.31120
|
cs.LGcs.AI
|
Binglun Wang, Do\u{g}a Y{\i}lmaz, Niloy Mitra, Edmond S. L. Ho, He Wang |
Deep reinforcement learning (RL) has driven rapid progress in physically-based motion generation, yet synthesizing robust motion policies in dense fluid environments (e.g. swimming) remains unsolved. Unlike land motion, where environmental impact is sparse and...Deep reinforcement learning (RL) has driven rapid progress in physically-based motion generation, yet synthesizing robust motion policies in dense fluid environments (e.g. swimming) remains unsolved. Unlike land motion, where environmental impact is sparse and can be coarsely modeled (e.g. gravity, normal reaction), swimming requires continuous, full-body coordination under pressure and flow forces across the entire body surface; fully-coupled rigid-fluid simulation is far too slow for the millions of interactions RL requires. We propose SWIM, an RL framework for physically-based human swimming learned from a single reference motion. Its core is a compact, low-dimensional environment representation: a dual-branch graph-convolutional VQ-VAE that tokenizes per-link body-water forces and torques, informative enough for control yet robust to the rapidly changing force exchanges that would otherwise destabilize training. We pair this with a GPU-based Lagrangian solver with rigid-fluid coupling that yields 10-15x faster training. Across hundreds of training runs and several hundred zero-shot conditions, SWIM generalizes to unseen goals and trajectories, fluids with varying density, external flows and perturbations, altered body morphologies, and all four competitive swimming styles. Against imitation-learning baselines (MimicKit-DeepMimic, MimicKit-AMP, ADD) and alternative force models (a PINN world model, MuJoCo's inertial fluid model, and an underwater-robotics simulator), SWIM achieves better stability, goal satisfaction, and physical realism: a 63.9% success rate on held-out generalization versus 36.8% for the strongest baseline.
|
| 1539 |
Characterizing How Complex Agentic AI Systems Handle General Tasks: A Trace-Based Simulation Study
2606.01725
|
cs.LGcs.AI
|
Donghwan Kim, Prakhar Singh, Younghoon Min, Jongryool Kim, Jongse Park |
Agentic AI completes tasks through iterative planning, tool use, and reasoning based on observed outcomes. Despite its popularity, its system-level behavior remains poorly understood, particularly for complex datasets and agent architectures-owing to highly no...Agentic AI completes tasks through iterative planning, tool use, and reasoning based on observed outcomes. Despite its popularity, its system-level behavior remains poorly understood, particularly for complex datasets and agent architectures-owing to highly non-deterministic execution, prohibitive evaluation costs, and limited visibility into proprietary models. This paper presents GAIATrace, the first token-level trace dataset of two state-of-the-art agentic systems (MiroThinker and OWL) running GAIA, a benchmark composed of a heterogeneous mix of general-purpose tasks. Unlike prior trace datasets, GAIATrace captures full reasoning tokens, task-level structures, and activities of every major participating LLMs, enabling in-depth systems research. Complementing the dataset, we present Vidur-Agent, a trace-driven simulator that can replay GAIATrace to perform reproducible, low-cost system evaluation across diverse simulated environments. Using both artifacts, we characterize how modern agentic systems handle general tasks and how various system design choices shape their behavior, yielding several unique findings.
|
| 1540 |
All you need to break LLMs are Black-Box, Adapting, Efficient, Transferable, Harmful, Applicable ... Attacks
2606.03647
|
cs.LGcs.AI
|
Vincent Limbach, Jonas Dornbusch, David L\"udke, Stephan G\"unnemann, Leo Schwinn |
Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense comparison unreliable. Historically, standardized attacks such as AutoAttack have l...Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment risk assessment and defense comparison unreliable. Historically, standardized attacks such as AutoAttack have largely resolved this for image classifiers, providing a reliable evaluation baseline for systematic comparison across defenses. However, no equivalent exists for LLM jailbreak evaluation yet, where designing such an attack is considerably more difficult. A reliable attack must, among other things, be black-box compatible, applicable to arbitrary defense pipelines, and efficient, which no existing method jointly satisfies. We introduce Indirect Harmfulness Optimization (IHO), a masked diffusion language model attacker trained via iterative preference optimization against a harmfulness judge, requiring only black-box access to the target. The same method can be used without modification as a strong adapting attack on individual behaviors, or as an efficient amortized policy that transfers to held-out behaviors and unseen target models without fine-tuning. Even against layered defenses, such as GLM 5.3-flash served through a black-box API deploying an auxiliary detector, IHO improves attack success considerably over state-of-the-art approaches, without any defense-specific adaptation. Our results position IHO as a practical step toward the kind of standardized jailbreak evaluation that has improved reliability in the past. Code and models are available on GitHub and Hugging Face.
|
| 1541 |
Beyond Objective Equivalence: Constraint Injection for LLM-Based Optimization Modeling on Vehicle Routing Problems
2606.04816
|
cs.LGcs.AI
|
Xizi Luo, Changhong He, Dongdong Geng, Chenggong Shi, Yu Mei |
Large language models (LLMs) can generate executable solver code from natural-language descriptions of optimization problems. However, existing verification signals focus on whether the generated formulation produces the correct objective value. Such signals c...Large language models (LLMs) can generate executable solver code from natural-language descriptions of optimization problems. However, existing verification signals focus on whether the generated formulation produces the correct objective value. Such signals can overlook constraint-level errors: incorrect formulations may still produce the same optimum when extra or missing constraints do not affect the tested instance. We propose constraint injection, a verification method that directly tests whether the generated code implements the intended constraints. We use feasible solutions to detect incorrectly added constraints and one-constraint-violating solutions to detect omissions. Together with the optimal value check, these tests verify both the objective and the constraint set. We evaluate this approach on vehicle routing problems (VRPs), which provide a challenging testbed with diverse, tightly coupled operational constraints. We develop VRPCoder, an 8B model for generating Gurobi code from natural-language VRP descriptions, and VRPBench, an expert-verified benchmark spanning 21 VRP variants across four subsets. The verifier serves as a rejection-sampling criterion for synthetic data construction and as a rollout-level reward for reinforcement learning (RL). On VRPBench, VRPCoder-RL reaches 93% overall Pass@1, outperforms Gemini-3.1-Pro Preview on three subsets, exceeds Claude-Sonnet-4.5 by 28 percentage points, and exceeds the strongest prior OR-LLM by 78 percentage points.
|
| 1542 |
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
2606.11063
|
cs.LGcs.AI
|
Joachim Schaeffer, Alexander Panfilov, Thomas Jiralerspong, Guillaume Lajoie, Jonas Geiping |
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the acting model detects such an intervention, it ma...AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the acting model detects such an intervention, it may infer properties of the monitor and adapt to evade the control protocol. We introduce CIAware-Bench, a benchmark for measuring control intervention (CI) awareness across frontier models. CIAware-Bench tests whether models can distinguish their own trajectories from those modified by a control intervention. The benchmark comprises four task domains (essay writing, BigCodeBench, Bash Arena, and SHADE-Arena), with options to vary trajectory watermarking, side-task presence, and the control protocol. Evaluating thirteen acting models with eight intervening models shows substantial variation between settings and model pairs. CI awareness rises sharply for GPT-6 Astra and the Claude 5 models (Fable 5 and Opus 5). When explicitly probed, Astra reaches mean AUROC of 0.90 on Essay, 0.91 on BigCodeBench, 0.86 on Bash Arena and 0.77 on SHADE-Arena. Fable 5 and Opus 5 both reach 0.77 on Essay, with less consistent gains in the other environments. On SHADE-Arena, we demonstrate that telling an acting model that an action was replaced and asking it to adapt leads to moderate improvements in monitor evasion rates. In summary, our results highlight that control evaluations should already assume perfect CI awareness for conservative safety estimates, and that protocol design should explore countermeasures that make interventions harder to detect.
|
| 1543 |
Modelling magnetic material properties with uncertainty-aware neural networks
2606.11870
|
cs.LG
|
Clemens Wager, Heisam Moustafa, Alexander Kovacs, Qais Ali, Harald Oezelt |
Machine learning is increasingly used to accelerate materials discovery across large compositional and structural design spaces, but limited and heterogeneous data make reliable uncertainty estimation essential. In this work, we investigate uncertainty quantif...Machine learning is increasingly used to accelerate materials discovery across large compositional and structural design spaces, but limited and heterogeneous data make reliable uncertainty estimation essential. In this work, we investigate uncertainty quantification for permanent magnet modeling across three complementary studies. First, on a public Curie temperature surrogate dataset, we benchmark Gaussian process regression, random forest bagging, and dropout-based Bayesian neural networks and show that calibration, sharpness, and confidence curve diagnostics reveal differences in uncertainty quality that are not visible from point-prediction metrics alone. Second, we transfer this framework to the prediction of intrinsic magnetic properties in Nd2 Fe14 B-based magnets. Third, we extend the same uncertainty-aware perspective to coercivity prediction from microstructural information using a graph neural network. Together, these studies show that uncertainty quantification improves the trustworthiness of magnetic material property predictions and can be transferred from composition-based surrogate models to more complex structure-sensitive learning tasks.
|
| 1544 |
Quantum Statistical Memory Advantage Reveals Predictive Structure in Chaotic Invariant Measures
2606.13422
|
cs.LG
|
Maida Wang, Xiao Xue, Minh Chung, Peter V. Coveney |
Long-horizon prediction of a chaotic system is governed by fidelity to its invariant measure, and which parts of that measure a learned predictor needs is a physical question not known in advance. We establish a quantum statistical memory advantage for interro...Long-horizon prediction of a chaotic system is governed by fidelity to its invariant measure, and which parts of that measure a learned predictor needs is a physical question not known in advance. We establish a quantum statistical memory advantage for interrogating it. A quantum statistical prior (Q-Prior), trained once on classical data, stores an efficiently preparable $k$-point marginal in polynomially many circuit parameters against exponentially many for explicit tabulation. Collective Bell measurements on two copies then estimate any post hoc Pauli-expectation magnitude at a copy cost independent of register size, with a worst-case exponential separation from single-copy protocols. A single Bell dataset supports an entire family of candidate observables at a cost growing only logarithmically in the family size, so an exponentially large candidate space stays open for later analysis. We implement the protocol on IQM superconducting processors with two-copy registers of up to 54 physical-qubit chips. We apply this to turbulent channel flow and ERA5 forecasting, resolving invariant structure by statistical sector and order. Turbulent phase correlations persist through eighth order, while most predictive gains arise from low-order constraints; in ERA5, planetary-wave phase coherence complements covariance regularisation, identifying distinct low-order sectors with predictive value. Quantum statistical memory is therefore both a near-term computational resource and an instrument for identifying which invariant structures matter for prediction.
|
| 1545 |
Learning Topological Representations of Protein Structure and Dynamics
2606.14737
|
cs.LG
|
Dominik Geng, Florian Graf, Martin Uray, Roland Kwitt |
Modern protein representation models support tasks such as enzyme design and drug discovery, but their reliance on static data such as sequence and native structure limits their ability to capture the conformational dynamics that drive protein function. We inv...Modern protein representation models support tasks such as enzyme design and drug discovery, but their reliance on static data such as sequence and native structure limits their ability to capture the conformational dynamics that drive protein function. We investigate whether persistent homology (PH) can provide descriptors shared across diverse proteins that retain global structure, fine-grained conformational variability, and kinetically relevant information without large-scale pretraining. We introduce the masked Flood complex, i.e., an adaptation of a recently proposed simplicial complex construction, that incorporates domain knowledge to emphasize inter-residue structure at low computational cost. We then use it to compute PH on molecular dynamics (MD) sampled structures, vectorize the persistence diagrams into a shared coordinate system, and probe the capacity of these representations in terms of the aforementioned aspects. To assess the amount of kinetic information, we learn low-dimensional embeddings from time-lagged observations and evaluate Markov state models (MSMs) estimated from them. Using these MSMs to guide training of the recent marsfm generative framework improves several ensemble statistics relative to the original model. After finetuning on lower-temperature MD data and adapting the sampling procedure, the resulting model also shows promising transfer to fast folding proteins.
|
| 1546 |
Task-Error Residual Learning for Real-Robot Five-Ball Juggling
2606.16978
|
cs.LG
|
Kai Ploeger, Alap Kshirsagar, Jan Peters |
For residual learning that refines existing behavior, sample efficiency depends on two things: how much information each rollout returns, and how efficiently the learner uses that information. Reinforcement learning's standard scalar reward carries far less in...For residual learning that refines existing behavior, sample efficiency depends on two things: how much information each rollout returns, and how efficiently the learner uses that information. Reinforcement learning's standard scalar reward carries far less information than the directional task error that defines the task. Random exploration further discards whatever information each rollout returns. Through residual learning with directional task-error supervision and a task error model that drives sample selection, we achieve stable three-, four-, and five-ball juggling on anthropomorphic Barrett WAM arms. Despite planning and controlling through a simple, idealized stack, the system converges from the second attempt. The first attempt drops, after which task error decreases monotonically without further failures. In comparison, five-ball juggling typically takes humans years of practice. We compare residual learners across two ternary axes, the directional information in the learning feedback and the commitment of the analytic prior, spanning Newton-style Jacobian updates, Composite Bayesian Optimization, and stochastic search methods. Both axes prove necessary for fast convergence: neither directional feedback nor an informative prior alone matches their combination, and the simplest method that combines them, a fixed-Jacobian Newton update, is the fastest and fully reliable. The learned residual tolerates substantial prior misalignment and degraded joint tracking, affecting mainly convergence speed. The bottleneck for residual learning on real robots is therefore the information content of the supervision signal and how the learner uses it, not the accuracy of the surrounding stack. Video documentation of all experiments is available at https://kai-ploeger.com/residual-juggling.
|
| 1547 |
Multi-Level Distributional Entropy from Flow Summary Statistics for Explainable Network Intrusion Detection
2606.29797
|
cs.LGcs.AI
|
Mohamed Aly Bouke, Md Shohel Sayeed, Swee-Huay Heng, Azizol Abdullah, Mohamed Othman |
Machine learning network intrusion detection systems (IDS) operate on aggregate flow statistics that discard the distributional structure of traffic, and although information-theoretic measures capture that structure, established entropy estimators require raw...Machine learning network intrusion detection systems (IDS) operate on aggregate flow statistics that discard the distributional structure of traffic, and although information-theoretic measures capture that structure, established entropy estimators require raw packet sequences that pre-aggregated flow datasets do not contain. No prior method derives entropy from the summary statistics those records already hold. We introduce Multi-Level Distributional Entropy (MDE), which computes interpretable information-theoretic features analytically from flow-level summary statistics at three levels, within-flow Gaussian differential entropy, cross-directional Jensen-Shannon divergence (JSD), and Transmission Control Protocol (TCP) flag-incidence Shannon entropy, with closed-form properties and no raw packet access; only imputation medians and score bounds are fitted on the training split. We pair the features with a leakage-free, fold-local evaluation protocol that reports the full operational metric suite across cross-validation, temporal, pseudo-live, cross-dataset, and unseen-attack-family settings, on four benchmarks (NSL-KDD, CICIDS-2017, CICIDS-2018, UNSW-NB15) with tree-ensemble classifiers and SHAP. The protocol exposes failure modes that aggregate weighted F1 conceals: on CICIDS-2018 an F1 of 0.73 hides a detection rate (DR) of 0.44, on held-out attack families F1 exceeds 0.998 while DR falls to zero, and a 703K-flow pseudo-live replay reveals a threshold-ranking divergence in which score ranking is largely preserved (area under the ROC curve, AUC, 0.84 to 0.86) while fixed-threshold detection collapses (DR 0.08). The entropy features match conventional features within 0.1 percentage points of F1 and receive reproducible SHAP attributions (Spearman 0.84 to 0.94), so their contribution is a grounded, interpretable representation and an evaluation methodology rather than an accuracy gain.
|
| 1548 |
Optimal Stabilizer Testing and Learning with Limited Quantum Memory
2607.02444
|
cs.LG
|
Srinivasan Arunachalam, Louis Schatzki |
We study stabilizer state testing and learning with limited coherent quantum memory. Here an algorithm sequentially receives copies of an unknown $n$-qubit state, but may keep only $k$ qubits of coherent quantum memory between measurements. With unrestricted m...We study stabilizer state testing and learning with limited coherent quantum memory. Here an algorithm sequentially receives copies of an unknown $n$-qubit state, but may keep only $k$ qubits of coherent quantum memory between measurements. With unrestricted memory, seminal work of Gross, Nezami and Walter showed how to test $n$-qubit stabilizer states using $6$ copies, which is dimension independent, unlike the learning complexity of $\Theta(n)$. We show that this testing-vs-learning separation is lost under memory constraints. More concretely we show that (1) The sample complexity of testing stabilizer states in the $k$-qubit memory framework is $\Theta(n-k)$. Our upper bound goes via a novel connection to the hidden shift problem and the lower bound is proven using a novel approach to average case bounds on likelihood ratios via combinatorics of the stochastic orthogonal group. (2) The sample complexity of learning stabilizer states with $k$ qubits of memory, in the non-adaptive framework, is $\Theta(n^2/k)$. As a further application of our techniques, we prove an exponential lower bound for purity testing even when the memory may be left coherent throughout the protocol. Our main results identify coherent quantum memory as the resource enabling the usual separation between stabilizer testing and learning. In particular, even with $k=0.99n$ qubits of memory, there is no constant-copy stabilizer tester; furthermore for $k=cn$ qubits of memory (for $0< c < 1$), stabilizer testing is as hard as learning, with both requiring $\Theta(n)$ copies.
|
| 1549 |
Comparison of a Parametric Physics-Informed Neural Network and a Tensorial Reduced-Order Model for the Shallow-Water Dam-Break Problem
2607.27433
|
cs.LG
|
Anton Myshak, Md Rezwan Bin Mizan, Ilya Timofeyev |
We develop two parametric data-driven reduced models: a physics-informed neural network (PINN) and a non-intrusive tensorial reduced-order model (TROM), and apply both approaches to the parametrized one-dimensional shallow-water dam-break problem. Neither redu...We develop two parametric data-driven reduced models: a physics-informed neural network (PINN) and a non-intrusive tensorial reduced-order model (TROM), and apply both approaches to the parametrized one-dimensional shallow-water dam-break problem. Neither reduced model requires time integration: both learn a direct parameter-to-solution map from space, time, and dam-break parameters to the physical state, with the PINN providing predictions at arbitrary times and the TROM reconstructing solutions at the stored snapshot times. In addition, we demonstrate that it is essential to introduce shock-aware collocation to improve the robustness of the PINN model.
|
| 1550 |
Evaluation Resolution Confounds Learning-Rule Comparisons in Model-Brain RSA of Early Visual Cortex
2608.12408
|
cs.LG
|
Nils Leutenegger |
Studies of biologically plausible learning rules (feedback alignment, predictive coding, STDP) train small networks on $32\times32$ images and compare them with brain responses to naturalistic stimuli presented at much higher resolution. We find that a common ...Studies of biologically plausible learning rules (feedback alignment, predictive coding, STDP) train small networks on $32\times32$ images and compare them with brain responses to naturalistic stimuli presented at much higher resolution. We find that a common result in this setting, that untrained or locally trained networks rival or beat backpropagation at early visual cortex, depends on the resolution at which networks are evaluated. In representational similarity analysis (RSA) against human V1, the gap between an untrained and a backprop-trained network grows monotonically from $-0.001$ (95% CI $[-0.011, 0.009]$) at the 32 px training resolution to $+0.030$ $[0.018, 0.042]$ at 224 px (six resolutions, 5 seeds; per-subject RSA on cross-run stimulus pairs). The growth reproduces for three further learning rules, directionally in single-seed macaque electrophysiology, along training, and for an ImageNet ResNet-50 and a Swin-Tiny transformer, which also align best at low resolution although trained at 224 px. Gabor and pixel structure, the normalization state of the untrained baseline, and a global brightness statistic do not account for it. Against a criterion fixed before the analysis, limiting image content to 32 px while the input still grows does not shrink the V1 gap ($+0.035$ $[0.023, 0.046]$ vs. $+0.030$ at 224 px; ratio 1.14 $[1.03, 1.34]$): it arises once the input exceeds the training resolution, through a decline of backprop. A single luminance value per image reaches $\rho = 0.054$ against V1, matching the untrained network (0.053), which bounds what RSA on globally pooled features can resolve at V1 in this dataset. The one learning effect that holds across resolution is backprop above untrained at LOC. Comparisons of learning rules or architectures at early visual cortex need to control, and report, the evaluation resolution.
|
| 1551 |
Difference-of-Convex Regularization for Graph Learning by Differentiable Programming
2608.12757
|
cs.LG
|
Liping Tao, Chee Wei Tan |
Laplacian-regularized minimization is fundamental in signal processing and machine learning, but explicit construction of the dense graph Laplacian pseudoinverse can be computationally expensive. To address this issue, we consider the setting where the graph L...Laplacian-regularized minimization is fundamental in signal processing and machine learning, but explicit construction of the dense graph Laplacian pseudoinverse can be computationally expensive. To address this issue, we consider the setting where the graph Laplacian is given and propose a Difference-of-Convex Regularizer (DCR) graph learning framework that learns a reusable approximation of the pseudoinverse action through regularized Maximum Likelihood Estimation (MLE) and shrinkage-regularized Convex--Concave Procedure (CCCP) iterations. For Laplacian-Regularized Nonnegative Least Squares (LR-NNLS), dual and Karush-Kuhn-Tucker (KKT) analysis motivates separating graph-dependent pseudoinverse learning from instance-specific optimization. The learned pseudoinverse is embedded into a differentiable primal optimization procedure as a graph-aware preconditioner for nonnegative reconstruction and can be reused across LR-NNLS instances sharing the same graph. We establish stability and fixed-point existence guarantees for the proposed pseudoinverse-learning iteration. Experiments across different graph topologies and scales demonstrate near-reference accuracy, competitive time-to-moderate-accuracy behavior, and effective cross-instance reuse across independently generated problem instances sharing the same graph.
|
| 1552 |
Composing Learned Robot Behaviors with Temporal Logic at Runtime
2608.13678
|
cs.LG
|
Moritz Zoellner, Anastasios Manganaris, Ahmed H. Qureshi, Rohan Paleja |
Executing Linear Temporal Logic (LTL) instructions with learned robot policies faces two practical challenges: demonstrations may cover individual behaviors without containing the temporal compositions requested at deployment, and semantic success predicates o...Executing Linear Temporal Logic (LTL) instructions with learned robot policies faces two practical challenges: demonstrations may cover individual behaviors without containing the temporal compositions requested at deployment, and semantic success predicates often provide little useful motor guidance through quantitative robustness. We address these challenges by separating behavior learning from temporal composition. From the same offline demonstrations, we train an unconditioned multimodal diffusion policy and a semantic predictor that estimates which semantic outcomes are likely to follow a proposed action sequence. These predictions provide a learned guidance signal without requiring hand-designed robustness measures for semantic task objectives. At deployment, an LTL automaton tracks instruction progress and scores the policy's proposed actions according to their predicted semantic outcomes. Neither learned component receives the specification during training, allowing demonstrated behaviors to be reused in longer, previously unseen temporal compositions without retraining. For additional runtime safety constraints with informative continuous margins, an optional robustness-based controller locally refines the selected actions. Experiments in a navigation environment, CALVIN manipulation, and on a real robot demonstrate reliable execution of complex temporal instructions and additional safety constraints, with substantial improvements over temporal-logic baselines.
|
| 1553 |
Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning
2608.25100
|
cs.LGcs.AI
|
Xuzhong Wang, Maiqi Jiang, Tejal Nair, Girija Bhusal, Yanfu Zhang |
Large Language Models (LLMs) are powerful but limited by static parametric knowledge that becomes outdated once pretraining ends. Knowledge editing addresses this problem by updating model behavior on target facts without full retraining. In particular, in-con...Large Language Models (LLMs) are powerful but limited by static parametric knowledge that becomes outdated once pretraining ends. Knowledge editing addresses this problem by updating model behavior on target facts without full retraining. In particular, in-context knowledge editing has gained attention because it is training-free and readily applicable to black-box LLMs. Recent reinforcement learning (RL)-based approaches improve over fixed retrieval strategies by adapting prompt construction to the quantity-quality trade-off. Despite initial success, they fail to model the prompt as a structured entity under the distinct and often competing objectives of reliability, generality, and specificity. Previous methods largely optimize a single objective and make decisions over only part of the prompt construction process, thereby overlooking both the balance of different objectives and the global organization of demonstrations. We propose Multi-Objective In-context Knowledge Editing (MO-IKE), a multi-objective RL algorithm that formulates prompt construction for in-context knowledge editing as a Constrained Markov Decision Process. MO-IKE trains a dynamic retriever to optimize competing objectives in knowledge editing, enabling more balanced and globally coherent prompt construction. On Llama-3.2, MO-IKE improves edit success (reliability) from 85.0% to 92.0%, paraphrase consistency (generality) from 77% to 79%, while increasing retention rate (specificity) by 23.0% compared to prior RL-based methods.
|
| 1554 |
Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale
2608.25735
|
cs.LGcs.AI
|
Peichun Hua, Danyang Chen, Junan Zhang, Haifeng Sun, Jingyu Wang |
Hosted retrieval-augmented generation (RAG) and semantic search allow users to query valuable provider-held corpora, raising two competing demands: to hide each query and chosen result, yet reveal only the documents that the user is authorized to receive. Exis...Hosted retrieval-augmented generation (RAG) and semantic search allow users to query valuable provider-held corpora, raising two competing demands: to hide each query and chosen result, yet reveal only the documents that the user is authorized to receive. Existing cryptographic approaches either make this costly by processing the entire corpus for every query, or sacrifice quality for efficiency by scanning a few clusters. We repurpose learned deep hashing as a private filter: a randomized binary code points the provider to a short candidate list, while encrypted reranking and oblivious key transfer protect the precise query and final selection. This shortlist short-circuits full-corpus cryptographic search without sacrificing retrieval quality: with 200-500 candidates, it closely matches full-corpus retrieval across five zero-shot corpora spanning 25K to 5.4M documents. On the full 2.68M-passage NQ corpus over a 10-Gbps link, our protocol only adds 0.73 seconds, or 10 percent, to a 128-token Qwen3-32B RAG pipeline. The released code satisfies directional metric differential privacy (DP) and substantially reduces embedding-inversion and property-inference leakage, demonstrating that a carefully learned shortlist can make private dense retrieval both accurate and practical.
|
| 1555 |
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
2608.27831
|
cs.LGcs.AI
|
Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim, Kyuhong Shim, Sunjae Lee |
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characteriz...Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap, we define a six-category information taxonomy and four dimensions of linguistic style, and apply them to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro. We find that requests carrying only a problem statement, alone or with limited additional context, account for 88% of real prompts but just 7% of benchmark problems. Furthermore, 87% of real prompts are casually written whereas 94% of benchmark problems are formal. Guided by these observations, we introduce RealSWE, 381 multi-variant task families derived from SWE-bench Verified and Pro. Variants within each family share the same underlying task and gold patch while differing only in information composition and linguistic style. Evaluating seven contemporary LLMs with RealSWE, we find that i) realistic inputs reduce resolution rates by 6.4 pp on average and can change model rankings. Controlled analysis further shows that ii) including Desired Behavior and Motivation significantly affects performance, whereas Environment Information and Reproduction Steps merely add tokens without measurable benefit; iii) linguistic style has only small, model-dependent effects. These findings provide actionable guidance for users and agents: explicitly stating the desired behavior and motivation, which most real prompts omit, substantially improves the LLM's software engineering performance.
|
| 1556 |
The PUR-1 Cyber-Physical Digital Twin
2608.30186
|
cs.LG
|
Vasileios Theos, Jonah Lau, Konstantinos Gkouliaras, Zachery Dahm, Konstantinos Vasili |
Digital twin technologies have the potential to improve operational flexibility and responsiveness capabilities of nuclear systems. To provide decision support, cyber event characterization, state estimation, predictive control, and real-time dynamic processin...Digital twin technologies have the potential to improve operational flexibility and responsiveness capabilities of nuclear systems. To provide decision support, cyber event characterization, state estimation, predictive control, and real-time dynamic processing of operational data, however, an efficient digital twin needs to integrate multiple models (data-driven as well as physics-based) with explainability while at the same time maintain two-way synchronization with the physical facility at a time constant less than its operational cycle. In this work, we present the Purdue University Reactor One Digital Twin (PUR-1 DT), a cyber-physical digital twin with a complete high-fidelity physics-based and AI-driven virtual model stack (neutronics, thermal-hydraulics, point kinetics) which provides closed-loop explainable diagnostics, forecasting, predictive control, and action recommendation back to the reactor via two-way communications and a cyber-physical testbed. We demonstrate real-time synchronized state estimation and short-term forecasting over a full reactor operational cycle and conduct a series of benchmarking experiments to validate accuracy and latency. Our results show good agreement with experimental results and lay the groundwork for further development and experimental demonstration of DT-enabled functionalities in real-world facilities.
|
| 1557 |
Capability-Gated Language Models: Security Composes, Utility Does Not
2609.00445
|
cs.LGcs.AI
|
Patrikas Vanagas, Augustas Ma\v{c}ijauskas, Laurynas Lopata |
Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets the same ...Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets the same model configuration. This motivates us to define capability-gated deployment: per-principal access control inside one set of weights, whose configurations form a lattice - meets accumulate a principal's restrictions and joins pool a coalition's reach. We instantiate it by sparse rank gating over an existing nested-factorisation mechanism, guide profile search with one-pass attribution, and read every result once from a pre-registered held-out split. Security approximately composes: provably exactly at meets under a monotone-elicitation assumption we falsify pointwise. In two lineages the median held-out meet deepens suppression; the one effect surviving correction strengthens it. Utility does not: individually harmless profiles can compose to retention and fluency damage, and no compositional bound exists.
|
| 1558 |
Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
2609.01552
|
cs.LGcs.AI
|
Yiming Huang, Ziche Liu, Junxia Cui, Zhuohang Wu, Yiqian Wang |
Scientific law discovery has long been central to scientific progress, proceeding through iterative cycles of generating hypotheses, testing them against empirical evidence, and refining them under scientific constraints. As large language models (LLMs) become...Scientific law discovery has long been central to scientific progress, proceeding through iterative cycles of generating hypotheses, testing them against empirical evidence, and refining them under scientific constraints. As large language models (LLMs) become increasingly involved in scientific research, whether they can discover scientific laws and how to evaluate this ability remain open questions. A central evaluation challenge is to move beyond familiar published equations while keeping discovery tasks grounded in scientific data and constraints. We introduce SciLaws-Bench, a curated collection of scientific task packages grounded in the source literature, each linking a scientific problem, supporting data, published reference equations, and scientific-validity rubrics. Through agent-assisted curation and human verification, we assemble 118 problems spanning six disciplines, drawing on 381 papers, 291 candidate laws, and roughly 8M data points. Each problem supports two complementary evaluation settings. SciLaws-Real uses fixed scientific data to evaluate proposed laws for held-out predictive fit and scientific validity. SciLaws-Parallel evaluates recovery of a newly synthesized structural variant of a published equation through active queries to a simulator calibrated to the source data. Our evaluation reveals three limitations: good predictive fit need not imply scientific validity, recovering a published formula does not establish recovery of its new structural terms, and candidate selection remains a bottleneck in scientific law discovery. Project page: https://yiyihum.github.io/SciLaws-Bench
|
| 1559 |
Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings
2609.03376
|
cs.LG
|
Peichun Hua, Yunming Xiao |
Retrieval-Augmented Generation (RAG) has made dense retrieval over large document collections a standard building block. Organizations increasingly outsource vector indexes to untrusted clouds, exposing proprietary corpora and user queries. Cryptographic prote...Retrieval-Augmented Generation (RAG) has made dense retrieval over large document collections a standard building block. Organizations increasingly outsource vector indexes to untrusted clouds, exposing proprietary corpora and user queries. Cryptographic protection is challenging because each query searches corpus-scale state, causing computation, correlated randomness, and communication to grow with the corpus. At million-document scale, a naive secure implementation takes minutes and about 90 GB of communication per query. Even recent optimized systems require 10--22 seconds. We propose Spruce (Scalable Private Outsourced Retrieval Using Compact Embeddings), which co-designs representations with the cryptographic protocol. Spruce learns compact binary codes that preserve candidates for full-precision reranking, replacing corpus-wide embedding scoring with efficient Hamming-distance computation under two-server multi-party computation (MPC). A corpus-calibrated fixed-radius protocol avoids multi-round candidate selection while preserving retrieval quality. Spruce also provides private cluster pruning, which trades minor quality loss for substantially less computation, and a one-core owner-operated dealer that removes cloud OT preprocessing bottlenecks. Across four corpora containing 383K--5.42M documents, Spruce preserves the original search quality with median candidate sets of only 382--1,952. At 10 Gbps inter-server bandwidth, full scans take 0.21--2.97 seconds, $4.8$--$6.7\times$ faster than the closest measured prior work. Private pruning takes 0.06--1.09 seconds, achieves $13.1$--$22.9\times$ speedups, and retains $93.9\%$--$97.3\%$ of full-float NDCG. On the largest corpus, pruning and the dealer jointly improve sustained throughput by $31.5\times$ at 1 Gbps per link.
|
| 1560 |
Optimal Low-Rank Quantum State Tomography with Bounded-Sample Joint Measurements
2609.10514
|
cs.LG
|
Ashwin Nayak, Xingyu Zhou |
We determine the optimal sample complexity of low-rank quantum state tomography when each measurement may act jointly on at most $t$ samples. For sufficiently small $\varepsilon$, estimating an unknown state on $\mathbb{C}^d$ of rank at most $r$ to trace norm ...We determine the optimal sample complexity of low-rank quantum state tomography when each measurement may act jointly on at most $t$ samples. For sufficiently small $\varepsilon$, estimating an unknown state on $\mathbb{C}^d$ of rank at most $r$ to trace norm error $\varepsilon$ with constant success probability requires, and is achievable with, $$\Theta\left(\frac{dr}{\varepsilon^2}\mathop{\mathrm{max}}\left\{1,\frac{r}{\sqrt{t}}\right\}\right)$$ samples. The lower bound allows the protocol to choose each joint measurement adaptively using all previous classical outcomes; the matching upper bound is nonadaptive. Thus joint measurements on at most $t$ samples improve the complexity of algorithms making single-sample measurements by at most a factor $\sqrt{t}$. Further, measuring order $r^2$ samples jointly is necessary and sufficient to attain the unrestricted collective rate. For the lower bound, we vary the support of a state with fixed uniform spectrum and bound the Fisher information trace of every joint measurement on $t$ samples. The adaptive Fisher chain rule and the van Trees inequality then give the trace norm lower bound. For the upper bound, we construct and analyze a nonadaptive tomography protocol based on a Gaussian joint measurement. An explicit second moment identity and a conditional Gaussian law outside the state's support give a rank-dependent error analysis, yielding the matching rate.
|
| 1561 |
TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
2609.14762
|
cs.LGcs.AI
|
Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary |
Operational logs create a need for private, resource-efficient incident analysis, but aggregate detection scores can conceal severe prediction bias. We present TriCalRAG, a reproducible benchmark for log-anomaly detection with generated root-cause and remediat...Operational logs create a need for private, resource-efficient incident analysis, but aggregate detection scores can conceal severe prediction bias. We present TriCalRAG, a reproducible benchmark for log-anomaly detection with generated root-cause and remediation outputs across BGL, HDFS, Thunderbird, and OpenStack. The primary evaluation compares Qwen2.5-14B and Mistral-Small-22B, served through vLLM on one NVIDIA RTX PRO 6000 GPU (96 GB), under zero-shot, few-shot, and retrieval-augmented generation (RAG) prompting across three data-sampling seeds. We report F1, bootstrap confidence intervals, predicted-positive rates, throughput, and memory use, with DeepLog as a held-out classical baseline. Mistral-Small attains a higher macro-averaged F1 than Qwen2.5-14B (0.644 versus 0.560), whereas Qwen provides approximately twice the throughput. A separate log-probability evaluation compares raw decisions with Contextual Calibration (CC) and Batch Calibration (BC): neither correction consistently improves prediction-balance diagnostics across prompting strategies. Supplementary single-run comparisons extend evaluation to 4-bit Llama-3.1-70B via local Ollama and Claude Haiku 4.5 via Anthropic's cloud API; RAG improves F1 on all four datasets for both models. Claude's reported aggregate F1 increases from 0.566 to 0.695, while the estimated API cost rises from \$0.94 to \$2.15 per 1,000 incidents. Local deployment ablations show approximately 41-fold throughput scaling with batching and 20% lower latency with 4-bit quantization on the tested workload. These findings support retrieval as useful context for anomaly decisions, while the supplementary protocols, unvalidated explanation quality, and prediction-balance diagnostics limit broader claims about RCA accuracy and probabilistic calibration.
|
| 1562 |
Bias-Induced Crossover in Absolute Capacity of Dense Associative Memory
2609.17477
|
cs.LG
|
Yuto Sakurai, Takeaki Shimokawa, Kazunori Iwata, Kazushi Mimura |
The absolute capacity of dense associative memory has mainly been analyzed for unbiased patterns. Here we examine the effect of bias in centered binary patterns under the Krotov-Hopfield single-site criterion $P_{\mathrm{error}}=1/N$, where $P_{\mathrm{error}}...The absolute capacity of dense associative memory has mainly been analyzed for unbiased patterns. Here we examine the effect of bias in centered binary patterns under the Krotov-Hopfield single-site criterion $P_{\mathrm{error}}=1/N$, where $P_{\mathrm{error}}$ is the probability that a single-site flip lowers the energy of a stored pattern and $N$ is the number of neurons. Each pattern component takes $1-q$ with probability $q$ and $-q$ otherwise, where $0<q\le1/2$. For polynomial interactions of order $n$, a signal-to-noise analysis gives an absolute capacity of order $N^{n-1}/\ln N$ at $q=1/2$. For fixed $q<1/2$, however, the capacity is $O(N^{n/2})$ for even $n\ge4$ and $O(N^{(n+1)/2})$ for odd $n\ge5$. For $n=3$, both the unbiased and fixed-bias capacities remain $O(N^2/\ln N)$. For $n\ge4$, these different asymptotic forms imply a nonuniform large-$N$ limit near $q=1/2$. Asymptotic matching predicts a bias-induced crossover in the region $1-2q=O(\ln N/N^{\lfloor n/2\rfloor-1})$. The crossover originates from a bias-dependent crosstalk mean that reduces the stability of sites carrying the more frequent value $-q$. Computer simulations are compared with the finite-size conditioned-Gaussian predictions. An activity-dependent control potential that cancels the conditional crosstalk mean restores the $N^{n-1}/\ln N$ capacity for fixed $0<q<1/2$ within the conditioned-Gaussian approximation.
|
| 1563 |
Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
2609.19963
|
cs.LG
|
Lishang Xu, Guodong Ma, Pengcheng Weng, Zixuan Xia |
Exploration in centralized serial-dictatorship matching bandits must use complete matchings, so learning one player-arm pair can impose regret on others. We study this externality under a known common priority order and Gaussian rewards with unit variance. We ...Exploration in centralized serial-dictatorship matching bandits must use complete matchings, so learning one player-arm pair can impose regret on others. We study this externality under a known common priority order and Gaussian rewards with unit variance. We show that the matching-level Graves-Lai constraints reduce to finitely many pairwise exploration quotas and, at top-choice-separated instances, yield a polynomial-size marginal linear program. At these instances, the exact attainable set of expected logarithmic regret coefficients is $G(\theta)\mathcal{X}(\theta)$, where $\mathcal{X}$ is the feasible matching-allocation set and $G$ maps allocations to player regret. The usual upper-closed Graves-Lai region can be strictly larger despite having the same Pareto-minimal boundary. We further show that identical exploration quotas can induce very different regret through their scheduling. Finally, we construct estimate-solve-track policies, uniformly good on the full row-strict class, that attain every fixed positively weighted optimum without assuming optimizer uniqueness. Every Pareto-minimal point is pointwise attainable, possibly through an instance-calibrated target.
|
| 1564 |
Automated Physics-Informed Neural-Networks-Based Calibration of Highly Segmented Silicon Telescopes
2609.20868
|
cs.LG
|
M. Rejmund, A. Lemasson, P. Morfouace, D. Ramos, J. Taieb |
Transfer and multi-nucleon transfer reactions are essential tools for probing nuclear structure and reaction dynamics, requiring precise determination of the identity, energy, and emission angles of reaction products. The increasing granularity of modern silic...Transfer and multi-nucleon transfer reactions are essential tools for probing nuclear structure and reaction dynamics, requiring precise determination of the identity, energy, and emission angles of reaction products. The increasing granularity of modern silicon telescope arrays enhances experimental capabilities but challenges detector calibration, as conventional channel-by-channel approaches become inefficient and difficult to scale. In this work, we present a fully automated, physics-informed calibration framework based on neural networks, specifically designed for highly segmented silicon detector arrays. The method formulates calibration as a global optimization problem, in which detector gains and geometrical corrections are determined simultaneously by minimizing the width of the reconstructed excitation energy under two-body kinematics constraints. The approach relies exclusively on experimental data and well-established physical principles, without requiring explicit modeling of detector response. A distinctive feature is the use of multiple neural network sub-models sharing a common loss function with embedded physics constraints, enabling coherent and self-consistent calibration across all detector channels. This strategy ensures scalability, robustness, and reproducibility, making it particularly suitable for next-generation detector systems with increasing complexity. The performance of the method is demonstrated using experimental data from the Particle-Identification Silicon-Telescope Array (PISTA) in high-resolution fission studies in inverse kinematics. The results show excellent agreement with theoretical kinematics, high-quality particle identification, and a significant improvement in calibration efficiency. The proposed framework provides a general and adaptable solution for the calibration of complex detector systems in modern nuclear physics experiments.
|
| 1565 |
Graph Learning for Cross-Subject, Cross-Population EEG Emotion Decoding and Model-Derived Spatial-Spectral Neural Signatures
2609.22103
|
cs.LG
|
Dongyi He, Bin Jiang, Xiangkai Wang, Yun Zhao, Hongjie Yan |
Cross subject emotion decoding from electroencephalography EEG requires representations that accommodate individual variability while preserving spatial spectral structure for interpretation. This study introduces EmoDiPyraTrans, a differential graph Transform...Cross subject emotion decoding from electroencephalography EEG requires representations that accommodate individual variability while preserving spatial spectral structure for interpretation. This study introduces EmoDiPyraTrans, a differential graph Transformer that integrates adaptive graph recurrence, differential attention, pyramid fusion and distribution regularization over sequential relative power spectral density graphs. Across SEED, FACED, MAHNOB HCI, DEAP and DREAMER, the model achieved the highest participant mean accuracy and positive class F1 among the evaluated methods, with accuracy and F1 both reaching 0.928 on SEED. On DEP EEG, positive versus neutral accuracy reached 0.802 within healthy controls and 0.704 within participants with depression, compared with 0.591 under healthy to depression transfer and 0.581 with mixed population development. Complementary SEED analyses identified distributed spatial weighting and an alpha centred spectral preference, while configurations averaging six channels retained near full performance. These findings link generalization assessment with model derived candidate signatures to support interpretable EEG emotion decoding, with code available at https://github.com/hdy6438/EmoDiPyraTrans.
|
| 1566 |
Stochastic Flow Map for Count Data
2609.23290
|
cs.LG
|
Ganchao Wei |
High-dimensional count data are common in scientific applications, but most diffusion and flow models are designed for continuous or categorical data, and generation often requires many sequential model evaluations. We propose Count Flow Map, a generative mode...High-dimensional count data are common in scientific applications, but most diffusion and flow models are designed for continuous or categorical data, and generation often requires many sequential model evaluations. We propose Count Flow Map, a generative model that learns finite-time transitions directly in count space for one- or few-step generation. Our model directly learns stochastic transitions over finite time intervals, using Poisson births and Binomial deaths to preserve nonnegative integer counts without a predefined maximum. These transition models are trained to match the underlying local birth--death dynamics and to maintain consistency across step sizes. We characterize the connection between local dynamics and finite-time transition consistency and derive a bound on the generation error. After validating Count Flow Map in several simulations, including a high-dimensional, high-count setting, we apply it to single-cell drug perturbation prediction and neural population forecasting, where it captures perturbation responses and supports forecasts of high-activity events with only one or a few model evaluations. Together, these experiments demonstrate that Count Flow Map enables high-quality generation directly in count space across inference budgets, from one-step to few-step generation, using a single trained model.
|
| 1567 |
PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
2609.24840
|
cs.LG
|
Lei Ye, Haibo Gao, Yitang Li, Peng Xu, Zetong Jing |
Diffusion models provide a flexible framework for motion generation, but turning this flexibility into closed-loop humanoid control remains challenging. Hierarchical generator-tracker systems steer motion through reference trajectories, yet these references ma...Diffusion models provide a flexible framework for motion generation, but turning this flexibility into closed-loop humanoid control remains challenging. Hierarchical generator-tracker systems steer motion through reference trajectories, yet these references may exceed the capabilities of the downstream tracker, leaving physical feasibility and disturbance recovery largely to a separate control module. Action-only diffusion avoids this separation by directly generating executable actions, but provides no explicit future-state trajectory that can be steered toward test-time motion objectives. Joint state-action diffusion offers a natural alternative, but existing controllers often rely on privileged full-body states, while learned behavior selection and test-time motion steering remain only partially integrated. We present PredActor, a predictive action diffusion policy that unifies both steering modes in one directly executed policy using proprioception alone. Given proprioceptive history and optional task context, PredActor jointly predicts actions and an internal future-state trajectory that enables guidance: classifier-free guidance strengthens text-conditioned motion, while classifier guidance steers future states toward test-time objectives. Only actions are executed, requiring neither a motion-reference tracker nor privileged full-body states. In simulation, PredActor reaches 44 of 45 destination targets and achieves a text retrieval score of 0.539 versus 0.424 for conditional action diffusion, with similar disturbance survival. Rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, within the 20 ms control period. Deployed on a Unitree G1, PredActor demonstrates text-conditioned motion, disturbance response, joystick control, and semantic interpolation in simulation and hardware.
|
| 1568 |
SAGE: Semantic Audio Generative Encoder
2609.32755
|
cs.LGcs.SDeess.AS
|
Francesco Brigante, Luca Cerovaz, Davide Marincione, Giorgio Strano, Luca Zhou |
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space...Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.
|
| 1569 |
Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
2609.33875
|
cs.LGcs.AI
|
Yuanhao Li, Hongbo Wang, Xuhong Chen, Yiming Cao, Xunzhu Tang |
Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable execut...Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.
|
| 1570 |
Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges
2609.33878
|
cs.LGcs.AI
|
Donghao Huang, Jinling Pei, Zhaoxia Wang |
Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels con...Merchant matching resolves a noisy payment descriptor to a retrieved merchant entity or returns no match. A key challenge in curating training labels is distinguishing teacher abstention from evidence that no acceptable entity exists: false no-match labels contaminate pseudo-labeled data, while conservative labeling reduces coverage. We investigate whether agreement between two local large language model judges improves pseudo-label reliability. A label is retained only when the judges agree, with separate ordered thresholds for selections and abstentions that guarantee disjoint positive and negative label sets. Retrospective replay on 2,000 expert-annotated queries shows that higher selection thresholds can improve positive-label purity, whereas higher abstention thresholds increase false no-match labels. At thresholds (0.86, 0.80), Muse Glimmer 30B and Gemma 4 31B jointly label 1,633 queries (81.7% coverage) at 96.88% purity; positive and negative purities are 99.47% and 93.38%. This exceeds either constituent model at the same thresholds by more than two percentage points, with lower coverage. A split-half check finds only 0.14 percentage points of threshold-selection optimism. A symmetric threshold of 0.86 adds 40 erroneous no-match labels, while 46 false abstentions persist even with no confidence threshold. Across five matched within-model comparisons, higher reasoning effort yields no clear F0.5 gain and increases median latency by 1.8-5.0 times. These results motivate separate thresholding and auditing for positive and negative pseudo-labels. The study establishes label purity, not student utility; fresh-data curation and student fine-tuning remain necessary to demonstrate downstream value.
|
| 1571 |
ReplayLens: Auditing Agents' Use of Outcomes
2609.34177
|
cs.LGcs.AI
|
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu |
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black...When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored history at a time, holds the remaining interface fixed, and measures the resulting decision. Four interventions target four relationships. Outcome reassignment swaps which scores belong to which actions. Pair transport moves intact action-score pairs to new record slots. Consistent renaming relabels actions in both history and menu. Key-slot reassignment changes both score attachment and position. A constructive separation shows why the audit is needed: two memory writers with identical endpoint accuracy respond differently to the same replay, so conventional evaluation cannot resolve the underlying dependence. On black-box LLM interfaces, swapping scores changes decisions while moving intact pairs does not, separating score attachment from record order. A bounded-memory study exposes ingestion-order sensitivity that endpoint comparison misses. In sequential experiment planning, altered historical scores redirect exploration and reduce final utility despite fresh measurements. A code-debugging agent with sealed hidden tests shows the same pattern outside model selection. ReplayLens provides a relationship-level audit for deciding whether logged experience can be merged, reordered, or reindexed safely.
|
| 1572 |
Into the danger zone: stable extrapolation in high-dimensional function and operator learning
2609.36709
|
cs.LG
|
Ben Adcock, Simone Brugiapaglia, Xuemeng Wang |
Out-of-distribution (OOD) generalization is a central challenge in scientific machine learning. We study regression problems in which the test distribution differs from the training distribution and ask: under what assumptions on the target function or operato...Out-of-distribution (OOD) generalization is a central challenge in scientific machine learning. We study regression problems in which the test distribution differs from the training distribution and ask: under what assumptions on the target function or operator is stable extrapolation possible, and how far beyond the training domain can one extrapolate? Existing theory controls the test error through additive penalties measuring the discrepancy between the training and test distributions. Such guarantees show robustness to small distribution shifts, but can be very pessimistic in comparison to OOD performance observed empirically. We identify classes of holomorphic functions and operators for which the OOD generalization error converges at algebraic rates even in the presence of large distribution shifts. This phenomenon stems from the increasing smoothness of higher-index coordinates, leading to what we term a `blessing of high dimensionality'. For learning with either polynomials, deep neural networks or deep neural operators, we derive explicit rates for arbitrary test measures supported on suitable domains and quantify how the admissible domain depends on the underlying regularity of the function or operator. Our extrapolation guarantees are independent of the test distribution, depending only on its support. We also present a series of numerical experiments across a range of functions and operators that support the main theoretical findings.
|
| 1573 |
Steepest Guidance: A Practical and Principled Approach to Inference-Time Alignment of Flow and Diffusion-based Models
2609.39091
|
cs.LG
|
Shokichi Takakura, Akifumi Wachi, Rei Higuchi, Kohei Miyaguchi, Taiji Suzuki |
Inference-time alignment of flow and diffusion-based models is critical for achieving flexible generative modeling. Theoretically, Doob's $h$-transform provides an elegant solution to this problem, and most existing methods are based on this principle. However...Inference-time alignment of flow and diffusion-based models is critical for achieving flexible generative modeling. Theoretically, Doob's $h$-transform provides an elegant solution to this problem, and most existing methods are based on this principle. However, in practice, estimating the optimal guidance derived from Doob's $h$-transform at inference time is challenging. To deal with this issue, we regard inference-time alignment as a sequential optimization problem in the space of probability measures and propose a novel framework called *Steepest Guidance*, based on the principle of maximizing local improvement in the objective. We provide a theoretical analysis of the proposed method and demonstrate its effectiveness through extensive experiments.
|
| 1574 |
Minimax Additive Regression under Unknown Dependent Designs
2609.39212
|
cs.LG
|
Baptiste Ferrere, Fabrice Gamboa, Jean-Michel Loubes |
We study additive regression under an unknown and potentially non product design distribution, allowing the number of covariates to grow with the sample size. We consider a coupled class that separately controls the smoothness of the marginal densities and of ...We study additive regression under an unknown and potentially non product design distribution, allowing the number of covariates to grow with the sample size. We consider a coupled class that separately controls the smoothness of the marginal densities and of each additive component multiplied by the corresponding marginal density. Under joint-density bounds that hold uniformly in the dimension and suitable dimension-growth conditions, we establish matching minimax bounds for prediction. With known marginal densities, the classical additive rate is attainable. When the marginals are unknown, this rate is preserved if the densities are at least as smooth as the weighted components. Otherwise, marginal-density smoothness determines the minimax rate over the coupled class. Finally, we recover all additive components with total squared error of the same order as the prediction error.
|
| 1575 |
Generalized Geometry Block Proximal Linearized Method for Multiblock Nonconvex and Nonsmooth Optimization
2609.39301
|
cs.LG
|
Weifeng Yang |
This paper considers a class of multiblock nonconvex and nonsmooth optimization problems arising in many applications. Existing methods construct proximal linearized operators or their variants within standard Euclidean geometry to solve this class of problems...This paper considers a class of multiblock nonconvex and nonsmooth optimization problems arising in many applications. Existing methods construct proximal linearized operators or their variants within standard Euclidean geometry to solve this class of problems, forcing their block variable updates to rely on the standard inner product and its induced norm. Nevertheless, this construction fails to capture the geometric structure of the target problem, leading to low numerical efficiency. To overcome these drawbacks, we propose a generalized geometry proximal linearized operator for updating block variables, and develop the Generalized Geometry Block Proximal Linearized (GGBPL) method based on this operator. Compared with existing proximal linearized operators, the proposed operator allows the block surrogate functions to be constructed using arbitrary inner products and general admissible metrics, thereby enabling the GGBPL method to adapt its updates to the geometric structure of various problems. We also introduce the inertial version of GGBPL, named the inertial GGBPL (iGGBPL) method. We further establish a new unified convergence framework under this generalized geometry, within which we prove that our methods guarantee convergence of the objective function values, establish global convergence of the generated sequence to a critical point, and derive the convergence rate of our methods. We also establish an $\mathcal{O}(\varepsilon^{-2})$ iteration complexity bound for obtaining an $\varepsilon$-stationary point. We apply our methods to two nonconvex and nonsmooth problems: sparse nonnegative matrix factorization with $\ell_0$-constraints and sparse nonnegative CP decomposition with $\ell_0$-constraints. Numerical results demonstrate the superior numerical performance of our proposed methods over several state-of-the-art methods.
|
| 1576 |
Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization
2609.39402
|
cs.LGcs.AI
|
Yun Kim, Nojun Kwak |
Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance ...Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.
|
| 1577 |
Safety of Latent Communication in Multi-Agent Systems
2609.39788
|
cs.LGcs.AI
|
Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer |
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to ma...Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety
|
| 1578 |
MiDShip: Multimodal Dataset of Ship Cargo Hold Structures for Engineering Design
2610.02214
|
cs.LG
|
Noah J. Bagazinski, Md. Ferdous Alam, Jaya Manideep Rebbagondla, Faez Ahmed |
Ship structures govern vessel strength, safety, and manufacturability, but their design must satisfy hundreds of classification society requirements, making the process complex and iterative. Data-driven approaches are limited by the lack of structured dataset...Ship structures govern vessel strength, safety, and manufacturability, but their design must satisfy hundreds of classification society requirements, making the process complex and iterative. Data-driven approaches are limited by the lack of structured datasets linking design geometry, structural performance, and rule-based constraints. This paper presents MiDShip, a multimodal dataset of 12,753 synthetic cargo-hold structural designs: 6,020 random, 496 generated by an SGLD-inspired procedure, and 6,237 generated by an equation-informed repair procedure. Each design includes parametric data, full and mesh-ready 3D geometry, engineering drawings and annotations, a bill of materials, and preliminary structural evaluations. Twenty-five constraints derived from a subset of ABS MVR are also evaluated. None of the random designs satisfies all constraints. Among the SGLD-inspired designs, 322 (64.9%) were fully compliant, with an average of 0.409 violations, 82.7% below the seed mean and 96.9% below the random-design mean. The repair procedure, developed through LLM-assisted code analysis, produced 4,952 fully compliant designs (79.4%), averaging 0.296 violations, 97.1% below the paired-source mean. In equal-size comparisons, mean nearest-neighbor distances in the scaled 120-parameter space were 3.495 for repaired designs, 1.144 for SGLD batches, and 3.729 for random designs. The primary contribution is the synchronized dataset and its generation and evaluation infrastructure; the generation studies demonstrate its utility rather than proposing new optimization algorithms. MiDShip supports machine learning, generative design, and automated rule-based evaluation for ship structures.
|
| 1579 |
Generalization Bounds for Flow-matching Generative Models for Intrinsically Low-dimensional Data
2610.02663
|
cs.LGcs.AI
|
Saptarshi Chakraborty, Quentin Berthet, Peter L. Bartlett |
Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to...Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to reflect the intrinsic low-dimensional structure common in real data, such as natural images and molecular geometries. In this work, we study the statistical generalization of flow-matching models for learning an unknown distribution $P_{\mathrm{data}}$ from finitely many samples. We derive finite-sample error bounds on the learned generative distribution, measured in the Wasserstein-$p$ distance, for all $p\geq 1$. Specifically, given $n$ i.i.d. samples from $P_{\mathrm{data}}$, we show that, for every $d>d_p^\ast(P_{\mathrm{data}})$ and appropriately chosen network architectures and hyperparameters, the learned distribution $\widehat{P}^{\mathrm{FM}}$ satisfies $ \mathbb{W}_p(\widehat{P}^{\mathrm{FM}},P_{\mathrm{data}}) \lesssim n^{-1/d}+n^{-1/(2p)}\bigl(\log(1/\xi)\bigr)^{1/(2p)}$ with probability at least $1-\xi$, where $d_p^\ast(P_{\mathrm{data}})$ denotes the Wasserstein-$p$ dimension of the target measure. Our results demonstrate that flow matching naturally adapts to the intrinsic geometry of data and mitigates the curse of dimensionality, as the convergence exponent depends on the intrinsic rather than ambient dimension. These guarantees remain meaningful in high-dimensional regimes and provide a theoretical explanation for the empirical success of flow matching on structured data distributions under substantially more relaxed assumptions than those in existing analyses.
|
| 1580 |
Amortized Structured Stochastic Variational Inference for Gaussian Process Latent Variable Models
2610.03647
|
cs.LG
|
Maksym Tretiakov, Sarah Filippi, Vincent Fortuin, Ruth Misener, Ruby Sedgwick |
Many machine learning methods aim to approximate the lower-dimensional manifold on which the data lives. A desirable feature of such methods is that they should capture the epistemic uncertainty of this learned manifold. One model that achieves this is the Gau...Many machine learning methods aim to approximate the lower-dimensional manifold on which the data lives. A desirable feature of such methods is that they should capture the epistemic uncertainty of this learned manifold. One model that achieves this is the Gaussian Process Latent Variable Model, in which a Gaussian Process (GP) mapping from the latent space provides an estimate of the uncertainty of the manifold. However, the effectiveness of this uncertainty estimation is limited by the mean-field variational approximation between the GP inducing points and the latent variables. In this work, we apply Amortized Structured Stochastic Variational Inference to allow the variational posterior for the latent space to be conditionally dependent on the value of the inducing points. We demonstrate that this more flexible variational posterior improves several metrics relating to the reconstruction of points on the data manifold.
|
| cs.MM 2 papers | ||||
| 1993 |
FATE: Frame-Level Audio-Visual Temporal Embedding
2608.01310
|
cs.MM
|
Kaisi Guan, Bingzi Zhang, Xihua Wang, Ying Ba, Xin Cheng |
When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short...When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding models match semantic but lose temporal information; synchronization models capture temporal offsets but lack semantic understanding. To bridge this gap, we propose FATE, Frame-level Audio-visual Temporal Embedding. Unlike prior embedding models that pool each modality into a single embedding and discard temporal information, FATE retains frame-level sequences, aligns them on the physical timeline, and computes similarity over strictly aligned frame pairs. Unlike synchronization models that output only an offset prediction, FATE encodes synchronization in a reusable embedding space, trained with a joint objective combining cross-video semantic and within-video temporal contrastive learning to capture both what sounds and when it occurs. Across three tasks, FATE surpasses the strongest baseline on temporal and semantic retrieval by a large margin, matches fully supervised methods on event localization in a zero-shot setting, and achieves the best correlation with human judgments as a generation evaluation metric. The source code can be found at \texttt{https://github.com/guankaisi/FATE}.
|
| 1994 |
Emotion Understanding in Streaming Video with Trajectory-Aware Reliability
2608.26786
|
cs.MM
|
Qingsong Wang, Qigong Lei, Zitong Wang, Bohan Yu, Zhiang Dong |
Video emotion understanding is commonly studied as an offline classification problem, where the complete video segment is available before prediction. Real-time interaction, however, requires emotion decisions from incomplete and evolving evidence. This paper ...Video emotion understanding is commonly studied as an offline classification problem, where the complete video segment is available before prediction. Real-time interaction, however, requires emotion decisions from incomplete and evolving evidence. This paper studies streaming video emotion understanding as a reliability-aware decision process over evolving emotion beliefs. In this setting, a single confident prefix prediction can still be unreliable when the underlying belief trajectory is unstable or repeatedly switches across emotion classes. We propose TRACE, a trajectory-aware reliability framework that forms low-latency emotion beliefs from streaming audio prefixes, estimates reliability from confidence, entropy, stability, and class-switching patterns, and selectively invokes contextual belief reinterpretation with visual, textual, and neighboring-utterance evidence. TRACE keeps stable cases in the low-latency online pathway while allocating stronger multimodal reasoning to uncertain cases that remain ambiguous. Experiments on StreamMER, MELD, and MER2024 show that TRACE improves the accuracy-cost trade-off, retaining most full-context gains while reducing unnecessary contextual reasoning.
|
| cs.SD 24 papers | ||||
| 1958 |
Hallucination Reduction for LLM-Based Audio Understanding via Multimodal Direct Preference Optimization
2610.04004
|
cs.SD
|
Bebe Cosgrove, Aaron Isidore Grace, Weiran Wang |
Large Audio Language Models (LALMs) are prone to hallucinating and over-relying on text priors when simultaneously presented with audio and text inputs. To mitigate these hallucinations, we propose utilizing the multimodal Direct Preference Optimization (mDPO)...Large Audio Language Models (LALMs) are prone to hallucinating and over-relying on text priors when simultaneously presented with audio and text inputs. To mitigate these hallucinations, we propose utilizing the multimodal Direct Preference Optimization (mDPO) objective, which forces the model to ground its generation in the acoustic input by contrasting intact and distorted audio counterparts. We extend this preference learning framework to the audio domain by applying a variety of acoustic perturbations. Evaluating the Qwen2-Audio backbone across the DCASE 2025 Challenge and AH Existence datasets, we demonstrate that extending mDPO to LALMs significantly enhances temporal reasoning in the complex DCASE dataset, and improves performance on basic existence verification in the AH benchmark. We identify temporal reversal, frequency masking, and random noise as the most effective perturbations. Ultimately, our approach achieves an absolute accuracy improvement of 14.0% on the DCASE 2025 dataset and 27.4% on AH Existence.
|
| 1959 |
Temporal Anchors and Editing Sensitivity in Partial Speech Spoofing: A Controlled Study
2610.04479
|
cs.SD
|
Xiaosu Su, Yun Cao, Yiping Ni, Xiaowei Yi |
Partial-spoof detectors must reject synthetic content while accepting benign edits. We diagnose temporal anchors and boundary-consistency training using frozen WavLM features, nonlinguistic unit controls, and source-locked evaluation. Genuine-genuine and genui...Partial-spoof detectors must reject synthetic content while accepting benign edits. We diagnose temporal anchors and boundary-consistency training using frozen WavLM features, nonlinguistic unit controls, and source-locked evaluation. Genuine-genuine and genuine-fake splices contrast editing false alarms with synthetic-content misses; they do not isolate a unique causal artifact. With layer-6 features, phone units have higher PartialSpoof evaluation frame EER than soft frames (11.18% versus 9.59%). In 500 retrospective PS-eval cases, augmentation reduces genuine-splice false alarms; synthetic-core misses increase at 5% PS-development FPR but not demonstrably at 1%. A 200-case Llama A follow-up retains the false-alarm reduction but does not establish a miss-rate increase. Constructed-case ranking can improve while fixed-threshold misses rise. Consistency gives no uniform gain, and encoder-layer rankings change across corpora. We report frame and event localization with three seeds, including word units, and treat PS evaluation as retrospective. Audit code, selected training scripts, and result summaries are available at https://github.com/mysxs/partial-spoof-diagnostics.
|
| 1960 |
VoiceWeaver: Staged Learning of Structured Controls for Expressive Speech and Sound-Event Generation
2610.04500
|
cs.SD
|
Xiaosu Su, Yun Cao, Yiping Ni, Xiaowei Yi |
Adding expressive and environmental controls to a speech generator requires learning heterogeneous attributes without losing earlier capabilities. VoiceWeaver addresses this problem with structured label prefixes and staged emotion-tone-event training in a sha...Adding expressive and environmental controls to a speech generator requires learning heterogeneous attributes without losing earlier capabilities. VoiceWeaver addresses this problem with structured label prefixes and staged emotion-tone-event training in a shared text-audio model. Replay-based distillation retains earlier predictions, while embedding decorrelation and attribute dropout regularize conditioning. Evaluation separates single-attribute correctness, joint emotion-event generation, and three-attribute correctness. Rounded to whole percentage points, emotion accuracy is 86% in Chinese and 83% in English. Chinese emotion-event joint accuracy is about 70%, versus 56% for mixed-task training and 48% for Ming-omni-tts; English joint accuracy is 66%. Three-attribute accuracy is 59% in Chinese and 55% in English on 500 samples each (57% pooled). Text error rates exceed those of the external TTS baselines. Audio samples are available at https://anonymous.4open.science/api/repo/VoiceWeaver1-4D88/file/index.html.
|
| 1961 |
"Spectral Harmony" Towards the Unification of Timbre and Speech: Resolution of the Young Mahler Schoenberg Dream
2610.04570
|
cs.SD
|
Yusei TAMURA, Shigekazu ISHIHARA, Ken ITO |
In this article, we present a method of spectral serialization envisioned by G. Mahler and A. Schoenberg between the late 19th and early 20th centuries that enables the selection of timbres rooted in musical structure. Based on information geometry, we evaluat...In this article, we present a method of spectral serialization envisioned by G. Mahler and A. Schoenberg between the late 19th and early 20th centuries that enables the selection of timbres rooted in musical structure. Based on information geometry, we evaluate the Wasserstein distance between polyphonic timbre spectra, as demonstrated in our previous paper, in terms of the second moment of spectral chords comprising three or more notes, expressed in squared hertz. By embedding these quantities in a Banach space, we analyze specific passages from works by R. Strauss, Mahler, A. Berg, and A. Webern from the perspective of timbre, showing how the same chord exhibits different affinities and divergences depending on orchestration. Taking the wah wah mute as an example, we propose a model that treats instrumental timbre and speech within the same Banach space, and we attempt to quantify the "vowel triangle" in Hz. These efforts constitute a partial but integrated solution to the historical three problems of Schoenberg. Serialism, Klangfarbenmelodie, and Sprechgesang which neither Mahler nor Schoenberg in their lifetimes, nor the postwar generation of composers including Nono, Boulez, and Stockhausen, was able to resolve fully. We aim to reconsider musical thought as a whole from the perspective of timbre and to open new possibilities for interpretation, performance, and composition. All tools are publicly available on GitHub, so that anyone can reproduce and extend these analyses.
|
| 1962 |
Prompt-Consistency Inference for Zero-Shot Flow-Matching Text-to-Speech Models
2610.04757
|
cs.SD
|
Vasily Zadorozhnyy, Can Goksen, Kazuhito Koishida, Dung Tran |
In recent years, flow-matching models have produced significant improvements in zero-shot text-to-speech synthesis. Conditioned on an audio prompt and text, these models learn a velocity field and generate speech by iteratively solving an ODE. During inference...In recent years, flow-matching models have produced significant improvements in zero-shot text-to-speech synthesis. Conditioned on an audio prompt and text, these models learn a velocity field and generate speech by iteratively solving an ODE. During inference, the solver evolves a single state spanning both the prompt and the region to be generated, although only the generated region is ultimately retained. The discarded prompt state, however, still matters; its intermediate values influence generation through the velocity field that couples the two regions. As sampling continues, this state can drift away from the prescribed conditional path, introducing a discrepancy into subsequent generation updates. Unlike the unknown generated trajectory, the prompt path is available in closed form from the reference audio and the initial noise. We exploit this observation with Prompt-Consistency Inference (PCI), a training-free rule that restores the prompt block to its analytic value before each velocity evaluation, while leaving the generated block unchanged. PCI improves speaker similarity and intelligibility across the evaluated flow-matching TTS backbones without additional network evaluations. Our ablation studies further show that PCI keeps post-step prompt discrepancies smaller and that corrections covering the later sampling stages recover much of the observed similarity gain.
|
| 1963 |
EchoChat: Structured Cognitive Reasoning in Empathetic Spoken Dialogue
2610.04826
|
cs.SD
|
Dingdong Wang, Shujie Liu, Yayue Deng, Yuxuan Hu, Yunrui Cai |
Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user's latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-re...Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user's latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to "superficially warm" but emotionally hollow interactions. In addition, since empathy relies on a multi-stage process with strong inter-step dependency, errors at any intermediate step can cascade through subsequent steps and lead to inappropriate responses, while existing training paradigms lack mechanisms to precisely localize and improve such errors. In this work, we propose EchoChat, a unified framework that reformulates empathetic spoken dialogue as a structured cognitive reasoning process integrating perception, mental-state reasoning, and response generation. To support this paradigm, we first construct EchoDialogue-400K, an acoustically rich dataset for multi-stage empathetic supervision. During the SFT stage, we strengthen acoustic grounding through proposed Acoustic-Anchored Attention (AAA). During the RL stage, we further introduce a novel stage-aware optimization objective with Step-Decomposed Credit Assignment (SDCA) to localize reasoning errors and mitigate cascaded error propagation. In addition, we introduce EchoEval, an expert-annotated benchmark for multi-dimensional empathy evaluation. Extensive experiments demonstrate that EchoChat achieves state-of-the-art performance in perception, reasoning, and response alignment. Project page: https://github.com/dingdongwang/EchoChat
|
| 1964 |
A Multidimensional Model for Quantifying Tonal Strength: A Continuous Framework of Analyzing Tonal Evolution Beyond Tonal-Atonal Binary Classification
2610.04871
|
cs.SDcs.MM
|
Yuliang Li, Nan Nan, Meilian Gu, Xiaohong Guan |
Tonality is a fundamental organizing principle of Western music. However, the gradual transition from common-practice tonality to atonality around the turn of the twentieth century cannot be captured by binary labels. To characterize intermediate tonal states ...Tonality is a fundamental organizing principle of Western music. However, the gradual transition from common-practice tonality to atonality around the turn of the twentieth century cannot be captured by binary labels. To characterize intermediate tonal states and trace this evolution, we propose a multidimensional framework for quantifying tonal strength, grounded in tonal theory and mechanisms underlying tonal dissolution. This framework operationalizes the global statistical regularity, temporal coherence, and local perceptual alignment of tonal organization through three complementary features: pitch-class distribution uniformity, temporal stability of tonal centers, and local tonal clarity, respectively. To validate the model, we introduce two new datasets: a binary dataset comprising 196 tonal and atonal pieces, and a historical MIDI corpus of 1,561 compositions by 21 composers spanning the Baroque era to the early twentieth century. The proposed model achieves high discriminative performance on the binary classification task (F_1=0.985), and quantitatively reveals a continuous historical trajectory from established tonality through progressive weakening to eventual dissolution. Computational analysis of representative composers further confirms significant declines in tonal strength throughout the careers of late-Romantic and modernist figures. Additionally, analysis of EMOPIA dataset reveals a weak but statistically significant positive correlation between higher tonal strength values and positive emotional valence. This work establishes a unified quantitative model for analyzing tonal evolution and provides a continuous measure for future behavioral and neurophysiological investigations of the relationships among tonal organization, musical emotion, and cognition.
|
| 1965 |
TS-SP: Learning Speaker-Preserving Representations in Audio Large Language Models
2610.04887
|
cs.SD
|
Junjie Li, Zheng Liang, Zhe Li, Tianchi Liu, Kong Aik Lee |
Audio large language models (ALLMs) can understand speech content, yet their ability to use speaker identity for verification remains limited. We propose TS-SP (Two-Stage Speaker Preservation), a parameter-efficient framework for learning speaker-preserving re...Audio large language models (ALLMs) can understand speech content, yet their ability to use speaker identity for verification remains limited. We propose TS-SP (Two-Stage Speaker Preservation), a parameter-efficient framework for learning speaker-preserving representations and making them accessible to an ALLM's language-model component. We instantiate and evaluate TS-SP on Qwen2.5-Omni-7B. First, we adapt the audio encoder with speaker identity supervision. We then freeze the adapted encoder and train the language model to compare speakers. Both stages use low-rank adaptation (LoRA), keeping the pretrained base weights fixed. On Vox1-O, TS-SP reduces the equal error rate (EER) from 7.01\% for the Paired Loss Adaptation Baseline to 4.37\%. EER remains within 4.31--4.79\% under unseen prompts. Cross-domain evaluation on CN-Celeb yields a similar EER to the baseline, but lower accuracy at the native decision threshold. These findings support two-stage adaptation for improving speaker verification on the evaluated backbone.
|
| 1966 |
NeuMark-Native: Robust Text-to-Speech-Native Watermarking Through Full Utilization of Neural Audio Codec Latent Space
2610.05215
|
cs.SD
|
Annan Wu, Wen-Chin Huang, Tomoki Toda |
Speech watermarking offers proactive traceability for synthetic speech, yet most existing models operate only after text-to-speech (TTS) synthesis by adding a watermark perturbation to the generated waveform. This post-hoc design leaves watermarking as an exte...Speech watermarking offers proactive traceability for synthetic speech, yet most existing models operate only after text-to-speech (TTS) synthesis by adding a watermark perturbation to the generated waveform. This post-hoc design leaves watermarking as an external step that can be omitted or bypassed and restricts the watermark to a shallow waveform representation. We propose NeuMark-Native, a TTS-native watermarking framework for neural codec-based synthesis. It embeds payload information into every generated codec-latent layer before waveform decoding, improving watermark persistence under downstream digital signal processing (DSP) and neural codec resynthesis. NeuMark-Native keeps the pretrained TTS model and the neural codec frozen, while optimizing only the watermark modules on generated codec tokens. Experiments on two corpora under 11 DSP attacks and 9 neural-codec attacks demonstrate robust watermark detection while preserving naturalness, intelligibility, and speech quality close to synthetic speech.
|
| 1967 |
SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
2610.05336
|
cs.SD
|
Junyan Jiang, Ruibin Yuan, Jiahao Pan, Wei Xue, Yike Guo |
Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequenc...Transcribing music into a human-readable score requires a coherent understanding of rhythm, harmony, melody, and form. Two obstacles limit this goal: annotated recordings are scarce, and accurate local predictions can still produce inconsistent musical sequences. We present SheetSage2, a unified music transcription framework that combines synthetic data, task-specific structured decoding, and autoregressive distillation. Automatically annotated MIDI, rendered into audio, provides scalable supervision across music understanding tasks. Task-specific structured decoders integrate complementary musical cues and their temporal dependencies to produce musically coherent scores. Autoregressive distillation further retains transcription accuracy without task-specific dynamic programming at inference. Across eight benchmark collections, a single SheetSage2-AR model exceeds the listed prior systems on 12 of 15 benchmark--metric pairs in our evaluation, substantially improving over SheetSage1 and surpassing task-specific models on several benchmarks. Model weights and inference code are publicly available.
|
| 1968 |
AudioGAR: Bridging Reconstruction and Generation in Latent Audio Generative Models
2610.05691
|
cs.SD
|
Xianghong Fang, Geeyang Tay, Wentao Ma, Tim G. J. Rudner, Dehan Kong |
Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents...Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents but deployed on generator-produced latents at inference time. Across diverse datasets and latent generative models, we observe clear reconstruction-generation gaps under both FD and FAD, showing that strong reconstruction quality does not necessarily translate into strong end-to-end generation quality. A natural remedy is to adapt the decoder on generation-produced latents, but generated latents lack correspondence with source audio and therefore cannot directly provide the paired supervision used for decoder fine-tuning. We introduce \textbf{AudioGAR}, which constructs intermediate latents by perturbing encoder latents and denoising them through the frozen latent diffusion model. These latents form a trajectory from reconstruction toward generation, with lower-noise latents retaining source correspondence and supporting paired decoder fine-tuning. We fine-tune only the codec decoder on these latents, while keeping the codec encoder and latent generative model frozen. When applied to AudioX, AudioGAR substantially improves generative performance. It requires only 1.5\% of the original training audio hours and 0.26\% of the original training cost.
|
| 1969 |
Smorph: Playable Sound Morphing with Diffusion Models
2610.06478
|
cs.SDeess.AS
|
Annie Chu, Hugo Flores Garc\'ia, Johannes Imort, Oriol Nieto, Bryan Pardo |
Sound morphing, generating intermediate sounds that transition from one sonic identity to another, can be a powerful tool for musical sound design. Existing diffusion-based morphing approaches entangle temporal structure and timbral identity, offering no mecha...Sound morphing, generating intermediate sounds that transition from one sonic identity to another, can be a powerful tool for musical sound design. Existing diffusion-based morphing approaches entangle temporal structure and timbral identity, offering no mechanism to hold one fixed while transforming the other. We present smorph, a training-free guidance framework that preserves how a sound behaves over time while transforming what the sound is, allowing users to morph, for instance from brass to strings at a fixed pitch. We demonstrate across three morphing modes: prompt-to-prompt, audio-to-prompt, and audio-to-audio. Evaluations across diverse datasets show that smorph effectively produces smooth morph trajectories while substantially improving temporal-structure and source preservation over baselines, albeit with more conservative target-ward transformation in some settings. In an exploratory case study, musicians found smorph trajectories to be expressive and playable, suggesting structural anchoring can serve as a productive constraint for instrumental interaction.
|
| 1970 |
AuraSE: Low-Hallucination Generative Speech Enhancement via Multimodal Flow Matching and Inference Policy Optimization
2610.06632
|
cs.SD
|
Yingda Shen, Yao Qian, Yuxuan Hu, Junan Zhang, Yuxiang Wang |
Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a f...Generative speech enhancement models can produce cleaner and more natural-sounding speech than conventional discriminative approaches, but may hallucinate by changing speech content or speaker identity, even with transcript conditioning. We present AuraSE, a flow-matching framework that addresses hallucination through complementary modality and inference designs. First, a double-stream-to-single-stream multimodal Diffusion Transformer (MMDiT) allows transcript and acoustic representations to interact while preserving a dedicated pathway for the degraded input. Second, we find that the best decoder configuration, governed by guidance scale, sampling temperature, and step count, varies substantially across utterances. This observation motivates Inference Policy Optimization (IPO), an online, on-policy preference optimization method. IPO generates multiple candidates from the current model under different inference configurations, ranks them with a multi-objective reward, and learns from their relative preferences. AuraSE-IPO ranks first on 11 of 12 metrics across the synthetic test sets and obtains the highest DNSMOS and blind-listening scores among the evaluated systems on the real DNS blind test set. At deployment, it uses a fixed $10$-step ODE decoder without classifier-free guidance (CFG) or per-utterance configuration search.
|
| 1971 |
HelixWorld: A Real-time Interactive Audio-Visual World Model
2609.38123
|
cs.SDcs.MM
|
Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang, Haoyuan Huang |
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic d...World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.
|
| 1972 |
Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction
2610.04235
|
cs.SD
|
Weizhi Liu, Yue Li, Hui Tian, Zhaoxia Yin |
Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can suppo...Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can support by embedding multi-bit identifiers directly into synthesized speech. Once released, however, speech may undergo heterogeneous learned transformations during distribution and editing, with reconstruction objectives that can preserve speech utility while affecting watermark recoverability differently. We find that no single watermark carrier remains consistently reliable across reconstruction models, as its survival depends jointly on the embedded structure, reconstruction mechanism, and observation representation. To this end, we propose Thrive, a multi-bit generative speech watermarking framework for modern autoregressive TTS, covering both discrete-token and continuous-representation generation under reconstruction attacks. Specifically, Rise synchronizes watermark injection into intermediate representations with its continued integration into subsequent generation, while Care combines waveform and spectral experts using bit-wise reliability selection. Experiments on both autoregressive paradigms show that Thrive preserves synthesis fidelity, achieves 87.6% average recovery accuracy under reconstruction attacks, and supports source attribution over candidate sets of up to 10,000 identities.
|
| 1973 |
Revisiting Frame-Wise Saliency for Audio Moment Retrieval
2610.05737
|
cs.SDeess.AScs.MM
|
Tatsuya Munakata, Hokuto Munakata |
This paper revisits frame-wise saliency for audio moment retrieval (AMR). We show that the frame-wise saliency sequence, conventionally used only as an auxiliary output in DETR-based AMR models, can itself serve as an effective source of moment predictions. We...This paper revisits frame-wise saliency for audio moment retrieval (AMR). We show that the frame-wise saliency sequence, conventionally used only as an auxiliary output in DETR-based AMR models, can itself serve as an effective source of moment predictions. We convert the saliency sequence into ranked moments using a simple SED-inspired segmentation rule with no learned parameters, enabling moment retrieval directly from frame-wise temporal information. On the CASTELLA dataset, saliency-based prediction consistently outperforms decoder-based prediction from the same model across all 18 runs of QD-DETR and CG-DETR. For QD-DETR, simply replacing the inference output improves R1@0.7 from 21.0 to 36.1. The advantage remains 7-15 points when the two outputs are evaluated at their independently selected best epochs. The same tendency extends to TaskWeave and UVCOM, whereas TR-DETR shows the opposite behavior, suggesting that how saliency construction may matter. The performance gap is especially pronounced for short moments: for queries whose annotated moments average at most 2 s, R1@0.7 improves from 6.3 to 27.8 with QD-DETR. Decoder supervision nevertheless benefits saliency-based prediction, indicating that its role during training differs from the utility of its inference output.
|
| 1974 |
Character Identity is not Speaker Identity: KyaraBench and KyaraEmbed for Character Verification
2610.06013
|
cs.SDeess.AS
|
Joonyong Park, Jerry Li |
A dubbed character keeps its identity while the voice actor changes, so character identity and speaker identity are distinct properties of one recording, yet speaker verification measures only the latter. To address this gap, we propose KyaraBench, a benchmark...A dubbed character keeps its identity while the voice actor changes, so character identity and speaker identity are distinct properties of one recording, yet speaker verification measures only the latter. To address this gap, we propose KyaraBench, a benchmark that scores character voice directly instead of speaker voice, built from 85 human-audited identities in a dubbed anime corpus. It poses two challenging conditions: one that swaps the performer under a fixed character, and one that fixes the performer under changing characters. Listening studies with 78 participants provide human reference scores for cross-performer verification and same-actor discrimination. Speaker-verification baselines show increased errors under these character-specific conditions. We then train KyaraEmbed, a compact encoder using multilingual character supervision, same-actor negatives, and a language-alignment term. The model achieves the best performance on all character-specific conditions in the main comparison. We release the benchmark, protocol, and encoder publicly.
|
| 1975 |
Ensemble-Based Perceptual Audio Quality Assessment with Confidence Intervals
2610.06569
|
cs.SDeess.AS
|
Pablo M. Delgado, Andreas Brendel, Konstantin Schmidt, J\"urgen Herre |
Objective audio quality metrics typically provide point estimates, whereas listening tests yield score distributions from which mean opinion scores (MOS), confidence intervals (CIs), and significance decisions are derived. We propose a lightweight intrusive me...Objective audio quality metrics typically provide point estimates, whereas listening tests yield score distributions from which mean opinion scores (MOS), confidence intervals (CIs), and significance decisions are derived. We propose a lightweight intrusive metric that combines a PEAQ-style perceptual front-end (ITU-R BS.1387) with a bagging ensemble of regressors. Calibration with subjective data aligns ensemble outputs with listener scores. The resulting item-dependent score distributions enable uncertainty assessment, panel-size-matched CIs, and identification of less conclusive predictions. Calibration improves agreement with subjective distributions and CI coverage across all evaluated datasets while preserving MOS accuracy. Using only 11 fixed PEAQ features, low-capacity regressors, and public training data, the method performs comparably to more data-intensive end-to-end approaches. Its output can support uncertainty-aware assessment and target listening tests towards uncertain conditions. The distributions also enable approximate pairwise comparisons, but not yet reliable significance inference.
|
| 1976 |
A mathematical model of the vowel space
2111.00868
|
cs.SDeess.AS
|
Fr\'ed\'eric Berthommier (GIPSA-PCMD) |
The articulatory-acoustic relationship is many-to-one and non linear and this is a great limitation for studying speech production. A simplification is proposed to set a bijection between the vowel space (f1, f2) and the parametric space of different vocal tra...The articulatory-acoustic relationship is many-to-one and non linear and this is a great limitation for studying speech production. A simplification is proposed to set a bijection between the vowel space (f1, f2) and the parametric space of different vocal tract models. The generic area function model is based on mixtures of cosines allowing the generation of main vowels with two formulas. Then the mixture function is transformed into a coordination function able to deal with articulatory parameters. This is shown that the coordination function acts similarly with the Fant's model and with the 4-Tube DRM derived from the generic model.
|
| 1977 |
Re-purposing Multimodal Large Language Models for Audio-Text Retrieval
2602.18010
|
cs.SD
|
Jilan Xu, Carl Thom\'e, Danijela Horak, Weidi Xie, Andrew Zisserman |
Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentally limited by the capacity of small-scale encoders. Specifically, the text encode...Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentally limited by the capacity of small-scale encoders. Specifically, the text encoders struggle to understand complex queries that require reasoning or world knowledge. In this paper, we propose AuroLA, a novel contrastive language-audio pre-trained model that re-purposes Multimodal Large Language Models (MLLMs) as a unified backbone for audio-text retrieval. Specifically, we make the following contributions: (i) we construct a scalable data pipeline that curates diverse audio from multiple sources and generates multi-granular captions, ranging from long descriptions to structured tags, via automated annotation; (ii) we adapt an MLLM for retrieval by prompting it to summarise the audio/text input and using the hidden state of a special token as audio/text embeddings. (iii) extensive experiments demonstrate that AuroLA consistently outperforms state-of-the-art dual-encoder models, including the recent PE-AV. This validates the effectiveness of MLLM as a unified backbone for audio-text retrieval.
|
| 1978 |
AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration
2602.19409
|
cs.SD
|
Henry Zhong, J\"org M. Buchholz, Julian Maclaren, Simon Carlile, Richard F. Lyon |
Manual annotation and clustering of audio datasets is labour intensive. We introduce AuditoryHuM, a training-free framework for the unsupervised discovery and clustering of auditory scene labels using human-Multimodal Large Language Model (MLLM) collaboration....Manual annotation and clustering of audio datasets is labour intensive. We introduce AuditoryHuM, a training-free framework for the unsupervised discovery and clustering of auditory scene labels using human-Multimodal Large Language Model (MLLM) collaboration. Leveraging MLLMs, our framework generates contextually relevant labels for audio data. To ensure label quality and mitigate hallucinations, zero-shot learning (Human-CLAP) quantifies the alignment between generated text labels and raw audio. A targeted human-in-the-loop intervention, refines only the lowest aligned pairs. The discovered labels form an interpretable alignment vector to group audio into cohesive clusters. The framework was evaluated across three auditory scene datasets (ADVANCE, AHEAD-DS, and TAU 2019), achieving a 96.1\% reduction in human labour for ADVANCE during testing. Downstream models trained on our clusters exhibit enhanced classification accuracy vs the baseline (0.79 to 0.84) in clean acoustic environments, though performance scales down in dense, complex soundscapes. The project page: \url{https://github.com/Australian-Future-Hearing-Initiative}
|
| 1979 |
Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement
2603.02641
|
cs.SD
|
Szu-Wei Fu, Rong Chao, Xuesong Yang, Sung-Feng Huang, Ryandhimas E. Zezario |
Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curatio...Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Model weights are available for download at: https://huggingface.co/nvidia/RE-USE.
|
| 1980 |
One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications
2606.25621
|
cs.SD
|
Szu-Wei Fu, Rong Chao, Xuesong Yang, Sung-Feng Huang, Ante Juki\'c |
Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhancement models for each scenario. In this paper, we propose a one-for-all, real-time universal speech enhancement model that provides explicit contr...Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhancement models for each scenario. In this paper, we propose a one-for-all, real-time universal speech enhancement model that provides explicit control over both algorithmic and computational latency. Algorithmic latency is flexibly adjusted via configurable look-ahead frames. To avoid learning inefficiency caused by varying padding configurations, we introduce parallel convolutional layers corresponding to different look-ahead settings. Computational latency is controlled through an early-exit mechanism, enabling inference at different network depths. To narrow the performance gap between specialized and flexible models, we propose a two-stage training strategy with a shared-to-multiple decoder transition. Overall, the proposed framework enables a single model to be deployed across diverse latency budgets without retraining separate models. Model weights are available for download at: https://huggingface.co/nvidia/Real-time_RE-USE
|
| 1981 |
Rubric-Based Optimization for Text-to-Music Generation
2610.03589
|
cs.SD
|
Ping Wang, Guang Yang, Shao-Rong Su, Junkai Wu, Pang Wei Koh |
Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as train...Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.
|
| eess.AS 11 papers | ||||
| 1982 |
MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations
2610.04180
|
eess.AS
|
Xiaomin He, Vishal Choudhari, Tristan J. Spratt, Aarya Raghavan, Richard T. Lee |
Auditory attention decoding (AAD) is often evaluated on static, simplified speech scenes that poorly match everyday listening. We introduce MOV-AAD, a large-scale dataset for studying auditory attention under moving, naturalistic conversations. MOV-AAD combine...Auditory attention decoding (AAD) is often evaluated on static, simplified speech scenes that poorly match everyday listening. We introduce MOV-AAD, a large-scale dataset for studying auditory attention under moving, naturalistic conversations. MOV-AAD combines 64-channel EEG with synchronized physiological recordings, including eye tracking, respiration, galvanic skin response, heart rate, peripheral oxygen saturation, body temperature, body motion, and photoplethysmography, enabling analysis of cross-modal neural and physiological markers of attention and listening effort. This dataset uses a more ecologically valid listening paradigm with dynamically moving conversational speech sources and behavioral measures of attentional engagement. MOV-AAD supports research on robust AAD in realistic spatial dynamics, multimodal attention modeling, listening effort, and intersubject neural responses, providing a resource for benchmarking selective auditory attention in naturalistic listening.
|
| 1983 |
Influence of Geometrical Acoustic Simulator Complexity on a Trained Multisource Localizer
2610.05055
|
eess.AS
|
Fabian Staub, Nils Meyer-Kahlen, Thomas Deppisch, Sergio de las Heras, Florian Klein |
Learning-based acoustic sound source localization and detection (SSLD) requires large labeled datasets covering diverse acoustic conditions. Since obtaining measured data is costly, training commonly uses simulated data, while practical devices must operate un...Learning-based acoustic sound source localization and detection (SSLD) requires large labeled datasets covering diverse acoustic conditions. Since obtaining measured data is costly, training commonly uses simulated data, while practical devices must operate under real-world conditions. However, higher simulation complexity increases data-generation cost, and the complexity required for reliable generalization remains unclear. In practice, SSLD methods often rely on efficient geometrical acoustic simulation, typically the image-source method. This work investigates how geometrical acoustic simulation complexity affects the real-world performance of a common multi-source SSLD model. We train the model with simulators ranging from anechoic conditions to high-order image-source simulations, optionally including diffuse reverberation, array simulation, and randomized image-source positions, and evaluate them on three measurement-based test datasets. Results show that anechoic simulation is insufficient, while medium-complexity image-source simulations already provide strong real-world performance. Further increases in complexity yield only marginal gains, with the best performance obtained using the highest tested image-source order, array simulation, and randomized image-source positions. For this configuration, measured-domain performance approaches within-domain simulated performance, suggesting limited benefit from further increasing simulation complexity. These findings show how the complexity-performance trade-off can be exploited in future data-driven SSLD, and highlight image-source randomization as an efficient way to improve generalization.
|
| 1984 |
UltraM2M: Leveraging Text Transcripts and Mixture Constraints for Weakly-Supervised Speech Enhancement
2610.05155
|
eess.AS
|
Liu, Jiachen, Wu, Fulin, Wang |
We propose UltraM2M, a weakly-supervised speech enhancement algorithm building upon the recent unsupervised mixture-to-mixture (M2M) algorithm. M2M realizes unsupervised speech enhancement by training deep neural networks on a set of real-recorded noisy-reverb...We propose UltraM2M, a weakly-supervised speech enhancement algorithm building upon the recent unsupervised mixture-to-mixture (M2M) algorithm. M2M realizes unsupervised speech enhancement by training deep neural networks on a set of real-recorded noisy-reverberant multi-channel mixture signals to estimate target speech and non-target signals. The two estimated signals are penalized by a so-called mixture-constraint (MC) loss, which constrains them to reconstruct the observed mixture signals. Although shown to be effective, the mixture constraint may be too weak to enable sufficient noise reduction. To deal with this, UltraM2M extends M2M by further leveraging text transcripts of real-recorded mixtures to design an automatic speech recognition (ASR) loss to penalize the estimated speech signal. The ASR loss can be viewed as a form of weak supervision that could help unsupervised enhancement. Evaluation results on the CHiME-4 dataset show the effectiveness of UltraM2M.
|
| 1985 |
Speaker Tracking: Segment-online Multi-talker Organization with a Varying Number of Speakers
2610.05693
|
eess.AS
|
Vahid A. Kalkhorani, Daniel Wong, Jacob Donley, Ashutosh Pandey, Buye Xu |
Speaker tracking is the task of separating and following multiple speakers over time. It must address overlapped speech, speech onset and offset, talker identity, and time-varying speaker count. We propose a segment-online, modular framework for single- and mu...Speaker tracking is the task of separating and following multiple speakers over time. It must address overlapped speech, speech onset and offset, talker identity, and time-varying speaker count. We propose a segment-online, modular framework for single- and multi-channel speaker tracking. The proposed system first performs speaker separation in each segment and computes speech activity via voice activity detection (VAD). To generate speaker tracks over time, we introduce a two-stage sequential organization strategy: Overlap-based stitching for continuous grouping and memory-based speaker verification for discontinuous grouping. For speaker separation, we employ complex spectral mapping to estimate the real and imaginary spectrograms of underlying speakers. The proposed system achieves state-of-the-art segment-online tracking performance on the LibriCSS and AMI datasets. Our framework significantly reduces diarization error rate (DER) and concatenated minimum-permutation word error rate (cpWER) compared to other methods.
|
| 1986 |
Enhancing Pathological Speech through Articulatory Bottlenecks
2610.05944
|
eess.AS
|
Ail\'in Pollio San Pedro, Olivier Perrotin, Thomas Hueber |
Dysarthric speech reconstruction (DSR) typically relies on linguistic or phonetic representations extracted from impaired speech to generate a more intelligible waveform. We investigate a complementary approach that instead intervenes in a representation relat...Dysarthric speech reconstruction (DSR) typically relies on linguistic or phonetic representations extracted from impaired speech to generate a more intelligible waveform. We investigate a complementary approach that instead intervenes in a representation related to speech production. Starting from a neural analysis--synthesis framework, we introduce a residual mapper that modifies an articulatory-aligned latent space while preserving speaker and prosodic information. The mapper is pretrained on parallel synthetic healthy and artificially dysarthric speech, then adapted to natural dysarthric speech using phoneme-guided and adversarial objectives. We further compare the articulatory bottleneck with a dimension-matched unsupervised representation to assess the benefit of explicit articulatory supervision.
|
| 1987 |
Preemptive defense against re-identification attacks on voice anonymization via adversarial perturbation
2610.06074
|
eess.AS
|
Michele Panariello, Yibo Bai, Massimiliano Todisco, Nicholas Evans |
Voice anonymization (VA) is used to conceal the voice identities in speech data. Performance is usually estimated using automatic speaker verification (ASV) as a proxy to judge the capability of an attacker to re-identify speakers after anonymization. The defe...Voice anonymization (VA) is used to conceal the voice identities in speech data. Performance is usually estimated using automatic speaker verification (ASV) as a proxy to judge the capability of an attacker to re-identify speakers after anonymization. The defender anonymizes test utterances; the attacker compares them to equally anonymized reference utterances to infer voice identity, with degraded ASV performance indicating successful anonymization. Thus, the defender's protection is one-sided and only applied to test utterances via VA. Though unexplored so far, there is an opportunity to enhance anonymization by protecting reference utterances too. We present a new, preemptive VA paradigm: reference utterances are protected using adversarial noise to degrade their potential to infer voice identity. Preemptive protection does not degrade the perceived quality of reference utterances and is independent of the specific ASV system used for identity inference. By combining preemptive protection and regular test-side VA, ASV equal error rates can be increased from 14\% to 45\% (near-perfect privacy).
|
| 1988 |
Lyric: Wave-Domain Computing for Efficient Spoken-Digit Recognition
2610.06433
|
eess.AS
|
Jeeven Balasubramaniam, Nakul Garg |
Low-power speech recognition requires reducing audio acquisition and feature-processing costs as well as neural inference. We present Lyric, a speech-recognition front end that uses wave-domain computing to extract features before digitization. Our proposed sy...Low-power speech recognition requires reducing audio acquisition and feature-processing costs as well as neural inference. We present Lyric, a speech-recognition front end that uses wave-domain computing to extract features before digitization. Our proposed system uses passive acoustic resonators to separate speech by frequency and supplies their slowly varying envelopes to a compact temporal neural network. This moves spectral filtering into the acoustic structure, reducing the data and processing needed for recognition. We build a prototype and evaluate spoken-digit recognition on a speakerdisjoint AudioMNIST split across nine classifier families. The front end acquires 32 times fewer scalar samples than a 16-kHz waveform. In the EdgeSpeechNet-A comparison on a Raspberry Pi 4, we reduce preprocessing-plus-inference latency and estimated processor energy per inference by 98.6%, with a 3.58-percentage-point decrease in accuracy to 95.70%.
|
| 1989 |
When Layer Selection Misleads Speech Depression Detection
2610.06465
|
eess.AS
|
Paula A. Perez-Toro, David Gimeno-G\'omez, Daniel R\"uckert, Andreas Maier |
Pretrained speech representations are increasingly used for depression detection, but selecting the best encoder layer on the same data used for evaluation biases reported performance. On DAIC-WOZ, with repeated nested cross-validation across five deep encoder...Pretrained speech representations are increasingly used for depression detection, but selecting the best encoder layer on the same data used for evaluation biases reported performance. On DAIC-WOZ, with repeated nested cross-validation across five deep encoder families, naive best-of-25-layer probing inflates AUC by up to $0.09$. On shuffled labels over the \emph{real} latent representations, naive best-of-25 still reaches AUC $0.59$ versus $0.50$ under nested selection, and this bias grows with the number of probed layers and with smaller samples, isolating selection, not signal, as the cause. The effect replicates on a second clinical corpus and language (Androids, Italian). Under leakage-free selection, a fair comparison across representation families identifies a compact, task-aligned affect--prosody model as competitive with substantially more complex SSL, deep-learning, audio-LLM, and codec-based approaches, while using $\sim\!10^{3}\times$ fewer trainable downstream parameters. Extending the analysis to individual PHQ-8 symptoms shows that the same selection bias persists at this finer-grained level. We argue nested model-selection protocols should be standard when probing encoder layer representations on small clinical speech datasets.
|
| 1990 |
Multiplexing Neural Audio Watermarks with Adaptive Routing
2511.02278
|
eess.AS
|
Zheqi Yuan, Yucheng Huang, Guangzhi Sun, Zengrui Jin, Chao Zhang |
Audio watermarking supports speech authenticity verification. We study whether combining heterogeneous watermarks can retain at least one detectable provenance signal when their failure modes differ. This Any-survival objective concerns complementary evidence ...Audio watermarking supports speech authenticity verification. We study whether combining heterogeneous watermarks can retain at least one detectable provenance signal when their failure modes differ. This Any-survival objective concerns complementary evidence retention, not simultaneous survival of all constituent marks or recovery of all payloads. To our knowledge, this is the first systematic study of neural audio watermark multiplexing, with a scoped benchmark covering five released systems, parallel and sequential baselines, and 14 evaluation conditions. The benchmark shows complementary failure modes, but naive composition does not reliably turn them into system-level robustness under the Any-survival objective. We therefore formulate multiplexing as watermark allocation and study perceptual-adaptive time-frequency multiplexing (PA-TFM), a training-free routing method, and MaskNet, a learned time-domain router for separately trained systems with native detectors. We use the five-system scoped benchmark to characterize multiplexing behavior and the AudioSeal-PerTh pair as a representative heterogeneous pair for adaptive watermark routing. Compared with direct parallel composition, MaskNet improves average TPR@1%FPR from 0.75 to 0.88 and SNR from 15.20 dB to 25.36 dB; compared with PA-TFM, it gives a clear system-level robustness gain under the Any-survival objective while retaining similar high-fidelity behavior. A SpeechTokenizer-aware case study further improves average TPR@1%FPR to 0.91 and SpeechTokenizer robustness from 0.20 to 0.60, showing that channel-adapted constituents can enter the same routing framework.
|
| 1991 |
DASHA: Diarization-Aware Speech-LLM ASR for Healthcare in Code-Switched Indic Doctor-Patient Conversations
2603.06373
|
eess.AS
|
S\'everin Baroudi, Yanis Labrak, Shashi Kumar, Joonas Kalda, Sergio Burdisso |
Extracting patients' medical conditions from recorded primary-care conversations requires knowing who said what, which is difficult when speech is far-field, spontaneous, overlapping and code-switched. We study this problem on DISPLACE-M, a corpus of Hindi--En...Extracting patients' medical conditions from recorded primary-care conversations requires knowing who said what, which is difficult when speech is far-field, spontaneous, overlapping and code-switched. We study this problem on DISPLACE-M, a corpus of Hindi--English (Hinglish) doctor--patient conversations, and propose DASHA, a modular cascade of diarization, speaker-attributed ASR (SA-ASR) and LLM-based extraction. An end-to-end neural diarization with vector clustering (EEND-VC) system built on a multilingual w2v-BERT~2.0 encoder reduces DER from 9.31% to 7.76%. A Qwen3-ASR model adapted to Hindi and the target domain, with Devanagari normalization, reduces tcpWER from 26.78% to 19.27%, and to 18.59% with optional LLM correction. A module-swap analysis shows that better diarization improves extraction only when paired with the adapted ASR. On code-switched words, the multilingual encoder roughly halves speaker confusion. DASHA sets a new state of the art for two-speaker, code-switched Hindi SA-ASR and ranked first among 25 participants in the DISPLACE-M challenge.
|
| 1992 |
Interpretable Destination-Aware Synthesizer Modulation Recovery
2610.00642
|
eess.AS
|
David Liu, Giulio Cengarle, David Cooper, Mark Vinton, Haici Yang |
Synthesizer programming is a challenging task, particularly in configuring modulation: the time-varying control of parameters such as pitch, filter cutoff, and oscillator level by modulators such as Low-Frequency Oscillators (LFOs). Prior work has shown that m...Synthesizer programming is a challenging task, particularly in configuring modulation: the time-varying control of parameters such as pitch, filter cutoff, and oscillator level by modulators such as Low-Frequency Oscillators (LFOs). Prior work has shown that modulation can be reconstructed from the clean audio through curve matching, yet the recovered parameters are often untransferable to modern synthesizers. We present an interpretable destination-aware modulation and waveform recovery pipeline that mirrors how musicians often recreate sounds on a synthesizer. First, the system identifies the modulated destinations; then it recovers the LFO shape associated with each destination, and estimates the oscillator waveform to match the source timbre. The training is supported by a differentiable synthesizer that incorporates a noise oscillator, and more modulation options than previous work. We carefully train the models for optimal perceptual quality, with extensive exploration of perceptual losses and adversarial training. Through objective and subjective evaluation, we show that predicting the destinations correctly is essential, our model excels in the modulation focused inverse synthesis tasks, and our Gammatone loss and CQT discriminator significantly outperform the traditionally-used MSS loss in improving perceptual quality. We provide audio samples from the subjective test.
|